Year over year, performance per engineer rose 117%.
Quarter over quarter, the change can’t be told apart from zero.
Cloudflare · Vercel · OpenAI · Google · Meta · Microsoft
Two notes before the data.
We sell a measurement layer for engineering work, so we have a commercial interest in the conclusion that such work can be measured at all. This report makes only descriptive and correlational claims: it reports what moved, not why. Sample boundaries and known limitations are in Section 7 and Appendix A.
Every quarter is recomputed. Re-scoring, contributor-role reclassification, and late-arriving commit attribution move historical quarters between snapshots, so this edition recomputes the whole window rather than carrying figures forward. On this snapshot the Q1 2025 → Q1 2026 open-cohort change reads +112%, where No. 01 printed +116%. Read the tables here as one internally consistent series; comparisons should be drawn within a snapshot, not line by line across editions.
What one more quarter of data did to the previous reading.
The two figures sit on different baselines — No. 01 measured Q1 2025 to Q1 2026, this edition measures Q2 2025 to Q2 2026 — and on different snapshots. They are not a quarter’s worth of change, and subtracting them would not produce one. What did change is the shape: the level shift that arrived in Q1 2026 held, and the quarter that followed it is flat enough that its confidence interval contains zero.
The horizon you pick decides the answer you get.
Mean ETV per qualifying engineer moved +13.6% between Q1 2026 and Q2 2026 on the open cohort, with a 95% interval of [−3.5%, +32.7%]. That interval contains zero. On the fixed panel of the same 388 engineers the same quarter moved +1.4%.
The year-over-year reading is not ambiguous in the same way: +117% open and +81% fixed, both intervals clear of zero. That distinction is load-bearing for everything below — the data supports a year-over-year statement about scored output per engineer, and does not support a quarter-over-quarter one.
The Shape
A level shift in Q1 2026, then a flat quarter on top of it.
The open cohort takes every engineer qualifying in the quarter; the fixed panel takes the 388 present in all six, which removes cohort composition as a source of trend at the cost of survivorship bias. Across the first four quarters the two series track each other closely and both stay near flat. The level shift arrives in Q1 2026 and is present in both — the pattern that rules out pure composition, because it survives holding the population constant.
Q2 2026 is where the two separate. The open cohort adds +13.6% while the fixed panel adds +1.4%, so roughly a tenth of the quarterly move survives once the population is held fixed. Over the year the correspondence is much stronger, +117% open against +81% fixed.
The bootstrap intervals behind these lines are wide enough that single-quarter jumps should be read cautiously; the multi-quarter slope is the signal of interest. It matters particularly here, because the headline quarter-over-quarter interval contains zero.
The Composition
The mix rotates out of maintenance and into tests.
Output splits across the five categories the scorer emits. Classification is file-scoped, so a commit that adds a feature, extends its test suite, and updates a README splits across Features, Tests, and Docs rather than booking whole to one category.
Over the full window Maintenance falls from 29.7% of scored output to 17.3% and Tests rises from 18.7% to 26.8%, now the second-largest category. In the last quarter alone Tests moves +6.1pp against Features at −2.9pp; the other three move less than 2.4pp between them.
| Share of output | Q1 '25 | Q2 '25 | Q3 '25 | Q4 '25 | Q1 '26 | Q2 '26 | Window Δ | QoQ Δ |
|---|---|---|---|---|---|---|---|---|
| Features | 32.5% | 33.6% | 34.3% | 32.6% | 39.0% | 36.1% | +3.6pp | −2.9pp |
| Maintenance | 29.7% | 28.9% | 26.9% | 24.4% | 19.4% | 17.3% | −12.4pp | −2.1pp |
| Tests | 18.7% | 20.1% | 20.1% | 23.2% | 20.7% | 26.8% | +8.1pp | +6.1pp |
| Docs | 4.9% | 4.5% | 4.6% | 3.9% | 3.2% | 4.4% | −0.5pp | +1.2pp |
| Fixes | 14.2% | 12.8% | 14.1% | 15.9% | 17.7% | 15.3% | +1.1pp | −2.4pp |
These are shifts in composition, not statements about intent, and a falling share is not falling output. Maintenance is the clearest case: its share fell 12.4pp across the window while total output per engineer rose 140%, so absolute maintenance output per engineer per month rose +40%, from 0.25 to 0.35 ETV. Less of a larger total, not less maintenance.
The Spread
Every organization is positive on both horizons — in a different order.
Mean ETV per qualifying engineer · % Δ
Microsoft leads the year at +133.5% while OpenAI leads the quarter at +24.5%. Meta is last on the annual horizon at +41.7%, and Cloudflare last on the quarterly one at +3.6%. Reading the two together separates organizations on a sustained annual trajectory from those whose gain is concentrated in the most recent quarter.
OpenAI · held out of the scale above
Its Q2 2025 cohort of 18 SWEs sits below the 20-SWE sample floor, so the year-over-year change is measured from a baseline this report does not consider reportable on its own. Rebased to Q3 2025, its first at-floor quarter, the change to Q2 2026 is +199%.
Quarter over quarter, on a baseline at floor: +24.5%
Cohort size fell at all six organizations between Q1 2026 and Q2 2026. Where a cohort contracts and mean output per remaining engineer rises, the aggregate can move without any individual engineer changing output — the fixed panel above is the control for exactly that.
Per-organization work-mix panels, the quarter-by-quarter trajectories, and the reason Cloudflare’s documentation share runs several times the aggregate are in the white paper (Sections 5 and 5.1, Figures 4a and 4b).
Fewer commits, each carrying more.
The aggregate decomposes into frequency and intensity: commits per engineer and performance per commit. Quarter over quarter they move in opposite directions — commits per engineer −13.4% against performance per commit +31.2%. Year over year both rise, by +12.7% and +92.8%.
Engineers in the qualifying cohort landed fewer merged commits than in Q1 2026. Consistent with larger units of work, and equally consistent with a shift in commit-granularity practice such as heavier squashing before merge.
Each commit carried more scored complexity. Over the full window per-commit intensity is up +122% — the component that accounts for most of the headline.
Note: unlike No. 01, the two components multiply to the headline exactly, because the decomposition is an identity rather than an approximation — all three statistics are pooled quarterly ratios over the same commit set, so the commit count cancels. Falling frequency times rising intensity gives 0.866 × 1.312 = 1.136, the +13.6% above. What the identity cannot do is attribute the movement: both factors are cohort-level ratios, so either can move because the cohort changed rather than because any individual engineer did.
Same engine as No. 01, reported on five categories instead of three.
Five sub-scores, one scalar. For each merged commit the engine produces one sub-score per work category — Features, Maintenance, Tests, Docs, Fixes. Each is assembled per file and summed across the commit, so classification is file-scoped. Their sum is the Engineering Throughput Value. The scalar is unchanged from No. 01; only the reported breakdown moved from three categories to the scorer’s native five.
Sample. 137,129 qualifying commits to 65 public repositories across six organizations over a 78-week window. A contributor is an engineer when the role classifier assigns Software Engineer and they have commit activity in at least 10 of those 78 weeks; they join a given quarter’s cohort by authoring at least one scored commit in it. Bots are excluded by pattern match on email and display name.
Organization selection. Not a random sample. Cloudflare, Vercel, and OpenAI were chosen because each makes frequent public claims about AI productivity gains in its engineering organization; Google, Meta, and Microsoft as incumbents of substantially larger scale with strong public reputations for engineering talent density. Cross-organization comparisons reflect that.
Sample floor. Per-organization results are reported only where that organization’s qualifying cohort in a quarter reaches 20 engineers. Quarters below the floor are suppressed; where a below-floor quarter is nonetheless used as a comparison baseline, as with OpenAI above, it is marked in place.
Uncertainty. 95% intervals by bootstrap over 1,000 seeded iterations, resampling within each quarter. Quoted percentage-change intervals are endpoint intervals: they reflect sampling error in the endpoint quarter only and understate total uncertainty in the change, since the baseline quarter carries its own.
Temporal alignment. Calendar quarters. Commits are attributed to the quarter in which they were authored, not the quarter they merged, so long-lived branches distribute across the window rather than concentrating at merge. Q2 2026 closed 2026-07-01, nearly eight weeks before this snapshot. Q3 2026 was still open at publication and is excluded rather than reported as a partial quarter.
Limitations. Public repositories only. No view into private work, code review depth, incident response, planning, or mentorship — anything not captured as a merged commit to an in-scope repository. Monorepo versus polyrepo workflow, squash-merge policy, and public-versus-private code mix are not controlled.
Interpretation. The window coincides with broad production adoption of AI coding assistants across these organizations. The study does not quantify what share of the level shift, if any, is attributable to that adoption, and a single quarter of deceleration in a noisy series is not yet evidence of a plateau.
12 pages · five-way work classification · per-org detail in Section 5.1 · 65 repositories enumerated in Appendix A
Key findings
Year over year, scored output per engineer is up 117% and the finding survives holding the engineer population constant. Quarter over quarter, the change cannot be told apart from zero. The mix rotated into tests, the quarter’s movement came entirely from per-commit intensity, and whether Q1 2026’s level shift was a step or the start of a slope still needs another quarter to answer.
Whether any of this output was aimed at the right things is a separate question that commit history cannot answer. Navigara’s Alignment concept handles it by connecting to Jira or Linear. That’s outside this study.
Cloudflare · Vercel · OpenAI · Google · Meta · Microsoft
The series
Older edition
No. 01 · Q1 2026
Engineering performance measured at scale
+116% · 30 April 2026
Read the editionNext edition
Published quarterly. The next edition revisits the same cohort one quarter on.