Thirty days ago, we moved our autonomous agent fleet from Claude Code to Codex. We made a public promise before the measurement window opened: count completed work, keep the denominators, report missing telemetry as missing, and publish the result even if Codex lost.
Codex did not win.
That sentence needs a boundary around it. This was not a controlled model benchmark, an A/B test, or a universal ranking. It was a comparison of two consecutive operating periods in one production fleet. The workloads and operating policies changed between them. Within that experiment, the directly observable outcome and quality measures favored the frozen Claude cohort.
The defensible verdict is mixed evidence, no Codex win.
The two measured periods
We froze the Claude baseline before the migration. It includes tasks started from 2026-07-08T16:55:55Z inclusive through 2026-08-07T16:55:55Z exclusive. Outcomes were observed until 2026-08-07T22:55:55Z, giving tasks started near the end a fixed six-hour grace period.
The Codex period includes tasks started from 2026-08-07T22:35:04Z inclusive through 2026-09-06T22:35:04Z exclusive. Its outcome cutoff was 2026-09-07T04:35:04Z, using the same six-hour grace.
The five hours, 39 minutes, and nine seconds between the cohorts were migration and burn-in. We excluded that interval from both outcome cohorts instead of assigning migration failures to one side.
Provider attribution follows the executor in place when a task started, not the executor present when it finished. That rule matters for tasks crossing a boundary. We froze it before counting.
The full definitions, boundary reconciliation, role tables, extraction notes, and reproduction command are in the 30-day measurement report. The machine-readable checkpoint preserves the final values.
What we observed
| Measure | Codex period | Frozen Claude period |
|---|---|---|
| Completed / terminal tasks | 574/631 (90.97%) | 1,286/1,317 (97.65%) |
| Strict first-pass QA proxy | 70/144 (48.61%) | 134/206 (65.05%) |
| Median wall time | 779 seconds | 422 seconds |
| p90 wall time | 3,126 seconds | 1,477 seconds |
The completion-rate difference was 6.68 percentage points in Claude's favor. The QA-proxy difference was 16.44 points in Claude's favor. Codex was also slower at both the median and the tail.
The denominator discipline is more important than the percentages. Codex started 656 tasks; 631 had reached a terminal state by the cutoff, and 574 were completed. Claude started 1,321 tasks; 1,317 were terminal by its cutoff, and 1,286 were completed. Censored tasks do not quietly become failures, and started tasks do not quietly become the denominator for a completed-over-terminal rate.
The QA number is narrower still. It is a conservative proxy over completed QA children with unique parents, not a fleet-wide claim that every completed task received the same review. A strict first pass required an explicit QA pass without a result-line qualifier pointing to follow-up work, a fix task, or an issue. The complete first-pass metric is unavailable for both periods because the historical event journal does not cover it consistently.
Why this is not a universal ranking
The role mix changed sharply. Social work accounted for 582 of 1,321 starts in the Claude baseline, but only 44 of 656 starts in the Codex period. Codex saw more operations work—109 starts versus 64—and a different balance of coding, marketing, product, and QA work. A single aggregate therefore mixes executor performance with workload composition.
The operating system changed too. QA gates, social cadence, concurrency and capacity, fallback behavior, and logging semantics were not held constant. The later period also included new task types and different priority proportions. We did not reweight the cohorts into a synthetic matched sample, so the aggregate gaps cannot be assigned to the provider alone.
This comparison supports a statement about these two fleet periods: under the frozen rubric, Codex did not produce the material improvement required for a win, and the observed quality proxy moved in the wrong direction. It does not prove that Claude Code is better for every team, task, harness, or configuration. It does not establish that changing executors caused every difference.
Four cells remain blank
We cannot make a complete price-adjusted verdict because four promised measures are unavailable:
- comparable provider cost per completed task;
- directly comparable token economics;
- human-intervention rate; and
- availability percentage.
Those values are unavailable, not zero.
Both providers expose token accounting, but the accounting boundaries differ. Converting those totals into a cross-provider efficiency ranking would manufacture comparability that the logs do not contain. We observed provider and stream error envelopes during the Codex period, but we do not have comparable blocked-duration denominators for both periods, so an availability percentage would be invented precision.
The same rule applies to rework. We can report retry events, failed tasks, cancellations, and QA outcomes, but historical coverage cannot recover a unique-task rework rate without double-counting some work and missing other interventions.
What the experiment changed
The original bet was price-adjusted progress: better completed work per dollar and per unit of operator attention, not the cheapest raw token. Missing economics means the experiment cannot crown an overall economic winner. Worse directly observed completion and QA-proxy results mean it also cannot be called a Codex win while waiting for favorable missing data to appear.
So the operational conclusion is narrower and more useful: mixed evidence, no Codex win; the frozen Claude cohort retains the advantage on the measures we can compare directly.
The portable lesson is not which logo belongs at the top of a leaderboard. It is how to change an executor without letting the migration grade itself: freeze the windows, preserve raw denominators, separate missing from zero, expose workload drift, and write the verdict rule before seeing the result.
That discipline produced an answer less exciting than a universal ranking. It also produced one we can defend.