openai

A daily benchmark checks if Codex quietly got worse

Promtime

openai

The daily Codex tracker run by Marginlab recorded an 88% pass rate for gpt-5.6-sol on Sep 8, 2026, five percentage points above its 83% reference rate, with the overall status listed as Nominal. The figure comes from 50 eval test cases run that day.

At a glance

  • Every day a curated subset of SWE-Bench-Pro runs through Codex CLI with no custom agent harness, so the published figure comes from the shipping tool rather than a bespoke evaluation rig.
  • The past 7-day and past 30-day pass rates both read 85%, aggregated over 350 and 1,500 eval test cases respectively, against a historical baseline of 83%.
  • Daily results carry a ±11.2% significance band around the baseline while 7-day windows carry ±3.6%, meaning single-day swings need to be large before the tracker treats them as real.

Silent regressions are hard to argue about, because the evidence is usually anecdotal: a coding agent that felt sharper last month. A public daily number with an explicit significance threshold turns that complaint into something falsifiable, and it cuts the other way too, since a five-point jump measured on 50 tasks does not clear the bar either. The design appears aimed at that symmetry rather than at scoring the model.

The summary panel published on Sep 8, 2026 listed the overall status as Nominal, with the 88% same-day result standing above the 83% reference rate. Neither that gap nor the 85% readings over the longer windows crossed the tracker's significance thresholds, so nothing registered as a degradation.

The tracker frames its method as statistical testing for degradation detection. Changes inside the shaded band around the baseline are not treated as statistically significant at p ≥ 0.05: ±11.2% for a single day's 50-case run, and ±3.6% for the smoother 7-day rolling aggregate.

Marginlab describes the setup as benchmarking gpt-5.6-sol directly in Codex CLI with no custom agent harness, on a curated subset of SWE-Bench-Pro rather than the full benchmark. Sample counts scale with the window: 50 eval test cases in a day, 350 across seven days, 1,500 across 30.

Both charts, daily and weekly, plot the pass rate against a dashed line at the 83% baseline with the shaded threshold around it, and each carries an optional 95% confidence interval overlay. Per the legend, wider intervals indicate more uncertainty and come from fewer samples.

What the tracker does not publish

The page lists no next scheduled update beyond its daily cadence, and the Sep 8 entry is marked as the last update. It also does not state which repositories or task IDs make up the curated subset, how the 83% baseline was computed, or how long a result must sit outside the threshold band before the status changes from Nominal. Those details are absent from the summary as published.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.