Skip to content

openai

Luna 6 reports 32x fewer reasoning tokens on ChatGPT Pro

Promtime

Send the same prompt to the same model through the same Codex client, and the reported reasoning token count can come out about 32 times different depending on how you signed in. A community experiment published on Github found that Luna 6 at Medium averaged 96.30 reported reasoning tokens per response with an API key and 3.025 with a ChatGPT Pro sign-in, across 40 completed responses each.

At a glance

  • The authors collected 960 completed responses from Luna 6 and Sol 6.1 at three thinking levels, comparing API-key access with ChatGPT Pro sign-in inside the same Codex client.
  • On the bookstore task, four gaps held up under multiple-comparison correction in both batches, yet every one of the 480 bookstore decision sets was correct on both routes.
  • The setup used one API credential, one Pro account and two prompts, the counts reflect reported usage rather than internal reasoning, and the authors say they still do not know the cause.

If you have not been following, this is not the first report of its kind. Codex issue #49757 compares Astra through an Enterprise sign-in with the public API, and the authors' related-work review also covers a Luna report, bugs in the effort setting and studies that found different results. None of them, the authors note, establishes the cause of what they measured.

Luna 6 at Low averaged 0.000 reported reasoning tokens through ChatGPT Pro

The bookstore task, which asks the model to diagnose duplicate emails from a bookstore, produced the clearest gaps. For Luna 6 at Low, the API route averaged 68.550 reported reasoning tokens per response against 0.000 through the ChatGPT Pro sign-in. For Luna 6 at Medium the figures were 96.300 and 3.025, the roughly 32x gap in the headline.

Higher up, the gap shrinks in relative terms. Luna 6 at High averaged 148.525 tokens through the API and 85.200 through the subscription, and Sol 6.1 at High averaged 216.675 against 138.700. Each row covers 40 completed responses per sign-in, and these four differences passed the multiple-comparison correction in both the original study and the follow-up.

All 480 bookstore decision sets were correct on both routes

More reported reasoning did not translate into better answers where the token gap was clearest. All 480 bookstore decision sets came back correct, whichever way the Codex client signed in, so on that task the API route's extra tokens bought nothing measurable.

The second task, choosing projects within a budget, is where accuracy did move. Across the full dataset, Luna 6 at Low scored 35/40 on the project task through the API and 18/40 through the subscription. That pooled difference passed the planned conservative correction with p = 0.0083, but the authors still label it exploratory.

The follow-up batch on its own looks weaker. There, Luna 6 at Low on the project task scored 19/20 through the API versus 12/20 through the subscription, a gap that did not pass its correction (p = 0.371).

The study ran 960 responses in two batches of 480

The design is a grid: two tasks, two models (Luna 6 and Sol 6.1), three thinking levels (Low, Medium, High) and two sign-ins, with 40 completed responses for each combination. The original 480 responses were followed by a fixed, separately analyzed batch of another 480, which is why the authors can say the four bookstore gaps held in both.

The follow-up hit two capacity failures on the subscription route. The authors preserved both, documented the manual continuations and completed the original schedule. The statistics describe completed responses, a runtime amendment explains the interruptions, and incorrect answers stay in the dataset rather than being filtered out.

A local proxy keeps each pair byte-identical, so only the sign-in differs

Each response starts a fresh Codex session. A local proxy sits between the client and the server and makes sure both members of a pair send exactly the same model input, byte for byte, with the same requested model and thinking level. What differs is authentication and the service endpoint the request goes to.

Think of it as mailing the same sealed letter to two addresses of one company and comparing the receipts that come back. The proxy leaves real server responses unchanged, and the authors record the server's answers and usage counts, then check them against Codex's own records.

Anyone with Python 3.12 or later can recompute the published results from the included data without making model calls, by running python3 -m experiment verify and then python3 -m experiment analyze. The report lands in output/reports/followup/report.md, and the repository uses the MIT license.

The limits are the authors' own: two fixed prompts, one API credential, one Pro account, and counts that describe reported usage rather than the model's internal reasoning budget. In our view, the Luna 6 Low row deserves the closest look, since an average of 0.000 reported reasoning tokens on the subscription route reads as more than noise. Whether it reflects how usage is counted or how much the model actually reasons, the published data cannot say.

What independent reruns could settle

The authors are asking for reproductions with different accounts, plans and collection dates, and they say runs that find no difference are just as useful. Their investigation plan lists possible explanations to test, but no cause has been identified and no timeline for one has been given. Until other accounts repeat the comparison, nobody knows how widely the gap occurs.

Related stories

  1. GPT-6 Astra draws ducks on some accounts, a test finds
  2. Your ChatGPT Plus plan can now pay for other apps' AI
  3. Codex reset hunters lose Day 1 to a 50% speedup
  4. Left free to choose, OpenAI's Dots picked Cloudflare 8 of 8
  5. OpenAI's Dots run free until they open a Codex task
  6. Codex opens up to Kimi K3 and GLM-5.3 Flash via Baseten

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.