OpenAI's Decisions API beats its own engine by up to 10x

GPT-6 Luna runs inside OpenAI's new Decisions API. Even so, OpenAI's announcement on its developer forum says the API reaches a decision up to 10x faster than you would by asking Luna the same question through the Responses API. Since October 6, 2026, any developer can try it in public beta.
At a glance
- OpenAI opened the Decisions API to all developers in public beta and describes it as a way for an app to pick the right model, tool or action in near real time.
- It accepts text and image inputs and answers in one of three structured forms: a probability that a statement is true, a pick from fixed options with confidence scores, or a score.
- The catch is model choice. According to OpenAI, gpt-6-luna is the only model the Decisions API currently supports, and the announcement does not explain where the speed gain of up to 10x comes from.
If you have not been building routers, here is the usual approach. To get a model to make a small call, such as which tool to run or whether a message needs a human, developers send a full prompt through a general endpoint and parse the text that comes back. For OpenAI models that endpoint is the Responses API, the same one the Decisions API is measured against.
The Decisions API answers in three formats: predicates, choices and scores
The Decisions API does not return free text. It returns one of three output types. A predicate is an estimate of the probability that a statement is true. A choice is a selection from options you define in advance, and it comes back with confidence scores.
A score is an evaluation of an input. The published excerpt of the announcement cuts off at that point, so we cannot tell you what the input is scored against or what range the score takes. Inputs can be text or images, so the same kind of call can judge a screenshot as well as a sentence.
OpenAI pitches the whole package for routing. The idea is that an app decides in near real time which model to call, which tool to use or which action to take.
The Decisions API runs on GPT-6 Luna and is up to 10x faster than Luna via Responses
The model behind the Decisions API is GPT-6 Luna. According to OpenAI, gpt-6-luna is also the only model currently available for it. You cannot swap in a larger model for harder decisions or a different one for a specific domain. The public beta opened on October 6, 2026, and every developer has access, so there is no invite list to wait on.
The speed claim compares Luna with itself. OpenAI says the Decisions API makes decisions up to 10x faster than GPT-6 Luna does through the Responses API. The words "up to" mean 10x is the best case, not the typical one. The announcement excerpt gives no latency in milliseconds, no description of the workload and no input size behind the figure.
Each of the three outputs turns a decision into a number or a pick
In plain terms, the Decisions API asks the model a closed question. A predicate is a yes-or-no question answered with a probability, for example how likely it is that a support ticket is about billing. A choice is a multiple-choice question in which each option comes back with a confidence score. A score rates a single input.
Compare an essay exam with a multiple-choice sheet: whoever grades the second only needs to see which box is ticked. Your app gets back something it can check against a threshold. For example, it can send a request to a heavier model only when confidence falls below a level you set, and nobody has to parse prose to get there.
OpenAI does not say where the speed comes from, how the confidence scores are calibrated or what the beta costs. The figure of up to 10x also comes without a benchmark you could rerun yourself. In our view, the single-model limit is the tighter constraint. A router that can only consult gpt-6-luna is likely fine for routine calls, but it gives you no way to pay for more accuracy on the hard ones.
When Decisions gets more models. OpenAI has not given a date for the Decisions API to leave beta, and the announcement does not say whether models other than gpt-6-luna will be added. Until it does, you can run the test yourself. Send the same routing question through the Decisions API and through the Responses API, and see how close your own workload gets to 10x.
Related stories
- OpenAI's Decisions API picks an answer in 150 milliseconds
- HTTPX2 becomes the default in OpenAI's Python SDK
- ChatGPT's new Mac memory watches what you do, without screenshots
- ChatGPT takes meeting notes, but only in the Mac app
- OpenAI speeds up GPT-6 models by about 50% in ChatGPT
- ChatGPT Auto-review goes free for ChatGPT sign-ins
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
