OpenAI's Decisions API picks an answer in 150 milliseconds

OpenAI says its new Decisions API picks an answer from a fixed list in 150 milliseconds, while GPT-6 Luna, the model it is built on, would take 1.6 seconds. In the announcement on Twitter, OpenAI calls the service lightning fast constrained decision making powered by Luna. It is built to make choices, and it does not chat.
At a glance
- OpenAI unveiled the Decisions API at its annual DevDay conference, and it is already open in limited preview, with a broad release planned for the coming days.
- You supply a question, a set of allowed answers and some context, including images, and it returns each answer with a confidence score instead of generating chat-style prose.
- OpenAI has not said what a call costs, how many candidate answers one request can take, or whether developers can tune the model on their own data.
If you have not been following, The New Stack reports that decision models became very popular with the sudden rise of TypeSafe's Jev. The outlet describes them as explicitly not chat models: they return one of a set of predefined answers along with a confidence score. OpenAI has done the scoring part before, since its Moderation API has long returned per-category scores rather than prose.
OpenAI puts a Decisions API call at 150 milliseconds, against 1.6 seconds for GPT-6 Luna
The speed claim leads the announcement. According to The New Stack, OpenAI says the Decisions API returns results in 150 milliseconds, while GPT-6 Luna would take 1.6 seconds. The announcement post itself is less specific and says the service is tuned to make decisions in less than a few hundred milliseconds end to end.
Underneath is Luna, which The New Stack describes as the smallest and most affordable model in OpenAI's current lineup. The API also supports visual inputs, so the context for a decision can be an image as well as text. The developer's side of the deal is short: provide the questions, the answers and the context.
The Decisions API is in limited preview, and OpenAI promises more details at broad rollout
OpenAI announced the API at its annual DevDay conference on Tuesday. For now it is available in limited preview, and the broad release is planned for the coming days. An OpenAI spokesperson told The New Stack that the company plans to share more at broad rollout.
The New Stack sees the launch as a likely reaction to TypeSafe and Jev, and says OpenAI probably rushed the announcement out ahead of DevDay. According to the outlet, what has been announced so far is all OpenAI is sharing for now. It lists the jobs this kind of model is built for: classifying content, routing requests and choosing an agent's next action from a limited set of choices.
The Decisions API sits between a prompted chat model and a trained classifier
Most teams today, The New Stack writes, handle this job with a regular chat model and a carefully worded prompt that asks it to pick from a list. If they are lucky, they read the token probabilities to get something that looks like a confidence score. Those scores, the outlet notes, are often a rough guess, and the model burns quite a few tokens to produce them.
The other route is a small trained classifier. It is fast and cheap, but it needs labeled data and a new training run every time the label set changes. A decision model takes new labels in the prompt, as a chat model does, yet returns a score a developer can work with. The New Stack says such models give more realistic confidence scores and return them extremely fast.
Think of an essay question versus a multiple-choice sheet. A chat model writes the essay and leaves you to guess how sure it is. A decision model ticks a box and writes a number next to each option. The Moderation API works in a similar way, but there OpenAI sets the categories itself, while here the developer writes them.
The missing details are the ones that matter most for adoption. Price per call, the limit on candidate answers and any way to tune the model on your own data are all undisclosed, and The New Stack says these will decide whether the API becomes a standard building block in agent frameworks or stays a niche tool. In our view, the 150-millisecond figure is hard to rely on without the conditions behind it, because the announcement itself only promises less than a few hundred milliseconds end to end.
What the broad rollout must answer. OpenAI has promised more details with the broad release, which is planned for the coming days. No exact date has been given. The first things to check will be the price per call and the maximum number of candidate answers in one request. Until those are public, anyone thinking about replacing a trained classifier is comparing a cost they know with one they don't.
Related stories
- OpenAI's half-price GPT-6 Sol and Luna aren't a promo
- HTTPX2 becomes the default in OpenAI's Python SDK
- ChatGPT's new Mac memory watches what you do, without screenshots
- OpenAI's $200 Pro plan comes back worth half the API spend
- Microsoft Copilot gets a Code mode on GitHub Copilot tech
- Cursor's new bot follows each pull request into production
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
