AWS's Strands Decider 2B picks from your list, never writes

AWS took Qwen3.5-2B, removed the part that writes text and attached a scoring head of just over a million parameters, so the model can only pick from the answers you give it. Strands Decider 2B came out Thursday, as The New Stack reports, as a downloadable model that comes with the data and scripts used to train it.
At a glance
- Decision models like TypeSafe's Jev give up free-form text and instead choose among options a developer supplies or return numerical scores. AWS is now the latest major vendor to release one of its own.
- AWS reports decisions under 100 milliseconds on an Nvidia RTX 3090 and around 150 milliseconds median for small tasks on an M3 MacBook. It also claims second place among roughly 2B public models on JevBench's public set.
- The closed option list stops the model from inventing answers but not from choosing a wrong one, and AWS has not said whether it will offer a hosted version for production use.
If you haven't been following, TypeSafe's Jev started the current wave of decision models a few weeks ago, and Kev, imajev, Laya and others followed. According to TypeSafe, Jev belongs to a class it calls the "System One Model", meaning models "built to make fast, structured decisions that software can use directly." TypeSafe also says it spent two years in stealth before releasing Jev in early access.
OpenAI shipped a hosted Decisions API on Tuesday, and AWS shipped weights on Thursday
The major AI vendors are now bringing out their own versions. On Tuesday OpenAI launched its Decisions API as a limited preview. It is a hosted API that aims its Luna model at questions with predefined answers. Two days later AWS went the other way and released a downloadable model along with the data and scripts used to train it.
All decision models make the same trade. They give up free-form text generation and either select from options the developer supplies or return numerical scores. That makes them useful for routing natural language requests, selecting tools, evaluating outputs and checking proposed actions. Conversation and more complex work stay with generative models.
A pointer head of just over a million parameters replaces the text generator
Strands Decider uses Qwen3.5-2B as its language-understanding base, which AWS calls the "torso." The team removed the language-model head, the layer that turns the model's understanding into generated text. In its place they put a pointer head that scores the answer options supplied with each request. That head has just over a million parameters. The backbone is tuned with a rank-16 LoRA (low-rank adaptation) adapter, a small set of extra trainable weights.
Think of a machine-graded multiple-choice exam. The student can only fill in one of the printed bubbles and can never write in a new answer, but the bubble they pick can still be wrong. Restricting the answer space works the same way. In return, developers get faster decisions and confidence scores they can act on.
AWS says it tried to balance accuracy, calibration and latency. Calibration here means how closely the model's confidence scores track how often it is actually right, so a high score should go with a high hit rate.
Before a weather tool runs, Decider checks whether the agent knows the city
In AWS's demo, built with its open-source Strands agent framework, a user asks for the weather without saying where. The agent guesses a city and proposes calling a weather tool. Before the call runs, Decider checks two things: whether the argument values are grounded in the conversation, and whether the agent has enough information to go ahead. The application then sends the agent back to ask which city the user meant.
The check goes through Strands' intervention system. That system gives developers four choices for a tool call: let it proceed, deny it, ask a human to confirm it, or send feedback to the agent. In the demo, Decider runs locally while the agent calls its generative model through Amazon Bedrock. AWS says it is also working on decision-model integration libraries.
Strands Decider was incubated at Strands Labs, AWS's home for experimental approaches to agentic AI, which launched earlier this year. It follows AWS's recent release of Strands Harness, which packages the tools and supporting machinery needed to run longer-lived agents.
On JevBench, AWS says Strands Decider ranks first among public models with a full training recipe
On JevBench's public set, AWS says Strands Decider ranks second among public models with roughly 2 billion parameters. It ranks first among public models that come with a full training recipe. AWS also says it answers every question in JevBench's easy tier correctly, and that tier is the kind of routine agent decision the model is built for.
AWS reports decisions in under 100 milliseconds on an Nvidia RTX 3090, with response times going up as tasks grow. On an M3 MacBook, the company says, the median for small tasks is around 150 milliseconds.
This release is the second major iteration of the architecture. AWS says an earlier head design performed significantly worse. Every earlier iteration is still in the repository, so developers can trace how the model evolved. Like Kev, Strands Decider is built on an open Qwen model.
Independent figures exist only for other decision models. According to AIMultiple, in a benchmark of 50 browser tasks Kev-9B completed 20 and Jev 1.13 completed 17. GPT-6 Astra completed 47, and Gemini 3.8 Flash completed 42 at low reasoning effort.
Every Strands Decider number so far comes from AWS itself. AWS reports a perfect easy-tier result but gives no matching detail for the harder tiers. In our view, hosting is the bigger gap: the demo already splits the work between a local Decider and Bedrock, and production apps will likely want a cloud version of both halves.
Whether Decider lands on Bedrock
AWS has not said whether this model, or a future version of it, will be offered as a hosted service in its cloud, and it has given no date. The nearer checkpoint is the decision-model integration libraries AWS says it is working on. AWS has not given a timeline for those either. OpenAI's Decisions API, for its part, is still a limited preview.
Related stories
- OpenAI's Decisions API picks an answer in 150 milliseconds
- TypeSafe AI's Jev answers with choices, never sentences
- OpenAI's MCP Extensions put plugins in the ChatGPT sidebar
- OpenAI opens ChatGPT's 1.2B weekly users to developers
- Claude Directory opens a submission flow for plugins
- Four Opus 5.5 API changes make Opus 5 requests fail
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
