Skip to content

openai

OpenAI's GPT-6 guide says drop blanket "always ask" rules

Promtime

Stop telling your coding agent to "always ask" before it acts. That is one of the most concrete tips in OpenAI's practical guide to the GPT-6 family. The guide recommends clear boundaries between what the model may do alone and what needs your sign-off, instead of blanket approval rules.

At a glance

  • OpenAI's guide sorts the GPT-6 family into GPT-6 Astra for the hardest reasoning, GPT-6.1 Sol for complex coding, research and computer use, and GPT-6 Luna for repeated, clearly scoped work.
  • The cost levers include prompt caching, where cached input tokens cost up to 95% less depending on the model. There is also compaction for long conversations, and reasoning effort that runs from Low to Extra high or Max.
  • The catch is maturity: multi-agent delegation in GPT-6.1 Sol, which hands independent subtasks to subagents on long jobs, is still in beta, and steering updates cannot undo completed actions.

If you have not been following, the GPT-6 generation arrived in stages. According to TechCrunch, OpenAI launched GPT-6 Astra in early September 2026. Updated Sol and Luna models followed on September 22 at half the API cost of the 5.6 series, a drop OpenAI puts down to better caching and inference. TechCrunch also reports that Anthropic released a new version of Opus 5.5 about 90 minutes before that release.

Astra, Sol and Luna split the GPT-6 family by how hard the task is

OpenAI treats the choice of model and reasoning level as a trade between intelligence and price. GPT-6 Astra is for the hardest reasoning, where you need maximum intelligence. GPT-6.1 Sol handles complex coding, research and computer use. GPT-6 Luna takes focused tasks at scale, such as extracting invoice fields, classifying requests or producing structured summaries.

In the API, you set reasoning effort per task. Low covers routine work like extracting facts or making small edits. Medium covers judgment calls such as planning a feature, and High is for difficult debugging or careful review. Where supported, test Extra high and Max when High falls short, and keep them only if the improvement justifies the added time and cost. In Codex, start at the model's default and adjust from there.

Speed is a separate dial. Fast mode in the API gives faster, more consistent response times at a higher per-token cost than Standard processing. Ultrafast works in Codex and the API, is available for GPT-6 Astra, and speeds up token generation independently of reasoning effort.

Cached input tokens cost up to 95% less, depending on the model

The production advice starts with context. Cut what the task does not need, but keep the evidence it does need. Where your application supports it, run independent tasks together so one slow step does not hold up unrelated work. For recurring work, reuse shared context through prompt caching: put stable instructions and reference material before the task details that change, and keep tool definitions consistent.

OpenAI points to a caching dashboard and a diagnostics guide that show where reuse breaks down. When you estimate the cost of a complete workflow, the guide says to include cache writes and any long-context rates. For longer conversations, compaction shrinks the context while keeping the state needed to continue.

Before you deploy, run representative tasks and measure task success, latency and cost per successful task. The guide also asks you to decide how you will monitor behavior and to review the data controls for your application.

OpenAI wants decision boundaries instead of blanket "always ask" rules

The prompting section opens with a line from Eric Provencher of Developer Experience at OpenAI: give the model a clear assignment. That means the result you want, who it is for, the relevant context and constraints, and what counts as done.

Four areas follow, summarized from OpenAI's companion piece on rethinking skills and prompts for GPT-6 Astra. Skill descriptions should be short and say when each skill runs. AGENTS.md should explain when particular documents and tests matter and explicitly allow safe routine workflows, such as running local tests with disposable data and no production access. Decision boundaries should state which actions can proceed alone and which need approval.

Persistence gets its own rule. "Done" should include implementing the change, running it, inspecting the result and fixing failures. The guide's example of a boundary is simple: the model may decide how to organize a summary, but it should check with you before changing the project's scope. A useful handoff covers what changed, what was checked and what still needs attention.

Tasks that span hours or days get steering, async tools and subagents

With the GPT-6 family, OpenAI says, you can take on tasks that span hours or days. In the API, mid-turn steering sends a correction through the Responses WebSocket API while the model works. Updates are queued, and they do not cancel running tools or undo completed actions. Asynchronous tool calling lets the model keep doing independent work while your app runs something slower, such as tests.

GPT-6.1 Sol supports multi-agent workflows in the Responses API. It assigns independent subtasks to subagents and combines what they find into one answer, and this feature is in beta. In Codex, GPT-6 Astra can ask for clarification mid-task, and you can tell it which work may continue while you decide.

Computer use works with Astra, Sol and Luna, including on apps that have no API. The guide's rule is to use an API or connected tool when one can do the job, and to use screens and clicks only when needed. In your own app, Playwright handles browsers and PyAutoGUI handles desktop apps.

Extra high and Max stay only if the gain covers the extra time and cost

Reasoning effort is easiest to picture as how long the model thinks before it answers. According to T-Minus AI, higher effort makes the model generate more internal thinking tokens before the visible answer. Those tokens are where the cost and latency come from, and they are billed at the same rate as output tokens. Think of a contractor who bills by the hour: a more careful inspection buys more hours whether or not it finds anything.

The returns are uneven. Digital Applied ran a benchmark in April 2026 with 900 task runs across five frontier models and three effort tiers. It found that turning up effort improved quality by 8 to 22 points, while fees grew 4-17× and latency 5-60×. The point where extra effort stopped paying depended on the task: high effort won AIME, medium won an Expert-SWE refactor and low won PR-scale review.

OpenAI has given this advice before. Its GPT-5 prompting guide said many workflows get consistent results at medium or even low reasoning_effort. It also recommended the Responses API for agentic flows, because reasoning carries over between tool calls.

The guide gives few numbers of its own. It tells you to compare pricing for each model but lists no prices, and the 95% caching discount is an upper limit that varies by model. Oddly, multi-agent delegation in GPT-6.1 Sol, the feature best suited to the multi-day jobs the guide talks about, is still in beta, so the longest workflows appear to depend on the least finished feature.

When subagents leave beta

OpenAI has not given a date for multi-agent workflows in the Responses API to leave beta. The guide also does not say whether subagent support will come to models other than GPT-6.1 Sol. Until then, the guide leaves the checkpoint to you: run representative tasks and score them on success, latency and cost per successful task before Extra high, Max or Ultrafast go into production.

Related stories

  1. GPT-6 caching update cuts cached input costs by up to 90%
  2. OpenAI's half-price GPT-6 Sol and Luna aren't a promo
  3. OpenAI's MCP Extensions put plugins in the ChatGPT sidebar
  4. Your ChatGPT Plus plan can now pay for other apps' AI
  5. OpenAI opens ChatGPT's 1.2B weekly users to developers
  6. Codex keeps coding in the cloud after your laptop closes

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.