Skip to content

openai

GPT-6 caching update cuts cached input costs by up to 90%

Promtime

If your app resends the same long system prompt or document on every call, you now pay less for the repeated part. OpenAI's prompt caching update for GPT-6, announced September 22, cuts cached input token costs by up to 90%.

OpenAI lists higher cache hit rates, new diagnostics, explicit breakpoints and controls aimed at lower latency and cost. Newsroom America also reports a caching dashboard.

Caching stores the already processed start of a prompt, so repeat calls skip that work. A breakpoint marks where the reusable part ends. OpenAI's summary doesn't say how much hit rates actually go up.

Related stories

  1. Code strings point to a $50 production plan for OpenAI's API
  2. Token costs drop out of Replit's new Free Mode
  3. OpenAI's half-price GPT-6 Sol and Luna aren't a promo
  4. libheif bugs reached OpenAI repos and GitHub Enterprise
  5. OpenAI's lawyers and recruiters now work through Codex
  6. Every chatgpt.com page carries 377 KB of flag decisions

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.