GPT-6 caching update cuts cached input costs by up to 90%

If your app resends the same long system prompt or document on every call, you now pay less for the repeated part. OpenAI's prompt caching update for GPT-6, announced September 22, cuts cached input token costs by up to 90%.
OpenAI lists higher cache hit rates, new diagnostics, explicit breakpoints and controls aimed at lower latency and cost. Newsroom America also reports a caching dashboard.
Caching stores the already processed start of a prompt, so repeat calls skip that work. A breakpoint marks where the reusable part ends. OpenAI's summary doesn't say how much hit rates actually go up.
Related stories
- Code strings point to a $50 production plan for OpenAI's API
- Token costs drop out of Replit's new Free Mode
- OpenAI's half-price GPT-6 Sol and Luna aren't a promo
- libheif bugs reached OpenAI repos and GitHub Enterprise
- OpenAI's lawyers and recruiters now work through Codex
- Every chatgpt.com page carries 377 KB of flag decisions
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
