Codex opens up to Kimi K3 and GLM-5.3 Flash via Baseten

If your company has already promised OpenAI a set amount of spending, the money it now spends on open models served by Baseten can count toward that promise, as a post on Twitter announcing the deal describes it. OpenAI enterprise customers get GLM-5.3 Flash and Kimi K3 natively inside Codex, with Baseten doing the serving.
At a glance
- Baseten is now an open-model inference provider in the OpenAI B2B Marketplace, which puts models such as GLM-5.3 Flash and Kimi K3 inside Codex for OpenAI enterprise customers.
- Baseten hosts the model weights and serves the answers, so you keep working in the Codex interface while an open model, not a GPT model, does the work behind it.
- Neither Baseten nor OpenAI has published per-model prices, and The State of AI reports that native access to open models in Codex currently runs through a waitlist.
If you have not been following, Codex could already talk to open models before this deal. According to Explainx, on June 17, 2026 Tibo Sottiaux of OpenAI Codex posted a reminder that the Codex App, CLI and SDK work with any open source model, not just OpenAI's, and the post reportedly hit 1.6M+ views within a day. The Baseten deal adds a hosted route that enterprise customers pay for through their OpenAI commitment.
Open-model spend on Baseten counts against an OpenAI enterprise commitment
Baseten has joined the OpenAI B2B Marketplace as an inference provider for open models. An inference provider hosts the model weights on its own hardware and serves the answers when you send a request. For OpenAI enterprise customers the deal does two things: it brings GLM-5.3 Flash and Kimi K3 natively into Codex, and it lets spend on those Baseten-served models count toward the existing OpenAI commitment.
According to Lavx, the commitment covers open models served by Baseten whether you use them within Codex or through the Responses API, and Baseten is among the first open-model inference providers in the marketplace. Lavx also reports that native Codex access goes through Baseten's waitlist, while other use cases get direct engineering consultations. Neither company has published per-model prices.
Baseten says its inference runs on US-based infrastructure with zero data retention for all prompts, and that deployments can be pinned to regions, with fine-grained AuthN/AuthZ and usage views across model, user or key. Blaxel, now part of Baseten, gives each agent a sandbox, according to the company.
How did teams run open models in Codex before Baseten?
According to Explainx, they did it themselves through Ollama: launch Codex with GLM-5.2 and Kimi-K2.7-Code, or run codex --oss. The setup requires wire_api = responses, and Ollama recommends a minimum context of 64k tokens. Explainx also lists gaps: the desktop model picker hides custom providers, and GPT-only features such as computer use and browser automation stay out of reach.
The ChatGPT Learn documentation covers the configuration side. Codex supports custom model providers and profiles, and from Codex 0.134.0 a profile lives in a file like ~/.codex/profile-name.config.toml. According to the same documentation, custom model providers do not carry over through managed configuration for Work Cloud.
What does Baseten actually do between Codex and the model?
Baseten's own answer, on its website, is routing. The company says businesses are shifting to a mix of models at different trade-offs between intelligence, cost, latency and capacity, with static or adaptive routers sending each task to the model with the best mix for the job. Think of a restaurant pass: the ticket looks the same, but the dish goes to whichever cook handles it best.
Baseten also argues that the cycle between new open and closed models is down to weeks, so no model stays the best tool for every job for long. Code generation, it says, is a hard case for inference, because multiple turns, long prompts over large repos and huge context windows tax every part of the stack.
On capacity, Baseten says it runs on more than 90 clusters across 20+ clouds with active-active deployments, so losing a provider or region reroutes traffic instead of stopping work. The company cites relationships with 200+ compute providers.
The gaps sit where you would look first. Without per-model prices, you cannot yet compare a Kimi K3 or GLM-5.3 Flash request against a GPT one, and the waitlist means access is not immediate. Oddly, for a deal whose headline feature is billing, the billing numbers are the missing piece; it is also unclear whether the limits Explainx lists for the DIY route, such as GPT-only computer use, carry over to the native one.
When the Baseten waitlist opens
No date has been given for when waitlisted OpenAI enterprise customers get native open-model access in Codex, and no per-model price list has been published. The model versions are worth watching too: the DIY route Explainx describes used GLM-5.2 and Kimi-K2.7-Code, while the Baseten route names GLM-5.3 Flash and Kimi K3. Baseten says it launches major open models on the day they ship, so the lineup may grow beyond two.
Related stories
- Codex keeps coding in the cloud after your laptop closes
- OpenAI cuts $200 Pro usage in half, adds a $500 tier
- Codex Security Cloud reviews commits with the laptop closed
- OpenAI puts 8x faster Codex behind a new Pro 500 plan
- GPT-6 Astra draws ducks on some accounts, a test finds
- AWS says its open agent runs 45% cheaper than Claude Code
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
