Typing "hi" into Claude Code crosses ten layers

When your agent runs a tool, the output never goes straight back to the model. It lands in the harness, which reads it, picks the next step and only then calls the API. That loop is one of ten layers mapped by The-agent-stack, a pixel-art explainer by Clifford and Gaurav at Nori Agentic that follows a single request, req_7f3a, from the box you type into down to the model weights on a GPU.
At a glance
- Nori Agentic built a ten-layer map of one agent request and split it in two: an application side you shape, and a provider side that starts at the API and ends in the model weights.
- Providers cache a prompt's prefix, and those caches usually live five minutes to an hour, so an early change to a system prompt or tool definition throws away everything cached after it.
- The explainer describes the stack only in general terms and gives no prices, latencies or provider-specific numbers, so it shows why a long pause costs you money but not how much.
If you haven't followed the history: according to Taskade, Claude Code began as one engineer's side project during his first month at Anthropic in September 2024. ScriptByAI reports that it entered research preview on February 24, 2025, reached general availability on May 22, and later expanded into IDEs, the web and Desktop. Taskade adds that, per public reporting, it was past a $1B annualized run rate by November 2025.
Four of the ten layers are yours, and the harness sits in the middle of them
The top layer is the interface, the part you type into and watch. It handles presentation, not execution, so the same agent can show up as a terminal, an editor panel, a native app or a chat bot. The Agent Client Protocol calls this role the client.
Next comes context: everything the model sees before it answers. Most of it is pulled off your file system, including AGENTS.md, skills, source files and the output of earlier tool calls, and the harness decides what to load and in what order. According to ScriptByAI, as of a Sep. 18 change, Claude Code reads CLAUDE.md and falls back to AGENTS.md when none exists.
The harness is the code that turns a model into an agent. Claude Code, Codex and Nori are each one. It runs a single loop: read context, call the model, evaluate the reply, run the tools it asked for, feed the results back in, repeat until the work is done. The tools are its hands: shell commands, file reads and edits, URL fetches, MCP servers.
Every tool call makes a round trip through the harness
This is the detail we would retell first. When the model asks for a shell command, the result does not go back to the model. It goes back to the harness, which reads the output, decides what to do next and only then calls the API again.
Each harness defines its own tool set and its own instructions for using those tools, so the same model behaves differently depending on what it was handed. Taskade describes Claude Code in the same loop terms: it searches files, makes edits, runs commands, reads the output and keeps going until the job is done or it needs approval.
One part of that loop has recently moved. Claude Code's Auto Mode uses a classifier, and on Sep. 19 the default classifier moved server-side for API, Enterprise and supported provider users, which removes the classifier overhead from the local side.
Past one HTTP POST, the provider controls everything, including your cache
The API is where your control stops. Past that HTTP POST you cannot see or change how the request is served. You keep two levers at the boundary: enterprise plan policy applied across a whole team, and an AI gateway, a stand-in API that meters spend per user, limits which models are reachable and issues one key for many providers.
On the other side, the request lands in a provider region at a hyperscaler, gets routed to a machine and is served by two neighboring services, a cache and the inference service. The cache keeps a prompt's prefix, so a change near the start invalidates everything after it. The explainer says this is why your choice of harness and long gaps between turns show up directly on your bill. According to GitHub release notes, since the Aug. 13 release Claude Code's forked subagents inherit the current conversation and prompt cache by default.
The second token is much cheaper than the first
Inference reads the whole prompt in one pass, then generates one token at a time and batches your request with other people's to keep the hardware busy. Each new token would normally mean reading back over everything before it. Instead, the service stores what that reading produced, a key vector and a value vector for every token so far. That store is the KV cache.
Think of it as notes in the margin of a long book. Before writing the next sentence you check the notes instead of rereading every chapter. The notes still take up room, so a long conversation costs memory as well as time, because the cache grows with every token in it.
A trillion-parameter model can do the work of a far smaller one
A data center GPU packs a few hundred small processors. Each has tensor cores that multiply one tile of a matrix per instruction, and they sit next to high-bandwidth memory. Generation speed is limited by how fast the weights move out of that memory, not by raw arithmetic.
The weights are the model. A token passes through layer after layer of matrix multiplies until one output token comes out. A dense model sends every token through every weight. In a mixture of experts, each layer is split into many blocks and a small router picks a few of them per token, leaving the rest untouched. That is why, the explainer argues, parameter count stopped predicting cost.
The explainer is a map, not a measurement. It gives no prices, cache hit rates, latencies or model sizes for any provider, and it does not say which descriptions are specific to Claude Code and which apply to harnesses in general. In our view, the most useful figure for daily users is the five-minutes-to-an-hour cache lifetime, and it is also the vaguest one, since the source does not say which provider sits at which end of that range.
Pricing the gap between your turns
The explainer gives no update schedule and no numbers to check its claims against. It raises one question for Claude Code users without answering it: how much a five-minute or an hour-long pause actually adds to the cost of a session on your plan. Until providers publish cache lifetimes next to diagrams like this one, you will have to work that number out from your own usage logs.
Related stories
- Claude Code shipped its own TypeScript in npm source maps
- Claude Code Projects admits its whole Pro and Max waitlist
- In Claude Code, a "block" hook could let commands through
- A Claude Code mod that reverts its own failed ideas
- Claude went from bug hunt to abuse report on New Year's Eve
- Claude Code Projects threads can now run on your machine
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
