Back home

Runtime Anatomy / Claude Code source notes

Agents are not chatbots. They are runtimes that keep reconnecting the loop.

I recently spent time reading through the source of Claude Code-style agent tooling. After reading it, I cared less about whether the model can answer and more about how the runtime turns the model’s next intent into action: call a tool, collect the result, reshape the context, and continue reasoning. This article focuses on three mechanisms: tool use, context management, and looped reasoning.

1. Model

The model proposes the next intent

It may emit text, or it may emit a structured tool_use block. The runtime treats that as an action request, not as the final answer.

2. Tool

The tool layer touches the real world

Tool definitions handle schema validation, permission checks, concurrency policy, error wrapping, and conversion back into tool_result.

3. Context

The context is rebuilt

Assistant messages, tool results, attachments, memory, and compaction boundaries are assembled into the next model request.

4. Loop

Reasoning continues until no tool is needed

As long as the turn produces tool calls, the runtime continues. The turn ends only when there is no tool_use, blocking hook, or recovery retry.

01 / Loop

The life of one agentic turn

A normal chat model feels like one user message followed by one assistant answer. An agent runtime works differently. After the user gives a goal, the runtime enters a loop. Each iteration sends the current messages, system prompt, tool definitions, and context state to the model. If the model returns a tool call, the runtime pauses the answer, executes the tool, then feeds the result back into the next model request.

query loop sketchsrc/query.ts
let state = {
  messages: initialMessages,
  toolUseContext: initialToolUseContext,
  turnCount: 1,
}

while (true) {
  const { messages, toolUseContext, turnCount } = state

  messagesForQuery = compactAndPrepare(messages)
  assistantMessages = streamModel(messagesForQuery)

  if (!assistantMessages.hasToolUse) {
    return completed
  }

  toolResults = runTools(assistantMessages.toolUseBlocks)
  attachments = collectMemoryAndNotifications()

  const nextTurnCount = turnCount + 1

  if (maxTurns && nextTurnCount > maxTurns) {
    return maxTurnsReached
  }

  const next = {
    messages: [
      ...messagesForQuery,
      ...assistantMessages,
      ...toolResults,
      ...attachments,
    ],
    toolUseContext: toolUseContext,
    turnCount: nextTurnCount,
    transition: 'next_turn',
  }
  state = next
}

The crucial line is state = next. Agent continuity comes from runtime discipline. Every observation is fed back into the next input, so each model call starts from a freshly prepared scene assembled by the runtime.

02 / Nested query

A sub-agent is another query loop inside a tool

Once the main loop is clear, sub-agents become much easier to understand. AgentTool does not start a totally different runtime. For synchronous sub-agents, it behaves like a special tool: the outer query pauses on that tool_use, enters runAgent, and runAgent calls the same query() function again. The inner query performs its own model calls, tool executions, and context updates, then returns a result as the AgentTool tool_result.

That is what “nested query” means. The work stays inside the current call stack, with a complete agent loop nested inside it. The outer query owns the main task; the inner query owns the subtask. They share the same engine, but the parameters differ: the sub-agent has its own messages, system prompt, tool allowlist, agentId, and queryTracking depth.

This keeps the architecture conservative. A sub-agent is not a mysterious new abstraction. It still follows the same tool protocol, permission protocol, and context governance. The difference is that it is invoked as a tool, completes a higher-level packet of work, and hands the result back to the main agent.

1Outer query

The user goal enters the outer loop, and the model chooses the next action.

2AgentTool

The model delegates to a sub-agent, so the outer query pauses on this tool_use.

3runAgent

The runtime builds the sub-agent prompt, messages, toolUseContext, and tool set.

4Inner query

The same query() runs again while the sub-agent reads, searches, and uses tools.

5tool_result

When the sub-agent finishes, its result returns to the outer query as tool_result.

A normal Bash tool usually executes once and returns; AgentTool starts a full agent loop.

The sub-agent can keep reasoning, call tools again, and observe results until it finishes or hits maxTurns.

Isolation comes from parameters and context: the sub-agent does not automatically inherit the whole main conversation, and tools can be allowlisted.

queryTracking.depth tells the runtime which layer of the nested chain is currently running.

03 / Tool use

A tool call is a controlled execution protocol

From the model’s point of view, a tool looks like a JSON function. From the runtime’s point of view, it is a controlled execution protocol. It must be discovered, validated, authorized, routed through hooks, and only then allowed to touch files, shell commands, MCP servers, or network resources.

1

Discover tool

2

Parse input

3

Validate schema

4

Check permission

5

Execute / coordinate

6

Wrap tool_result

7

Feed context back

04 / Permission

Permission checks are the gate between tool_use and execution

A model-produced tool_use means “I want to do this.” It does not mean the runtime immediately does it. Before execution, the runtime parses the model arguments into structured input and validates them against the tool schema. If the input shape is wrong, the call never reaches the permission layer.

After the input is valid, permission becomes layered. A tool can declare whether it is read-only, concurrency-safe, and how it wants permissions checked. Rules, classifiers, hooks, and UI confirmation can then decide whether this specific call should run. Reading a file, editing a file, running a harmless grep, and executing a state-changing shell command should not share the same risk path.

The important design point is that permission is not one boolean check. It is a pre-execution pipeline. Any layer can allow, rewrite input, ask for confirmation, or block the call. A block is still useful: the runtime can feed it back into context so the model can choose a safer next step.

Schema

First make sure the model JSON can actually satisfy the tool schema, so malformed calls never enter execution.

Tool policy

Tools declare read-only status, concurrency safety, permission checks, and result shape; defaults stay conservative.

Rules / classifier

Rules and classifiers turn natural-language intent into risk decisions for edits, shell, network, and MCP calls.

Hooks / UI

PreToolUse hooks can block or add feedback; UI confirmation shows the user who wants to do exactly what.

Feedback

Blocks, errors, and denials are wrapped as tool_result or context feedback so the next model turn can adapt.

05 / Context

Context is a workbench that keeps getting rebuilt.

In a long task, the context window is usually consumed by tool results, file contents, attachments, memory, and intermediate traces. Reading Claude Code’s source makes this clear: context management is a layered system of result budgeting, micro-compaction, auto-compaction, reactive compaction, and session memory. Saving tokens matters, but the deeper goal is preserving the scene needed for the next step.

Context window as a workbench

Conceptual distribution, not measured token usage
System Prompt14%
User / Assistant Turns34%
Tool Results24%
Attachments / Memory16%
Safety Buffer12%

Compaction is more than shortening chat history

Auto-compaction behaves a lot like memory management. The runtime estimates token usage, compares it with the effective model window, reserved output space, and buffer, then folds history into a summary when the threshold is crossed. If compaction keeps failing, a circuit breaker prevents every future turn from wasting another doomed API call. The hard part is keeping the compacted context useful for the next action: preserve the goal, constraints, files already inspected, key errors, important diffs, and open tasks while clearing low-value stdout, repeated logs, and stale intermediate reasoning.

Result budget

Tool results cannot flow back forever. Long output needs truncation, summarization, or persistence.

Micro compact

Local cleanup handles an oversized result or noisy fragment before it pollutes the whole context.

Auto compact

When the full context nears the window threshold, history is folded into a shorter working summary.

Reactive compact

Prompt-too-long and related failures can trigger compression as a recovery path.

Circuit breaker

Repeated compaction failures need to stop, or every future turn wastes another doomed model call.

Preserved scene

The goal is not a pretty summary. It is preserving paths, conclusions, constraints, errors, and remaining work.

06 / Reasoning

Looped reasoning is observation-driven

Every agent iteration does the same thing: choose the next action from the latest observation. Tool results may disprove an assumption, reveal a new file, expose a new error, or hit a permission boundary. The runtime does not need to know the entire plan upfront. It only needs to feed observations back correctly so the model can decompose the task into verifiable steps.

Plan

Turn the goal into the next executable action

Act

Call Bash, file editing, MCP, sub-agents, and other tools

Observe

Feed stdout, diffs, errors, and attachments back into context

Architecture judgments I took from reading Claude Code

First, the agent product experience is mostly shaped by the runtime. The model may propose tool calls, but clear permissions, safe concurrency, recoverable errors, and faithful context compression are engineering responsibilities.

Second, tool results must be managed as prompt assets. Long outputs need truncation or persistence, important summaries must survive, and low-value noise must be replaced before the agent drowns in its own observations.

Third, a sub-agent is best understood as another query loop started from inside a tool. With that mental model, isolated context, tool allowlists, maxTurns, and depth tracking all become easier to reason about.

Fourth, permission should be a pipeline, not a boolean switch. Schema, tool policy, rules, classifiers, hooks, and UI confirmation each address a different layer of risk.

Fifth, looped reasoning needs explicit stop conditions. No tool call can mean completion; blocking hooks can trigger retries; token and prompt-too-long errors may be recoverable, but recovery must be bounded. Otherwise autonomy becomes a spin loop.

Architecture Breakdown Path

I did not want this piece to become a source-code tour, so I am not walking through file paths one by one. This path is closer to what I took from the reading: split the runtime into layers, then look at how those layers carry one agent call.

Main loop

Expands one user request into multiple model calls, tool executions, and state updates.

Start here to understand why an agent is not a single request-response exchange.

Tool protocol

Defines input schemas, permission checks, read-only/destructive classification, and result shape.

This layer decides what the model can touch, how it touches it, and how failures return.

Sub-agent

Enters query() again during AgentTool execution, using isolated parameters to finish one subtask.

This explains why “delegating to a helper” can still reuse the same runtime.

Tool orchestration

Runs concurrency-safe read tools in batches while serializing tools that mutate state.

This layer keeps speed and safety from undermining each other.

Permission pipeline

Runs schema checks, tool policy, rules, classifiers, hooks, and UI confirmation before real execution.

The model says what it wants; the runtime decides whether and how that action may happen.

Single tool execution

Handles validation, permissions, hooks, progress events, error wrapping, and tool_result creation.

This is the last runtime gate before model intent reaches the real system.

Context governance

Controls context growth through result budgets, summarization, recovery retries, and session memory.

Long-running tasks mostly depend on how faithfully this layer preserves state.