1 Foundations of self-improving agents
The chapter introduces self-improving agents as systems that get better without retraining the underlying model. The core idea is that most of the useful improvement lives in the surrounding harness and other editable artifacts, not in the frozen weights. That shift matters because production agents are usually improved manually, through slow cycles of user complaints, prompt edits, redeployments, and more complaints, which is both inefficient and noisy.
To replace that workflow, the chapter presents a closed improvement loop: the agent acts, a signal measures the gap between output and desired behavior, search proposes a better variant, and the winner is applied. This loop can run offline with a deployment gate for safety, or online for fast, limited adaptations. The chapter emphasizes that self-improvement is not just self-evaluation; the score must lead to search and then to an actual change in the running system.
The rest of the chapter organizes the space into two main design choices and four artifact layers. The three key dials are signal, search, and artifact type, while the editable artifacts include context, memory, metacognitive scaffolds, and code or tools. Signals range from precise but scarce ground truth to widely available but noisier judge-based methods, and search methods range from cheap hill-climbing to expensive evolutionary or meta-search approaches. Together, these ideas provide the foundation for building agents that improve their prompts, memories, reasoning routines, and tool logic in a disciplined, versioned way.
The harness and the agent shows the boundary and purpose for the harness with respect to the agent worker.
The manual improvement workflow, as many agents are improved today: ship, wait for user feedback, review it by hand, adjust the prompt, redeploy, and wait again.
A conceptual schematic, not a precise empirical chart. Individual reasoning capacity barely moves on biological timescales, while collective capability compounds through accumulated infrastructure. A self-improving agent improves the infrastructure curve, not the substrate one.
One turn of the improvement loop, run offline on the support-agent task. The agent acts, a judge signals the gap, a search proposes prompt variants, and the judge picks the winner, and the winner passes a deploy gate before reaching the running agent on the next task.
The three dials of the improvement loop. The signal dial sets how the gap is measured, the search dial sets how a better variant is found, and both belong to the L0 toolkit, which is reused at every layer. The artifact dial sets what is being changed, and it is the dial that the book turns as we get deeper into the book.
The same support-agent loop, run online. The cycle is the same four steps and the same closed return, but the deploy gate dissolves and improvements reach the running agent continuously. Online loops can only safely change cheap, reversible artifacts, prompts and memory, not reasoning scaffolds or code.
shows a comparison between the offline and online self-improvement loop where the significant difference is the deployment gate that manages deployments of improved variants
The harness artifact layers of self-improvement. The frozen model (LLM) serves as the unchanging substrate below. The four artifact layers stacked on it are shared infrastructure, inherited by every instance of the agent. Layer zero is the toolkit, the engineering work that measures, searches, and applies, reused at every layer.
The signal family arranged by precision against availability. Ground-truth and environment signals are the most precise but the rarest; judge-based signals are available almost everywhere but are noisier, and process and reflection signals assess how the agent worked rather than only what it produced.
The signal-quality ceiling. With only an intrinsic signal, an agent reflecting on its own output, improvement rises for two or three rounds and then flattens. Adding fresh information from outside the loop, ground truth, an execution result, or new evidence, raises the ceiling and lets the loop keep climbing.
The search family arranged by cost against improvement reach. Hill-climbing is the cheapest and most limited, pairwise and population methods cost more and reach further, reinforcement-learning searches fit policy-shaped artifacts, and evolutionary archive search is the most expensive and reaches the furthest. Meta-search sits apart, since it changes the search itself rather than spending more effort on the same curve.
Summary
- Most agents in production are improved by a slow, manual loop in which a human reads user feedback, edits a prompt by hand, redeploys, and waits for the next piece of feedback.
- That manual loop is fragile because user feedback can be wrong, and chasing incorrect feedback can quietly make an agent worse.
- A self-improving agent replaces the noisy human-complaint signal with structured evaluation the agent can produce on its own, so the loop runs continuously rather than waiting for a person.
- A self-improving agent improves the infrastructure, the scaffolding (harness) around a frozen model, never the model itself, so no fine-tuning, RLHF, or weight updates are involved.
- An agent harness or just the harness is the scaffolding software that wraps the LLM and encapsulates the agent.
- Agent intelligence is better understood as societal than individual, the way humanity advanced by building shared infrastructure rather than by growing bigger brains; a single agent runs as many shared instances and improves the same way.
- The infrastructure the agent harness improves is made of artifacts, organized into four layers: L1 text, L2 memory, L3 metacognition, and L4 code and tools.
- Artifacts compose across layers, so a higher-layer artifact like a metacognitive scaffold is usually built out of lower-layer prompts and memory references, and improving a layer can mean swapping the artifact or a component it contains.
- Shallow layers are cheap and safe to change; deep layers are powerful but harder to undo, so the book turns the artifact dial from shallow to deep.
- Self-improvement runs as a closed loop with four steps: act, signal, search, apply, then repeat.
- Closing the loop is what separates self-improvement from self-evaluation, which only measures the gap without acting on it.
- The loop has three independent dials: the signal that measures the gap, the search that finds a better variant, and the artifact being changed.
- The signal and search dials are the general-purpose L0 toolkit, reused at every layer, while the artifact dial is the one the book deliberately turns.
- The loop runs in two cadences: offline loops gate improvements through deploy review and can touch any artifact, while online loops apply improvements live but stay on cheap, reversible artifacts like text and memory.
- The signal family sorts into three groups that trade precision against availability: ground-truth and environment signals (labeled data, test outcomes, tool grounding), judge-based signals (absolute scoring, pairwise preference, contrastive judging), and process and reflection signals (Process Reward Models, verbal self-critique).
- A loop can only improve an agent as far as its signal can see, and intrinsic self-judgment saturates after a few rounds because of the generation-verification gap.
- The search family sorts into five groups that trade cost against improvement reach: hill-climbing (Karpathy autoresearch), pairwise and population methods (SPO, GEPA, MIPROv2, AFlow), reinforcement-learning searches (MemRL, GRPO), evolutionary archive search (DGM, AlphaEvolve, ShinkaEvolve), and meta-search (HyperAgents).
- Most published search techniques generalize across artifact layers when the artifact can be represented, mutated, evaluated, versioned, and rolled back.
Self-Improving Agents ebook for free