2 Getting started improving agents
This chapter covers
- Building the agent and the text prompt artifact to improve
- Measuring the output signal using LLM-as-Judge
- Searching for a better text prompt variant with Self-Supervised Prompt Optimization (SPO)
- Completing the offline improvement loop and running self-improvement
In the last chapter, we laid the foundations and the mental model for what a self-improving agent (harness) is and does. In this chapter, we put that model to work on the simplest artifact there is: a RAG (retrieval-augmented generation) agent's prompt instruction. The whole job is to take a baseline instruction and let an offline loop improve it for us.
To get there, we start with the basics of building agents in the Helix framework and the provided harness. From there, we compose the loop one block at a time, beginning with the signal: the component that scores the agent's output with an LLM-as-Judge. Then we move to the next block, search, where SPO proposes better prompt variants for the judge to pick from.
Finally, we wire the pieces together and run the full offline improvement loop, so the agent’s prompt is improved end-to-end. By the end of the chapter, you will have built and exercised that loop yourself, with Helix, an LLM-as-Judge, and SPO handling the search.