7 The Observability Service: seeing what your AI does
This chapter covers
- The observability data model: sessions, traces, spans, generations, and scores
- Structured logging that captures AI-specific context
- Following a single user request across the platform with distributed tracing
- Attaching quality scores to production traces
- Cost attribution and budget tracking
Every platform service we've built so far produces valuable signals. The Model Service records token counts and latency on every request. The Session Service tracks conversation lengths and context window utilization. The Data Service measures retrieval relevance scores. The Guardrails Service logs every policy evaluation and its outcome. But these signals exist in isolation. When Sarah's patient intake assistant takes four seconds to respond instead of the usual one second, she can't tell whether the delay came from a slow model call, an expensive vector search, a guardrail evaluation that triggered a secondary classification, or a tool execution that timed out. The data exists somewhere in each service's local logs, but nothing ties it together into a coherent story.
This chapter builds the Observability Service, which transforms isolated signals into unified visibility. It collects logs, metrics, traces, and quality scores from every platform component and makes them queryable through a single interface. It answers operational questions: What happened? How long did it take? How much did it cost? How good was it?