chapter ten
10 Deploying and monitoring
This chapter covers
- Seeing how LLMOps differs from traditional software operations
- Choosing between hosted APIs and self-hosted models
- Building hybrid deployment architectures that optimize for both cost and capability
- Implementing model-native monitoring systems that track response quality, user satisfaction, and business impact
- Designing automated quality assurance pipelines to maintain output standards at scale
At 3:04 AM, an alert arrives that no one wants to see:
URGENT: AI chatbot billing alert – $47,000 this month. System failing.
Just days before, the company’s new large language model (LLM)-powered support assistant had been a success story in the making. It sailed through internal testing, impressed executives, and promised to reduce support costs dramatically. Now it’s producing unpredictable results, racking up massive expenses, and creating more confusion than value.