11 Managing session state and chat history
This chapter covers
After using your chatbot for a while, you may notice something unexpected. During a long conversation—say, twenty or thirty exchanges—the AI’s responses start to degrade. It repeats itself, contradicts earlier statements, and takes longer to respond. Your chatbot is not broken. It has hit the context window limit.
This chapter explains why that happens and how to fix it. You’ll learn in depth how Streamlit’s session state works, why LLMs are stateless, what the context window is, and how to manage conversation history to keep your chatbot performing well.
11.1 The problem: LLMs are stateless
Before we look at Streamlit, we will start with the AI model itself. Stateless means the server does not automatically remember what happened in earlier interactions. An LLM does not remember previous exchanges. Each call to ollama.chat() is independent. If you want the model to respond as if it remembers the conversation, your app must send the relevant conversation history with each new request.