chapter eleven

11 Managing session state and chat history

 

This chapter covers

  • How st.session_state works and why Streamlit needs it
  • Why LLMs do not actually remember conversations
  • What the context window is and what happens when you exceed it
  • Trimming conversation history
  • Adding a robust reset button with confirmation

After using your chatbot for a while, you may notice something unexpected. During a long conversation—say, twenty or thirty exchanges—the AI’s responses start to degrade. It repeats itself, contradicts earlier statements, and takes longer to respond. Your chatbot is not broken. It has hit the context window limit.

This chapter explains why that happens and how to fix it. You’ll learn in depth how Streamlit’s session state works, why LLMs are stateless, what the context window is, and how to manage conversation history to keep your chatbot performing well.

11.1 The problem: LLMs are stateless

Before we look at Streamlit, we will start with the AI model itself. Stateless means the server does not automatically remember what happened in earlier interactions. An LLM does not remember previous exchanges. Each call to ollama.chat() is independent. If you want the model to respond as if it remembers the conversation, your app must send the relevant conversation history with each new request.

11.2 Understanding Streamlit session state

11.2.1 The script rerun problem

11.2.2 Session state in action

11.2.3 What you can store in session state

11.2.4 When session state resets

11.3 How LLMs use conversation history

11.4 Context window management

11.4.1 What happens when you exceed the context window

11.4.2 Implementing history trimming