chapter ten

10 Building a voice-enabled AI chat application

 

This chapter covers

  • The architecture of a voice AI pipeline
  • Connecting MLX Whisper to Streamlit
  • Refactoring the app into functions
  • Adding a text fallback so the app also works with keyboard input
  • Testing the full voice conversation loop
  • Understanding what makes this application challenging and how to extend it

This is the chapter where everything comes together. In the previous chapters, you learned how to use the terminal, install Ollama, pull AI models, write Python, call the Ollama API, build web interfaces with Streamlit, and transcribe speech with MLX Whisper. In this chapter, you will combine all of those skills into a voice-­enabled AI chat application: speak into your microphone, watch your words become text, and receive a streaming AI response, all running locally on your Mac, with complete privacy. You will continue working in the same my-ai-chatbot folder you used earlier; this chapter will add voice_chat.py in a ch10 subfolder, while voice_input.py file from chapter 9 remains in the project root.

10.1 The voice AI pipeline

10.2 Setting up the voice chat application project

10.3 Building the application step by step

10.3.1 Stage 1: Transcription in Streamlit

10.3.2 Why use a temporary WAV file?

10.3.3 Stage 2: Connecting transcription to the LLM

10.3.4 Stage 3: The complete application

10.4 The complete voice chat application

10.4.1 Running the application

10.4.2 Updating the Ollama model menu

10.5 Understanding the complete voice chat application code

10.5.1 Browser page configuration

10.5.2 The sidebar: Controlling two models

10.5.3 transcribe_audio(): The bridge between Streamlit and MLX Whisper

10.5.4 stream_response(): A reusable streaming function

10.5.5 handle_user_message(): The unified entry point