chapter nine

9 Recording audio and transcribing speech with MLX Whisper

This chapter covers

  • What speech recognition is and how Whisper achieves human-level accuracy
  • Why MLX Whisper runs faster on Apple silicon
  • How digital audio works
  • How to record from your microphone using Python
  • Transcribing speech to text with a single function call
  • Choosing the right Whisper model size
  • Building a reusable voice transcription script

You have spent the past eight chapters building the text component of your voice AI: the LLM backend, the Python layer, and the Streamlit web interface. In this chapter, you will add the audio component. You will create a standalone Python script that listens to your microphone, converts your speech to text, and prints the transcription. That capability will be the foundation for the voice-enabled chatbot application you will build in chapter 10.

9.1 What is speech recognition?

Speech recognition, also called automatic speech recognition (ASR), is the task of converting audio into the text that the speech represents. It sounds simple, but it is one of the hardest problems in AI: speakers vary in accent, pace, and pronunciation; background noise corrupts recordings; words run together; and the same sound can mean different things depending on the context.

9.2 Why use MLX Whisper on Apple silicon?

9.3 Installing MLX Whisper

9.4 Understanding digital audio

9.4.1 Audio sample rate

9.4.2 Audio channels

9.4.3 Sample bit depth and data type

9.5 Recording from the microphone

9.5.1 sd.rec(): The recording call

Exercises