4 Downloading an LLM and having your first conversation
This chapter covers
- Downloading your first AI model
- Having an interactive conversation with a local LLM
- Model sizes and their RAM requirements
- Comparing different model families (Gemma, Llama, Qwen, Mistral)
- Managing models with essential Ollama commands
At this point, Ollama is installed and the ollama serve process is running in a terminal window. Leave that terminal open and use a second terminal for the commands in this chapter. You have the music player, but no song is loaded yet. In this chapter, you will download an actual AI model, talk to it, and learn how to manage multiple models on your machine.
4.1 Downloading your first model
You will start with one small, reliable model before comparing the alternatives. This will keep your first download manageable, and it will give you a known, good baseline for later chapters.
4.1.1 Choosing a starting model
There are dozens of open source models available through Ollama [1]. For your first download, you will use Gemma 3 [2] with 4 billion parameters, a lightweight model created by Google DeepMind.
NOTE Google released Gemma 4 after this chapter’s tested workflow was prepared. Gemma 4 is newer and more capable and is available under the Apache 2.0 license. This chapter keeps gemma3:4b as the main path so that the commands and results remain consistent on modest Mac hardware; if you want to compare newer models, you should check the current Gemma tags in the Ollama library.