chapter four

4 Downloading an LLM and having your first conversation

 

This chapter covers

  • Downloading your first AI model
  • Having an interactive conversation with a local LLM
  • Model sizes and their RAM requirements
  • Comparing different model families (Gemma, Llama, Qwen, Mistral)
  • Managing models with essential Ollama commands

At this point, Ollama is installed and the ollama serve process is running in a terminal window. Leave that terminal open and use a second terminal for the commands in this chapter. You have the music player, but no song is loaded yet. In this chapter, you will download an actual AI model, talk to it, and learn how to manage multiple models on your machine.

4.1 Downloading your first model

You will start with one small, reliable model before comparing the alternatives. This will keep your first download manageable, and it will give you a known, good baseline for later chapters.

4.1.1 Choosing a starting model

There are dozens of open source models available through Ollama [1]. For your first download, you will use Gemma 3 [2] with 4 billion parameters, a lightweight model created by Google DeepMind.

4.1.2 Running the pull command

4.1.3 Understanding the model name format

4.2 Interactive chat

4.2.1 Starting interactive chat mode

4.2.2 A sample conversation

4.2.3 Exiting chat mode

4.3 Understanding model sizes and parameters

4.3.1 What are parameters?

4.3.2 The size, capability, and memory tradeoff

4.3.3 What “quantization” means