Part 7 · 1 chapters · ~8 min
Running a Model Locally
Choosing a model by memory, quantisation formats (GGUF Q4_K_M and friends), predicting speed from memory bandwidth, llama.cpp, Ollama and MLX, OpenAI-compatible local APIs, embedding models locally for private search, and the trade-offs against hosted models.
8
Commands and arithmetic
code
# not run on this machine (no local runtime installed); commands as documented by each project
brew install ollama && ollama serve &
ollama run llama3.1:8b "Explain a KV cache in two sentences"
curl http://localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"hello"}]}' # OpenAI-compatible
# llama.cpp directly
llama-server -m Llama-3.1-8B-Instruct-Q4_K_M.gguf -c 8192 --port 8080
# MLX on Apple silicon
pip install mlx-lm && mlx_lm.generate --model mlx-community/Meta-Llama-3.1-8B-Instruct-4bit --prompt "hello"
# will it fit on this 18 GB machine? 8B Q4 ≈ 4.5-5 GB + 1 GiB KV at 8k → yes
# 70B Q4 ≈ 35+ GB → no
# expected speed ceiling: 150 GB/s ÷ ~4.5 GB ≈ 33 tokens/s (real numbers land below the ceiling)RUNNING A MODEL LOCALLY
what fits, how fast, with which tool
swipe the figure sideways, or tap expand for full screen
1/4
will it fit?
Weights at the chosen precision plus KV cache plus runtime overhead must fit in memory: an 18 GB machine runs 8B models in 4-bit comfortably, not 70B.
weights + KV + overhead8B yes, 70B no