Part 7 · 1 chapters · ~8 min

Running a Model Locally

Choosing a model by memory, quantisation formats (GGUF Q4_K_M and friends), predicting speed from memory bandwidth, llama.cpp, Ollama and MLX, OpenAI-compatible local APIs, embedding models locally for private search, and the trade-offs against hosted models.

8

Commands and arithmetic

code
# not run on this machine (no local runtime installed); commands as documented by each project
brew install ollama && ollama serve &
ollama run llama3.1:8b "Explain a KV cache in two sentences"
curl http://localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"hello"}]}'     # OpenAI-compatible

# llama.cpp directly
llama-server -m Llama-3.1-8B-Instruct-Q4_K_M.gguf -c 8192 --port 8080

# MLX on Apple silicon
pip install mlx-lm && mlx_lm.generate --model mlx-community/Meta-Llama-3.1-8B-Instruct-4bit --prompt "hello"

# will it fit on this 18 GB machine?  8B Q4 ≈ 4.5-5 GB + 1 GiB KV at 8k  → yes
#                                      70B Q4 ≈ 35+ GB                    → no
# expected speed ceiling: 150 GB/s ÷ ~4.5 GB ≈ 33 tokens/s (real numbers land below the ceiling)
RUNNING A MODEL LOCALLY
what fits, how fast, with which tool
memory18 GB laptop (this machine): 8B4-bit fits; 70B 4-bit (~35 GB)does not.speed boundBandwidth ÷ model bytes: ~33tokens/s ceiling for 8B 4-bit at150 GB/s.llama.cppC/C++ inference, GGUF quantisedfiles, Metal and CUDA backends.OllamaA friendly wrapper: ollama run, anHTTP API on :11434.MLXApple's array framework; fast onApple silicon.why localPrivacy, offline use, no per-tokencost, experimentation.
swipe the figure sideways, or tap expand for full screen
1/4
will it fit?
Weights at the chosen precision plus KV cache plus runtime overhead must fit in memory: an 18 GB machine runs 8B models in 4-bit comfortably, not 70B.
weights + KV + overhead8B yes, 70B no