Originally published on the Macyou blog. Disclosure up front: I run Macyou, we rent dedicated Apple Silicon Macs for AI work. This post is about the hardware math, which is the same whether the Mac sits on your desk or in a rack.
The short answer: any Apple Silicon Mac with 16 GB of unified memory runs 7B to 14B models well, a 64 GB M4 Pro runs 70B class models at usable speed, and a 256 GB Mac Studio M3 Ultra runs 200B class models that no single consumer GPU can hold. Generation speed is set almost entirely by memory bandwidth, so the chip tier matters more than the year it came out.
Which Mac Runs Which Model: The Memory Rule
A model has to fit in unified memory at the quantization you choose, with a few gigabytes left for the context window and macOS. Once it fits, tokens per second scale with memory bandwidth. The base M4 and M5 Pro rows below are measured on our own machines (methodology and raw JSON, CC BY 4.0); the other rows come from the same formula, which holds within 4.6% on both measured machines.
| Mac | Bandwidth | Comfortable models (Q4) | What to expect |
| M4 Mac mini, 16 GB | 120 GB/s | 3B to 14B | Measured: Llama 3.2 3B 46.7 tok/s, Llama 3.1 8B 21.2, Qwen 2.5 14B 11.7 |
| M4 Mac mini, 24 to 32 GB | 120 GB/s | 14B comfortably, 32B at the edge | Same speeds as 16 GB: extra memory buys model size, not tok/s |
| M4 Pro Mac mini, 48 to 64 GB | 273 GB/s | 32B comfortably, 70B Q4 at 64 GB | About 2x the base M4 at equal model size, 70B at 5 to 6 tok/s |
| M5 Pro, 48 GB | 307 GB/s | 32B comfortably | Measured: Qwen3 8B 47.8 tok/s, Qwen 2.5 14B 29.4 |
| M4 Max Mac Studio, 128 GB | 546 GB/s | 70B Q8, 123B Q4 | About 3.5x the base M4 at equal model size, 70B Q4 at about 11 tok/s |
| M3 Ultra Mac Studio, 256 GB | 819 GB/s | 200B class Q4, 70B FP16 | Largest single box option; 405B still needs clustering |
The RAM Math: How Much Memory a Model Needs
Weights
Weights in GB = parameters (billions) × bits per weight / 8, plus about 15% runtime overhead. At Q4_K_M (about 4.85 bits per weight):
| Model size | Weights at Q4_K_M | In memory with overhead |
| 8B | 4.9 GB | about 6 GB |
| 32B | about 19 GB | about 22 GB |
| 70B | about 42 GB | about 49 GB |
| 123B | about 75 GB | about 86 GB |
Q8_0 roughly doubles those numbers, FP16 roughly quadruples them.
Context (the KV cache)
The context window costs memory on top of the weights, and this is where most "it fit yesterday" surprises come from. For Llama 3 8B the KV cache at fp16 is:
32 layers × 8 KV heads × 128 dims × 2 bytes × 2 (K and V)
= 131,072 bytes = 128 KB per token
32,768 tokens × 128 KB = 4 GB
So an 8B model that needs 6 GB at a short prompt needs about 10 GB at 32k context. That is why 16 GB tops out at 14B, 64 GB is the 70B threshold, and 128 GB is where 100B+ dense models become practical.
Why Speed Depends on Bandwidth, Not GPU Cores
Generating one token means reading every active weight once. A 4.9 GB model on a 120 GB/s bus can never exceed about 24 tok/s, and we measured 21.2. Fitting eight dense runs on two machines gives one line that predicts all of them:
seconds per token = weights_GB / (bandwidth_GB_s × 0.9075) + 0.00325
Example: Llama 3.1 8B Q4 on a base M4 (4.87 GB of weights)
4.87 / (120 × 0.9075) + 0.00325 = 0.0480 s → 20.9 tok/s predicted, 21.2 measured
The first number is how much of the theoretical bandwidth a real run gets; the second is a fixed cost per token that does not depend on model size. The constants were fitted on the base M4 first. The M5 Pro was not in that fit, and at two and a half times the bandwidth the same line held within 4.1%.
Two consequences:
- More GPU cores on the same bandwidth barely help generation.
- Mixture of experts models read only their active experts per token, so they run far faster than their total parameter count suggests. They do read more than the active weights alone: on Qwen3 30B A3B we measured 1.3x the active size crossing the bus, 78 tok/s where the naive rule said 91.
Prompt processing is the exception: it is compute bound, and it varies about 2x between model families at equal size. We measured Qwen 2.5 7B at 1,130 prompt tok/s against Llama 3.1 8B at 587. If your workload is long prompt, short answer (RAG, classification), that gap matters more than generation speed.
Ollama vs LM Studio vs MLX vs llama.cpp
Ollama
The default for anything headless or scripted: one command to pull a model, a REST API on port 11434, and an OpenAI compatible endpoint. It runs llama.cpp underneath.
LM Studio
The best GUI: model browser, chat window, and a local server that speaks the OpenAI API. Same engine class as Ollama, so the same speeds.
MLX
Apple's own array framework for Apple Silicon. Fastest on some models and the natural choice for fine tuning on a Mac. It is a Python library, not an app.
llama.cpp
The raw engine when you want every flag, the newest quant formats, or a C/C++ embed.
Five Minute Setup: Run Llama 3.1 8B with Ollama
# Install and start the Ollama server (Homebrew)
brew install ollama
ollama serve &
# Download the model (about 4.9 GB at Q4) and ask it something
ollama pull llama3.1:8b
ollama run llama3.1:8b "Explain unified memory in two sentences."
The server also exposes an OpenAI compatible endpoint, so existing code works without changes:
# Same request format as the OpenAI Chat Completions API
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"Hi"}]}'
From Python, point the official OpenAI SDK at the local server. The API key is required by the SDK but ignored by Ollama:
from openai import OpenAI
# base_url points at the local Ollama server instead of api.openai.com
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
reply = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Why does memory bandwidth limit tok/s?"}],
)
print(reply.choices[0].message.content)
Before pulling a bigger model, do the RAM math: a 70B build on a 16 GB machine downloads 40 GB and then fails to load. The Mac LLM calculator does it for any model and chip, context included.
Common Mistakes When Running LLMs on a Mac
- Buying GPU cores instead of memory. A 24 GB Mac with more GPU cores runs the same 8B model no faster than a 16 GB one on the same chip; the next tier of bandwidth is what changes speed.
- Ignoring the context window. A model that "fits" with 1 GB to spare will swap and crawl the moment you paste a long document. Leave 2 to 4 GB free, more for RAG.
- Running production on a laptop. Thermal throttling, sleep and a residential uplink turn a 21 tok/s machine into an unreliable one. Anything that needs to be up around the clock belongs on a desktop class Mac with a real network connection, yours or rented; the trade off is laid out in buying a Mac mini vs renting one.
FAQ: Local LLMs on Apple Silicon
Can a 16 GB Mac run a 70B model?
No. A 70B model needs about 42 GB for weights alone at Q4. On 16 GB the practical ceiling is 14B at Q4 with a moderate context.
Is an M4 Pro worth it over a base M4 for LLMs?
For generation speed, yes: 273 GB/s against 120 GB/s is about 2.3x the bandwidth, and generation follows bandwidth. For memory it depends on the configuration you buy, not the chip.
Does the chip generation matter more than the tier?
Only as much as the bandwidth changes. Compare chips by GB/s, not by year: moving up a tier (base to Pro to Max to Ultra) changes bandwidth far more than a new generation of the same tier does.
Why is my measured speed lower than the estimate?
The formula assumes nothing else is competing for the memory bus. A browser with many tabs, a build running in the background, or swapping from an oversized context will all pull the number down.
Ollama or MLX?
Ollama for serving models to apps and scripts; MLX when you want to fine tune or work in Python directly. Test your specific model on both if speed matters, results vary by architecture.
Conclusion
On a Mac, unified memory decides which models you can run and memory bandwidth decides how fast they run.
A 16 GB machine is a solid 7B to 14B box; 64 GB is the entry point for 70B; 128 GB and up opens 100B+ dense models.
The speed formula, fitted on a base M4 and confirmed on an M5 Pro within 4.1%, predicts tokens per second from weight size and bandwidth alone.
Mixture of experts models run faster than their size suggests, but read about 1.3x their active weights, not just the active weights.
Budget memory for the KV cache: 32k of context adds about 4 GB to an 8B model.
Ollama is the quickest path to a working OpenAI compatible endpoint; MLX and llama.cpp are there when you need more control.
Questions about a specific model or Mac: ask in the comments, I'll answer with numbers where we have them.