Skip to main content

Lab · LLM VRAM · Apple

Computed

What LLMs can the Apple M3 Ultra (512GB) run?

62 of 62 rostered models have a usable on-device fit. The table keeps all 62 visible, including 0 below-context fits and 0 that exceed this device, so “not listed” never masquerades as an answer.

Published memory

512 GB

LPDDR5 unified

Assumed usable

384.0 GiB

75% of capacity

Peak bandwidth

819 GB/s

theoretical device spec

Usable fits

62 / 62

largest: 298.8B at Q8_0

Every model, solved on this accelerator

Usable rows fit weights plus at least 4k tokens, capped at the model’s own trained window. Tight rows hold weights but fall below that floor. The speed column is a memory-bandwidth ceiling, not a benchmark.

62 usable · 0 tight · 0 no fit

ModelParametersOn-device resultSelected weightsMax contextDecode ceiling · read per token
Hy3HY V3298.8BQuantized fitQ8_0 · 295.8 GiBpublished file256k35 tok/sreads 21.6 GiB/token
Laguna S 2.1Laguna117.6BFull precision · comfortableFP16 / BF16 · 219.0 GiBpublished file1.0M55 tok/sreads 13.9 GiB/token
gpt-oss 120Bgpt-oss116.8BFull precision · comfortableAs released · 60.8 GiBpublished file128k234 tok/sreads 3.3 GiB/token
Qwen2.5 72B InstructQwen72.7BFull precision · comfortableFP16 / BF16 · 135.4 GiBpublished file32k6 tok/sreads 137.9 GiB/token
DeepSeek-R1-Distill-Llama 70BDeepSeek70.6BFull precision · comfortableFP16 / BF16 · 131.4 GiBpublished file128k6 tok/sreads 133.9 GiB/token
Llama 3.1 70B InstructLlama70.6BFull precision · comfortableFP16 / BF16 · 131.4 GiBpublished file128k6 tok/sreads 133.9 GiB/token
Llama 3.1 Nemotron 70B InstructNemotron70.6BFull precision · comfortableFP16 / BF16 · 131.4 GiBpublished file128k6 tok/sreads 133.9 GiB/token
Llama 3.3 70B InstructLlama70.6BFull precision · comfortableFP16 / BF16 · 131.4 GiBpublished file128k6 tok/sreads 133.9 GiB/token
Mixtral 8x7B InstructMistral46.7BFull precision · comfortableFP16 / BF16 · 87.0 GiBpublished file32k31 tok/sreads 25.0 GiB/token
Hermes 4.3 36BSeed OSS36.2BFull precision · comfortableFP16 / BF16 · 67.3 GiBpublished file512k11 tok/sreads 69.3 GiB/token
KAT Coder V2.5 DevQwen3 5 MOE Text34.7BFull precision · comfortableFP16 / BF16 · 64.6 GiBpublished file256k128 tok/sreads 5.9 GiB/token
Qwen AgentWorld 35B A3BQwen3 5 MOE Text34.7BFull precision · comfortableFP16 / BF16 · 64.6 GiBpublished file256k128 tok/sreads 5.9 GiB/token
Yi 1.5 34B ChatYi34.4BFull precision · comfortableFP16 / BF16 · 64.1 GiBpublished file4k12 tok/sreads 65.0 GiB/token
Laguna XS 2.1Laguna33.4BFull precision · comfortableFP16 / BF16 · 62.3 GiBpublished file256k147 tok/sreads 5.2 GiB/token
DeepSeek-R1-Distill-Qwen 32BDeepSeek32.8BFull precision · comfortableFP16 / BF16 · 61.0 GiBpublished file128k12 tok/sreads 63.0 GiB/token
Qwen2.5 32B InstructQwen32.8BFull precision · comfortableFP16 / BF16 · 61.0 GiBpublished file32k12 tok/sreads 63.0 GiB/token
Qwen2.5-Coder 32B InstructQwen32.8BFull precision · comfortableFP16 / BF16 · 61.0 GiBpublished file32k12 tok/sreads 63.0 GiB/token
QwQ 32BQwen32.8BFull precision · comfortableFP16 / BF16 · 61.0 GiBpublished file40k12 tok/sreads 63.0 GiB/token
Qwen3 32BQwen32.8BFull precision · comfortableFP16 / BF16 · 61.0 GiBpublished file40k12 tok/sreads 63.0 GiB/token
Qwen3 30B-A3BQwen30.5BFull precision · comfortableFP16 / BF16 · 56.9 GiBpublished file40k109 tok/sreads 7.0 GiB/token
Gemma 3 27B InstructGemma27.4BFull precision · comfortableFP16 / BF16 · 51.1 GiBpublished file128k15 tok/sreads 52.1 GiB/token
Gemma 2 27B InstructGemma27.2BFull precision · comfortableFP16 / BF16 · 50.7 GiBpublished file8k14 tok/sreads 53.6 GiB/token
Dolphin Mistral 24B Venice EditionMistral24BFull precision · comfortableFP16 / BF16 · 44.7 GiBpublished file128k17 tok/sreads 46.0 GiB/token
Mistral Small 24B InstructMistral23.6BFull precision · comfortableFP16 / BF16 · 43.9 GiBpublished file32k17 tok/sreads 45.2 GiB/token
gpt-oss 20Bgpt-oss21.5BFull precision · comfortableAs released · 12.8 GiBpublished file128k277 tok/sreads 2.8 GiB/token
DeepSeek-R1-Distill-Qwen 14BDeepSeek14.8BFull precision · comfortableFP16 / BF16 · 27.5 GiBpublished file128k26 tok/sreads 29.0 GiB/token
Qwen2.5 14B InstructQwen14.8BFull precision · comfortableFP16 / BF16 · 27.5 GiBpublished file32k26 tok/sreads 29.0 GiB/token
Qwen3 14BQwen14.8BFull precision · comfortableFP16 / BF16 · 27.5 GiBpublished file40k27 tok/sreads 28.8 GiB/token
Phi-4 14BPhi14.7BFull precision · comfortableFP16 / BF16 · 27.3 GiBpublished file16k26 tok/sreads 28.9 GiB/token
OLMo 2 13B InstructOLMo13.7BFull precision · comfortableFP16 / BF16 · 25.5 GiBpublished file4k27 tok/sreads 28.7 GiB/token
Mistral Nemo 12B InstructMistral12.2BFull precision · comfortableFP16 / BF16 · 22.8 GiBpublished file128k32 tok/sreads 24.1 GiB/token
Gemma 3 12B InstructGemma12.2BFull precision · comfortableFP16 / BF16 · 22.7 GiBpublished file128k32 tok/sreads 23.5 GiB/token
Qwythos 9B Claude Mythos 5 1MQwen3 5 Text9.4BFull precision · comfortableFP16 / BF16 · 17.5 GiBpublished file1.0M41 tok/sreads 18.5 GiB/token
Gemma 2 9B InstructGemma9.2BFull precision · comfortableFP16 / BF16 · 17.2 GiBpublished file8k38 tok/sreads 19.8 GiB/token
Qwen3 8BQwen8.2BFull precision · comfortableFP16 / BF16 · 15.3 GiBpublished file40k47 tok/sreads 16.4 GiB/token
Granite 3.3 8B InstructGranite8.2BFull precision · comfortableFP16 / BF16 · 15.2 GiBpublished file128k46 tok/sreads 16.5 GiB/token
DeepSeek-R1-Distill-Llama 8BDeepSeek8BFull precision · comfortableFP16 / BF16 · 15.0 GiBpublished file128k48 tok/sreads 16.0 GiB/token
Llama 3.1 8B InstructLlama8BFull precision · comfortableFP16 / BF16 · 15.0 GiBpublished file128k48 tok/sreads 16.0 GiB/token
DeepSeek-R1-Distill-Qwen 7BDeepSeek7.6BFull precision · comfortableFP16 / BF16 · 14.2 GiBpublished file128k52 tok/sreads 14.6 GiB/token
Qwen2.5 7B InstructQwen7.6BFull precision · comfortableFP16 / BF16 · 14.2 GiBpublished file32k52 tok/sreads 14.6 GiB/token
Qwen2.5-Coder 7B InstructQwen7.6BFull precision · comfortableFP16 / BF16 · 14.2 GiBpublished file32k52 tok/sreads 14.6 GiB/token
Falcon3 7B InstructFalcon7.5BFull precision · comfortableFP16 / BF16 · 13.9 GiBpublished file32k52 tok/sreads 14.8 GiB/token
Mistral 7B Instruct v0.3Mistral7.2BFull precision · comfortableFP16 / BF16 · 13.5 GiBpublished file32k53 tok/sreads 14.5 GiB/token
Gemma 3 4B InstructGemma4.3BFull precision · comfortableFP16 / BF16 · 8.0 GiBpublished file128k92 tok/sreads 8.3 GiB/token
Nanbeige4.2 3BNanbeige4.2BFull precision · comfortableFP16 / BF16 · 7.8 GiBpublished file256k90 tok/sreads 8.5 GiB/token
Qwen3 4BQwen4BFull precision · comfortableFP16 / BF16 · 7.5 GiBpublished file40k89 tok/sreads 8.6 GiB/token
Phi-4-mini 3.8B InstructPhi3.8BFull precision · comfortableFP16 / BF16 · 7.1 GiBpublished file128k94 tok/sreads 8.1 GiB/token
Phi-3.5-mini 3.8B InstructPhi3.8BFull precision · comfortableFP16 / BF16 · 7.1 GiBpublished file128k75 tok/sreads 10.1 GiB/token
Llama 3.2 3B InstructLlama3.2BFull precision · comfortableFP16 / BF16 · 6.0 GiBpublished file128k111 tok/sreads 6.9 GiB/token
Qwen2.5 3B InstructQwen3.1BFull precision · comfortableFP16 / BF16 · 5.7 GiBpublished file32k127 tok/sreads 6.0 GiB/token
Qwen3 1.7BQwen2BFull precision · comfortableFP16 / BF16 · 3.8 GiBpublished file40k164 tok/sreads 4.7 GiB/token
DeepSeek-R1-Distill-Qwen 1.5BDeepSeek1.8BFull precision · comfortableFP16 / BF16 · 3.3 GiBpublished file128k216 tok/sreads 3.5 GiB/token
SmolLM2 1.7B InstructSmolLM1.7BFull precision · comfortableFP16 / BF16 · 3.2 GiBpublished file8k163 tok/sreads 4.7 GiB/token
Qwen2.5 1.5B InstructQwen1.5BFull precision · comfortableFP16 / BF16 · 2.9 GiBpublished file32k247 tok/sreads 3.1 GiB/token
Llama 3.2 1B InstructLlama1.2BFull precision · comfortableFP16 / BF16 · 2.3 GiBpublished file128k299 tok/sreads 2.6 GiB/token
TinyLlama 1.1B ChatTinyLlama1.1BFull precision · comfortableFP16 / BF16 · 2.0 GiBpublished file2k365 tok/sreads 2.1 GiB/token
MiniCPM5 1BLlama1.1BFull precision · comfortableFP16 / BF16 · 2.0 GiBpublished file128k347 tok/sreads 2.2 GiB/token
Gemma 3 1B InstructGemma1000MFull precision · comfortableFP16 / BF16 · 1.9 GiBpublished file32k401 tok/sreads 1.9 GiB/token
Qwen3 0.6BQwen752MFull precision · comfortableFP16 / BF16 · 1.4 GiBpublished file40k335 tok/sreads 2.3 GiB/token
Qwen2.5 0.5B InstructQwen494MFull precision · comfortableFP16 / BF16 · 942 MiBpublished file32k752 tok/sreads 1.0 GiB/token
SmolLM2 360M InstructSmolLM362MFull precision · comfortableFP16 / BF16 · 690 MiBpublished file8k773 tok/sreads 1010 MiB/token
SmolLM2 135M InstructSmolLM135MFull precision · comfortableFP16 / BF16 · 257 MiBpublished file8k1,789 tok/sreads 437 MiB/token

Device specification and assumption

Memory
512 GB LPDDR5 unified
Bandwidth
819 GB/s
Usable budget
75% → 384.0 GiB

Capacity and bandwidth come from the vendor specification. The usable fraction is an explicit planning assumption, not a device specification.

What the table does — and does not — claim

Weight sizes use published checkpoint or GGUF files where the roster has one, otherwise the documented bits-per-weight calculation. Context comes from each model’s layer and KV-head geometry. The decode figure is the card’s peak bandwidth divided by the bytes one token reads — the weights it routes through plus one pass over the KV cache, printed under each ceiling — so it is a roofline bound, not measured application throughput. Multi-GPU splitting and host-memory offload are outside this on-device table.

Questions about this card

What LLMs can the Apple M3 Ultra (512GB) run?

62 of the 62 open models in this roster have an on-device configuration with usable context. The largest by exact parameter count is Hy3 at Q8_0, with room for 256k tokens. “Largest” describes parameter count, not model quality or task performance.

How much memory is usable on the Apple M3 Ultra (512GB)?

The device publishes 512 GB of LPDDR5 unified. This calculator budgets 75%, or 384.0 GiB, for model weights and KV cache; the remainder is an explicit allowance for the runtime, driver or operating system, workspace, and display. It is an assumption rather than a vendor specification.

Are the speed figures benchmarks for the Apple M3 Ultra (512GB)?

No. They are bandwidth-bound roofline ceilings: 819 GB/s divided by the bytes one decoded token actually reads, which the table prints beside every ceiling. That is not the size of the file on disk — a token reads the weights it routes through, all of them for a dense model but only the selected experts for a mixture-of-experts, plus one pass over the KV cache. gpt-oss 120B holds 60.8 GiB resident at As released yet reads 3.3 GiB/token, which is what produces its 234 tok/s. Real throughput is lower because kernels, cache traffic, prompt processing, scheduling, and runtime overhead also consume time.