Lab · On-device AI · LLM VRAM
ComputedWill it fit?
Weight bytes, KV-cache growth and the bandwidth ceiling on decode speed, worked out for 68 open models against 22 real accelerators. Every architecture figure is read from the model's own published config; nothing is estimated and nothing is written by a language model.
Pick a pairing
68 models × 22 accelerators. Every combination below is a page that already exists and is already solved — the panel names the exact pairing it resolved to, never a nearby one.
Resolved pairing
DeepSeek-R1-Distill-Llama 8B on GeForce RTX 4090
Comfortable
The unquantized weights fit and leave room for a long context.
- Quantization
- FP16 / BF16
- largest that fits with usable context
- Weights
- 15.0 GiB
- of 22.08 GiB usable
- Max context
- 58,347 tokens
- Decode ceiling
- 59 tok/s
The fit map
All 68 models against the selected accelerator. Horizontal is the weight bytes of the quantization each model is drawn at; vertical is the context those weights leave room for. Both axes are logarithmic — this set spans 0.09 GiB to 557 GiB of weights. The vertical rule is the memory a runtime can actually allocate on that device; the horizontal rule is the 4,096-token floor this lane counts as usable context.
ComfortableFitsQuantized onlyToo tightNo fit
A dot's colour is the solved verdict for that pair, never which side of a rule it landed on. The two disagree on purpose: the usable-context floor is capped at each model's own maximum position count, so a model whose entire window is 2,048 tokens is a complete fit while sitting below the 4,096-token line. Models with no usable fit are drawn at their smallest quantization — the leftmost point they can reach — which is why they cluster against the vertical rule.
The short answer, by size
Largest quantization that fits with at least 4k of context. A model that needs more than a card holds is not unusable — it splits across devices or spills into host RAM — but the speed cost of doing that is severe.
| Model | Q4_K_M | RTX 3060 12GB 12 GB | RTX 4060 Ti 16GB 16 GB | RTX 4090 24 GB | RTX 5090 32 GB | M4 Max (128GB) 128 GB | H100 80GB (SXM5) 80 GB |
|---|---|---|---|---|---|---|---|
| Gemma 3 1B Instruct | 576 MiB | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 |
| Qwen2.5 3B Instruct | 1.8 GiB | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 |
| DeepSeek-R1-Distill-Llama 8B | 4.6 GiB | Q8_0 | Q8_0 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 |
| OLMo 2 13B Instruct | 7.8 GiB | Q4_K_M | Q6_K | Q8_0 | FP16 / BF16 | FP16 / BF16 | FP16 / BF16 |
| Qwen3 32B | 18.4 GiB | no fit | no fit | Q4_K_M | Q6_K | FP16 / BF16 | FP16 / BF16 |
| DeepSeek-R1-Distill-Llama 70B | 39.6 GiB | no fit | no fit | no fit | no fit | Q8_0 | Q8_0 |
Every accelerator
| Accelerator | Memory | Bandwidth | Largest model that fits at Q4 |
|---|---|---|---|
| Apple M3 Ultra (512GB) | 512 GB LPDDR5 unified | 819 GB/s | Nex N2.5 Pro (396.8B) |
| H200 141GB (SXM) | 141 GB HBM3e | 4800 GB/s | Laguna S 2.1 (117.6B) |
| Apple M4 Max (128GB) | 128 GB LPDDR5X unified | 546 GB/s | Laguna S 2.1 (117.6B) |
| A100 80GB (SXM) | 80 GB HBM2e | 2039 GB/s | Laguna S 2.1 (117.6B) |
| H100 80GB (SXM5) | 80 GB HBM3 | 3350 GB/s | Laguna S 2.1 (117.6B) |
| Jetson AGX Orin (64GB) | 64 GB LPDDR5 unified | 204.8 GB/s | Qwen2.5 72B Instruct (72.7B) |
| RTX 6000 Ada Generation | 48 GB GDDR6 ECC | 960 GB/s | Qwen2.5 72B Instruct (72.7B) |
| L40S | 48 GB GDDR6 ECC | 864 GB/s | Qwen2.5 72B Instruct (72.7B) |
| Apple M4 Pro (48GB) | 48 GB LPDDR5X unified | 273 GB/s | Mixtral 8x7B Instruct (46.7B) |
| A100 40GB (SXM) | 40 GB HBM2 | 1555 GB/s | Mixtral 8x7B Instruct (46.7B) |
| GeForce RTX 5090 | 32 GB GDDR7 | 1792 GB/s | Mixtral 8x7B Instruct (46.7B) |
| GeForce RTX 4090 | 24 GB GDDR6X | 1008 GB/s | Hermes 4.3 36B (36.2B) |
| GeForce RTX 3090 | 24 GB GDDR6X | 936 GB/s | Hermes 4.3 36B (36.2B) |
| Radeon RX 7900 XTX | 24 GB GDDR6 | 960 GB/s | Hermes 4.3 36B (36.2B) |
| Apple M4 (24GB) | 24 GB LPDDR5X unified | 120 GB/s | Qwen3 30B-A3B (30.5B) |
| GeForce RTX 5080 | 16 GB GDDR7 | 960 GB/s | Dolphin Mistral 24B Venice Edition (24B) |
| GeForce RTX 5070 Ti | 16 GB GDDR7 | 896 GB/s | Dolphin Mistral 24B Venice Edition (24B) |
| GeForce RTX 4080 SUPER | 16 GB GDDR6X | 736 GB/s | Dolphin Mistral 24B Venice Edition (24B) |
| GeForce RTX 4070 Ti SUPER | 16 GB GDDR6X | 672 GB/s | Dolphin Mistral 24B Venice Edition (24B) |
| GeForce RTX 4060 Ti 16GB | 16 GB GDDR6 | 288 GB/s | Dolphin Mistral 24B Venice Edition (24B) |
| GeForce RTX 3060 12GB | 12 GB GDDR6 | 360 GB/s | gpt-oss 20B (20.9B) |
| Jetson Orin Nano Super (8GB) | 8 GB LPDDR5 unified | 102 GB/s | Ornith 1.5 9B (9.7B) |
Every model
68 models across 22 families, each solved against all 22 accelerators. Open a family to see its models and their Q4 weight size; pick one in the panel above to see its fit.
Current-model discovery comes from a ranked Hugging Face text-generation snapshot captured September 22, 2026.
Qwen17 models
- Qwen2.5 72B Instruct44.2 GiB
- Qwen2.5 32B Instruct18.5 GiB
- Qwen2.5-Coder 32B Instruct18.5 GiB
- QwQ 32B18.4 GiB
- Qwen3 32B18.4 GiB
- Qwen3 30B-A3B17.3 GiB
- Qwen2.5 14B Instruct8.4 GiB
- Qwen3 14B8.4 GiB
- Qwen3 8B4.7 GiB
- Qwen2.5 7B Instruct4.4 GiB
- Qwen2.5-Coder 7B Instruct4.4 GiB
- Qwen3 4B2.3 GiB
- Qwen2.5 3B Instruct1.8 GiB
- Qwen3 1.7B1.1 GiB
- Qwen2.5 1.5B Instruct940 MiB
- Qwen3 0.6B433 MiB
- Qwen2.5 0.5B Instruct379 MiB
Llama7 models
- Llama 3.1 70B Instruct39.6 GiB
- Llama 3.3 70B Instruct39.6 GiB
- Llama 3.1 8B Instruct4.6 GiB
- Llama 3.2 3B Instruct1.9 GiB
- MiniCPM5 2B1.4 GiB
- Llama 3.2 1B Instruct770 MiB
- MiniCPM5 1B622 MiB
DeepSeek6 models
- DeepSeek-R1-Distill-Llama 70B39.6 GiB
- DeepSeek-R1-Distill-Qwen 32B18.5 GiB
- DeepSeek-R1-Distill-Qwen 14B8.4 GiB
- DeepSeek-R1-Distill-Llama 8B4.6 GiB
- DeepSeek-R1-Distill-Qwen 7B4.4 GiB
- DeepSeek-R1-Distill-Qwen 1.5B1.0 GiB
Gemma6 models
- Gemma 3 27B Instruct15.4 GiB
- Gemma 2 27B Instruct15.5 GiB
- Gemma 3 12B Instruct6.9 GiB
- Gemma 2 9B Instruct5.4 GiB
- Gemma 3 4B Instruct2.4 GiB
- Gemma 3 1B Instruct576 MiB
Mistral5 models
- Mixtral 8x7B Instruct26.3 GiB
- Dolphin Mistral 24B Venice Edition13.5 GiB
- Mistral Small 24B Instruct13.3 GiB
- Mistral Nemo 12B Instruct7.0 GiB
- Mistral 7B Instruct v0.34.1 GiB
Qwen3 5 MOE Text4 models
- Nex N2.5 Pro223.1 GiB
- Ornith 1.5 35B A3B20.2 GiB
- KAT Coder V2.5 Dev19.5 GiB
- Qwen AgentWorld 35B A3B20.6 GiB
Phi3 models
- Phi-4 14B8.4 GiB
- Phi-4-mini 3.8B Instruct2.3 GiB
- Phi-3.5-mini 3.8B Instruct2.2 GiB
SmolLM3 models
- SmolLM2 1.7B Instruct1007 MiB
- SmolLM2 360M Instruct258 MiB
- SmolLM2 135M Instruct101 MiB
gpt-oss2 models
- gpt-oss 120B58.5 GiB
- gpt-oss 20B10.8 GiB
Laguna2 models
- Laguna S 2.166.8 GiB
- Laguna XS 2.119.1 GiB
Qwen3 5 Text2 models
- Ornith 1.5 9B5.4 GiB
- Qwythos 9B Claude Mythos 5 1M5.3 GiB
Falcon1 model
- Falcon3 7B Instruct4.3 GiB
Granite1 model
- Granite 3.3 8B Instruct4.6 GiB
HY V31 model
- Hy3169.6 GiB
K2 Horizon1 model
- K2 Horizon 7B5.1 GiB
Nanbeige1 model
- Nanbeige4.2 3B2.3 GiB
Nemotron1 model
- Llama 3.1 Nemotron 70B Instruct39.6 GiB
OLMo1 model
- OLMo 2 13B Instruct7.8 GiB
Seed OSS1 model
- Hermes 4.3 36B20.3 GiB
Spark2 51 model
- Spark X2.5 4B2.3 GiB
TinyLlama1 model
- TinyLlama 1.1B Chat633 MiB
Yi1 model
- Yi 1.5 34B Chat19.2 GiB
How these are computed
Weights
The byte size of the real published GGUF file, wherever one exists — measured, not multiplied out. Where none does, the exact parameter count from the Hub's index over the checkpoint's tensor shapes times the published llama.cpp bits-per-weight. Not "7B" — 7,615,616,512. Every row says which basis it used, because the nominal figure is sound in the middle and wrong at both ends.
KV cache
2 · layers · kv_heads · head_dim bytes per token per element, summed over layers with each sliding-window layer capped at its window. It scales with key/value heads, which is why parameter count predicts it badly: 65 of these 68 models use grouped-query attention and the rest do not.
Speed ceiling
Decode is memory-bound, so tok/s ≤ bandwidth ÷ bytes read per token. It is a bound, not a benchmark — real runtimes reach roughly 60–80% of it and none exceed it. For mixture-of-expert models only the routed experts are read, and that share is derived from the config rather than assumed.
Model data fetched from Hugging Face on 2026-09-22 (scripts/llm/fetch-model-configs.mjs); accelerator capacity and bandwidth come from vendor specifications, and every bandwidth figure that derives from a memory bus is checked against the bus width and data rate that produce it. The one assumption applied throughout is how much of a device's memory a runtime can actually allocate; it is printed on every page that uses it. The picker and the fit map read solved answers computed at build time — the browser positions and looks up, and does no arithmetic of its own.