Skip to main content

Lab · On-device AI · LLM VRAM

Computed

Will it fit?

Weight bytes, KV-cache growth and the bandwidth ceiling on decode speed, worked out for 68 open models against 22 real accelerators. Every architecture figure is read from the model's own published config; nothing is estimated and nothing is written by a language model.

Pick a pairing

68 models × 22 accelerators. Every combination below is a page that already exists and is already solved — the panel names the exact pairing it resolved to, never a nearby one.

Resolved pairing

DeepSeek-R1-Distill-Llama 8B on GeForce RTX 4090

Comfortable

The unquantized weights fit and leave room for a long context.

Quantization
FP16 / BF16
largest that fits with usable context
Weights
15.0 GiB
of 22.08 GiB usable
Max context
58,347 tokens
Decode ceiling
59 tok/s
Open this fit as a card →

The fit map

All 68 models against the selected accelerator. Horizontal is the weight bytes of the quantization each model is drawn at; vertical is the context those weights leave room for. Both axes are logarithmic — this set spans 0.09 GiB to 557 GiB of weights. The vertical rule is the memory a runtime can actually allocate on that device; the horizontal rule is the 4,096-token floor this lane counts as usable context.

0.10.31310301003001k4k16k64k256k4k context floorweights, GiB (log)context tokens (log)

ComfortableFitsQuantized onlyToo tightNo fit

A dot's colour is the solved verdict for that pair, never which side of a rule it landed on. The two disagree on purpose: the usable-context floor is capped at each model's own maximum position count, so a model whose entire window is 2,048 tokens is a complete fit while sitting below the 4,096-token line. Models with no usable fit are drawn at their smallest quantization — the leftmost point they can reach — which is why they cluster against the vertical rule.

The short answer, by size

Largest quantization that fits with at least 4k of context. A model that needs more than a card holds is not unusable — it splits across devices or spills into host RAM — but the speed cost of doing that is severe.

ModelQ4_K_MRTX 3060 12GB
12 GB
RTX 4060 Ti 16GB
16 GB
RTX 4090
24 GB
RTX 5090
32 GB
M4 Max (128GB)
128 GB
H100 80GB (SXM5)
80 GB
Gemma 3 1B Instruct576 MiBFP16 / BF16FP16 / BF16FP16 / BF16FP16 / BF16FP16 / BF16FP16 / BF16
Qwen2.5 3B Instruct1.8 GiBFP16 / BF16FP16 / BF16FP16 / BF16FP16 / BF16FP16 / BF16FP16 / BF16
DeepSeek-R1-Distill-Llama 8B4.6 GiBQ8_0Q8_0FP16 / BF16FP16 / BF16FP16 / BF16FP16 / BF16
OLMo 2 13B Instruct7.8 GiBQ4_K_MQ6_KQ8_0FP16 / BF16FP16 / BF16FP16 / BF16
Qwen3 32B18.4 GiBno fitno fitQ4_K_MQ6_KFP16 / BF16FP16 / BF16
DeepSeek-R1-Distill-Llama 70B39.6 GiBno fitno fitno fitno fitQ8_0Q8_0

Every accelerator

AcceleratorMemoryBandwidthLargest model that fits at Q4
Apple M3 Ultra (512GB)512 GB LPDDR5 unified819 GB/sNex N2.5 Pro (396.8B)
H200 141GB (SXM)141 GB HBM3e4800 GB/sLaguna S 2.1 (117.6B)
Apple M4 Max (128GB)128 GB LPDDR5X unified546 GB/sLaguna S 2.1 (117.6B)
A100 80GB (SXM)80 GB HBM2e2039 GB/sLaguna S 2.1 (117.6B)
H100 80GB (SXM5)80 GB HBM33350 GB/sLaguna S 2.1 (117.6B)
Jetson AGX Orin (64GB)64 GB LPDDR5 unified204.8 GB/sQwen2.5 72B Instruct (72.7B)
RTX 6000 Ada Generation48 GB GDDR6 ECC960 GB/sQwen2.5 72B Instruct (72.7B)
L40S48 GB GDDR6 ECC864 GB/sQwen2.5 72B Instruct (72.7B)
Apple M4 Pro (48GB)48 GB LPDDR5X unified273 GB/sMixtral 8x7B Instruct (46.7B)
A100 40GB (SXM)40 GB HBM21555 GB/sMixtral 8x7B Instruct (46.7B)
GeForce RTX 509032 GB GDDR71792 GB/sMixtral 8x7B Instruct (46.7B)
GeForce RTX 409024 GB GDDR6X1008 GB/sHermes 4.3 36B (36.2B)
GeForce RTX 309024 GB GDDR6X936 GB/sHermes 4.3 36B (36.2B)
Radeon RX 7900 XTX24 GB GDDR6960 GB/sHermes 4.3 36B (36.2B)
Apple M4 (24GB)24 GB LPDDR5X unified120 GB/sQwen3 30B-A3B (30.5B)
GeForce RTX 508016 GB GDDR7960 GB/sDolphin Mistral 24B Venice Edition (24B)
GeForce RTX 5070 Ti16 GB GDDR7896 GB/sDolphin Mistral 24B Venice Edition (24B)
GeForce RTX 4080 SUPER16 GB GDDR6X736 GB/sDolphin Mistral 24B Venice Edition (24B)
GeForce RTX 4070 Ti SUPER16 GB GDDR6X672 GB/sDolphin Mistral 24B Venice Edition (24B)
GeForce RTX 4060 Ti 16GB16 GB GDDR6288 GB/sDolphin Mistral 24B Venice Edition (24B)
GeForce RTX 3060 12GB12 GB GDDR6360 GB/sgpt-oss 20B (20.9B)
Jetson Orin Nano Super (8GB)8 GB LPDDR5 unified102 GB/sOrnith 1.5 9B (9.7B)

Every model

68 models across 22 families, each solved against all 22 accelerators. Open a family to see its models and their Q4 weight size; pick one in the panel above to see its fit.

Current-model discovery comes from a ranked Hugging Face text-generation snapshot captured September 22, 2026.

Qwen17 models
  • Qwen2.5 72B Instruct44.2 GiB
  • Qwen2.5 32B Instruct18.5 GiB
  • Qwen2.5-Coder 32B Instruct18.5 GiB
  • QwQ 32B18.4 GiB
  • Qwen3 32B18.4 GiB
  • Qwen3 30B-A3B17.3 GiB
  • Qwen2.5 14B Instruct8.4 GiB
  • Qwen3 14B8.4 GiB
  • Qwen3 8B4.7 GiB
  • Qwen2.5 7B Instruct4.4 GiB
  • Qwen2.5-Coder 7B Instruct4.4 GiB
  • Qwen3 4B2.3 GiB
  • Qwen2.5 3B Instruct1.8 GiB
  • Qwen3 1.7B1.1 GiB
  • Qwen2.5 1.5B Instruct940 MiB
  • Qwen3 0.6B433 MiB
  • Qwen2.5 0.5B Instruct379 MiB
Llama7 models
  • Llama 3.1 70B Instruct39.6 GiB
  • Llama 3.3 70B Instruct39.6 GiB
  • Llama 3.1 8B Instruct4.6 GiB
  • Llama 3.2 3B Instruct1.9 GiB
  • MiniCPM5 2B1.4 GiB
  • Llama 3.2 1B Instruct770 MiB
  • MiniCPM5 1B622 MiB
DeepSeek6 models
  • DeepSeek-R1-Distill-Llama 70B39.6 GiB
  • DeepSeek-R1-Distill-Qwen 32B18.5 GiB
  • DeepSeek-R1-Distill-Qwen 14B8.4 GiB
  • DeepSeek-R1-Distill-Llama 8B4.6 GiB
  • DeepSeek-R1-Distill-Qwen 7B4.4 GiB
  • DeepSeek-R1-Distill-Qwen 1.5B1.0 GiB
Gemma6 models
  • Gemma 3 27B Instruct15.4 GiB
  • Gemma 2 27B Instruct15.5 GiB
  • Gemma 3 12B Instruct6.9 GiB
  • Gemma 2 9B Instruct5.4 GiB
  • Gemma 3 4B Instruct2.4 GiB
  • Gemma 3 1B Instruct576 MiB
Mistral5 models
  • Mixtral 8x7B Instruct26.3 GiB
  • Dolphin Mistral 24B Venice Edition13.5 GiB
  • Mistral Small 24B Instruct13.3 GiB
  • Mistral Nemo 12B Instruct7.0 GiB
  • Mistral 7B Instruct v0.34.1 GiB
Qwen3 5 MOE Text4 models
  • Nex N2.5 Pro223.1 GiB
  • Ornith 1.5 35B A3B20.2 GiB
  • KAT Coder V2.5 Dev19.5 GiB
  • Qwen AgentWorld 35B A3B20.6 GiB
Phi3 models
  • Phi-4 14B8.4 GiB
  • Phi-4-mini 3.8B Instruct2.3 GiB
  • Phi-3.5-mini 3.8B Instruct2.2 GiB
SmolLM3 models
  • SmolLM2 1.7B Instruct1007 MiB
  • SmolLM2 360M Instruct258 MiB
  • SmolLM2 135M Instruct101 MiB
gpt-oss2 models
  • gpt-oss 120B58.5 GiB
  • gpt-oss 20B10.8 GiB
Laguna2 models
  • Laguna S 2.166.8 GiB
  • Laguna XS 2.119.1 GiB
Qwen3 5 Text2 models
  • Ornith 1.5 9B5.4 GiB
  • Qwythos 9B Claude Mythos 5 1M5.3 GiB
Falcon1 model
  • Falcon3 7B Instruct4.3 GiB
Granite1 model
  • Granite 3.3 8B Instruct4.6 GiB
HY V31 model
  • Hy3169.6 GiB
K2 Horizon1 model
  • K2 Horizon 7B5.1 GiB
Nanbeige1 model
  • Nanbeige4.2 3B2.3 GiB
Nemotron1 model
  • Llama 3.1 Nemotron 70B Instruct39.6 GiB
OLMo1 model
  • OLMo 2 13B Instruct7.8 GiB
Seed OSS1 model
  • Hermes 4.3 36B20.3 GiB
Spark2 51 model
  • Spark X2.5 4B2.3 GiB
TinyLlama1 model
  • TinyLlama 1.1B Chat633 MiB
Yi1 model
  • Yi 1.5 34B Chat19.2 GiB

How these are computed

Weights

The byte size of the real published GGUF file, wherever one exists — measured, not multiplied out. Where none does, the exact parameter count from the Hub's index over the checkpoint's tensor shapes times the published llama.cpp bits-per-weight. Not "7B" — 7,615,616,512. Every row says which basis it used, because the nominal figure is sound in the middle and wrong at both ends.

KV cache

2 · layers · kv_heads · head_dim bytes per token per element, summed over layers with each sliding-window layer capped at its window. It scales with key/value heads, which is why parameter count predicts it badly: 65 of these 68 models use grouped-query attention and the rest do not.

Speed ceiling

Decode is memory-bound, so tok/s ≤ bandwidth ÷ bytes read per token. It is a bound, not a benchmark — real runtimes reach roughly 60–80% of it and none exceed it. For mixture-of-expert models only the routed experts are read, and that share is derived from the config rather than assumed.

Model data fetched from Hugging Face on 2026-09-22 (scripts/llm/fetch-model-configs.mjs); accelerator capacity and bandwidth come from vendor specifications, and every bandwidth figure that derives from a memory bus is checked against the bus width and data rate that produce it. The one assumption applied throughout is how much of a device's memory a runtime can actually allocate; it is printed on every page that uses it. The picker and the fit map read solved answers computed at build time — the browser positions and looks up, and does no arithmetic of its own.