Lab · LLM VRAM · NVIDIA
ComputedWhat LLMs can the H200 141GB (SXM) run?
62 of 62 rostered models have a usable on-device fit. The table keeps all 62 visible, including 0 below-context fits and 0 that exceed this device, so “not listed” never masquerades as an answer.
Published memory
141 GB
HBM3e
Assumed usable
129.7 GiB
92% of capacity
Peak bandwidth
4800 GB/s
theoretical device spec
Usable fits
62 / 62
largest: 298.8B at Q3_K_M
Every model, solved on this accelerator
Usable rows fit weights plus at least 4k tokens, capped at the model’s own trained window. Tight rows hold weights but fall below that floor. The speed column is a memory-bandwidth ceiling, not a benchmark.
62 usable · 0 tight · 0 no fit
| Model | Parameters | On-device result | Selected weights | Max context | Decode ceiling · read per token |
|---|---|---|---|---|---|
| Hy3HY V3 | 298.8B | Quantized fit | Q3_K_M · 128.1 GiBpublished file | 5k | 452 tok/sreads 9.9 GiB/token |
| Laguna S 2.1Laguna | 117.6B | Quantized fit | Q8_0 · 116.4 GiBpublished file | 282k | 588 tok/sreads 7.6 GiB/token |
| gpt-oss 120Bgpt-oss | 116.8B | Full precision · comfortable | As released · 60.8 GiBpublished file | 128k | 1,373 tok/sreads 3.3 GiB/token |
| Qwen2.5 72B InstructQwen | 72.7B | Quantized fit | Q8_0 · 72.0 GiBpublished file | 32k | 60 tok/sreads 74.5 GiB/token |
| DeepSeek-R1-Distill-Llama 70BDeepSeek | 70.6B | Quantized fit | Q8_0 · 69.8 GiBpublished file | 128k | 62 tok/sreads 72.3 GiB/token |
| Llama 3.1 70B InstructLlama | 70.6B | Quantized fit | Q8_0 · 69.8 GiBpublished file | 128k | 62 tok/sreads 72.3 GiB/token |
| Llama 3.1 Nemotron 70B InstructNemotron | 70.6B | Quantized fit | Q8_0 · 69.8 GiBpublished file | 128k | 62 tok/sreads 72.3 GiB/token |
| Llama 3.3 70B InstructLlama | 70.6B | Quantized fit | Q8_0 · 69.8 GiBpublished file | 128k | 62 tok/sreads 72.3 GiB/token |
| Mixtral 8x7B InstructMistral | 46.7B | Full precision · comfortable | FP16 / BF16 · 87.0 GiBpublished file | 32k | 179 tok/sreads 25.0 GiB/token |
| Hermes 4.3 36BSeed OSS | 36.2B | Full precision · comfortable | FP16 / BF16 · 67.3 GiBpublished file | 250k | 64 tok/sreads 69.3 GiB/token |
| KAT Coder V2.5 DevQwen3 5 MOE Text | 34.7B | Full precision · comfortable | FP16 / BF16 · 64.6 GiBpublished file | 256k | 752 tok/sreads 5.9 GiB/token |
| Qwen AgentWorld 35B A3BQwen3 5 MOE Text | 34.7B | Full precision · comfortable | FP16 / BF16 · 64.6 GiBpublished file | 256k | 753 tok/sreads 5.9 GiB/token |
| Yi 1.5 34B ChatYi | 34.4B | Full precision · comfortable | FP16 / BF16 · 64.1 GiBpublished file | 4k | 69 tok/sreads 65.0 GiB/token |
| Laguna XS 2.1Laguna | 33.4B | Full precision · comfortable | FP16 / BF16 · 62.3 GiBpublished file | 256k | 862 tok/sreads 5.2 GiB/token |
| DeepSeek-R1-Distill-Qwen 32BDeepSeek | 32.8B | Full precision · comfortable | FP16 / BF16 · 61.0 GiBpublished file | 128k | 71 tok/sreads 63.0 GiB/token |
| Qwen2.5 32B InstructQwen | 32.8B | Full precision · comfortable | FP16 / BF16 · 61.0 GiBpublished file | 32k | 71 tok/sreads 63.0 GiB/token |
| Qwen2.5-Coder 32B InstructQwen | 32.8B | Full precision · comfortable | FP16 / BF16 · 61.0 GiBpublished file | 32k | 71 tok/sreads 63.0 GiB/token |
| QwQ 32BQwen | 32.8B | Full precision · comfortable | FP16 / BF16 · 61.0 GiBpublished file | 40k | 71 tok/sreads 63.0 GiB/token |
| Qwen3 32BQwen | 32.8B | Full precision · comfortable | FP16 / BF16 · 61.0 GiBpublished file | 40k | 71 tok/sreads 63.0 GiB/token |
| Qwen3 30B-A3BQwen | 30.5B | Full precision · comfortable | FP16 / BF16 · 56.9 GiBpublished file | 40k | 639 tok/sreads 7.0 GiB/token |
| Gemma 3 27B InstructGemma | 27.4B | Full precision · comfortable | FP16 / BF16 · 51.1 GiBpublished file | 128k | 86 tok/sreads 52.1 GiB/token |
| Gemma 2 27B InstructGemma | 27.2B | Full precision · comfortable | FP16 / BF16 · 50.7 GiBpublished file | 8k | 83 tok/sreads 53.6 GiB/token |
| Dolphin Mistral 24B Venice EditionMistral | 24B | Full precision · comfortable | FP16 / BF16 · 44.7 GiBpublished file | 128k | 97 tok/sreads 46.0 GiB/token |
| Mistral Small 24B InstructMistral | 23.6B | Full precision · comfortable | FP16 / BF16 · 43.9 GiBpublished file | 32k | 99 tok/sreads 45.2 GiB/token |
| gpt-oss 20Bgpt-oss | 21.5B | Full precision · comfortable | As released · 12.8 GiBpublished file | 128k | 1,622 tok/sreads 2.8 GiB/token |
| DeepSeek-R1-Distill-Qwen 14BDeepSeek | 14.8B | Full precision · comfortable | FP16 / BF16 · 27.5 GiBpublished file | 128k | 154 tok/sreads 29.0 GiB/token |
| Qwen2.5 14B InstructQwen | 14.8B | Full precision · comfortable | FP16 / BF16 · 27.5 GiBpublished file | 32k | 154 tok/sreads 29.0 GiB/token |
| Qwen3 14BQwen | 14.8B | Full precision · comfortable | FP16 / BF16 · 27.5 GiBpublished file | 40k | 155 tok/sreads 28.8 GiB/token |
| Phi-4 14BPhi | 14.7B | Full precision · comfortable | FP16 / BF16 · 27.3 GiBpublished file | 16k | 155 tok/sreads 28.9 GiB/token |
| OLMo 2 13B InstructOLMo | 13.7B | Full precision · comfortable | FP16 / BF16 · 25.5 GiBpublished file | 4k | 156 tok/sreads 28.7 GiB/token |
| Mistral Nemo 12B InstructMistral | 12.2B | Full precision · comfortable | FP16 / BF16 · 22.8 GiBpublished file | 128k | 186 tok/sreads 24.1 GiB/token |
| Gemma 3 12B InstructGemma | 12.2B | Full precision · comfortable | FP16 / BF16 · 22.7 GiBpublished file | 128k | 190 tok/sreads 23.5 GiB/token |
| Qwythos 9B Claude Mythos 5 1MQwen3 5 Text | 9.4B | Full precision · comfortable | FP16 / BF16 · 17.5 GiBpublished file | 898k | 241 tok/sreads 18.5 GiB/token |
| Gemma 2 9B InstructGemma | 9.2B | Full precision · comfortable | FP16 / BF16 · 17.2 GiBpublished file | 8k | 225 tok/sreads 19.8 GiB/token |
| Qwen3 8BQwen | 8.2B | Full precision · comfortable | FP16 / BF16 · 15.3 GiBpublished file | 40k | 273 tok/sreads 16.4 GiB/token |
| Granite 3.3 8B InstructGranite | 8.2B | Full precision · comfortable | FP16 / BF16 · 15.2 GiBpublished file | 128k | 271 tok/sreads 16.5 GiB/token |
| DeepSeek-R1-Distill-Llama 8BDeepSeek | 8B | Full precision · comfortable | FP16 / BF16 · 15.0 GiBpublished file | 128k | 280 tok/sreads 16.0 GiB/token |
| Llama 3.1 8B InstructLlama | 8B | Full precision · comfortable | FP16 / BF16 · 15.0 GiBpublished file | 128k | 280 tok/sreads 16.0 GiB/token |
| DeepSeek-R1-Distill-Qwen 7BDeepSeek | 7.6B | Full precision · comfortable | FP16 / BF16 · 14.2 GiBpublished file | 128k | 306 tok/sreads 14.6 GiB/token |
| Qwen2.5 7B InstructQwen | 7.6B | Full precision · comfortable | FP16 / BF16 · 14.2 GiBpublished file | 32k | 306 tok/sreads 14.6 GiB/token |
| Qwen2.5-Coder 7B InstructQwen | 7.6B | Full precision · comfortable | FP16 / BF16 · 14.2 GiBpublished file | 32k | 306 tok/sreads 14.6 GiB/token |
| Falcon3 7B InstructFalcon | 7.5B | Full precision · comfortable | FP16 / BF16 · 13.9 GiBpublished file | 32k | 303 tok/sreads 14.8 GiB/token |
| Mistral 7B Instruct v0.3Mistral | 7.2B | Full precision · comfortable | FP16 / BF16 · 13.5 GiBpublished file | 32k | 308 tok/sreads 14.5 GiB/token |
| Gemma 3 4B InstructGemma | 4.3B | Full precision · comfortable | FP16 / BF16 · 8.0 GiBpublished file | 128k | 540 tok/sreads 8.3 GiB/token |
| Nanbeige4.2 3BNanbeige | 4.2B | Full precision · comfortable | FP16 / BF16 · 7.8 GiBpublished file | 256k | 529 tok/sreads 8.5 GiB/token |
| Qwen3 4BQwen | 4B | Full precision · comfortable | FP16 / BF16 · 7.5 GiBpublished file | 40k | 519 tok/sreads 8.6 GiB/token |
| Phi-4-mini 3.8B InstructPhi | 3.8B | Full precision · comfortable | FP16 / BF16 · 7.1 GiBpublished file | 128k | 549 tok/sreads 8.1 GiB/token |
| Phi-3.5-mini 3.8B InstructPhi | 3.8B | Full precision · comfortable | FP16 / BF16 · 7.1 GiBpublished file | 128k | 442 tok/sreads 10.1 GiB/token |
| Llama 3.2 3B InstructLlama | 3.2B | Full precision · comfortable | FP16 / BF16 · 6.0 GiBpublished file | 128k | 652 tok/sreads 6.9 GiB/token |
| Qwen2.5 3B InstructQwen | 3.1B | Full precision · comfortable | FP16 / BF16 · 5.7 GiBpublished file | 32k | 741 tok/sreads 6.0 GiB/token |
| Qwen3 1.7BQwen | 2B | Full precision · comfortable | FP16 / BF16 · 3.8 GiBpublished file | 40k | 959 tok/sreads 4.7 GiB/token |
| DeepSeek-R1-Distill-Qwen 1.5BDeepSeek | 1.8B | Full precision · comfortable | FP16 / BF16 · 3.3 GiBpublished file | 128k | 1,267 tok/sreads 3.5 GiB/token |
| SmolLM2 1.7B InstructSmolLM | 1.7B | Full precision · comfortable | FP16 / BF16 · 3.2 GiBpublished file | 8k | 954 tok/sreads 4.7 GiB/token |
| Qwen2.5 1.5B InstructQwen | 1.5B | Full precision · comfortable | FP16 / BF16 · 2.9 GiBpublished file | 32k | 1,445 tok/sreads 3.1 GiB/token |
| Llama 3.2 1B InstructLlama | 1.2B | Full precision · comfortable | FP16 / BF16 · 2.3 GiBpublished file | 128k | 1,752 tok/sreads 2.6 GiB/token |
| TinyLlama 1.1B ChatTinyLlama | 1.1B | Full precision · comfortable | FP16 / BF16 · 2.0 GiBpublished file | 2k | 2,137 tok/sreads 2.1 GiB/token |
| MiniCPM5 1BLlama | 1.1B | Full precision · comfortable | FP16 / BF16 · 2.0 GiBpublished file | 128k | 2,032 tok/sreads 2.2 GiB/token |
| Gemma 3 1B InstructGemma | 1000M | Full precision · comfortable | FP16 / BF16 · 1.9 GiBpublished file | 32k | 2,347 tok/sreads 1.9 GiB/token |
| Qwen3 0.6BQwen | 752M | Full precision · comfortable | FP16 / BF16 · 1.4 GiBpublished file | 40k | 1,965 tok/sreads 2.3 GiB/token |
| Qwen2.5 0.5B InstructQwen | 494M | Full precision · comfortable | FP16 / BF16 · 942 MiBpublished file | 32k | 4,409 tok/sreads 1.0 GiB/token |
| SmolLM2 360M InstructSmolLM | 362M | Full precision · comfortable | FP16 / BF16 · 690 MiBpublished file | 8k | 4,532 tok/sreads 1010 MiB/token |
| SmolLM2 135M InstructSmolLM | 135M | Full precision · comfortable | FP16 / BF16 · 257 MiBpublished file | 8k | 10,485 tok/sreads 437 MiB/token |
Device specification and assumption
- Memory
- 141 GB HBM3e
- Bandwidth
- 4800 GB/s
- Usable budget
- 92% → 129.7 GiB
Capacity and bandwidth come from the vendor specification. The usable fraction is an explicit planning assumption, not a device specification.
What the table does — and does not — claim
Weight sizes use published checkpoint or GGUF files where the roster has one, otherwise the documented bits-per-weight calculation. Context comes from each model’s layer and KV-head geometry. The decode figure is the card’s peak bandwidth divided by the bytes one token reads — the weights it routes through plus one pass over the KV cache, printed under each ceiling — so it is a roofline bound, not measured application throughput. Multi-GPU splitting and host-memory offload are outside this on-device table.
Questions about this card
What LLMs can the H200 141GB (SXM) run?
62 of the 62 open models in this roster have an on-device configuration with usable context. The largest by exact parameter count is Hy3 at Q3_K_M, with room for 5k tokens. “Largest” describes parameter count, not model quality or task performance.
How much memory is usable on the H200 141GB (SXM)?
The device publishes 141 GB of HBM3e. This calculator budgets 92%, or 129.7 GiB, for model weights and KV cache; the remainder is an explicit allowance for the runtime, driver or operating system, workspace, and display. It is an assumption rather than a vendor specification.
Are the speed figures benchmarks for the H200 141GB (SXM)?
No. They are bandwidth-bound roofline ceilings: 4800 GB/s divided by the bytes one decoded token actually reads, which the table prints beside every ceiling. That is not the size of the file on disk — a token reads the weights it routes through, all of them for a dense model but only the selected experts for a mixture-of-experts, plus one pass over the KV cache. gpt-oss 120B holds 60.8 GiB resident at As released yet reads 3.3 GiB/token, which is what produces its 1,373 tok/s. Real throughput is lower because kernels, cache traffic, prompt processing, scheduling, and runtime overhead also consume time.