Lab · LLM VRAM · NVIDIA
ComputedWhat LLMs can the Jetson Orin Nano Super (8GB) run?
30 of 62 rostered models have a usable on-device fit. The table keeps all 62 visible, including 2 below-context fits and 30 that exceed this device, so “not listed” never masquerades as an answer.
Published memory
8 GB
LPDDR5 unified
Assumed usable
6.0 GiB
75% of capacity
Peak bandwidth
102 GB/s
theoretical device spec
Usable fits
30 / 62
largest: 9.4B at Q4_K_M
Every model, solved on this accelerator
Usable rows fit weights plus at least 4k tokens, capped at the model’s own trained window. Tight rows hold weights but fall below that floor. The speed column is a memory-bandwidth ceiling, not a benchmark.
30 usable · 2 tight · 30 no fit
| Model | Parameters | On-device result | Selected weights | Max context | Decode ceiling · read per token |
|---|---|---|---|---|---|
| Qwythos 9B Claude Mythos 5 1MQwen3 5 Text | 9.4B | Quantized fit | Q4_K_M · 5.3 GiBarchitecture calculation | 6k | 16 tok/sreads 6.0 GiB/token |
| Gemma 2 9B InstructGemma | 9.2B | Quantized fit | Q3_K_M · 4.4 GiBpublished file | 5k | 16 tok/sreads 6.0 GiB/token |
| Qwen3 8BQwen | 8.2B | Quantized fit | Q4_K_M · 4.7 GiBpublished file | 9k | 16 tok/sreads 5.8 GiB/token |
| Granite 3.3 8B InstructGranite | 8.2B | Quantized fit | Q4_K_M · 4.6 GiBpublished file | 9k | 16 tok/sreads 5.9 GiB/token |
| DeepSeek-R1-Distill-Llama 8BDeepSeek | 8B | Quantized fit | Q5_K_M · 5.3 GiBpublished file | 5k | 16 tok/sreads 6.0 GiB/token |
| Llama 3.1 8B InstructLlama | 8B | Quantized fit | Q5_K_M · 5.3 GiBpublished file | 5k | 16 tok/sreads 6.0 GiB/token |
| DeepSeek-R1-Distill-Qwen 7BDeepSeek | 7.6B | Quantized fit | Q5_K_M · 5.1 GiBpublished file | 17k | 17 tok/sreads 5.5 GiB/token |
| Qwen2.5 7B InstructQwen | 7.6B | Quantized fit | Q5_K_M · 5.1 GiBpublished file | 17k | 17 tok/sreads 5.5 GiB/token |
| Qwen2.5-Coder 7B InstructQwen | 7.6B | Quantized fit | Q5_K_M · 5.1 GiBpublished file | 17k | 17 tok/sreads 5.5 GiB/token |
| Falcon3 7B InstructFalcon | 7.5B | Quantized fit | Q5_K_M · 5.0 GiBpublished file | 10k | 16 tok/sreads 5.8 GiB/token |
| Mistral 7B Instruct v0.3Mistral | 7.2B | Quantized fit | Q5_K_M · 4.8 GiBpublished file | 10k | 16 tok/sreads 5.8 GiB/token |
| Gemma 3 4B InstructGemma | 4.3B | Quantized fit | Q8_0 · 4.3 GiBarchitecture calculation | 84k | 21 tok/sreads 4.5 GiB/token |
| Nanbeige4.2 3BNanbeige | 4.2B | Quantized fit | Q8_0 · 4.1 GiBarchitecture calculation | 22k | 20 tok/sreads 4.8 GiB/token |
| Qwen3 4BQwen | 4B | Quantized fit | Q8_0 · 4.0 GiBpublished file | 14k | 19 tok/sreads 5.1 GiB/token |
| Phi-4-mini 3.8B InstructPhi | 3.8B | Quantized fit | Q8_0 · 3.8 GiBarchitecture calculation | 18k | 20 tok/sreads 4.8 GiB/token |
| Phi-3.5-mini 3.8B InstructPhi | 3.8B | Quantized fit | Q8_0 · 3.8 GiBpublished file | 6k | 16 tok/sreads 6.0 GiB/token |
| Llama 3.2 3B InstructLlama | 3.2B | Quantized fit | Q8_0 · 3.2 GiBpublished file | 26k | 23 tok/sreads 4.1 GiB/token |
| Qwen2.5 3B InstructQwen | 3.1B | Full precision · limited context | FP16 / BF16 · 5.7 GiBpublished file | 7k | 16 tok/sreads 6.0 GiB/token |
| Qwen3 1.7BQwen | 2B | Full precision · limited context | FP16 / BF16 · 3.8 GiBpublished file | 20k | 20 tok/sreads 4.7 GiB/token |
| DeepSeek-R1-Distill-Qwen 1.5BDeepSeek | 1.8B | Full precision · comfortable | FP16 / BF16 · 3.3 GiBpublished file | 98k | 27 tok/sreads 3.5 GiB/token |
| SmolLM2 1.7B InstructSmolLM | 1.7B | Full precision · comfortable | FP16 / BF16 · 3.2 GiBpublished file | 8k | 20 tok/sreads 4.7 GiB/token |
| Qwen2.5 1.5B InstructQwen | 1.5B | Full precision · comfortable | FP16 / BF16 · 2.9 GiBpublished file | 32k | 31 tok/sreads 3.1 GiB/token |
| Llama 3.2 1B InstructLlama | 1.2B | Full precision · comfortable | FP16 / BF16 · 2.3 GiBpublished file | 118k | 37 tok/sreads 2.6 GiB/token |
| TinyLlama 1.1B ChatTinyLlama | 1.1B | Full precision · comfortable | FP16 / BF16 · 2.0 GiBpublished file | 2k | 45 tok/sreads 2.1 GiB/token |
| MiniCPM5 1BLlama | 1.1B | Full precision · comfortable | FP16 / BF16 · 2.0 GiBpublished file | 128k | 43 tok/sreads 2.2 GiB/token |
| Gemma 3 1B InstructGemma | 1000M | Full precision · comfortable | FP16 / BF16 · 1.9 GiBpublished file | 32k | 50 tok/sreads 1.9 GiB/token |
| Qwen3 0.6BQwen | 752M | Full precision · comfortable | FP16 / BF16 · 1.4 GiBpublished file | 40k | 42 tok/sreads 2.3 GiB/token |
| Qwen2.5 0.5B InstructQwen | 494M | Full precision · comfortable | FP16 / BF16 · 942 MiBpublished file | 32k | 94 tok/sreads 1.0 GiB/token |
| SmolLM2 360M InstructSmolLM | 362M | Full precision · comfortable | FP16 / BF16 · 690 MiBpublished file | 8k | 96 tok/sreads 1010 MiB/token |
| SmolLM2 135M InstructSmolLM | 135M | Full precision · comfortable | FP16 / BF16 · 257 MiBpublished file | 8k | 223 tok/sreads 437 MiB/token |
| Mistral Nemo 12B InstructMistral | 12.2B | Below usable context | Q3_K_M · 5.7 GiBpublished file | 2k | 16 tok/sreads 6.0 GiB/token |
| Gemma 3 12B InstructGemma | 12.2B | Below usable context | Q3_K_M · 5.5 GiBarchitecture calculation | 2k | 16 tok/sreads 6.0 GiB/token |
| Hy3HY V3 | 298.8B | No on-device fit | Q3_K_M · 128.1 GiBpublished file | — | Not resident— |
| Laguna S 2.1Laguna | 117.6B | No on-device fit | Q3_K_M · 50.3 GiBpublished file | — | Not resident— |
| gpt-oss 120Bgpt-oss | 116.8B | No on-device fit | Q3_K_M · 58.3 GiBpublished file | — | Not resident— |
| Qwen2.5 72B InstructQwen | 72.7B | No on-device fit | Q3_K_M · 35.1 GiBpublished file | — | Not resident— |
| DeepSeek-R1-Distill-Llama 70BDeepSeek | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
| Llama 3.1 70B InstructLlama | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
| Llama 3.1 Nemotron 70B InstructNemotron | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
| Llama 3.3 70B InstructLlama | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
| Mixtral 8x7B InstructMistral | 46.7B | No on-device fit | Q3_K_M · 21.3 GiBarchitecture calculation | — | Not resident— |
| Hermes 4.3 36BSeed OSS | 36.2B | No on-device fit | Q3_K_M · 16.5 GiBarchitecture calculation | — | Not resident— |
| KAT Coder V2.5 DevQwen3 5 MOE Text | 34.7B | No on-device fit | Q3_K_M · 15.8 GiBarchitecture calculation | — | Not resident— |
| Qwen AgentWorld 35B A3BQwen3 5 MOE Text | 34.7B | No on-device fit | Q3_K_M · 15.5 GiBpublished file | — | Not resident— |
| Yi 1.5 34B ChatYi | 34.4B | No on-device fit | Q3_K_M · 15.5 GiBpublished file | — | Not resident— |
| Laguna XS 2.1Laguna | 33.4B | No on-device fit | Q3_K_M · 14.5 GiBpublished file | — | Not resident— |
| DeepSeek-R1-Distill-Qwen 32BDeepSeek | 32.8B | No on-device fit | Q3_K_M · 14.8 GiBpublished file | — | Not resident— |
| Qwen2.5 32B InstructQwen | 32.8B | No on-device fit | Q3_K_M · 14.8 GiBpublished file | — | Not resident— |
| Qwen2.5-Coder 32B InstructQwen | 32.8B | No on-device fit | Q3_K_M · 14.8 GiBpublished file | — | Not resident— |
| QwQ 32BQwen | 32.8B | No on-device fit | Q3_K_M · 14.9 GiBarchitecture calculation | — | Not resident— |
| Qwen3 32BQwen | 32.8B | No on-device fit | Q3_K_M · 14.9 GiBarchitecture calculation | — | Not resident— |
| Qwen3 30B-A3BQwen | 30.5B | No on-device fit | Q3_K_M · 13.7 GiBpublished file | — | Not resident— |
| Gemma 3 27B InstructGemma | 27.4B | No on-device fit | Q3_K_M · 12.5 GiBarchitecture calculation | — | Not resident— |
| Gemma 2 27B InstructGemma | 27.2B | No on-device fit | Q3_K_M · 12.5 GiBpublished file | — | Not resident— |
| Dolphin Mistral 24B Venice EditionMistral | 24B | No on-device fit | Q3_K_M · 10.9 GiBarchitecture calculation | — | Not resident— |
| Mistral Small 24B InstructMistral | 23.6B | No on-device fit | Q3_K_M · 10.7 GiBpublished file | — | Not resident— |
| gpt-oss 20Bgpt-oss | 21.5B | No on-device fit | Q3_K_M · 10.7 GiBpublished file | — | Not resident— |
| DeepSeek-R1-Distill-Qwen 14BDeepSeek | 14.8B | No on-device fit | Q3_K_M · 6.8 GiBpublished file | — | Not resident— |
| Qwen2.5 14B InstructQwen | 14.8B | No on-device fit | Q3_K_M · 6.8 GiBpublished file | — | Not resident— |
| Qwen3 14BQwen | 14.8B | No on-device fit | Q3_K_M · 6.8 GiBpublished file | — | Not resident— |
| Phi-4 14BPhi | 14.7B | No on-device fit | Q3_K_M · 6.9 GiBpublished file | — | Not resident— |
| OLMo 2 13B InstructOLMo | 13.7B | No on-device fit | Q3_K_M · 6.3 GiBpublished file | — | Not resident— |
Device specification and assumption
- Memory
- 8 GB LPDDR5 unified
- Bandwidth
- 102 GB/s
- Usable budget
- 75% → 6.0 GiB
Capacity and bandwidth come from the vendor specification. The usable fraction is an explicit planning assumption, not a device specification.
What the table does — and does not — claim
Weight sizes use published checkpoint or GGUF files where the roster has one, otherwise the documented bits-per-weight calculation. Context comes from each model’s layer and KV-head geometry. The decode figure is the card’s peak bandwidth divided by the bytes one token reads — the weights it routes through plus one pass over the KV cache, printed under each ceiling — so it is a roofline bound, not measured application throughput. Multi-GPU splitting and host-memory offload are outside this on-device table.
Questions about this card
What LLMs can the Jetson Orin Nano Super (8GB) run?
30 of the 62 open models in this roster have an on-device configuration with usable context. The largest by exact parameter count is Qwythos 9B Claude Mythos 5 1M at Q4_K_M, with room for 6k tokens. “Largest” describes parameter count, not model quality or task performance.
How much memory is usable on the Jetson Orin Nano Super (8GB)?
The device publishes 8 GB of LPDDR5 unified. This calculator budgets 75%, or 6.0 GiB, for model weights and KV cache; the remainder is an explicit allowance for the runtime, driver or operating system, workspace, and display. It is an assumption rather than a vendor specification.
Are the speed figures benchmarks for the Jetson Orin Nano Super (8GB)?
No. They are bandwidth-bound roofline ceilings: 102 GB/s divided by the bytes one decoded token actually reads, which the table prints beside every ceiling. That is not the size of the file on disk — a token reads the weights it routes through, all of them for a dense model but only the selected experts for a mixture-of-experts, plus one pass over the KV cache. SmolLM2 135M Instruct holds 257 MiB resident at FP16 / BF16 yet reads 437 MiB/token, which is what produces its 223 tok/s. Real throughput is lower because kernels, cache traffic, prompt processing, scheduling, and runtime overhead also consume time.