Skip to main content

Lab · LLM VRAM · NVIDIA

Computed

What LLMs can the Jetson Orin Nano Super (8GB) run?

30 of 62 rostered models have a usable on-device fit. The table keeps all 62 visible, including 2 below-context fits and 30 that exceed this device, so “not listed” never masquerades as an answer.

Published memory

8 GB

LPDDR5 unified

Assumed usable

6.0 GiB

75% of capacity

Peak bandwidth

102 GB/s

theoretical device spec

Usable fits

30 / 62

largest: 9.4B at Q4_K_M

Every model, solved on this accelerator

Usable rows fit weights plus at least 4k tokens, capped at the model’s own trained window. Tight rows hold weights but fall below that floor. The speed column is a memory-bandwidth ceiling, not a benchmark.

30 usable · 2 tight · 30 no fit

ModelParametersOn-device resultSelected weightsMax contextDecode ceiling · read per token
Qwythos 9B Claude Mythos 5 1MQwen3 5 Text9.4BQuantized fitQ4_K_M · 5.3 GiBarchitecture calculation6k16 tok/sreads 6.0 GiB/token
Gemma 2 9B InstructGemma9.2BQuantized fitQ3_K_M · 4.4 GiBpublished file5k16 tok/sreads 6.0 GiB/token
Qwen3 8BQwen8.2BQuantized fitQ4_K_M · 4.7 GiBpublished file9k16 tok/sreads 5.8 GiB/token
Granite 3.3 8B InstructGranite8.2BQuantized fitQ4_K_M · 4.6 GiBpublished file9k16 tok/sreads 5.9 GiB/token
DeepSeek-R1-Distill-Llama 8BDeepSeek8BQuantized fitQ5_K_M · 5.3 GiBpublished file5k16 tok/sreads 6.0 GiB/token
Llama 3.1 8B InstructLlama8BQuantized fitQ5_K_M · 5.3 GiBpublished file5k16 tok/sreads 6.0 GiB/token
DeepSeek-R1-Distill-Qwen 7BDeepSeek7.6BQuantized fitQ5_K_M · 5.1 GiBpublished file17k17 tok/sreads 5.5 GiB/token
Qwen2.5 7B InstructQwen7.6BQuantized fitQ5_K_M · 5.1 GiBpublished file17k17 tok/sreads 5.5 GiB/token
Qwen2.5-Coder 7B InstructQwen7.6BQuantized fitQ5_K_M · 5.1 GiBpublished file17k17 tok/sreads 5.5 GiB/token
Falcon3 7B InstructFalcon7.5BQuantized fitQ5_K_M · 5.0 GiBpublished file10k16 tok/sreads 5.8 GiB/token
Mistral 7B Instruct v0.3Mistral7.2BQuantized fitQ5_K_M · 4.8 GiBpublished file10k16 tok/sreads 5.8 GiB/token
Gemma 3 4B InstructGemma4.3BQuantized fitQ8_0 · 4.3 GiBarchitecture calculation84k21 tok/sreads 4.5 GiB/token
Nanbeige4.2 3BNanbeige4.2BQuantized fitQ8_0 · 4.1 GiBarchitecture calculation22k20 tok/sreads 4.8 GiB/token
Qwen3 4BQwen4BQuantized fitQ8_0 · 4.0 GiBpublished file14k19 tok/sreads 5.1 GiB/token
Phi-4-mini 3.8B InstructPhi3.8BQuantized fitQ8_0 · 3.8 GiBarchitecture calculation18k20 tok/sreads 4.8 GiB/token
Phi-3.5-mini 3.8B InstructPhi3.8BQuantized fitQ8_0 · 3.8 GiBpublished file6k16 tok/sreads 6.0 GiB/token
Llama 3.2 3B InstructLlama3.2BQuantized fitQ8_0 · 3.2 GiBpublished file26k23 tok/sreads 4.1 GiB/token
Qwen2.5 3B InstructQwen3.1BFull precision · limited contextFP16 / BF16 · 5.7 GiBpublished file7k16 tok/sreads 6.0 GiB/token
Qwen3 1.7BQwen2BFull precision · limited contextFP16 / BF16 · 3.8 GiBpublished file20k20 tok/sreads 4.7 GiB/token
DeepSeek-R1-Distill-Qwen 1.5BDeepSeek1.8BFull precision · comfortableFP16 / BF16 · 3.3 GiBpublished file98k27 tok/sreads 3.5 GiB/token
SmolLM2 1.7B InstructSmolLM1.7BFull precision · comfortableFP16 / BF16 · 3.2 GiBpublished file8k20 tok/sreads 4.7 GiB/token
Qwen2.5 1.5B InstructQwen1.5BFull precision · comfortableFP16 / BF16 · 2.9 GiBpublished file32k31 tok/sreads 3.1 GiB/token
Llama 3.2 1B InstructLlama1.2BFull precision · comfortableFP16 / BF16 · 2.3 GiBpublished file118k37 tok/sreads 2.6 GiB/token
TinyLlama 1.1B ChatTinyLlama1.1BFull precision · comfortableFP16 / BF16 · 2.0 GiBpublished file2k45 tok/sreads 2.1 GiB/token
MiniCPM5 1BLlama1.1BFull precision · comfortableFP16 / BF16 · 2.0 GiBpublished file128k43 tok/sreads 2.2 GiB/token
Gemma 3 1B InstructGemma1000MFull precision · comfortableFP16 / BF16 · 1.9 GiBpublished file32k50 tok/sreads 1.9 GiB/token
Qwen3 0.6BQwen752MFull precision · comfortableFP16 / BF16 · 1.4 GiBpublished file40k42 tok/sreads 2.3 GiB/token
Qwen2.5 0.5B InstructQwen494MFull precision · comfortableFP16 / BF16 · 942 MiBpublished file32k94 tok/sreads 1.0 GiB/token
SmolLM2 360M InstructSmolLM362MFull precision · comfortableFP16 / BF16 · 690 MiBpublished file8k96 tok/sreads 1010 MiB/token
SmolLM2 135M InstructSmolLM135MFull precision · comfortableFP16 / BF16 · 257 MiBpublished file8k223 tok/sreads 437 MiB/token
Mistral Nemo 12B InstructMistral12.2BBelow usable contextQ3_K_M · 5.7 GiBpublished file2k16 tok/sreads 6.0 GiB/token
Gemma 3 12B InstructGemma12.2BBelow usable contextQ3_K_M · 5.5 GiBarchitecture calculation2k16 tok/sreads 6.0 GiB/token
Hy3HY V3298.8BNo on-device fitQ3_K_M · 128.1 GiBpublished fileNot resident
Laguna S 2.1Laguna117.6BNo on-device fitQ3_K_M · 50.3 GiBpublished fileNot resident
gpt-oss 120Bgpt-oss116.8BNo on-device fitQ3_K_M · 58.3 GiBpublished fileNot resident
Qwen2.5 72B InstructQwen72.7BNo on-device fitQ3_K_M · 35.1 GiBpublished fileNot resident
DeepSeek-R1-Distill-Llama 70BDeepSeek70.6BNo on-device fitQ3_K_M · 31.9 GiBpublished fileNot resident
Llama 3.1 70B InstructLlama70.6BNo on-device fitQ3_K_M · 31.9 GiBpublished fileNot resident
Llama 3.1 Nemotron 70B InstructNemotron70.6BNo on-device fitQ3_K_M · 31.9 GiBpublished fileNot resident
Llama 3.3 70B InstructLlama70.6BNo on-device fitQ3_K_M · 31.9 GiBpublished fileNot resident
Mixtral 8x7B InstructMistral46.7BNo on-device fitQ3_K_M · 21.3 GiBarchitecture calculationNot resident
Hermes 4.3 36BSeed OSS36.2BNo on-device fitQ3_K_M · 16.5 GiBarchitecture calculationNot resident
KAT Coder V2.5 DevQwen3 5 MOE Text34.7BNo on-device fitQ3_K_M · 15.8 GiBarchitecture calculationNot resident
Qwen AgentWorld 35B A3BQwen3 5 MOE Text34.7BNo on-device fitQ3_K_M · 15.5 GiBpublished fileNot resident
Yi 1.5 34B ChatYi34.4BNo on-device fitQ3_K_M · 15.5 GiBpublished fileNot resident
Laguna XS 2.1Laguna33.4BNo on-device fitQ3_K_M · 14.5 GiBpublished fileNot resident
DeepSeek-R1-Distill-Qwen 32BDeepSeek32.8BNo on-device fitQ3_K_M · 14.8 GiBpublished fileNot resident
Qwen2.5 32B InstructQwen32.8BNo on-device fitQ3_K_M · 14.8 GiBpublished fileNot resident
Qwen2.5-Coder 32B InstructQwen32.8BNo on-device fitQ3_K_M · 14.8 GiBpublished fileNot resident
QwQ 32BQwen32.8BNo on-device fitQ3_K_M · 14.9 GiBarchitecture calculationNot resident
Qwen3 32BQwen32.8BNo on-device fitQ3_K_M · 14.9 GiBarchitecture calculationNot resident
Qwen3 30B-A3BQwen30.5BNo on-device fitQ3_K_M · 13.7 GiBpublished fileNot resident
Gemma 3 27B InstructGemma27.4BNo on-device fitQ3_K_M · 12.5 GiBarchitecture calculationNot resident
Gemma 2 27B InstructGemma27.2BNo on-device fitQ3_K_M · 12.5 GiBpublished fileNot resident
Dolphin Mistral 24B Venice EditionMistral24BNo on-device fitQ3_K_M · 10.9 GiBarchitecture calculationNot resident
Mistral Small 24B InstructMistral23.6BNo on-device fitQ3_K_M · 10.7 GiBpublished fileNot resident
gpt-oss 20Bgpt-oss21.5BNo on-device fitQ3_K_M · 10.7 GiBpublished fileNot resident
DeepSeek-R1-Distill-Qwen 14BDeepSeek14.8BNo on-device fitQ3_K_M · 6.8 GiBpublished fileNot resident
Qwen2.5 14B InstructQwen14.8BNo on-device fitQ3_K_M · 6.8 GiBpublished fileNot resident
Qwen3 14BQwen14.8BNo on-device fitQ3_K_M · 6.8 GiBpublished fileNot resident
Phi-4 14BPhi14.7BNo on-device fitQ3_K_M · 6.9 GiBpublished fileNot resident
OLMo 2 13B InstructOLMo13.7BNo on-device fitQ3_K_M · 6.3 GiBpublished fileNot resident

Device specification and assumption

Memory
8 GB LPDDR5 unified
Bandwidth
102 GB/s
Usable budget
75% → 6.0 GiB

Capacity and bandwidth come from the vendor specification. The usable fraction is an explicit planning assumption, not a device specification.

What the table does — and does not — claim

Weight sizes use published checkpoint or GGUF files where the roster has one, otherwise the documented bits-per-weight calculation. Context comes from each model’s layer and KV-head geometry. The decode figure is the card’s peak bandwidth divided by the bytes one token reads — the weights it routes through plus one pass over the KV cache, printed under each ceiling — so it is a roofline bound, not measured application throughput. Multi-GPU splitting and host-memory offload are outside this on-device table.

Questions about this card

What LLMs can the Jetson Orin Nano Super (8GB) run?

30 of the 62 open models in this roster have an on-device configuration with usable context. The largest by exact parameter count is Qwythos 9B Claude Mythos 5 1M at Q4_K_M, with room for 6k tokens. “Largest” describes parameter count, not model quality or task performance.

How much memory is usable on the Jetson Orin Nano Super (8GB)?

The device publishes 8 GB of LPDDR5 unified. This calculator budgets 75%, or 6.0 GiB, for model weights and KV cache; the remainder is an explicit allowance for the runtime, driver or operating system, workspace, and display. It is an assumption rather than a vendor specification.

Are the speed figures benchmarks for the Jetson Orin Nano Super (8GB)?

No. They are bandwidth-bound roofline ceilings: 102 GB/s divided by the bytes one decoded token actually reads, which the table prints beside every ceiling. That is not the size of the file on disk — a token reads the weights it routes through, all of them for a dense model but only the selected experts for a mixture-of-experts, plus one pass over the KV cache. SmolLM2 135M Instruct holds 257 MiB resident at FP16 / BF16 yet reads 437 MiB/token, which is what produces its 223 tok/s. Real throughput is lower because kernels, cache traffic, prompt processing, scheduling, and runtime overhead also consume time.