Lab · LLM VRAM · AMD
ComputedWhat LLMs can the Radeon RX 7900 XTX run?
54 of 62 rostered models have a usable on-device fit. The table keeps all 62 visible, including 0 below-context fits and 8 that exceed this device, so “not listed” never masquerades as an answer.
Published memory
24 GB
GDDR6
Assumed usable
22.1 GiB
92% of capacity
Peak bandwidth
960 GB/s
theoretical device spec
Usable fits
54 / 62
largest: 46.7B at Q3_K_M
Every model, solved on this accelerator
Usable rows fit weights plus at least 4k tokens, capped at the model’s own trained window. Tight rows hold weights but fall below that floor. The speed column is a memory-bandwidth ceiling, not a benchmark.
54 usable · 0 tight · 8 no fit
| Model | Parameters | On-device result | Selected weights | Max context | Decode ceiling · read per token |
|---|---|---|---|---|---|
| Mixtral 8x7B InstructMistral | 46.7B | Quantized fit | Q3_K_M · 21.3 GiBarchitecture calculation | 7k | 134 tok/sreads 6.7 GiB/token |
| Hermes 4.3 36BSeed OSS | 36.2B | Quantized fit | Q4_K_M · 20.3 GiBarchitecture calculation | 7k | 40 tok/sreads 22.1 GiB/token |
| KAT Coder V2.5 DevQwen3 5 MOE Text | 34.7B | Quantized fit | Q4_K_M · 19.5 GiBarchitecture calculation | 33k | 401 tok/sreads 2.2 GiB/token |
| Qwen AgentWorld 35B A3BQwen3 5 MOE Text | 34.7B | Quantized fit | Q4_K_M · 20.6 GiBpublished file | 19k | 385 tok/sreads 2.3 GiB/token |
| Yi 1.5 34B ChatYi | 34.4B | Quantized fit | Q4_K_M · 19.2 GiBpublished file | 4k | 44 tok/sreads 20.2 GiB/token |
| Laguna XS 2.1Laguna | 33.4B | Quantized fit | Q4_K_M · 19.1 GiBpublished file | 74k | 483 tok/sreads 1.8 GiB/token |
| DeepSeek-R1-Distill-Qwen 32BDeepSeek | 32.8B | Quantized fit | Q4_K_M · 18.5 GiBpublished file | 14k | 44 tok/sreads 20.5 GiB/token |
| Qwen2.5 32B InstructQwen | 32.8B | Quantized fit | Q4_K_M · 18.5 GiBpublished file | 14k | 44 tok/sreads 20.5 GiB/token |
| Qwen2.5-Coder 32B InstructQwen | 32.8B | Quantized fit | Q4_K_M · 18.5 GiBpublished file | 14k | 44 tok/sreads 20.5 GiB/token |
| QwQ 32BQwen | 32.8B | Quantized fit | Q4_K_M · 18.4 GiBarchitecture calculation | 15k | 44 tok/sreads 20.4 GiB/token |
| Qwen3 32BQwen | 32.8B | Quantized fit | Q4_K_M · 18.4 GiBarchitecture calculation | 15k | 44 tok/sreads 20.4 GiB/token |
| Qwen3 30B-A3BQwen | 30.5B | Quantized fit | Q5_K_M · 20.2 GiBpublished file | 20k | 301 tok/sreads 3.0 GiB/token |
| Gemma 3 27B InstructGemma | 27.4B | Quantized fit | Q6_K · 20.9 GiBarchitecture calculation | 9k | 41 tok/sreads 22.0 GiB/token |
| Gemma 2 27B InstructGemma | 27.2B | Quantized fit | Q5_K_M · 18.1 GiBpublished file | 8k | 43 tok/sreads 21.0 GiB/token |
| Dolphin Mistral 24B Venice EditionMistral | 24B | Quantized fit | Q6_K · 18.3 GiBarchitecture calculation | 24k | 46 tok/sreads 19.6 GiB/token |
| Mistral Small 24B InstructMistral | 23.6B | Quantized fit | Q6_K · 18.0 GiBpublished file | 26k | 46 tok/sreads 19.3 GiB/token |
| gpt-oss 20Bgpt-oss | 21.5B | Full precision · comfortable | As released · 12.8 GiBpublished file | 128k | 324 tok/sreads 2.8 GiB/token |
| DeepSeek-R1-Distill-Qwen 14BDeepSeek | 14.8B | Quantized fit | Q8_0 · 14.6 GiBpublished file | 40k | 55 tok/sreads 16.1 GiB/token |
| Qwen2.5 14B InstructQwen | 14.8B | Quantized fit | Q8_0 · 14.6 GiBpublished file | 32k | 55 tok/sreads 16.1 GiB/token |
| Qwen3 14BQwen | 14.8B | Quantized fit | Q8_0 · 14.6 GiBpublished file | 40k | 56 tok/sreads 15.9 GiB/token |
| Phi-4 14BPhi | 14.7B | Quantized fit | Q8_0 · 14.5 GiBpublished file | 16k | 56 tok/sreads 16.1 GiB/token |
| OLMo 2 13B InstructOLMo | 13.7B | Quantized fit | Q8_0 · 13.6 GiBpublished file | 4k | 54 tok/sreads 16.7 GiB/token |
| Mistral Nemo 12B InstructMistral | 12.2B | Quantized fit | Q8_0 · 12.1 GiBpublished file | 64k | 67 tok/sreads 13.4 GiB/token |
| Gemma 3 12B InstructGemma | 12.2B | Quantized fit | Q8_0 · 12.1 GiBarchitecture calculation | 128k | 69 tok/sreads 12.9 GiB/token |
| Qwythos 9B Claude Mythos 5 1MQwen3 5 Text | 9.4B | Full precision · comfortable | FP16 / BF16 · 17.5 GiBpublished file | 36k | 48 tok/sreads 18.5 GiB/token |
| Gemma 2 9B InstructGemma | 9.2B | Full precision · comfortable | FP16 / BF16 · 17.2 GiBpublished file | 8k | 45 tok/sreads 19.8 GiB/token |
| Qwen3 8BQwen | 8.2B | Full precision · comfortable | FP16 / BF16 · 15.3 GiBpublished file | 40k | 55 tok/sreads 16.4 GiB/token |
| Granite 3.3 8B InstructGranite | 8.2B | Full precision · comfortable | FP16 / BF16 · 15.2 GiBpublished file | 44k | 54 tok/sreads 16.5 GiB/token |
| DeepSeek-R1-Distill-Llama 8BDeepSeek | 8B | Full precision · comfortable | FP16 / BF16 · 15.0 GiBpublished file | 57k | 56 tok/sreads 16.0 GiB/token |
| Llama 3.1 8B InstructLlama | 8B | Full precision · comfortable | FP16 / BF16 · 15.0 GiBpublished file | 57k | 56 tok/sreads 16.0 GiB/token |
| DeepSeek-R1-Distill-Qwen 7BDeepSeek | 7.6B | Full precision · comfortable | FP16 / BF16 · 14.2 GiBpublished file | 128k | 61 tok/sreads 14.6 GiB/token |
| Qwen2.5 7B InstructQwen | 7.6B | Full precision · comfortable | FP16 / BF16 · 14.2 GiBpublished file | 32k | 61 tok/sreads 14.6 GiB/token |
| Qwen2.5-Coder 7B InstructQwen | 7.6B | Full precision · comfortable | FP16 / BF16 · 14.2 GiBpublished file | 32k | 61 tok/sreads 14.6 GiB/token |
| Falcon3 7B InstructFalcon | 7.5B | Full precision · comfortable | FP16 / BF16 · 13.9 GiBpublished file | 32k | 61 tok/sreads 14.8 GiB/token |
| Mistral 7B Instruct v0.3Mistral | 7.2B | Full precision · comfortable | FP16 / BF16 · 13.5 GiBpublished file | 32k | 62 tok/sreads 14.5 GiB/token |
| Gemma 3 4B InstructGemma | 4.3B | Full precision · comfortable | FP16 / BF16 · 8.0 GiBpublished file | 128k | 108 tok/sreads 8.3 GiB/token |
| Nanbeige4.2 3BNanbeige | 4.2B | Full precision · comfortable | FP16 / BF16 · 7.8 GiBpublished file | 167k | 106 tok/sreads 8.5 GiB/token |
| Qwen3 4BQwen | 4B | Full precision · comfortable | FP16 / BF16 · 7.5 GiBpublished file | 40k | 104 tok/sreads 8.6 GiB/token |
| Phi-4-mini 3.8B InstructPhi | 3.8B | Full precision · comfortable | FP16 / BF16 · 7.1 GiBpublished file | 119k | 110 tok/sreads 8.1 GiB/token |
| Phi-3.5-mini 3.8B InstructPhi | 3.8B | Full precision · comfortable | FP16 / BF16 · 7.1 GiBpublished file | 40k | 88 tok/sreads 10.1 GiB/token |
| Llama 3.2 3B InstructLlama | 3.2B | Full precision · comfortable | FP16 / BF16 · 6.0 GiBpublished file | 128k | 130 tok/sreads 6.9 GiB/token |
| Qwen2.5 3B InstructQwen | 3.1B | Full precision · comfortable | FP16 / BF16 · 5.7 GiBpublished file | 32k | 148 tok/sreads 6.0 GiB/token |
| Qwen3 1.7BQwen | 2B | Full precision · comfortable | FP16 / BF16 · 3.8 GiBpublished file | 40k | 192 tok/sreads 4.7 GiB/token |
| DeepSeek-R1-Distill-Qwen 1.5BDeepSeek | 1.8B | Full precision · comfortable | FP16 / BF16 · 3.3 GiBpublished file | 128k | 253 tok/sreads 3.5 GiB/token |
| SmolLM2 1.7B InstructSmolLM | 1.7B | Full precision · comfortable | FP16 / BF16 · 3.2 GiBpublished file | 8k | 191 tok/sreads 4.7 GiB/token |
| Qwen2.5 1.5B InstructQwen | 1.5B | Full precision · comfortable | FP16 / BF16 · 2.9 GiBpublished file | 32k | 289 tok/sreads 3.1 GiB/token |
| Llama 3.2 1B InstructLlama | 1.2B | Full precision · comfortable | FP16 / BF16 · 2.3 GiBpublished file | 128k | 350 tok/sreads 2.6 GiB/token |
| TinyLlama 1.1B ChatTinyLlama | 1.1B | Full precision · comfortable | FP16 / BF16 · 2.0 GiBpublished file | 2k | 427 tok/sreads 2.1 GiB/token |
| MiniCPM5 1BLlama | 1.1B | Full precision · comfortable | FP16 / BF16 · 2.0 GiBpublished file | 128k | 406 tok/sreads 2.2 GiB/token |
| Gemma 3 1B InstructGemma | 1000M | Full precision · comfortable | FP16 / BF16 · 1.9 GiBpublished file | 32k | 469 tok/sreads 1.9 GiB/token |
| Qwen3 0.6BQwen | 752M | Full precision · comfortable | FP16 / BF16 · 1.4 GiBpublished file | 40k | 393 tok/sreads 2.3 GiB/token |
| Qwen2.5 0.5B InstructQwen | 494M | Full precision · comfortable | FP16 / BF16 · 942 MiBpublished file | 32k | 882 tok/sreads 1.0 GiB/token |
| SmolLM2 360M InstructSmolLM | 362M | Full precision · comfortable | FP16 / BF16 · 690 MiBpublished file | 8k | 906 tok/sreads 1010 MiB/token |
| SmolLM2 135M InstructSmolLM | 135M | Full precision · comfortable | FP16 / BF16 · 257 MiBpublished file | 8k | 2,097 tok/sreads 437 MiB/token |
| Hy3HY V3 | 298.8B | No on-device fit | Q3_K_M · 128.1 GiBpublished file | — | Not resident— |
| Laguna S 2.1Laguna | 117.6B | No on-device fit | Q3_K_M · 50.3 GiBpublished file | — | Not resident— |
| gpt-oss 120Bgpt-oss | 116.8B | No on-device fit | Q3_K_M · 58.3 GiBpublished file | — | Not resident— |
| Qwen2.5 72B InstructQwen | 72.7B | No on-device fit | Q3_K_M · 35.1 GiBpublished file | — | Not resident— |
| DeepSeek-R1-Distill-Llama 70BDeepSeek | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
| Llama 3.1 70B InstructLlama | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
| Llama 3.1 Nemotron 70B InstructNemotron | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
| Llama 3.3 70B InstructLlama | 70.6B | No on-device fit | Q3_K_M · 31.9 GiBpublished file | — | Not resident— |
Device specification and assumption
- Memory
- 24 GB GDDR6
- Bandwidth
- 960 GB/s
- Bus
- 384-bit × 20 Gbps
- Usable budget
- 92% → 22.1 GiB
Capacity and bandwidth come from the vendor specification. The usable fraction is an explicit planning assumption, not a device specification.
What the table does — and does not — claim
Weight sizes use published checkpoint or GGUF files where the roster has one, otherwise the documented bits-per-weight calculation. Context comes from each model’s layer and KV-head geometry. The decode figure is the card’s peak bandwidth divided by the bytes one token reads — the weights it routes through plus one pass over the KV cache, printed under each ceiling — so it is a roofline bound, not measured application throughput. Multi-GPU splitting and host-memory offload are outside this on-device table.
Questions about this card
What LLMs can the Radeon RX 7900 XTX run?
54 of the 62 open models in this roster have an on-device configuration with usable context. The largest by exact parameter count is Mixtral 8x7B Instruct at Q3_K_M, with room for 7k tokens. “Largest” describes parameter count, not model quality or task performance.
How much memory is usable on the Radeon RX 7900 XTX?
The device publishes 24 GB of GDDR6. This calculator budgets 92%, or 22.1 GiB, for model weights and KV cache; the remainder is an explicit allowance for the runtime, driver or operating system, workspace, and display. It is an assumption rather than a vendor specification.
Are the speed figures benchmarks for the Radeon RX 7900 XTX?
No. They are bandwidth-bound roofline ceilings: 960 GB/s divided by the bytes one decoded token actually reads, which the table prints beside every ceiling. That is not the size of the file on disk — a token reads the weights it routes through, all of them for a dense model but only the selected experts for a mixture-of-experts, plus one pass over the KV cache. Laguna XS 2.1 holds 19.1 GiB resident at Q4_K_M yet reads 1.8 GiB/token, which is what produces its 483 tok/s. Real throughput is lower because kernels, cache traffic, prompt processing, scheduling, and runtime overhead also consume time.