Skip to main content

Lab · LLM VRAM · GPU lookup

Computed

I have a GPU. What can it run?

Choose an accelerator to see every one of the 62 rostered models ranked by on-device fit. A usable fit means the weights plus at least 4k tokens of context fit together, capped at the model’s own trained window; tight and no-fit cases stay visible rather than disappearing.

Choose your accelerator

Capacity and peak bandwidth are published device specifications. “Usable” reserves 8% of dedicated VRAM or 25% of unified memory for the runtime, OS, display, and workspace. Each result page prints both the raw spec and that explicit assumption.

Desktop GPUs

32 GB
Usable fits
54 / 62
Assumed usable
29.4 GiB
Bandwidth
1792 GB/s
Largest fit here
46.7B

Largest by parameter count: Mixtral 8x7B Instruct at Q4_K_M.

See all 62 models →
24 GB
Usable fits
54 / 62
Assumed usable
22.1 GiB
Bandwidth
1008 GB/s
Largest fit here
46.7B

Largest by parameter count: Mixtral 8x7B Instruct at Q3_K_M.

See all 62 models →
24 GB
Usable fits
54 / 62
Assumed usable
22.1 GiB
Bandwidth
936 GB/s
Largest fit here
46.7B

Largest by parameter count: Mixtral 8x7B Instruct at Q3_K_M.

See all 62 models →
16 GB
Usable fits
43 / 62
Assumed usable
14.7 GiB
Bandwidth
960 GB/s
Largest fit here
30.5B

Largest by parameter count: Qwen3 30B-A3B at Q3_K_M. 1 more only fit below the usable-context floor.

See all 62 models →
Usable fits
43 / 62
Assumed usable
14.7 GiB
Bandwidth
896 GB/s
Largest fit here
30.5B

Largest by parameter count: Qwen3 30B-A3B at Q3_K_M. 1 more only fit below the usable-context floor.

See all 62 models →
Usable fits
43 / 62
Assumed usable
14.7 GiB
Bandwidth
736 GB/s
Largest fit here
30.5B

Largest by parameter count: Qwen3 30B-A3B at Q3_K_M. 1 more only fit below the usable-context floor.

See all 62 models →
Usable fits
43 / 62
Assumed usable
14.7 GiB
Bandwidth
672 GB/s
Largest fit here
30.5B

Largest by parameter count: Qwen3 30B-A3B at Q3_K_M. 1 more only fit below the usable-context floor.

See all 62 models →
Usable fits
43 / 62
Assumed usable
14.7 GiB
Bandwidth
288 GB/s
Largest fit here
30.5B

Largest by parameter count: Qwen3 30B-A3B at Q3_K_M. 1 more only fit below the usable-context floor.

See all 62 models →
Usable fits
38 / 62
Assumed usable
11.0 GiB
Bandwidth
360 GB/s
Largest fit here
21.5B

Largest by parameter count: gpt-oss 20B at Q5_K_M. 2 more only fit below the usable-context floor.

See all 62 models →

Workstation GPUs

Datacenter accelerators

NVIDIA

L40S

48 GB
Usable fits
59 / 62
Assumed usable
44.2 GiB
Bandwidth
864 GB/s
Largest fit here
72.7B

Largest by parameter count: Qwen2.5 72B Instruct at Q3_K_M.

See all 62 models →
40 GB
Usable fits
59 / 62
Assumed usable
36.8 GiB
Bandwidth
1555 GB/s
Largest fit here
72.7B

Largest by parameter count: Qwen2.5 72B Instruct at Q3_K_M.

See all 62 models →

Apple silicon

Usable fits
58 / 62
Assumed usable
36.0 GiB
Bandwidth
273 GB/s
Largest fit here
70.6B

Largest by parameter count: DeepSeek-R1-Distill-Llama 70B at Q3_K_M. 1 more only fit below the usable-context floor.

See all 62 models →

Edge modules

Usable fits
30 / 62
Assumed usable
6.0 GiB
Bandwidth
102 GB/s
Largest fit here
9.4B

Largest by parameter count: Qwythos 9B Claude Mythos 5 1M at Q4_K_M. 2 more only fit below the usable-context floor.

See all 62 models →