Skip to main content

Playground · app-grounded instrument

On-device AI

LLM VRAM & KV-Cache Footprint Calculator

Exact memory math for open LLMs (DeepSeek-R1, Llama 3.3, Qwen 2.5, Mistral). Calculate weight sizes, KV-cache scaling up to 128K context, the runtime reserve, and hardware fit across Apple Silicon, RTX GPUs, and Jetson dev kits.

Already solved for 68 models and 22 accelerators. The calculator below runs the same solver on the same model configs — to see every pairing at once, browse the fit matrix.

Model & Architecture Configuration

Deterministic Math
Model, quantisation, context and accelerator ride in the URL.

70.6B parameters · 80 layers × 8 KV heads × 128 dims · trained context 131,072 · Q4_K_M weights: measured file.

Memory Footprint Breakdown

56.1 GiB VRAM
Model Weights Footprint (0.60 B/param):39.60 GiB
KV-Cache Residency (<8K context):2.50 GiB
Runtime & OS Reserve (25% of unified memory, assumed):14.03 GiB
Min Recommended VRAM Pool:56.13 GiB
Decode Speed Ceiling (546 GB/s ÷ 45.20 GB read per token):12.1 tok/sec

Hardware Target Compatibility Matrix

GeForce RTX 5090 (32GB)❌ Weights exceed 29 GiB usable
GeForce RTX 5080 (16GB)❌ Weights exceed 15 GiB usable
GeForce RTX 5070 Ti (16GB)❌ Weights exceed 15 GiB usable
GeForce RTX 4090 (24GB)❌ Weights exceed 22 GiB usable
GeForce RTX 4080 SUPER (16GB)❌ Weights exceed 15 GiB usable
GeForce RTX 4070 Ti SUPER (16GB)❌ Weights exceed 15 GiB usable
GeForce RTX 4060 Ti 16GB (16GB)❌ Weights exceed 15 GiB usable
GeForce RTX 3090 (24GB)❌ Weights exceed 22 GiB usable
GeForce RTX 3060 12GB (12GB)❌ Weights exceed 11 GiB usable
RTX 6000 Ada Generation (48GB)✅ Fits · ≤ 21 tok/s
L40S (48GB)✅ Fits · ≤ 19 tok/s
A100 40GB (SXM) (40GB)❌ Weights exceed 37 GiB usable
A100 80GB (SXM) (80GB)✅ Fits · ≤ 45 tok/s
H100 80GB (SXM5) (80GB)✅ Fits · ≤ 74 tok/s
H200 141GB (SXM) (141GB)✅ Fits · ≤ 106 tok/s
Radeon RX 7900 XTX (24GB)❌ Weights exceed 22 GiB usable
Apple M4 (24GB) (24GB)❌ Weights exceed 18 GiB usable
Apple M4 Pro (48GB) (48GB)❌ Weights exceed 36 GiB usable
Apple M4 Max (128GB) (128GB)✅ Fits · ≤ 12 tok/s
Apple M3 Ultra (512GB) (512GB)✅ Fits · ≤ 18 tok/s
Jetson Orin Nano Super (8GB) (8GB)❌ Weights exceed 6 GiB usable
Jetson AGX Orin (64GB) (64GB)✅ Fits · ≤ 5 tok/s

Anatomy of LLM Memory Allocation

Running an open LLM on consumer GPUs or Apple Silicon requires partitioning memory across three distinct structures:

1. Static Weights

Quantized Weight Tensor

The model parameters stored on disk. Quantization packs two-byte FP16 weights down to Q4_K_M (4.83 bits nominal, 0.60 B/param) or Q8_0 (8.5 bits, 1.06 B/param). Where a real GGUF file is published, its measured size is used instead — small models run well above the nominal figure. Stays resident in VRAM/Unified RAM continuously.

2. Dynamic KV-Cache

Key-Value Attention Cache

Stores intermediate key and value states per layer for every prompt token to avoid re-computing attention during generation. Grows linearly with context length and batch size.

3. Runtime Overhead

CUDA/Metal Runtime Context

GPU driver context buffers, scratch spaces for FlashAttention, activation tensors — and on a desktop, the display and the OS. This calculator holds back 8% of a dedicated card and 25% of unified memory for them. That is an assumption, not a measurement: a headless server card gives back more, a card driving two displays less.

Gear & Hardware Picks for Local LLMs

Tested studio gear for running large models on-device and high-throughput inference.

Build this lab

Local vs cloud inference bench

Jetson Orin Nano for local GPU inference + Pi 5 as an offload gateway — reproduce the client-vs-serverless tradeoff on real hardware before you commit to a cloud GPU bill.

Prices shown were retrieved from the Amazon Product Advertising API on 8 October 2026 and are indicative only — the price and availability on Amazon at the time of purchase apply.

Prices shown were each checked against the Amazon product listing between 8 August 2026 and 17 September 2026 and are indicative only — the price and availability on Amazon at the time of purchase apply.

Estimated total

$659

Prices from Amazon catalog cache · may change

Open primary listing ↗

Kit Total

—

Buy ↗

The Mathematics of Memory & Bandwidth

The card memory M_card a configuration needs is the weights plus the KV cache, divided by the share of the card a runtime can allocate:

Mcard=Wfile+2×L×Hkv×Dhead×Nctx×Nbatch×BkvfusableM_{\text{card}} = \frac{W_{\text{file}} + 2 \times L \times H_{\text{kv}} \times D_{\text{head}} \times N_{\text{ctx}} \times N_{\text{batch}} \times B_{\text{kv}}}{f_{\text{usable}}}

W_file is the byte size of the published GGUF file where one exists, and parameters × bits per weight ÷ 8 where none does. Sliding-window layers count at most their window in the cache term. f_usable is 0.92 on a dedicated card and 0.75 on unified memory — the reserve row above is the difference.

Autoregressive decoding reads the weights it routes through once per token, plus one pass over the KV cache, so its speed T_gen (tokens/sec) is bounded by bandwidth divided by bytes read per token:

Tgen≤BWmema×Wfile+KV(Nctx)T_{\text{gen}} \le \frac{\text{BW}_{\text{mem}}}{a \times W_{\text{file}} + \text{KV}(N_{\text{ctx}})}

a is the share of weights read per token: 1 for a dense model, the routed-expert share for a mixture of experts. It is a ceiling no runtime exceeds, and the results panel prints both terms of the division beside it.

Inference Startup & Allocation Command

Copyable startup template for vLLM / llama.cpp with exact memory ceilings.

# Python / vLLM / llama.cpp startup script generator
# Run with estimated memory ceiling
python3 -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-72B-Instruct-GGUF \
  --gpu-memory-utilization 0.90 \
  --max-model-len 32768 \
  --kv-cache-dtype fp8

Export · Soft gate

Export vLLM & llama.cpp Benchmark Config Script

Free watermarked file instantly. Unlock clean export via email — soft gate, no hard paywall. We store unlock in localStorage per sim.

File · llm-vram-kvcache-calculator.txt

llm-vram-kvcache-calculator.txttext/plain+ watermark line on free path

Free download adds a small footer: # Export from makerportal.ai — free watermarked build. Unlock c…Clean export removes footer. Both are generated fresh from your current sim tuning.

Privacy: email stays in your browser localStorage (mp_export_email_llm-vram-kvcache-calculator) + unlock flag (mp_export_unlock_llm-vram-kvcache-calculator). If Buttondown username is configured, we also POST to Buttondown (privacy-first mode, no tracking pixels per D-014). See privacy → affiliates & email.

Unlock clean export

Soft gate — no hard paywall, no Clerk. Email stays local unless you explicitly check the newsletter box. Unsubscribe anytime. RSS at /rss.xml.

Export → Fab bonusAfter export, your tuned stackup can be ordered via PCBWay/JLCPCB CTA (when live) — see /privacy#affiliates for live merchants.

Frequently Asked Questions

How is total LLM VRAM requirement calculated?↓

The weights and the KV cache together have to fit in the memory a runtime can actually allocate. Weights are the byte size of the published GGUF file wherever one exists, and parameters × bits per weight ÷ 8 where none does. The KV cache is 2 × layers × KV heads × head dim × context × batch × bytes per element, with any sliding-window layer capped at its window. This calculator — like the /lab/llm-vram fit map, which shares its solver — assumes 92% of a dedicated card and 75% of unified memory is allocatable, so the card you need is (weights + KV cache) ÷ that fraction. On Apple Silicon that pool is system RAM shared with macOS.

What is the KV-Cache memory impact at 128K context lengths?↓

At 131,072 tokens the KV cache is as large as the weights. Llama 3.3 70B Instruct keeps 40.0 GiB of FP16 cache for one request at that length — 1.01× its 39.6 GiB Q4_K_M weight file. An FP8 cache halves that to 20.0 GiB and a 4-bit cache quarters it to 10.0 GiB.

Why does Apple Silicon (M1–M4 Ultra) perform so well on large models?↓

Capacity, not speed. Apple M3 Ultra (512GB) moves 819 GB/s, less than the 1,792 GB/s of GeForce RTX 5090, but even at the 75% of unified memory this calculator assumes the GPU can use, it holds 384.0 GiB against 29.4 GiB on the 5090. Llama 3.3 70B Instruct's 131.4 GiB FP16 checkpoint fits there with room for 131,072 tokens of context, at a ceiling of 5.7 tok/s; on the 5090 even its 39.6 GiB Q4_K_M file does not fit.

What token generation speed (tokens/sec) can I expect?↓

Decoding is memory-bound: each token reads the weights it routes through once, plus one pass over the KV cache, so the ceiling is memory bandwidth divided by bytes read per token — not by the weight file alone. Llama 3.3 70B Instruct at Q4_K_M on Apple M4 Max (128GB) reads 45.20 GB per token at 8K context (a 42.52 GB weight file plus 2.68 GB of cache), so 546 GB/s gives a ceiling of 12.1 tok/s per stream. A mixture-of-experts model reads only its routed share of the weights. No runtime beats the ceiling, and a real one typically reaches 60–80% of it.