Skip to main content

Playground · app-grounded instrument

On-device AI

CoreML Model Size & Quantization Calculator

Exact on-device footprint math, no throughput guesswork. Enter parameter count or per-layer shapes, pick quantization, and get deterministic bytes — what Notiary checks before bundling a model for ANE.

Model definition

Quantization

Deterministic footprint

No latency guesses. Just bytes = params × bytesPerParam × overheadFactor. ANE tile = activation buffer.

fp32 — 4 B/param

1.56 GB (1483 MiB) weights + 64 MB act = 1.62 GB (1544 MiB) peak • of which vocab/overhead 8M ≈ 34.6 MB

1.56 GB (1483 MiB)

×1.00 vs fp32

fp16 — 2 B/param

777.6 MB weights + 64 MB act = 841.6 MB peak • of which vocab/overhead 8M ≈ 17.3 MB

777.6 MB

×2.00 vs fp32

int8 — 1 B/param

388.8 MB weights + 64 MB act = 452.8 MB peak • of which vocab/overhead 8M ≈ 8.6 MB

388.8 MB

×4.00 vs fp32

int4 — 0.5 B/param

202.5 MB weights + 64 MB act = 266.5 MB peak • of which vocab/overhead 8M ≈ 4.5 MB

202.5 MB

×7.68 vs fp32

Practical check (Thumbdash): 360M SmolLM2 at INT4 = 360M×0.5=180 MB raw weights. With 1.08× CoreML overhead and 0.96 packing → 202.5 MB on-disk. Add 64 MB activation tile → 267 MB peak. iOS app limit ~4 GB, extension limit ~ 1 GB, ANE SRAM per layer ~ 8 MB. This is why Thumbdash streams layers and uses 8-bit KV cache.

Formula

bytes = Σ(layers) params * bytesPerElement * overheadFactor / packingEfficiency
FP32: 4, FP16/BF16:2, INT8:1, INT4:0.5, INT2:0.25
Total on-disk ≈ weights + 3-12% CoreML protobuf + activation working set
ANE residency ≈ min(total, activationTile*2 + quantizedWeightsTile)
Compression ratio = fp32Size / quantizedSize

Swift — size check (Notiary)

import CoreML
let model = try MLModel(contentsOf: url)
let sizeBytes = try FileManager.default.attributesOfItem(atPath: url.path)[.size] as! Int
// Quantized: check weight descriptions
if let prog = try? MLModelStructure(contentsOf: url) { /* inspect */ }

Swift for your own project.

Why not latency? On-device AI latency depends on ANE vs GPU scheduling, thermal state, and model’s op mix (MatMul vs LayerNorm). Any online calculator that prints “12 ms” is fabricating. We only print what is deterministic: bytes. For a 360M model at INT4: 360M × 0.5 B = 180 MB raw, × 1.08 overhead ÷ 0.96 packing ≈ 202.5 MB on disk, plus a 64 MB activation tile ≈ 267 MB peak residency.

How to calculate a model's file size

A model file is its weights plus a little structure. The weights take one fixed number of bits each, so their size is the parameter count times the bits per weight, divided by 8 to get bytes. Then divide by 220 for MiB or 230 for GiB. The calculator above adds two settings on top, a metadata multiplier and, for sub-byte types, a packing divisor.

bytes=N×b8×foverheadeffpack,GiB=bytes230\text{bytes} = \frac{N \times b}{8} \times \frac{f_{\text{overhead}}}{\text{eff}_{\text{pack}}}, \qquad \text{GiB} = \frac{\text{bytes}}{2^{30}}

Here N is the parameter count, b the bits per weight, f the metadata multiplier and eff the packing efficiency (1 for every type except INT4, INT2 and binary). Both examples are solved when the page is built, by the same function that fills the results grid, and checked in the site's test suite against an independent bits-to-bytes calculation. These figures are in binary units (MiB, GiB); the grid above prints decimal MB and GB.

A · A 7-billion-parameter model, weights only

N = 7 billion parameters, metadata multiplier 1, packing 1.

7000M×168 B=13.04 GiB\frac{7000\text{M} \times 16}{8}\ \text{B} = 13.04\ \text{GiB}
FP16 weights (16 bits each)
13.04 GiB
4-bit weights
3.26 GiB
4-bit compression vs FP32
8

FP16 is 14 billion bytes, which a decimal-unit tool prints as 14 GB and a binary-unit tool as 13.04 GiB. Both are right; they are different units.

B · The 360M reference model at the page's defaults

N = 360 million parameters, metadata multiplier 1.08, packing 0.96 on 4-bit weights, 64 MB (decimal) activation buffer.

360M×48×1.080.96 B=193.1 MiB\frac{360\text{M} \times 4}{8} \times \frac{1.08}{0.96}\ \text{B} = 193.1\ \text{MiB}
FP16 weights
741.6 MiB
4-bit weights
193.1 MiB
4-bit weights plus activation buffer
254.2 MiB
4-bit compression vs FP32
7.68

The 4-bit weights alone are 180 MB before the multiplier and divisor, and 202.5 MB after (decimal, as in the grid above). The 1.08 and 0.96 are the page's defaults, not measured constants for every model.

Byte math is the only honest metric

iPhone apps don't die from FLOPs. They die from bytes. Exceed 4 GB resident and you get jetsam. Exceed app clip limit and you don't pass review. Put a 7B model in FP32 and you need 28 GB just for weights — game over before ANE even wakes up.

This calculator does one thing without lying: count parameters × bytes per parameter, then add the real-world taxes. CoreML protobuf overhead, packing inefficiency for sub-byte types, and the activation working set. Per-block quantization scales are not added separately; fold them into the overhead factor. No latency fantasy. Just arithmetic that matches ls -lh on device.

Base formula — no magic

bytes = params × bpe × f_overhead / eff_pack
params = total − overhead_vocab counted separately

bpe is bytes per element: FP32 4, FP16/BF16 2, INT8 1, INT4 0.5, INT2 0.25, binary 0.125. That's it. 360M×0.5=180MB360M × 0.5 = 180 MB raw INT4. Multiply by the 1.08 CoreML tax → 194.4 MB, then divide by the 0.96 packing efficiency → 202.5 MB. This page computes that deterministically. No ML needed to count bytes.

Quantization overhead — where 0.5 is not 0.5

INT4 group=32: scale FP16 =2B per block
effective bpe = 0.5 + 2/32 = 0.5625 B
Packing 0.96 → 0.5859 B realized

Sub-byte types need metadata. For each 32 weights you store one scale (and often a zero-point). That adds 2/32=0.06252/32 = 0.0625 B per param. Small group → better accuracy, worse size. Large group → opposite. The packing slider models alignment waste: CoreML aligns weight blobs to 16-byte boundaries, so 0.96 means about 4% waste (dividing by 0.96 adds 4.2%). The calculator applies the packing divisor but not this scale term: it prices INT4 at 0.5 B, so put the scale overhead into the overhead factor.

CoreML overhead factor f_overhead

.mlpackage = weights.bin + model.mlmodel + metadata
Typical f = 1.03 to 1.12 (3-12%)

The .mlmodel is a protobuf describing graph topology, op types, input shapes. For LLMs with 24 blocks, graph is ~5-10 MB. For MobileNet, graph is proportionally larger vs weights → overhead 1.12. For Llama 7B, overhead shrinks to 1.03. We default 1.08 matching Notiary's measured 360M model. Check with FileManager attributes, not guess.

ANE residency ≠ on-disk size

peak ≈ quantized_weights_tile + 2 × act_tile
act_tile = B × seq × hidden × bpe_act

ANE has ~8 MB SRAM per core, streams weights. You never hold whole model in SRAM. Peak DRAM residency = weights for current layered tile + activation double-buffer. For 360M INT4 at seq 2048, hidden 960: act = 2048×960×2B ≈ 3.9 MB per layer, ×2 → 8 MB + 20 MB weight tile ≈ 28 MB resident at once, but on-disk still 194 MB. App memory limit is about peak, not SRAM.

Playbook — how Notiary ships

  1. Count in Python first: sum(p.numel for p in model.parameters). That M number goes in this calculator. Don't trust HF config alone — vocab embeddings count.
  2. Pick quant target from memory budget: Want <200 MB on-disk for App Store cellular limit? 360M → need 0.55 B/param → INT4 group 32. 7B → need 7e9×0.5=3.5GB7e9 × 0.5 = 3.5 GB → too big, need INT2+ streaming.
  3. Measure packing: Convert with coremltools.optimize.coreml.quantize_weights, then ls -lh .mlpackage/Data/com.apple.CoreML/weights/weight.bin. Compare to $ params × bpe $ to derive eff_pack. Our default 0.96 came from real 360M INT4 export.
  4. Check act tile: Run Instruments → Allocations while prompt processing. Largest transient = act_tile. Add that to weights for jetsam calc. If peak > 1 GB in extension, use layer streaming loader like Thumbdash does.
  5. Don't trust latency field: CoreML predicts execution placement ANE vs GPU at compile time, but iOS can fallback under thermal pressure. Same model shows 18 ms then 90 ms. Bytes never lie, ms does.

Honesty — what this calculator does NOT do

  • It doesn't model KV-cache growth: 2×layers×seq×hidden×bpekv2 × layers × seq × hidden × bpe_kv — for 32 layers seq 4096 hidden 960 FP16 KV-cache ≈ 241 MB alone, larger than weights.
  • It doesn't predict accuracy drop. INT4 vs INT2 perplexity delta needs eval, not bytes.
  • It doesn't count tokenizer, embedding tables duplicated in some CoreML exports, or ANE compiler scratch.
  • It ignores iOS 17+ weight compression (palettization) that can make INT4 look like 0.35 B/param on-disk but expands to 0.5 in DRAM.
  • FP8 types (E4M3, E5M2) not yet in CoreML quantization — we list BF16 but it's really only useful for training, not ANE today.

Anatomy of the calculator

Six presets, seven quantization levels, three number fields, two sliders, one formula. Here is what each control actually computes.

The model definition panel

  1. 01

    Preset dropdown. Six reference models with pre-filled parameter counts, overhead estimates, and activation buffer sizes. SmolLM2 360M (Notiary\'s model) defaults to 360M params, 8M vocab overhead, 64 MB activation tile. Selecting a preset fills the sliders — you can then override individual values for custom quantization scenarios.

  2. 02

    Total parameters (M). The raw number of trainable weights — embeddings, attention weights, FFN layers, output projection. This is typically sourced from model.parameters() in PyTorch or the model card. The overhead slider separates non-matmul parameters (embedding tables, LayerNorm gammas) for more accurate footprint estimation.

  3. 03

    Quantization checkboxes. Seven types, each multiplying params × bytesPerElement. FP32 (4 B), FP16/BF16 (2 B), INT8 (1 B), INT4 (0.5 B), INT2 (0.25 B), Binary (0.125 B). Checked types appear in the results grid. The INT4 slider enables packing efficiency — unchecked types ignore it.

The deterministic footprint formula

bytes=params×bpe×foverheadeffpack\text{bytes} = \text{params} \times \text{bpe} \times \frac{f_{\text{overhead}}}{\text{eff}_{\text{pack}}}

params = the total parameter count as entered (the vocab-overhead figure is reported per row, not subtracted). bpe = bytes per element from the quantization table. f_overhead = CoreML protobuf tax (default 1.08×). eff_pack = alignment efficiency (default 0.96 for INT4). Peak DRAM = weights + activation tile.

Results, export, and the Swift snippet

  1. 01

    Results grid. One row per checked quantization. Shows raw weights size, total with activation buffer, overhead contribution, and compression ratio vs FP32 baseline. The ratio is fp32_bytes / quantized_bytes — INT4 typically gives ~7.5× compression. A practical check note updates with the current slider values for the SmolLM2 reference.

  2. 04

    Export JSON. Serializes the current state (params, overhead, activation, sliders, checked quants) to a downloadable JSON file. Useful for CI pipelines that need the same numbers — paste this into a build script assertion that checked model size < App Store cellular limit.

  3. 05

    Swift snippet. The Notiary production check: MLModel(contentsOf:) + FileManager.default.attributesOfItem to get actual on-disk bytes. This is the only number that matters — the calculator predicts it, FileManager verifies it.

Gear behind this build

Notiary stack · 25 picks

ML reference25

More gear across every app: the full Gear list →

Two gotchas worth knowing

KV-cache is missing from this calculator

For transformer-based models, the key-value cache grows during generation: 2×layers×seq_len×hidden_dim×bpe_kv2 \times \text{layers} \times \text{seq\_len} \times \text{hidden\_dim} \times \text{bpe\_kv}. At 32 layers, seq 4096, hidden 960, FP16 — that\'s 241 MB of KV cache alone, exceeding the weight footprint. This calculator intentionally omits KV-cache because it\'s runtime-dependent (grows with each generated token) and belongs in a separate memory budget, not the static model size.

Palettization ≠ true INT4

CoreML\'s weight palettization (iOS 17+) compresses weights further on disk by mapping clusters of weight values to a lookup table, making INT4 appear like 0.35 B/param on disk. But at inference time, weights are decompressed back to FP16 in DRAM — the ANE operates on 2-byte values. So your 180 MB on-disk INT4 model could use 360 MB in DRAM plus decompression latency. FileManager reports the on-disk size; Instruments shows the true DRAM cost.

Frequently asked questions

Why can't this calculator predict latency in milliseconds?

On-device AI latency depends on Apple Neural Engine vs GPU scheduling, the model's op mix (MatMul vs LayerNorm vs Softmax), thermal throttling state, and iOS resource contention. Any online calculator that prints "12 ms" is fabricating — CoreML itself can give 18 ms for the same model and then 90 ms under thermal pressure. Bytes are deterministic: parameter count × bytes per element × overhead factor. ms is a runtime variable. Notiary measures latency at runtime with Instruments, not at compile time.

How does INT4 quantization actually store 0.5 bytes per parameter?

Two 4-bit weights are packed into one byte. For a group size of 32, the quantizer stores one FP16 scale (2 bytes) per 32 weights, adding 2/32 = 0.0625 B/param overhead. The effective bytes per parameter is then 0.5 + 0.0625 = 0.5625. The calculator does not add that scale term itself: it prices INT4 at 0.5 B per parameter and divides by the packing efficiency (0.96 by default, which adds about 4.2%), so fold any scale overhead into the overhead factor. At the defaults, a 360M-parameter model comes to 202.5 MB of INT4 weights.

What's the difference between on-disk size and peak DRAM residency?

The .mlpackage file on disk includes weights, model graph topology (protobuf), and metadata. Peak DRAM residency adds the activation working set — the intermediate tensors computed during inference. For an LLM with seq_len = 2048 and hidden_dim = 960, each layer's activation tile is 2048 × 960 × 2 bytes (FP16) ≈ 3.9 MB. Double-buffered for pipeline overlap → ~8 MB per layer. The 64 MB slider in this calculator models the worst-case tile across all layers.

Can I run a 7B model on iPhone?

At FP16: 7B × 2 B = 14 GB weights alone — exceeds available DRAM on any iPhone (max ~6 GB). At INT4: 7B × 0.5 = 3.5 GB → still too large for the typical ~3 GB app extension limit but potentially fits in the main app with careful memory management. At INT4 with layer streaming (loading one transformer block at a time), peak residency stays under ~500 MB. Thumbdash uses exactly this approach for 360M on constrained devices. For 7B, INT2 (0.25 B/param → 1.75 GB) may be required.

What FP types does CoreML actually support for quantization?

As of iOS 18 / Core ML 6, native quantization supports FP16, INT8, and INT4 (palettization). BF16 is listed in this calculator for reference but is primarily a training format — Apple's ANE operates in FP16 natively. FP8 (E4M3, E5M2) is not yet in production CoreML as of mid-2025. INT2 and Binary (0.125 B/param) are shown for theoretical reference only — CoreML does not currently support them natively, though custom compute units could implement them via Metal Performance Shaders.

How do I calculate a model's file size from its parameter count?

Multiply the number of parameters by the bits stored per weight and divide by 8 to get bytes: size = parameters × bits ÷ 8. A 7-billion-parameter model at FP16 (16 bits) is 14 billion bytes, which is 13.04 GiB; at 4 bits it is 3.5 billion bytes, 3.26 GiB. That is weights only. A real .mlpackage also carries the graph and metadata, which this page models as a multiplier (default 1.08) and, for INT4, INT2 and binary weights, a packing-efficiency divisor (default 0.96). Both defaults are settings, not constants of Core ML: measure your own export and set them to match.

Is the size in MB or MiB?

The results grid prints decimal units (1 MB = 1,000,000 bytes, 1 GB = 1,000,000,000 bytes), and its activation-buffer field is in decimal MB. The worked examples print binary units: 1 MiB = 1,048,576 bytes and 1 GiB = 1,073,741,824 bytes. The same 7-billion-parameter FP16 model is 14 GB and 13.04 GiB. Above 1 GB the grid also prints the MiB count in brackets, as "GB (… MiB)".

Shareable still

The instrument, captured—not illustrated.

This 16:9 frame is rendered from the real browser instrument above. It is the page's canonical preview for image search, link unfurls, and posts that need to show what the tool actually does.

Download 1280 × 720 JPEG
CoreML Model Size & Quantization Calculator — live MakerPortal instrument screenshot
Canonical capture · real UI · no generated scientific artwork