iPhone apps don't die from FLOPs. They die from bytes. Exceed 4 GB resident and you get jetsam. Exceed app clip limit and you don't pass review. Put a 7B model in FP32 and you need 28 GB just for weights — game over before ANE even wakes up.
This calculator does one thing without lying: count parameters × bytes per parameter, then add the real-world taxes. CoreML protobuf overhead, packing inefficiency for sub-byte types, and the activation working set. Per-block quantization scales are not added separately; fold them into the overhead factor. No latency fantasy. Just arithmetic that matches ls -lh on device.
Base formula — no magic
bytes = params × bpe × f_overhead / eff_pack
params = total − overhead_vocab counted separately
bpe is bytes per element: FP32 4, FP16/BF16 2, INT8 1, INT4 0.5, INT2 0.25, binary 0.125. That's it. 360M×0.5=180MB raw INT4. Multiply by the 1.08 CoreML tax → 194.4 MB, then divide by the 0.96 packing efficiency → 202.5 MB. This page computes that deterministically. No ML needed to count bytes.
Quantization overhead — where 0.5 is not 0.5
INT4 group=32: scale FP16 =2B per block
effective bpe = 0.5 + 2/32 = 0.5625 B
Packing 0.96 → 0.5859 B realized
Sub-byte types need metadata. For each 32 weights you store one scale (and often a zero-point). That adds 2/32=0.0625 B per param. Small group → better accuracy, worse size. Large group → opposite. The packing slider models alignment waste: CoreML aligns weight blobs to 16-byte boundaries, so 0.96 means about 4% waste (dividing by 0.96 adds 4.2%). The calculator applies the packing divisor but not this scale term: it prices INT4 at 0.5 B, so put the scale overhead into the overhead factor.
CoreML overhead factor f_overhead
.mlpackage = weights.bin + model.mlmodel + metadata
Typical f = 1.03 to 1.12 (3-12%)
The .mlmodel is a protobuf describing graph topology, op types, input shapes. For LLMs with 24 blocks, graph is ~5-10 MB. For MobileNet, graph is proportionally larger vs weights → overhead 1.12. For Llama 7B, overhead shrinks to 1.03. We default 1.08 matching Notiary's measured 360M model. Check with FileManager attributes, not guess.
ANE residency ≠ on-disk size
peak ≈ quantized_weights_tile + 2 × act_tile
act_tile = B × seq × hidden × bpe_act
ANE has ~8 MB SRAM per core, streams weights. You never hold whole model in SRAM. Peak DRAM residency = weights for current layered tile + activation double-buffer. For 360M INT4 at seq 2048, hidden 960: act = 2048×960×2B ≈ 3.9 MB per layer, ×2 → 8 MB + 20 MB weight tile ≈ 28 MB resident at once, but on-disk still 194 MB. App memory limit is about peak, not SRAM.