Skip to main content

Library · Edge AI Radar · 2026-09-16

Daily

Which fresh models actually fit your board.

A daily, deterministic scan of new GGUF and ONNX releases on Hugging Face — verified file bytes in, datasheet memory ceilings out. Every fit verdict on this page is arithmetic you can check, not a benchmark we invented.

Wednesday, September 16, 2026 · 12 models × 10 boards · generated Sep 16, 05:42 PM UTC

Daily Snapshot · 2026-09-16

Real file sizes vs real hardware memory limits.

Quantized model weights vary dramatically by quantization scheme. Below is today's verified fit profile across microcontrollers, single-board computers, and edge GPU accelerators.

Scanned
12
HF Releases · 10 boards
Ceilings
7M–30G
RAM Span · MCU→Sigma
SBC+GPU Fits
12/12
Any ≥6.5 GiB board
16 GB+ Fits
12/12
Pi5 16GB / Orin NX · Sigma 12/12
Interactive Board Coverage Radar

Click any board vertex or pill below to inspect model headroom & RAM limits.

Fits (≤80%) Tight (≤100%)
Sigma 32GB12/12 fit (100%)Orin NX 16GB12/12 fit (100%)Pi 5 16GB12/12 fit (100%)Orange Pi 5+ 16GB12/12 fit (100%)ROCK 5B 16GB12/12 fit (100%)Orin Nano 8GB11/12 fit (92%)Pi 5 8GB11/12 fit (92%)Coral TPU1/12 fit (8%)Teensy 4.1 8MB1/12 fit (8%)ESP32-S3 N8R81/12 fit (8%)
Inspecting LattePanda Sigma (32 GB) — 12 of 12 scanned models fit
Interactive Simulator 01

Memory topology & context pressure

Simulate how model tensor weights, runtime scratchpads, and expanding KV cache context windows occupy memory across hardware targets — from 7 MiB microcontrollers to 30 GiB x86 edge servers (Pi 5 8GB/16GB, Orange Pi 5 Plus 16GB, Rock 5B 16GB, Orin Nano/NX, Sigma 32GB).

Live Context Simulator
Hardware RAM Containment & Context Pressure

Pick a real Hugging Face release (verified file bytes). The bar showsweights + 12% overhead + KV cache growing with context. MCU & Coral TPU rows show architectural class mismatch for 2–16 GiB LLMs — not a fabricated 98,000% overflow.

Weights KV Cache (context) Overhead Tight Exceeds
1. Select model (verified bytes)12 releases today
2. Context window2,048 tokens
5122k4k8k

KV cache grows linearly with tokens. At vs 512 baseline.Formula: max(runtime×ratio×seq/2048, 256 KiB). ratio 0.08 MCU, 0.04 SBC/GPU.

Active:onnx-community/vit-beans-demo-ONNX(onnx)
Tensors 327 MiBOverhead 39 MiBKV 16 MiBTotal 367 MiB

Single-Board

· 6.5–30 GiB model budget · 5 boards
Sigma 32GB32 GB LPDDR5 (x86)
Ceiling 30.0 GiB
Total 383 MiB (1%)KV 16 MiB
✅ Fits · 29.6 GiB free
Pi 5 16GB16 GB LPDDR4X
Ceiling 14.5 GiB
Total 383 MiB (3%)KV 16 MiB
✅ Fits · 14.1 GiB free
Orange Pi 5+ 16GB16 GB LPDDR4X (RK3588)
Ceiling 14.0 GiB
Total 383 MiB (3%)KV 16 MiB
✅ Fits · 13.6 GiB free
ROCK 5B 16GB16 GB LPDDR4X (RK3588)
Ceiling 14.0 GiB
Total 383 MiB (3%)KV 16 MiB
✅ Fits · 13.6 GiB free
Pi 5 8GB8 GB LPDDR4X
Ceiling 6.50 GiB
Total 383 MiB (6%)KV 16 MiB
✅ Fits · 6.13 GiB free

Edge GPU

· Unified RAM · CUDA/TensorRT · 2 boards
Orin NX 16GB16 GB LPDDR5 (100 TOPS)
Ceiling 14.5 GiB
Total 383 MiB (3%)KV 16 MiB
✅ Fits · 14.1 GiB free
Orin Nano 8GB8 GB LPDDR5 (unified)
Ceiling 6.50 GiB
Total 383 MiB (6%)KV 16 MiB
✅ Fits · 6.13 GiB free

AI Accelerator

· 8 MB cache · int8 graph · 1 board
Coral TPU8 MB on-chip model cache
Ceiling 8.0 MiB
Not in class · 8.0 MiB ceiling vs 327 MiB model
Edge TPU 8 MB cache — graph spill to host

Microcontroller

· 7 MiB arena · TinyML only · 2 boards
Teensy 4.1 8MB1 MB on-chip + 8 MB PSRAM
Ceiling 7.0 MiB
Not in class · 7.0 MiB ceiling vs 327 MiB model
MCU ceiling — needs <16 MiB TinyML model
ESP32-S3 N8R8512 KB SRAM + 8 MB PSRAM
Ceiling 7.0 MiB
Not in class · 7.0 MiB ceiling vs 327 MiB model
MCU ceiling — needs <16 MiB TinyML model
7 MiB MCU → 30 GiB 16/32 GB SBCs · KV linear with context · 12% overhead constantVerified bytes × 1.25 + KV(seq)
🟢 Weights + 🟦 Overhead + 🔷 KV Cache7M–30G MCU → Sigma 32GB512 to 8,192 token context scaling
Data Table 02

Today's fit matrix

Representative quant file per repo (Q4_K_M preferred), runtime estimate = verified bytes × 1.25, against each board's disclosed model-memory ceiling. Fits ≤ 80% of ceiling · Tight ≤ 100% · otherwise No.

Scan 2026-09-16
Model · representative fileRuntime est.Sigma 32GB30.00 GiB ceilingOrin NX 16GB14.50 GiB ceilingPi 5 16GB14.50 GiB ceilingOrange Pi 5+ 16GB14.00 GiB ceilingROCK 5B 16GB14.00 GiB ceilingOrin Nano 8GB6.50 GiB ceilingPi 5 8GB6.50 GiB ceilingCoral TPU8.0 MiB ceilingTeensy 4.1 8MB7.0 MiB ceilingESP32-S3 N8R87.0 MiB ceilingScore
onnx-community/vit-beans-demo-ONNXONNX · 327.5 MiB file · image-classification409.4 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/distilbert-sentiment-demo-ONNXONNX · 255.5 MiB file · text-classification319.4 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/japanese-roberta-base-ONNXONNX · 161.7 MiB file · fill-mask202.1 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/line-distilbert-base-japanese-ONNXONNX · 358.2 MiB file · fill-mask447.8 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/VibeVoice-Realtime-0.5B-OnnxONNX · 374.0 MiB file · text-to-speech467.5 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/rubert-tiny-toxicity-ONNXONNX · 45.0 MiB file · text-classification56.3 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/sarashina2.2-0.5b-instruct-v0.1-ONNXONNX · 1.1 MiB file · text-generation1.3 MiBFitsFitsFitsFitsFitsFitsFitsFitsFitsFits10/10
onnx-community/siglip-base-patch16-224-ONNXONNX · 213.5 MiB file · zero-shot-image-classification266.9 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/wav2vec2-base-10k-voxpopuli-ft-en-ONNXONNX · 360.3 MiB file · automatic-speech-recognition450.4 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
onnx-community/paraphrase-multilingual-MiniLM-L12-v2-ONNXONNX · 448.4 MiB file · sentence-similarity560.5 MiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
unsloth/Qwen3.8-Flash-Next-GGUFGGUF · Q4_K_M · 4.37 GiB file · image-text-to-text5.46 GiBFitsFitsFitsFitsFitsTightTightNoNoNo7/10
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUFGGUF · BF16 · 888.0 MiB file · image-text-to-text1.08 GiBFitsFitsFitsFitsFitsFitsFitsNoNoNo7/10
Analytics 03 & 04

Quantization & board fit analytics

Bits-per-weight distribution across today's models and breakdown of fit capacity per hardware board class.

Quantization Spectrum · Bits Per Weight (bpw)

Lower bpw fits smaller boards; 4.85 bpw (Q4_K_M) retains ~99% perplexity with 70% RAM savings.

12 models across 3 quants
2 bpwExtreme loss4.85 bpw (Q4_K_M)Sweet spot8.5 bpw (Q8_0)High fidelity16.0 bpwFP16 fullQ4_K_M (1)BF16 (1)ONNX (10)

Board Fit Breakdown · 12 Scanned Models

Percentage of today's models fitting within each board's datasheet memory limit.

Fits (≤80%) Tight (80-100%) No (>100%)
sbc30.0 GiB

Sigma 32GB

12 Fits0 No
edge-gpu14.5 GiB

Orin NX 16GB

12 Fits0 No
sbc14.5 GiB

Pi 5 16GB

12 Fits0 No
sbc14.0 GiB

Orange Pi 5+ 16GB

12 Fits0 No
sbc14.0 GiB

ROCK 5B 16GB

12 Fits0 No
edge-gpu6.5 GiB

Orin Nano 8GB

11 Fits1 Tight0 No
sbc6.5 GiB

Pi 5 8GB

11 Fits1 Tight0 No
accelerator8 MiB

Coral TPU

1 Fits11 No
mcu7 MiB

Teensy 4.1 8MB

1 Fits11 No
mcu7 MiB

ESP32-S3 N8R8

1 Fits11 No
Per-Model Cards 05

The arithmetic, per model

File bytes come from the repo's verified Hugging Face file tree. Parameter count is back-derived from those bytes ÷ the quant's published bits-per-weight. FLOPs/token is the dense-transformer 2·P rule.

  • onnx7/10 boards
    onnx-community/vit-beans-demo-ONNXonnx/model.onnx
    File
    327.5 MiB
    Runtime est.
    409.4 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB409.4 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB409 MiBonnx-community/vit-beans-demo-ONNX — 409 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/distilbert-sentiment-demo-ONNXonnx/model.onnx
    File
    255.5 MiB
    Runtime est.
    319.4 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB319.4 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB319 MiBonnx-community/distilbert-sentiment-demo-ONNX — 319 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/japanese-roberta-base-ONNXonnx/model_q4.onnx
    File
    161.7 MiB
    Runtime est.
    202.1 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB202.1 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB202 MiBonnx-community/japanese-roberta-base-ONNX — 202 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/line-distilbert-base-japanese-ONNXonnx/model.onnx
    File
    358.2 MiB
    Runtime est.
    447.8 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB447.8 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB448 MiBonnx-community/line-distilbert-base-japanese-ONNX — 448 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/VibeVoice-Realtime-0.5B-Onnxcpu_fp16/text_lm.onnx
    File
    374.0 MiB
    Runtime est.
    467.5 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB467.5 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB467 MiBonnx-community/VibeVoice-Realtime-0.5B-Onnx — 467 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/rubert-tiny-toxicity-ONNXonnx/model.onnx
    File
    45.0 MiB
    Runtime est.
    56.3 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB56.3 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB56 MiBonnx-community/rubert-tiny-toxicity-ONNX — 56 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx10/10 boards
    onnx-community/sarashina2.2-0.5b-instruct-v0.1-ONNXonnx/model.onnx
    File
    1.1 MiB
    Runtime est.
    1.3 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB1.3 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB1.3 MiBonnx-community/sarashina2.2-0.5b-instruct-v0.1-ONNX — 1.3 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/siglip-base-patch16-224-ONNXonnx/model_q4.onnx
    File
    213.5 MiB
    Runtime est.
    266.9 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB266.9 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB267 MiBonnx-community/siglip-base-patch16-224-ONNX — 267 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/wav2vec2-base-10k-voxpopuli-ft-en-ONNXonnx/model.onnx
    File
    360.3 MiB
    Runtime est.
    450.4 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB450.4 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB450 MiBonnx-community/wav2vec2-base-10k-voxpopuli-ft-en-ONNX — 450 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • onnx7/10 boards
    onnx-community/paraphrase-multilingual-MiniLM-L12-v2-ONNXonnx/model.onnx
    File
    448.4 MiB
    Runtime est.
    560.5 MiB
    ≈ Params
    FLOPs/token
    Footprint vs ceilings · log 1 MiB→32 GiB560.5 MiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB561 MiBonnx-community/paraphrase-multilingual-MiniLM-L12-v2-ONNX — 561 MiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • gguf · Q4_K_M7/10 boards
    unsloth/Qwen3.8-Flash-Next-GGUFMTP/mtp-Qwen3.8-Flash-Next-Q4_K_M.gguf
    File
    4.37 GiB
    Runtime est.
    5.46 GiB
    ≈ Params
    7.74B
    FLOPs/token
    15.5 GFLOP
    Footprint vs ceilings · log 1 MiB→32 GiB5.46 GiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB5.46 GiBunsloth/Qwen3.8-Flash-Next-GGUF — 5.46 GiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
  • gguf · BF167/10 boards
    ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUFmmproj-Qwen3.8-27B-BF16.gguf
    File
    888.0 MiB
    Runtime est.
    1.08 GiB
    ≈ Params
    466M
    FLOPs/token
    0.9 GFLOP
    Footprint vs ceilings · log 1 MiB→32 GiB1.08 GiB
    ESP32/Teensy 7MiBCoral TPUPi5/Nano 8GBOrange/ROCK 16GBPi5/OrinNX 16GBSigma 32GB1.08 GiBISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF — 1.08 GiB runtime — ceilings 7.0 MiB, 8.0 MiB, 6.50 GiB, 14.0 GiB, 14.5 GiB, 30.0 GiB
Historical Trend 06

Daily release & fit history

Tracking total Hugging Face releases vs edge-fittable models across daily radar snapshots.

Edge AI Release & Fit Trend · 55 Daily Snapshots

Daily count of fresh Hugging Face model releases and how many fit edge hardware ceilings.

Total Scanned Edge Fittable
2026-07-222026-07-292026-08-052026-08-142026-08-212026-08-282026-09-042026-09-112026-09-16
Hardware 07

The boards behind the ceilings

Prices shown were retrieved from the Amazon Product Advertising API on 19 July 2026 and are indicative only — the price and availability on Amazon at the time of purchase apply.

sbc30.00 GiB ceiling

Sigma 32GB

32 GB LPDDR5 (x86)

x86 edge server — 30 GiB model ceiling fits Gemma 4-26B-A4B Q4_K_M (15.8 GiB file / 19.7 GiB runtime) with 10 GiB headroom for 8k context.

edge-gpu14.50 GiB ceiling

Orin NX 16GB

16 GB LPDDR5 (100 TOPS)

Orin NX 16GB — double RAM vs Nano, fits 4-12B Q4_K_M with long context (8k) headroom for TensorRT-LLM.

sbc14.50 GiB ceiling

Pi 5 16GB

16 GB LPDDR4X

16GB flagship — fits 9-12B Q4_K_M (5-7 GiB files) with 2-4k context headroom. Same power envelope, double RAM for local LLM.

sbc14.00 GiB ceiling

Orange Pi 5+ 16GB

16 GB LPDDR4X (RK3588)

RK3588 SBC — cheaper Pi5-class alternative with PCIe NVMe for model storage; 14 GiB ceiling after RKLLM / llama.cpp overhead.

sbc14.00 GiB ceiling

ROCK 5B 16GB

16 GB LPDDR4X (RK3588)

ROCK 5B 16GB — solid mid-range for 7-12B Q4_K_M; active cooling recommended for sustained inference.

edge-gpu6.50 GiB ceiling

Orin Nano 8GB

8 GB LPDDR5 (unified)

Unified memory — CUDA/TensorRT runtime, desktop and display all share the same 8 GB pool.

sbc6.50 GiB ceiling

Pi 5 8GB

8 GB LPDDR4X

Full Linux host — llama.cpp / ONNX Runtime with headroom for the OS and KV cache growth at long context.

accelerator8.0 MiB ceiling

Coral TPU

8 MB on-chip model cache

Full-speed int8 only while the compiled model stays inside the 8 MB cache; larger graphs spill to host RAM. Host supplies its own memory via USB.

mcu7.0 MiB ceiling

Teensy 4.1 8MB

1 MB on-chip + 8 MB PSRAM

Ceiling assumes the PSRAM fitted. On-chip alone caps models near 0.75 MB. Linked card is the Teensy 4.0 sibling (same 600 MHz i.MX RT1062 core; the 4.1 adds the PSRAM pads this ceiling needs).

mcu7.0 MiB ceiling

ESP32-S3 N8R8

512 KB SRAM + 8 MB PSRAM

Model arena lives in octal PSRAM; on-chip SRAM stays free for the RTOS and tensor arena scratch.

Verification 08

Methodology — deterministic, checkable

Inputs (ground truth only)

  • Verified file bytes — each repo's public Hugging Face file tree, fetched at build time. No benchmark claims, no tokens/sec, no accuracy numbers.
  • Datasheet memory constants — ESP32-S3 N8R8 (512 KB + 8 MB PSRAM), Teensy 4.1 +8MB, Pi 5 8GB/16GB, Orange Pi 5+ 16GB, ROCK 5B 16GB, Coral TPU (8 MB cache), Orin Nano 8GB / NX 16GB, Sigma 32GB. Hardcoded in scripts/edgespec/radar-core.mjs with sources in comments.
  • Published quant formats — llama.cpp bits-per-weight values (Q4_K_M ≈ 4.85 bpw, Q8_0 ≈ 8.5, …).

Formulas (nothing else)

  • runtime = fileBytes × 1.25 (weights + KV/buffers)
  • ≈params = fileBytes × 8 ÷ bitsPerWeight (GGUF only)
  • FLOPs/token = 2 × params (dense transformer rule)
  • fits ≤ 80% ceiling · tight ≤ 100% · else no

Gate tests (scripts/edgespec/pipeline.test.mjs) run before every publish in CI. Days with zero grounded models publish nothing — the same anti-abuse posture as the Signals Journal.

Questions 09

Frequently Asked Questions

Why no tokens-per-second or accuracy numbers?+

Because we'd have to invent them. Throughput depends on your exact runtime, cooling, and build flags — a number we didn't measure on the named board would be fabrication. This radar answers only the question arithmetic can answer honestly: does the model's memory footprint fit the board's ceiling?

Why does a “fits” verdict stop at 80% of the ceiling?+

KV cache grows with context length, runtimes allocate scratch buffers, and OS/framework overhead varies. The 20% margin is the difference between “loads once in a demo” and “runs reliably on your bench.” Tight means it fits on paper — plan your context budget carefully.

The repo name says 27B but your table shows fewer params — why?+

We derive ≈params from the verified file bytes ÷ the quant's bits-per-weight, not from the repo's title. Some repos ship partial, draft, or experimental files whose real size doesn't match the name. The bytes are the ground truth; the name is marketing.

When does this page update?+

A daily GitHub Action (edgespec-digest.yml) rescans Hugging Face, re-runs the gate tests, and commits a new dated snapshot only when grounded models pass. This site itself makes zero runtime requests — the fetch happens at build time, per our privacy boundary.