<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet href="/rss-styles.xsl" type="text/xsl"?><rss version="2.0"><channel><title>Field Notes — MakerPortal</title><description>Notes on independent software, product decisions, interface craft, and what we learn while shipping.</description><link>https://makerportal.ai/</link><item><title>PrivateRedact limits 7B decode to 4.0 tok/s</title><link>https://makerportal.ai/blog/2026-08-11-itria-hackernews-49245161/</link><guid isPermaLink="true">https://makerportal.ai/blog/2026-08-11-itria-hackernews-49245161/</guid><description>PrivateRedact&apos;s auto-selection limits Ollama inference to 4.0 tok/s on an 8 GB Pi 5 by defaulting to CPU memory, ignoring unified bandwidth.</description><pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Agentic Coding Loops Beyond the LLM</title><link>https://makerportal.ai/blog/auralinter-agentic-dsp-verification/</link><guid isPermaLink="true">https://makerportal.ai/blog/auralinter-agentic-dsp-verification/</guid><description>How RAG-grounded retrieval and a clang++ verifier catch a hallucinated biquad formula a single LLM call ships — and why AuraLinter isn&apos;t fully on-device.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Calibrated Multichannel &amp; Binaural Audio Recording on iPhone</title><link>https://makerportal.ai/blog/binaural-head-headphone-measurement/</link><guid isPermaLink="true">https://makerportal.ai/blog/binaural-head-headphone-measurement/</guid><description>Class-compliant multichannel USB audio capture, a loaded 1 V / 103 dB SPL nominal reference at 1 kHz, and binaural recording analyzed on iPhone.</description><pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate></item><item><title>LLMKube&apos;s 66× speedup is actually 35.4×</title><link>https://makerportal.ai/blog/biquadia-github-1095330682/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1095330682/</guid><description>Two figures in LLMKube&apos;s GKE benchmark are wrong: prompt processing is 35.4×, not 66×, and token generation is 13.9×, not 17×.</description><pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Wax hybrid recall is 6.1 ms, not sub-ms</title><link>https://makerportal.ai/blog/biquadia-github-1138007869/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1138007869/</guid><description>Wax&apos;s own benchmark puts hybrid recall p95 at 6.1 ms — six times above the &quot;sub-millisecond&quot; label — and the figure omits which index engine ran.</description><pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Orion&apos;s ANE Decode vs CPU Baseline</title><link>https://makerportal.ai/blog/biquadia-github-1171688936/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1171688936/</guid><description>Driving Apple&apos;s Neural Engine through private frameworks: 172.4 tok/s on ANE vs 283 on CPU, and delta compilation that pays for itself in 1.21 steps.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Espresso bounds CoreML overhead at 168.7 µs</title><link>https://makerportal.ai/blog/biquadia-github-1173023018/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1173023018/</guid><description>Espresso&apos;s 3.41× ANE speedup over CoreML, read against a 70 µs dispatch floor, sets a 168.7 µs lower bound on CoreML per-token overhead.</description><pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Presspeech&apos;s 100ms Latency Overhead</title><link>https://makerportal.ai/blog/biquadia-github-1233899535/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1233899535/</guid><description>Bandwidth calculation for Presspeech&apos;s Parakeet model on M4 Max shows inference takes ~1% of latency, with pipeline dispatch consuming the rest.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate></item><item><title>ANEForge hits 2.28× speculative decode</title><link>https://makerportal.ai/blog/biquadia-github-1267734215/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1267734215/</guid><description>ANEForge reaches 16.8 tok/s on Qwen3-8B (up from 7.4), hitting the verify(K)≈verify(1) ceiling at 29.6 GB/s vs Orion&apos;s 42.8 GB/s floor.</description><pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate></item><item><title>M4 Max decode uses 99.8% of bandwidth</title><link>https://makerportal.ai/blog/biquadia-hackernews-49259339/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-hackernews-49259339/</guid><description>M4 Max llama.cpp inference consumes 99.0–99.8% of its 400 GB/s bandwidth, leaving no slack to absorb VM overhead — a direct constraint on near-native…</description><pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate></item><item><title>HM-10 vs nRF52: Not the Same BLE UART Profile</title><link>https://makerportal.ai/blog/blexar-ble-uart-bridge/</link><guid isPermaLink="true">https://makerportal.ai/blog/blexar-ble-uart-bridge/</guid><description>Why HM-10&apos;s single-characteristic UART and the Nordic UART Service aren&apos;t interchangeable, and the newline-framing gotcha that only appears at low MTU.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Designing the one-handed moment</title><link>https://makerportal.ai/blog/designing-the-one-handed-moment/</link><guid isPermaLink="true">https://makerportal.ai/blog/designing-the-one-handed-moment/</guid><description>A field note on glanceability, reach, and building interfaces for the imperfect moments where software is actually used.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>ElevenLabs &amp; Web Audio Streaming Latency</title><link>https://makerportal.ai/blog/elevenlabs-web-audio-streaming-latency/</link><guid isPermaLink="true">https://makerportal.ai/blog/elevenlabs-web-audio-streaming-latency/</guid><description>Browser latency ledger for ElevenLabs TTS, Web Audio scheduling, DSP, and output—plus mistakes that make a fast demo sound broken.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Zero-Latency Web Audio &amp; Voice AI Architecture</title><link>https://makerportal.ai/blog/elevenlabs-webassembly-audioworklet-streaming/</link><guid isPermaLink="true">https://makerportal.ai/blog/elevenlabs-webassembly-audioworklet-streaming/</guid><description>How to build sub-100ms real-time voice streaming applications in the browser using ElevenLabs API, WebAssembly decoders, and AudioWorklet ring buffers.</description><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Inside Biquadia: On-Device Audio Lab</title><link>https://makerportal.ai/blog/inside-biquadia/</link><guid isPermaLink="true">https://makerportal.ai/blog/inside-biquadia/</guid><description>How privacy, real-time constraints, and the character of physical sound shaped MakerPortal’s newest release.</description><pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate></item><item><title>watchOS LLM speed is capped by CPU bandwidth</title><link>https://makerportal.ai/blog/itria-github-1051796718/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-1051796718/</guid><description>ETOS-LLM-Studio pins watchOS inference to the CPU path, so bandwidth divided by model size sets the ceiling — 9.4 tok/s on a Pi 5 at 17 GB/s.</description><pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate></item><item><title>RTX 5090 Decode Bandwidth &amp; 25% Penalty</title><link>https://makerportal.ai/blog/itria-github-1164344011/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-1164344011/</guid><description>imp reports a 42-48% decode gap over llama.cpp on consumer Blackwell. The 5090&apos;s own 5090D comparison shows the gap is just a bandwidth model.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Qwen3.6-35B decodes ~11.2 B active per pass</title><link>https://makerportal.ai/blog/itria-github-725205304/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-725205304/</guid><description>Bandwidth-bound decode math on RTX 5090D measurements backs out that only ~32% of Qwen3.6-35B&apos;s nominal weights are active per forward pass.</description><pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Serial loading reclaims 4.2 GB on 8 GB devices</title><link>https://makerportal.ai/blog/itria-github-867433720/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-867433720/</guid><description>Two 7B Q4 models exceed an 8 GB ceiling when loaded concurrently. Serial residency via llama-swap reclaims 4.2 GB of KV cache headroom.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item><item><title>MoEspresso coder trades at 2.455:1 NLL</title><link>https://makerportal.ai/blog/itria-hackernews-49321813/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hackernews-49321813/</guid><description>Pruning DeepSeek-V4-Flash-0731 cuts safetensors by 32.63%, but each 0.1-unit code NLL gain costs 0.245 units of WikiText degradation.</description><pubDate>Mon, 17 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Qwen3.8-27B Cold Fusion: mixed attention</title><link>https://makerportal.ai/blog/itria-hf-davidau-qwen38-27b-cold-fusion-gain-v11-nm-dau-neo-max-mtp-gguf/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hf-davidau-qwen38-27b-cold-fusion-gain-v11-nm-dau-neo-max-mtp-gguf/</guid><description>GGUFs of a Qwen3-27B fine-tune with 75% linear attention layers and ~2/3 thinking-token reduction — benchmark gaps and unverified quality claims noted.</description><pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate></item><item><title>3.02× FastMTP requires draft depth ≥ 3</title><link>https://makerportal.ai/blog/itria-hf-hauhaucs-qwen38-27b-uncensored-hauhaucs-aggressive-mtp-gguf/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hf-hauhaucs-qwen38-27b-uncensored-hauhaucs-aggressive-mtp-gguf/</guid><description>A 3.02× speedup exceeds the 3.00× ceiling for depth-2 speculative decoding, placing FastMTP&apos;s draft depth at 3 or more with an implied acceptance rate of…</description><pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Benchmarking Edge AI SBCs &amp; Accelerators</title><link>https://makerportal.ai/blog/lattepanda-vs-jetson-orin-edge-ai-benchmarks/</link><guid isPermaLink="true">https://makerportal.ai/blog/lattepanda-vs-jetson-orin-edge-ai-benchmarks/</guid><description>Hardware benchmark comparing x86 SBCs, ARM CUDA Tensor Cores, and USB Edge TPUs for local vision, speech, and embedding inference.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate></item><item><title>LiteFS Multi-Region SQLite Latency Model</title><link>https://makerportal.ai/blog/litefs-multi-region-sqlite/</link><guid isPermaLink="true">https://makerportal.ai/blog/litefs-multi-region-sqlite/</guid><description>Mental model for LiteFS multi-region SQLite: single-primary writes, local replicas, fly-replay forwarding, and honest latency modeling.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>BEM HRTFs match measured: headphone VR only</title><link>https://makerportal.ai/blog/motionlink-arxiv-260816722v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/motionlink-arxiv-260816722v1/</guid><description>BEM-synthesised HRTFs match measured HRTFs on every polar localisation metric (N = 20), but torso-omission error at low rear elevations is untested.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate></item><item><title>CMHeadphoneMotionManager: What You Really Get</title><link>https://makerportal.ai/blog/motionlink-headphone-motion-api/</link><guid isPermaLink="true">https://makerportal.ai/blog/motionlink-headphone-motion-api/</guid><description>Quaternion attitude from the Headphone Motion API, why the reference frame resets on every launch, and the recentering pattern that fixes it.</description><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Minna: local doc search with grounded LLM</title><link>https://makerportal.ai/blog/notiary-hackernews-49362669/</link><guid isPermaLink="true">https://makerportal.ai/blog/notiary-hackernews-49362669/</guid><description>Minna indexes documents into a hybrid vector-plus-full-text store with citation-grounded chat, but requires macOS 26 beta and ships PostHog telemetry on…</description><pubDate>Thu, 20 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Running Semantic Search Entirely On-Device</title><link>https://makerportal.ai/blog/notiary-on-device-semantic-search/</link><guid isPermaLink="true">https://makerportal.ai/blog/notiary-on-device-semantic-search/</guid><description>Notiary&apos;s semantic brain: converting a sentence-transformer to CoreML, deterministic storage math, and the fixed-length-input problem long notes hit.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Why nymic Needs No Training Step</title><link>https://makerportal.ai/blog/nymic-knn-voice-conversion/</link><guid isPermaLink="true">https://makerportal.ai/blog/nymic-knn-voice-conversion/</guid><description>How kNN-VC voice conversion works: WavLM feature extraction, nearest-neighbor matching, HiFiGAN vocoding, and voice-bank coverage limits.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>ANE routes by expression, not arithmetic</title><link>https://makerportal.ai/blog/on-device-ai-arxiv-260822110v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-arxiv-260822110v1/</guid><description>A 64-shape primitive matrix and ANE counters reveal routing follows operation form, with a ~0.77 bytes-per-token law holding across fp16, int8, and 2-bit…</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Fun-ASR vLLM speedup: what conditions hide</title><link>https://makerportal.ai/blog/on-device-ai-github-1116516142/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-github-1116516142/</guid><description>Fun-ASR-Nano&apos;s 16x, 3–5x, and ~50% speedup figures are attached to different—or missing—conditions that the README never aligns.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>OBLITERATUS MMLU gaps: what conditions hide</title><link>https://makerportal.ai/blog/on-device-ai-hf-obliteratus-qwen38-27b-obliterated/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-hf-obliteratus-qwen38-27b-obliterated/</guid><description>V1&apos;s −6.0pp and V2&apos;s −0.3pp MMLU gaps come from sample sizes a factor of ten apart; V1&apos;s arithmetic doesn&apos;t reproduce — the figures are not…</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate></item><item><title>On-Device Vector Search in Swift</title><link>https://makerportal.ai/blog/on-device-vector-search-swift/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-vector-search-swift/</guid><description>How to build distraction-free local semantic search on iOS using CoreML feature extraction and native Swift HNSW index graphs, grounded in Notiary.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Running 3B LLMs on Edge SBCs &amp; MCUs</title><link>https://makerportal.ai/blog/running-3b-llms-on-microcontrollers/</link><guid isPermaLink="true">https://makerportal.ai/blog/running-3b-llms-on-microcontrollers/</guid><description>Why parameter count isn&apos;t the bottleneck for on-device LLMs: memory bandwidth, GGUF/INT4 quantization math, and ANE lessons from Thumbdash and itria.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate></item><item><title>buun-llama-cpp VBR uses 10 steps per layer</title><link>https://makerportal.ai/blog/thumb-dash-github-1192318297/</link><guid isPermaLink="true">https://makerportal.ai/blog/thumb-dash-github-1192318297/</guid><description>The step-count formula in buun-llama-cpp&apos;s examples reduces to exactly 10 degradation steps per attention layer — 6 codec tiers × 2 KV sides minus 2.</description><pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Two 7B Q4 models exceed DLLM&apos;s 8 GB VRAM</title><link>https://makerportal.ai/blog/thumb-dash-hackernews-49279500/</link><guid isPermaLink="true">https://makerportal.ai/blog/thumb-dash-hackernews-49279500/</guid><description>Measured Q4 GGUF sizes show two 7B models consume 8.4 GB, 0.4 GB over DLLM&apos;s ceiling, leaving no room for the embed model or KV-cache.</description><pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Acoustic Beamforming on iPhone: UMA-8</title><link>https://makerportal.ai/blog/uma-8-beamforming-iphone/</link><guid isPermaLink="true">https://makerportal.ai/blog/uma-8-beamforming-iphone/</guid><description>Real-time SRP-PHAT direction finding and MVDR superdirective beamforming with the miniDSP UMA-8 mic array on an iPhone in Biquadia.</description><pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Low-Latency WebAudio DSP for Voice AI</title><link>https://makerportal.ai/blog/webaudio-audioworklet-dsp-voice-ai/</link><guid isPermaLink="true">https://makerportal.ai/blog/webaudio-audioworklet-dsp-voice-ai/</guid><description>How to build zero-glitch AudioWorklet streaming, Schroeder reverb tails, and formant filtering for real-time speech and Voice AI in the browser.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Building Honest WebGPU Browser Benchmarks</title><link>https://makerportal.ai/blog/webgpu-benchmark-browser/</link><guid isPermaLink="true">https://makerportal.ai/blog/webgpu-benchmark-browser/</guid><description>Verification-first WebGPU benchmark methodology: warm-up, queue completion, CPU cross-checks, honest FLOP accounting, and cloud crossover math.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Why we ship small</title><link>https://makerportal.ai/blog/why-we-ship-small/</link><guid isPermaLink="true">https://makerportal.ai/blog/why-we-ship-small/</guid><description>The case for narrow products, strong opinions, and releases that can still surprise their makers.</description><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate></item></channel></rss>