<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet href="/rss-styles.xsl" type="text/xsl"?><rss version="2.0"><channel><title>Field Notes — MakerPortal</title><description>Notes on independent software, product decisions, interface craft, and what we learn while shipping.</description><link>https://makerportal.ai/</link><item><title>PrivateRedact limits 7B decode to 4.0 tok/s</title><link>https://makerportal.ai/blog/2026-08-11-itria-hackernews-49245161/</link><guid isPermaLink="true">https://makerportal.ai/blog/2026-08-11-itria-hackernews-49245161/</guid><description>PrivateRedact&apos;s auto-selection limits Ollama inference to 4.0 tok/s on an 8 GB Pi 5 by defaulting to CPU memory, ignoring unified bandwidth.</description><pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Agentic Coding Loops Beyond the LLM</title><link>https://makerportal.ai/blog/auralinter-agentic-dsp-verification/</link><guid isPermaLink="true">https://makerportal.ai/blog/auralinter-agentic-dsp-verification/</guid><description>How RAG-grounded retrieval and a clang++ verifier catch a hallucinated biquad formula a single LLM call ships — and why AuraLinter isn&apos;t fully on-device.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>CSAVocoder: mel adaptor beats pose 2.6×</title><link>https://makerportal.ai/blog/auralinter-arxiv-260825404v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/auralinter-arxiv-260825404v1/</guid><description>CSAVocoder renders binaural and FOA audio from mel-spectrograms and a 7D pose stream.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate></item><item><title>SwanWeave&apos;s &apos;spatial&apos; module is doing almost no spatial work</title><link>https://makerportal.ai/blog/auralinter-arxiv-260904975v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/auralinter-arxiv-260904975v1/</guid><description>The module named Spatial Edit supplies 91% of SwanWeave&apos;s fidelity win over SmartDJ but only 5% of the spatial one — pretraining does the spatial work.</description><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Calibrated Multichannel &amp; Binaural Audio Recording on iPhone</title><link>https://makerportal.ai/blog/binaural-head-headphone-measurement/</link><guid isPermaLink="true">https://makerportal.ai/blog/binaural-head-headphone-measurement/</guid><description>Class-compliant multichannel USB audio capture, a loaded 1 V / 103 dB SPL nominal reference at 1 kHz, and binaural recording analyzed on iPhone.</description><pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate></item><item><title>LLMKube&apos;s 66× speedup is actually 35.4×</title><link>https://makerportal.ai/blog/biquadia-github-1095330682/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1095330682/</guid><description>Two figures in LLMKube&apos;s GKE benchmark are wrong: prompt processing is 35.4×, not 66×, and token generation is 13.9×, not 17×.</description><pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Wax hybrid recall is 6.1 ms, not sub-ms</title><link>https://makerportal.ai/blog/biquadia-github-1138007869/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1138007869/</guid><description>Wax&apos;s own benchmark puts hybrid recall p95 at 6.1 ms — six times above the &quot;sub-millisecond&quot; label — and the figure omits which index engine ran.</description><pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Orion&apos;s ANE Decode vs CPU Baseline</title><link>https://makerportal.ai/blog/biquadia-github-1171688936/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1171688936/</guid><description>Driving Apple&apos;s Neural Engine through private frameworks: 172.4 tok/s on ANE vs 283 on CPU, and delta compilation that pays for itself in 1.21 steps.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Espresso bounds CoreML overhead at 168.7 µs</title><link>https://makerportal.ai/blog/biquadia-github-1173023018/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1173023018/</guid><description>Espresso&apos;s 3.41× ANE speedup over CoreML, read against a 70 µs dispatch floor, sets a 168.7 µs lower bound on CoreML per-token overhead.</description><pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Presspeech&apos;s 100ms Latency Overhead</title><link>https://makerportal.ai/blog/biquadia-github-1233899535/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1233899535/</guid><description>Bandwidth calculation for Presspeech&apos;s Parakeet model on M4 Max shows inference takes ~1% of latency, with pipeline dispatch consuming the rest.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate></item><item><title>ANEForge hits 2.28× speculative decode</title><link>https://makerportal.ai/blog/biquadia-github-1267734215/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1267734215/</guid><description>ANEForge reaches 16.8 tok/s on Qwen3-8B (up from 7.4), hitting the verify(K)≈verify(1) ceiling at 29.6 GB/s vs Orion&apos;s 42.8 GB/s floor.</description><pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate></item><item><title>8.75:1 hysteresis puts fan stop at 27°C</title><link>https://makerportal.ai/blog/biquadia-github-1349261167/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-github-1349261167/</guid><description>mac-tool-kit&apos;s 7.8°C hysteresis band splits 8.75:1, placing the fan stop at 27.0°C — a temperature inference or DSP sessions never return to mid-session.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate></item><item><title>M4 Max decode uses 99.8% of bandwidth</title><link>https://makerportal.ai/blog/biquadia-hackernews-49259339/</link><guid isPermaLink="true">https://makerportal.ai/blog/biquadia-hackernews-49259339/</guid><description>M4 Max llama.cpp inference consumes 99.0–99.8% of its 400 GB/s bandwidth, leaving no slack to absorb VM overhead — a direct constraint on near-native…</description><pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate></item><item><title>HM-10 vs nRF52: Not the Same BLE UART Profile</title><link>https://makerportal.ai/blog/blexar-ble-uart-bridge/</link><guid isPermaLink="true">https://makerportal.ai/blog/blexar-ble-uart-bridge/</guid><description>Why HM-10&apos;s single-characteristic UART and the Nordic UART Service aren&apos;t interchangeable, and the newline-framing gotcha that only appears at low MTU.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Designing the one-handed moment</title><link>https://makerportal.ai/blog/designing-the-one-handed-moment/</link><guid isPermaLink="true">https://makerportal.ai/blog/designing-the-one-handed-moment/</guid><description>A field note on glanceability, reach, and building interfaces for the imperfect moments where software is actually used.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>ElevenLabs &amp; Web Audio Streaming Latency</title><link>https://makerportal.ai/blog/elevenlabs-web-audio-streaming-latency/</link><guid isPermaLink="true">https://makerportal.ai/blog/elevenlabs-web-audio-streaming-latency/</guid><description>Browser latency ledger for ElevenLabs TTS, Web Audio scheduling, DSP, and output—plus mistakes that make a fast demo sound broken.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Zero-Latency Web Audio &amp; Voice AI Architecture</title><link>https://makerportal.ai/blog/elevenlabs-webassembly-audioworklet-streaming/</link><guid isPermaLink="true">https://makerportal.ai/blog/elevenlabs-webassembly-audioworklet-streaming/</guid><description>How to build sub-100ms real-time voice streaming applications in the browser using ElevenLabs API, WebAssembly decoders, and AudioWorklet ring buffers.</description><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Hy4 preview: 770B MoE, open weights, gaps</title><link>https://makerportal.ai/blog/field-note-hf-tencent-hy4-preview/</link><guid isPermaLink="true">https://makerportal.ai/blog/field-note-hf-tencent-hy4-preview/</guid><description>Tencent&apos;s Hy4 ships Apache 2.0 weights with 49B activated parameters and 1M-token context, but four unexplained residual streams block anything beyond…</description><pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Inside Biquadia: On-Device Audio Lab</title><link>https://makerportal.ai/blog/inside-biquadia/</link><guid isPermaLink="true">https://makerportal.ai/blog/inside-biquadia/</guid><description>How privacy, real-time constraints, and the character of physical sound shaped MakerPortal’s newest release.</description><pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate></item><item><title>iPhone 18 Pro vs iPhone 17 Pro: Pure Iteration</title><link>https://makerportal.ai/blog/iphone-18-pro-vs-iphone-17-pro-launch-iphone-18-pro/</link><guid isPermaLink="true">https://makerportal.ai/blog/iphone-18-pro-vs-iphone-17-pro-launch-iphone-18-pro/</guid><description>Apple&apos;s own spec pages for the iPhone 18 Pro and iPhone 17 Pro, diffed: 16 sections changed, 15 identical. Every figure copied from the two documents.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate></item><item><title>ShikumiMiner AST features fail cross-project</title><link>https://makerportal.ai/blog/itria-arxiv-260902789v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-arxiv-260902789v1/</guid><description>Clang LibTooling pipeline extracts AST and CFG features from C++ LLM inference code. In-sample macro F1 is 0.40; leave-one-project-out drops to 0.112.</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate></item><item><title>watchOS LLM speed is capped by CPU bandwidth</title><link>https://makerportal.ai/blog/itria-github-1051796718/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-1051796718/</guid><description>ETOS-LLM-Studio pins watchOS inference to the CPU path, so bandwidth divided by model size sets the ceiling — 9.4 tok/s on a Pi 5 at 17 GB/s.</description><pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate></item><item><title>OpenMed&apos;s 24–33× MLX speedup is compute, not memory</title><link>https://makerportal.ai/blog/itria-github-1069760430/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-1069760430/</guid><description>OpenMed&apos;s 24–33× MLX speedup is a compute advantage, not a memory one. Batch throughput (3.3× CPU, 2.2× MLX) shows both backends are partially…</description><pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate></item><item><title>RTX 5090 Decode Bandwidth &amp; 25% Penalty</title><link>https://makerportal.ai/blog/itria-github-1164344011/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-1164344011/</guid><description>imp reports a 42-48% decode gap over llama.cpp on consumer Blackwell. The 5090&apos;s own 5090D comparison shows the gap is just a bandwidth model.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Qwen3.6-35B decodes ~11.2 B active per pass</title><link>https://makerportal.ai/blog/itria-github-725205304/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-725205304/</guid><description>Bandwidth-bound decode math on RTX 5090D measurements backs out that only ~32% of Qwen3.6-35B&apos;s nominal weights are active per forward pass.</description><pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Serial loading reclaims 4.2 GB on 8 GB devices</title><link>https://makerportal.ai/blog/itria-github-867433720/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-github-867433720/</guid><description>Two 7B Q4 models exceed an 8 GB ceiling when loaded concurrently. Serial residency via llama-swap reclaims 4.2 GB of KV cache headroom.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item><item><title>MoEspresso coder trades at 2.455:1 NLL</title><link>https://makerportal.ai/blog/itria-hackernews-49321813/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hackernews-49321813/</guid><description>Pruning DeepSeek-V4-Flash-0731 cuts safetensors by 32.63%, but each 0.1-unit code NLL gain costs 0.245 units of WikiText degradation.</description><pubDate>Mon, 17 Aug 2026 00:00:00 GMT</pubDate></item><item><title>llama.cpp fork streams KV cache on 16 GB CUDA</title><link>https://makerportal.ai/blog/itria-hackernews-49511882/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hackernews-49511882/</guid><description>A llama.cpp fork adds block-granular KV cache streaming to the CUDA server path using pinned host memory and a bounded pool, enabling 262K context on a…</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate></item><item><title>gpt-oss-20b scores range 0% to 87% by harness</title><link>https://makerportal.ai/blog/itria-hackernews-49523381/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hackernews-49523381/</guid><description>The same gpt-oss-20b weights produce scores from ~0% to ~87% depending on template, tool format, wire API, and reasoning effort.</description><pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Qwen3.8-27B Cold Fusion: mixed attention</title><link>https://makerportal.ai/blog/itria-hf-davidau-qwen38-27b-cold-fusion-gain-v11-nm-dau-neo-max-mtp-gguf/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hf-davidau-qwen38-27b-cold-fusion-gain-v11-nm-dau-neo-max-mtp-gguf/</guid><description>GGUFs of a Qwen3-27B fine-tune with 75% linear attention layers and ~2/3 thinking-token reduction — benchmark gaps and unverified quality claims noted.</description><pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate></item><item><title>3.02× FastMTP requires draft depth ≥ 3</title><link>https://makerportal.ai/blog/itria-hf-hauhaucs-qwen38-27b-uncensored-hauhaucs-aggressive-mtp-gguf/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hf-hauhaucs-qwen38-27b-uncensored-hauhaucs-aggressive-mtp-gguf/</guid><description>A 3.02× speedup exceeds the 3.00× ceiling for depth-2 speculative decoding, placing FastMTP&apos;s draft depth at 3 or more with an implied acceptance rate of…</description><pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Tiel-Coder&apos;s gaps turn on unstated precision</title><link>https://makerportal.ai/blog/itria-hf-peculiar-ragdoll-tiel-coder-35b-a3b-gguf/</link><guid isPermaLink="true">https://makerportal.ai/blog/itria-hf-peculiar-ragdoll-tiel-coder-35b-a3b-gguf/</guid><description>Tiel-Coder&apos;s README places benchmark figures side by side across two quantization regimes, two chat templates, and unstated precision.</description><pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Benchmarking Edge AI SBCs &amp; Accelerators</title><link>https://makerportal.ai/blog/lattepanda-vs-jetson-orin-edge-ai-benchmarks/</link><guid isPermaLink="true">https://makerportal.ai/blog/lattepanda-vs-jetson-orin-edge-ai-benchmarks/</guid><description>Hardware benchmark comparing x86 SBCs, ARM CUDA Tensor Cores, and USB Edge TPUs for local vision, speech, and embedding inference.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate></item><item><title>LiteFS Multi-Region SQLite Latency Model</title><link>https://makerportal.ai/blog/litefs-multi-region-sqlite/</link><guid isPermaLink="true">https://makerportal.ai/blog/litefs-multi-region-sqlite/</guid><description>Mental model for LiteFS multi-region SQLite: single-primary writes, local replicas, fly-replay forwarding, and honest latency modeling.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>M3 Ultra decode: 9.8 to 139 tok/s measured</title><link>https://makerportal.ai/blog/local-decode-bench-2026-08-26/</link><guid isPermaLink="true">https://makerportal.ai/blog/local-decode-bench-2026-08-26/</guid><description>Decode, prefill and first-token measured on Apple M3 Ultra, 256 GB: 9.8–139 tok/s across 9 models, median of 3 runs, prompt cache defeated.</description><pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate></item><item><title>BEM HRTFs match measured: headphone VR only</title><link>https://makerportal.ai/blog/motionlink-arxiv-260816722v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/motionlink-arxiv-260816722v1/</guid><description>BEM-synthesised HRTFs match measured HRTFs on every polar localisation metric (N = 20), but torso-omission error at low rear elevations is untested.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate></item><item><title>CMHeadphoneMotionManager: What You Really Get</title><link>https://makerportal.ai/blog/motionlink-headphone-motion-api/</link><guid isPermaLink="true">https://makerportal.ai/blog/motionlink-headphone-motion-api/</guid><description>Quaternion attitude from the Headphone Motion API, why the reference frame resets on every launch, and the recentering pattern that fixes it.</description><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Minna: local doc search with grounded LLM</title><link>https://makerportal.ai/blog/notiary-hackernews-49362669/</link><guid isPermaLink="true">https://makerportal.ai/blog/notiary-hackernews-49362669/</guid><description>Minna indexes documents into a hybrid vector-plus-full-text store with citation-grounded chat, but requires macOS 26 beta and ships PostHog telemetry on…</description><pubDate>Thu, 20 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Running Semantic Search Entirely On-Device</title><link>https://makerportal.ai/blog/notiary-on-device-semantic-search/</link><guid isPermaLink="true">https://makerportal.ai/blog/notiary-on-device-semantic-search/</guid><description>Notiary&apos;s semantic brain: converting a sentence-transformer to CoreML, deterministic storage math, and the fixed-length-input problem long notes hit.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>XTTSv2 anonymizer hits 0.49 EER</title><link>https://makerportal.ai/blog/nymic-arxiv-260827360v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/nymic-arxiv-260827360v1/</guid><description>XTTSv2 reaches 0.49 EER on CommonVoice via speaker-embedding swaps, bounded by one ECAPA2 probe and seven of sixteen supported languages.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Why nymic Needs No Training Step</title><link>https://makerportal.ai/blog/nymic-knn-voice-conversion/</link><guid isPermaLink="true">https://makerportal.ai/blog/nymic-knn-voice-conversion/</guid><description>How kNN-VC voice conversion works: WavLM feature extraction, nearest-neighbor matching, HiFiGAN vocoding, and voice-bank coverage limits.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>ANE routes by expression, not arithmetic</title><link>https://makerportal.ai/blog/on-device-ai-arxiv-260822110v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-arxiv-260822110v1/</guid><description>A 64-shape primitive matrix and ANE counters reveal routing follows operation form, with a ~0.77 bytes-per-token law holding across fp16, int8, and 2-bit…</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate></item><item><title>PACodec&apos;s 30% reduction: two baselines</title><link>https://makerportal.ai/blog/on-device-ai-arxiv-260903363v1/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-arxiv-260903363v1/</guid><description>The 30% figure in the PACodec paper comes from two separate comparisons at different sample rates, baseline bitrates, and dataset sizes.</description><pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Fun-ASR vLLM speedup: what conditions hide</title><link>https://makerportal.ai/blog/on-device-ai-github-1116516142/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-github-1116516142/</guid><description>Fun-ASR-Nano&apos;s 16x, 3–5x, and ~50% speedup figures are attached to different—or missing—conditions that the README never aligns.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>CoreML-Models file sizes are comparable; the performance claims next to them are not</title><link>https://makerportal.ai/blog/on-device-ai-github-286898814/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-github-286898814/</guid><description>The CoreML-Models repo lists file sizes that are directly comparable across models, but the speed and parameter claims embedded in the same README…</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>AutoSubs star ratings hide different training</title><link>https://makerportal.ai/blog/on-device-ai-github-614149835/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-github-614149835/</guid><description>AutoSubs&apos; model table uses one star scale across models with different corpora, quantization levels, and language sets.</description><pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate></item><item><title>3.5-bit Qwen3.8-Flash-Next on two RTX 3090s beats BF16 on MMLU-Pro</title><link>https://makerportal.ai/blog/on-device-ai-hackernews-49703818/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-hackernews-49703818/</guid><description>A 3.5-bit GGUF quantization of the 180B Qwen3.8-Flash-Next model runs on 2× RTX 3090s and scores 2.85 points above its BF16 reference on MMLU-Pro.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>OBLITERATUS MMLU gaps: what conditions hide</title><link>https://makerportal.ai/blog/on-device-ai-hf-obliteratus-qwen38-27b-obliterated/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-ai-hf-obliteratus-qwen38-27b-obliterated/</guid><description>V1&apos;s −6.0pp and V2&apos;s −0.3pp MMLU gaps come from sample sizes a factor of ten apart; V1&apos;s arithmetic doesn&apos;t reproduce — the figures are not…</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate></item><item><title>On-Device Vector Search in Swift</title><link>https://makerportal.ai/blog/on-device-vector-search-swift/</link><guid isPermaLink="true">https://makerportal.ai/blog/on-device-vector-search-swift/</guid><description>How to build distraction-free local semantic search on iOS using CoreML feature extraction and native Swift HNSW index graphs, grounded in Notiary.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Running 3B LLMs on Edge SBCs &amp; MCUs</title><link>https://makerportal.ai/blog/running-3b-llms-on-microcontrollers/</link><guid isPermaLink="true">https://makerportal.ai/blog/running-3b-llms-on-microcontrollers/</guid><description>Why parameter count isn&apos;t the bottleneck for on-device LLMs: memory bandwidth, GGUF/INT4 quantization math, and ANE lessons from Thumbdash and itria.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate></item><item><title>buun-llama-cpp VBR uses 10 steps per layer</title><link>https://makerportal.ai/blog/thumb-dash-github-1192318297/</link><guid isPermaLink="true">https://makerportal.ai/blog/thumb-dash-github-1192318297/</guid><description>The step-count formula in buun-llama-cpp&apos;s examples reduces to exactly 10 degradation steps per attention layer — 6 codec tiers × 2 KV sides minus 2.</description><pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Two 7B Q4 models exceed DLLM&apos;s 8 GB VRAM</title><link>https://makerportal.ai/blog/thumb-dash-hackernews-49279500/</link><guid isPermaLink="true">https://makerportal.ai/blog/thumb-dash-hackernews-49279500/</guid><description>Measured Q4 GGUF sizes show two 7B models consume 8.4 GB, 0.4 GB over DLLM&apos;s ceiling, leaving no room for the embed model or KV-cache.</description><pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Task type splits MTP acceptance 2.05×</title><link>https://makerportal.ai/blog/thumbdash-hackernews-49434442/</link><guid isPermaLink="true">https://makerportal.ai/blog/thumbdash-hackernews-49434442/</guid><description>Document tasks yield acceptance 0.615 vs. 0.300 for reasoning — a 2.05× gap quant level and hardware cannot explain; both run Q8KP on RTX 6000 Ada.</description><pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Acoustic Beamforming on iPhone: UMA-8</title><link>https://makerportal.ai/blog/uma-8-beamforming-iphone/</link><guid isPermaLink="true">https://makerportal.ai/blog/uma-8-beamforming-iphone/</guid><description>Real-time SRP-PHAT direction finding and MVDR superdirective beamforming with the miniDSP UMA-8 mic array on an iPhone in Biquadia.</description><pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Low-Latency WebAudio DSP for Voice AI</title><link>https://makerportal.ai/blog/webaudio-audioworklet-dsp-voice-ai/</link><guid isPermaLink="true">https://makerportal.ai/blog/webaudio-audioworklet-dsp-voice-ai/</guid><description>How to build zero-glitch AudioWorklet streaming, Schroeder reverb tails, and formant filtering for real-time speech and Voice AI in the browser.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Building Honest WebGPU Browser Benchmarks</title><link>https://makerportal.ai/blog/webgpu-benchmark-browser/</link><guid isPermaLink="true">https://makerportal.ai/blog/webgpu-benchmark-browser/</guid><description>Verification-first WebGPU benchmark methodology: warm-up, queue completion, CPU cross-checks, honest FLOP accounting, and cloud crossover math.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Why we ship small</title><link>https://makerportal.ai/blog/why-we-ship-small/</link><guid isPermaLink="true">https://makerportal.ai/blog/why-we-ship-small/</guid><description>The case for narrow products, strong opinions, and releases that can still surprise their makers.</description><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate></item></channel></rss>