Correction / On-Device AI
Subtitles 0.15 RTF is a residual after 71× context amplification

Joshua HriskoPrincipal Engineer
5 min readSan Francisco, CA

Composed from the signals scanned on 2026-09-17.
The 0.15 real-time factor that Subtitles reports for its ANE transcription path is not a measure of model speed. It is the residual after a 71× context amplification factor has been absorbed by the engine. The Parakeet streaming export re-encodes 5.6 s of left context for every 80 ms of new audio, meaning the ANE must sustain roughly 470× real-time throughput of the raw model just to deliver the headline number.
What Subtitles is
Subtitles is a macOS menu-bar application that transcribes system audio in real time using NVIDIA Parakeet on the Apple Neural Engine via FluidAudio. It targets users who need live captions for any audio playing on their Mac — meetings, video, music — without sending audio to a cloud API. The core is Rust with a Swift macOS shell, requires Apple Silicon and macOS 14.2+, and ships a seven-day trial.
The numbers as stated
The README attaches these RTF figures to different conditions:
| RTF | Language | Feature flag | Hardware | Model variant |
|---|---|---|---|---|
| 0.08–0.11 | French | — | ANE | Not stated (headline names Parakeet) |
| 0.13–0.18 | Not stated | Speaker change off | ANE | Not stated |
| 0.27–0.33 | Not stated | Speaker change on | ANE | Not stated |
| 10.7–31.8 | Not stated | — | CPU | Not stated |
| ~0.01 | Not stated | Silero VAD only | ANE | Silero VAD |
The architectural detail that makes these numbers non-obvious is in the “How it works” section: “Parakeet’s streaming export re-encodes 5.6 s of left context per 80 ms chunk.”
Conditions around the comparison
scroll →- 0.15
- real-time factor
- 0.08–0.11
- measured real-time factor on French
- 160 ms
- chunk size
- 0.01 RTF
- measured real-time factor
The context amplification factor
For a streaming decoder that re-encodes seconds of context for every seconds of new audio, the effective RTF is:
where is the raw model throughput expressed in seconds-of-audio-processed per second-of-wall-clock. The ratio is the context amplification factor — the multiplier by which the model’s work exceeds the cost of transcribing only the new audio.
With and :
Solving for at each reported RTF:
| RTF | (× real-time) |
|---|---|
| 0.08 | 887.5 |
| 0.11 | 645.5 |
| 0.15 (headline) | 473.3 |
| 0.18 | 394.4 |
| 10.7 (CPU lower) | 6.64 |
| 31.8 (CPU upper) | 2.23 |
The ANE is sustaining 394–888× real-time raw throughput on this model. The CPU manages 2.2–6.6×. The ratio between the ranges is 60–400×, and the README’s “about 100× too slow” sits inside that band. It is not the geometric mean of those endpoints, which is ≈155×.
This reframes what the headline number means. A reader who sees “0.15 RTF” and thinks “the model runs 6.7× faster than real-time” is off by a factor of 71. The model is running hundreds of times faster than real-time in raw throughput; the 0.15 is what remains after the streaming export architecture re-feeds 71× more audio through the encoder than the user is hearing.
What this does not establish
The 5.6 s / 80 ms figures are attached to “Parakeet’s streaming export” specifically. The Nemotron variants (560, 1120, 2240 ms chunk sizes) and the Parakeet EOU variants (320, 1280, 160 ms) are different model configurations with different native chunk sizes. Their context re-encoding ratios are not stated in the document. The RTF measurements do not specify which model variant was active for each reading, so the values above are only valid for the Parakeet streaming export path.
What this means for a streaming ASR design
The context amplification factor, not the model’s per-token speed, is the binding constraint on whether on-device streaming transcription is viable. A model that is twice as fast per token but requires twice the context re-encoding yields the same RTF. If you are evaluating a model for live captioning, the ratio is the number to compare first; raw FLOPs or ms/token are secondary.
For Biquadia’s on-device neural DSP pipeline, this is directly relevant to any feature that processes streaming audio with lookback context — reverb estimation, source separation, or the neural enhancement models we already ship. The ANE routing paper we covered showed that the engine routes by expression shape, not arithmetic equivalence. Combined with the finding here, the design implication is that the cost of a streaming neural feature on ANE is governed by how much context you re-feed per new frame, and the ANE’s throughput headroom on small models is large enough to absorb very aggressive context windows. The constraint is not compute; it is the architectural decision of how much lookback you choose to re-encode.
Method: this note was drafted by qwen/qwen3.8-27b from two sources — the published README of daformat/subtitles and this studio’s own published measurements, linked above. Before publication an automated gate re-checked every extracted claim against the source document (33 claim(s) and 41 quantity(ies) verified) and every claim-shaped number in the draft against that evidence (14 derived from it, 16 row(s) supplied from our own tables). The studio has not re-run daformat/subtitles’s benchmarks; figures attributed to it are its own.
FAQ
Does the 71× factor mean the model is inefficient?
No. The 5.6 s left context is a design choice in Parakeet's streaming export to maintain accuracy across chunk boundaries. It is a correctness feature, not waste. The question is whether the ANE has enough headroom to absorb it, and at 394–888× raw throughput the answer is yes with substantial margin.
Can I use the RTF figures to predict performance on a different Mac?
Not directly. The RTF measurements do not state which Apple Silicon generation produced them. The ANE throughput scales with chip generation, and the context amplification factor is a property of the model architecture, not the hardware. You would need to re-measure $T$ on your target hardware.
Why does speaker change detection roughly double the RTF?
The README reports 0.13–0.18 without it and 0.27–0.33 with it, a factor of roughly 2×. It does not state the mechanism. Given that diarization "needs ~1 s of warmup and reports on a ~0.5 s cadence," the additional cost is likely from running a second model in parallel or in a tight loop, but this is not confirmed in the document.
Is the 0.01 RTF for Silero VAD meaningful in the same framework?
Silero VAD is a separate, much smaller model running at ~0.01 RTF. It does not re-encode context in the same way — it is a frame-level classifier. The 71× amplification framework applies to the Parakeet streaming export path, not to the VAD stage. The VAD's cost is negligible relative to the transcription path, so it does not meaningfully affect the end-to-end RTF.
Recommended Studio & Hardware Gear
Affiliate links support independent R&DTested studio equipment and reference hardware utilized for this build. Product images & pricing sourced from Amazon Creators API / SparkFun Electronics.
$14.98BookReal-Time Systems
Schedulability Theory: hyperbolic bound, SRP, and response-time analysis R_i = C_i + Σ⌈R_i/T_j⌉ C_j — formulas this scheduler evaluates to predict deadline misses before Gantt draws them.
$1,179.00PhoneApple iPhone 17 Pro, US Version, 512GB, eSIM, Silver- Unlocked (Renewed Premium)
The outgoing Pro generation, still the reference iOS device for on-device inference work here. Amazon Renewed unit — Apple no longer sells this model new, which is the same fact that retires its specification page (D-415).
$1,599.55Computer13-inch MacBook Air (M5): 32GB Memory, 512GB SSD - Midnight
32 GB of unified memory in the lightest Apple silicon body — enough to keep a quantized mid-size model resident instead of streaming it off SSD.
$6,999.00ComputerApple MacBook Pro Laptop with M5 Max, 18‑core CPU, 40‑core GPU: Standard 16.2-inch Display, 128GB Unified Memory, 2TB SSD Storage; Space Black
The portable 128 GB machine. Same resident-model argument as the Mac Studio, in a laptop you can profile on.
$329.00WearableApple Watch Series 11 [GPS 46mm] Smartwatch with Jet Black Aluminum Case with Black Sport Band - M/L. Sleep Score, Fitness Tracker, Health Monitoring, Always-On Display, Water Resistant
The watchOS target itself. Any on-device inference claim for the Watch is bounded by its CPU-accessible bandwidth, which Apple does not publish.
$4,999.99ComputerNVIDIA DGX Spark™ - Personal AI Desktop Supercomputer – Desktop GB10 Grace Blackwell Chip
128 GB of coherent unified memory on a GB10 Grace Blackwell chip. A 70B at Q4_K_M is ~33 GiB of weights, so this holds one resident with room for long context — and Q8 too.
Prices shown were retrieved from the Amazon Product Advertising API on 20 September 2026 and are indicative only — the price and availability on Amazon at the time of purchase apply.
Prices shown were each checked against the Amazon product listing between 8 August 2026 and 17 September 2026 and are indicative only — the price and availability on Amazon at the time of purchase apply.