Paper / Audio DSP
BEM HRTFs match measured: headphone VR only
BEM-synthesised HRTFs match measured HRTFs on every polar localisation metric (N = 20), but torso-omission error at low rear elevations is untested.

Joshua HriskoPrincipal Engineer
6 min readSan Francisco, CA

Composed from the signals scanned on 2026-08-18.
BEM-synthesised individual HRTFs match measured HRTFs on every polar localisation metric in a VR task (N = 20) and outperform a generic KEMAR baseline on both numerical fidelity and computational model predictions across 200 subjects — that is what this study establishes. What it does not establish is whether the numerical error pattern — elevated spectral distortion concentrated at low rear elevations due to omitted torso geometry — carries through to perceptual performance. Behavioural errors clustered at the front-back midline in all conditions, not at the numerically implicated angles.
Study design
The paper ran four arms against a common comparator set: BEM-synthesised HRTFs (Mesh2HRTF, individual head meshes, no torso geometry), individually measured HRTFs (ground truth), and a generic KEMAR dummy-head HRTF (current standard practice). Head meshes and paired measured HRTFs came from the Extended SONICOM dataset.
| Arm | What was measured | Baseline | Sample | Conditions |
|---|---|---|---|---|
| Numerical | ITD and ILD deviation from measured | KEMAR deviation from measured | 200 subjects | Mesh2HRTF BEM synthesis, full-head mesh only |
| Computational models | Predicted localisation error ranking | KEMAR predictions | Not stated | Two model architectures, all three HRTF types |
| Perceptual — localisation | Polar error metrics | KEMAR and measured | N = 20 | Head-tracked VR, headphone binaural rendering |
| Perceptual — SRM | Spatial release from masking | KEMAR and measured | N = 18 | Binaural masking paradigm |
The result transfers if you have a 3D head scan of sufficient quality for BEM meshing, you are rendering over headphones in a head-tracked context, and your application task is point-source localisation or spatial unmasking. A different mesh acquisition pipeline, loudspeaker rendering, crosstalk cancellation, or an elevation-specific task at low rear angles — treat the transfer as uncertain.
Results
Across all 200 subjects, the authors report synthetic HRTFs deviated less from individual measurements than KEMAR did on both interaural time difference and interaural level difference. Residual error in the synthetic set concentrated at low rear elevations; the authors attribute this to omission of torso geometry from the BEM pipeline.
Two computational models across all three HRTF conditions produced the same ranked ordering of predicted localisation error: measured best, synthetic intermediate, KEMAR worst — consistent with the numerical findings.
In the VR localisation task (N = 20), synthetic HRTFs matched measured on every polar metric; KEMAR was significantly worse. Behavioural errors clustered around the front-back midline in all three conditions, not at the low rear elevations flagged numerically and by the models. The paper characterises this as a discrepancy: BEM synthesis preserved the perceptually relevant cues even where aggregate spectral distortion was elevated.
The spatial release from masking task (N = 18) showed no effect of HRTF type. The paper does not report a minimum detectable effect size for this arm, so the null result cannot be read as equivalence — the sample may simply have been too small to resolve a difference of the magnitude present.
Where the result stops
The perceptual arms ran at N = 20 and N = 18. Both are standard for psychoacoustics laboratory work and adequate to detect large between-condition effects, but neither rules out small operationally meaningful gaps between synthetic and measured. Without variance estimates or confidence intervals in the abstract, the upper bound on any residual perceptual difference between synthetic and individually measured HRTFs is unconstrained.
The synthesis pipeline excluded torso geometry. Numerical and model arms both identify low rear elevations as the error-concentrated region. The VR localisation task found no perceptual consequences there — but it was not designed to probe rear-elevation accuracy, and the authors do not claim it clears that region. Any application depending on accurate elevation cues for sounds originating below and behind the listener sits outside the scope of what was tested and has not been cleared.
The Extended SONICOM dataset defines the population. Individuals with head morphologies outside that distribution — children, users with significant pinnae asymmetry, or individuals using prosthetics — are not covered. The study was not underpowered to see those populations; it did not sample them.
The computational model arm reports no sample size in the abstract. Whether it had adequate power to distinguish synthetic from measured — as opposed to distinguishing both from KEMAR — cannot be assessed from available information.
All perceptual results apply to headphone rendering. Loudspeaker binaural, room-acoustic simulation with synthetic HRTFs, and crosstalk-cancellation pipelines are not covered.
What to do differently
If your spatial audio renderer defaults to KEMAR or another generic HRTF, this study supports substituting BEM-synthesised individual HRTFs where head scan data is available. The authors report a clear numerical and perceptual advantage over KEMAR, with equivalence to measured HRTFs on every tested polar metric — validated here for headphone rendering in head-tracked VR localisation tasks.
If your application requires accurate elevation discrimination at low rear angles — a surround monitoring tool, a task involving detection of sounds from behind and below — do not treat these results as clearance. The torso-omission artefact is documented numerically; the behavioural task was not designed to probe it and cannot rule out a perceptual effect under the right conditions.
The null SRM result warrants no change to how you handle binaural masking tasks. At N = 18 with no reported power estimate, the absence of a detected effect says nothing about its magnitude. For spatial audio prototyping that incorporates on-device head tracking, the finding that synthetic personalised HRTFs match measured ones perceptually is relevant to the rendering layer; the ElevenLabs Web Audio Streaming Latency note covers the adjacent latency budget question for that stack.
Method: this note was drafted by us.anthropic.claude-sonnet-4-6 from a single source — the arXiv abstract for 2608.16722. Before publication an automated gate re-checked every extracted claim against the source document (7 claim(s) and 0 quantity(ies) verified) and every claim-shaped number in the draft against that evidence (0 derived from it, no rows from our own tables were supplied to the draft). The studio has not re-run 2608.16722’s benchmarks; figures attributed to it are its own.
FAQ
The VR task matched synthetic to measured at N = 20 — is that sample size convincing?
Convincing for the direction of the result — synthetic is not worse than measured, and both are better than KEMAR. Tightly bounding the magnitude of any residual gap between synthetic and measured is a different matter. No variance estimates appear in the abstract, so the study cannot rule out a small but operationally meaningful difference that N = 20 lacked the power to detect.
The SRM arm showed no effect — does that mean HRTF personalisation doesn't matter for spatial unmasking tasks?
Not from this data. A null result at N = 18 means the effect, if present, was smaller than what that sample could reliably detect. No minimum detectable effect size is reported for the SRM arm, so the result carries no information about the practical size of any gap. The hypothesis is not cleared.
Numerical errors concentrate at low rear elevations, but behavioural errors appear at the front-back midline — which one should I act on?
They measure different things. Numerical errors reflect spectral distortion relative to measured HRTFs in a specific angular region; behavioural errors reflect the task those 20 participants actually performed, which was not designed to stress-test rear-elevation accuracy. For front-back localisation in a head-tracked VR context, the perceptual result is the operative one. For elevation-specific judgements at low rear angles, the numerical result is the relevant signal — the perceptual task has nothing to say about it.
Does this transfer to archetype-based or interpolated synthetic HRTFs rather than fully individual BEM simulations?
This study does not address that question. All 200 subjects had individual BEM simulations run on their own head meshes; there is no within-paper comparison to clustered, averaged, or morphologically interpolated synthetic HRTFs. Any inference about archetype-based approaches requires other evidence.
Recommended Studio & Hardware Gear
Affiliate links support independent R&DTested studio equipment and reference hardware utilized for this build. Product images & pricing sourced from Amazon Creators API / SparkFun Electronics.
$74.50BookDigital Design and Computer Architecture: ARM Edition
From gates to RISC-V pipeline — Chapter 3 defines setup/hold T_su/T_h and clock skew that this sculptor checks live; Chapter 7 datapath mapping explains LUT vs MUX vs adder inference.
$530ApparatusSR3D® Dummy Head MKIII
Binaural dummy head acoustic fixture with anatomical silicone pinnae and Primo EM272 electret capsules for HRTF and headphone measurement.
Prices shown were retrieved from the Amazon Product Advertising API on 19 July 2026 and are indicative only — the price and availability on Amazon at the time of purchase apply.