Field note / On-Device Vector Search
On-Device Vector Search in Swift
How to build distraction-free local semantic search on iOS using CoreML feature extraction and native Swift HNSW index graphs, grounded in Notiary.

Joshua HriskoPrincipal Engineer
5 min readSan Francisco, CA

Building semantic search for local markdown notes or private documents requires running both vector embedding generation and nearest-neighbor retrieval 100% offline on-device.
In Notiary (App Store), our distraction-free iOS notes app, we execute local BERT/BGE embeddings on Apple Neural Engine (ANE) and search candidates using a native Hierarchical Navigable Small World (HNSW) vector graph in Swift.
1. Cosine Distance & Vector Normalization Math
Given a query vector and document embedding , the cosine similarity is:
If vectors are L2-normalized upon creation (), cosine similarity simplifies to a simple dot product:
In SIMD-accelerated Swift (Accelerate.framework), this dot product executes in less than 1.1 microseconds per 384-dimensional vector.
2. Flat Scan vs HNSW Recall Tradeoff
For small note archives (< 2,000 items), brute-force flat linear scans across normalized float arrays in SIMD are faster than traversing graph nodes due to zero pointer-indirection overhead:
| Corpus Size | Exact Flat Scan (SIMD) | HNSW Graph () | Memory Overhead |
|---|---|---|---|
| 500 Notes | 0.8 ms | 0.4 ms | None |
| 5,000 Notes | 7.2 ms | 1.1 ms | +20% Graph Nodes |
| 50,000 Notes | 71.0 ms | 2.8 ms (98% Recall) | +35% Graph Nodes |
3. Recommended Reading & Hardware Tools
To dive deeper into deep learning algorithms and spatial audio hardware:
- Deep Learning (Adaptive Computation Series) — Foundational textbook by Goodfellow, Bengio, and Courville.
- Apple AirPods Pro 3 Wireless Earbuds — Hardware testing for MotionLink’s head-tracking and spatial audio vector projections.
Try out our live interactive Vector Retrieval Recall Lab playground to measure recall loss, skipped candidates, and graph traversal performance live in your browser.
4. The Power of Local Semantics
By pairing Apple’s Neural Engine with native Swift implementations of algorithms like HNSW, we bypass network latency and privacy concerns entirely. For small to medium local knowledge bases, on-device semantic search isn’t just a gimmick—it’s a massive UX improvement, bringing the power of modern retrieval directly to the user’s pocket without sacrificing a millisecond of speed.
FAQ
Why is SIMD dot product faster than HNSW for less than 2,000 vectors?
HNSW graph traversal requires following pointer references across memory nodes. For fewer than 2,000 L2-normalized 384-dimensional vectors, contiguous array memory access with Swift's Accelerate vDSP SIMD instructions finishes in under 1.5ms with zero graph overhead.
How does CoreML target the Apple Neural Engine (ANE) for embeddings?
By setting MLModelConfiguration.computeUnits = .all, CoreML automatically offloads transformer encoder layers to the ANE coprocessor, freeing the GPU and main CPU cores for UI rendering.
Recommended Studio & Hardware Gear
Affiliate links support independent R&DTested studio equipment and reference hardware utilized for this build. Product images & pricing sourced from Amazon Creators API / SparkFun Electronics.
$199.99WearableApple AirPods Pro 3 Wireless Earbuds, Active Noise Cancellation, Live Translation, Heart Rate Sensing, Hearing Aid Feature, Bluetooth Headphones, Spatial Audio, High-Fidelity Sound, USB-C Charging
MotionLink's head-tracking feature depends on the Headphone Motion API, which requires AirPods Pro.
$61.00BookDeep Learning (Adaptive Computation and Machine Learning series)
Foundational deep-learning textbook referenced while building itria.
$49.50BookHands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
Practical ML reference used while building itria.
Prices shown were retrieved from the Amazon Product Advertising API on 19 July 2026 and are indicative only — the price and availability on Amazon at the time of purchase apply.