- Engineering
- Audio Processing
- Wake Word
- PocketSphinx
- YAMNet
- TensorFlow Lite
- WebSocket
- Performance
- Benchmarking
- Python
- Oremi
Oremi Ohunerin 4.0.0.b11: Turning Measurements into a Faster and More Reliable Audio Server
A benchmark-driven look at Oremi Ohunerin 4.0.0.b11, covering audio inference performance, WebSocket concurrency, wake-word reliability, decoder isolation, and the French pronunciation study.
Summarise this page with
your favorite AI assistant
I have just released Oremi Ohunerin 4.0.0.b11, a new beta release of the real-time audio detection component behind the Oremi Personal Assistant.
This release is less about adding shiny new features and more about something I find much more valuable for this kind of system: measuring what the server actually does, finding the real bottlenecks, and fixing the parts that matter.
Oremi Ohunerin receives raw PCM audio over WebSockets and runs two machine-learning pipelines:
- YAMNet / TensorFlow Lite for environmental sound detection.
- PocketSphinx for wake-word detection, currently in English and French.
The project is designed to run continuously, with multiple clients potentially streaming audio at the same time. That makes latency, CPU usage, concurrency and recognition reliability particularly important.
Version 4.0.0.b11 brings substantial improvements in all of these areas.
From assumptions to measurements
The starting point for this release was a proper performance audit.
Instead of optimizing individual functions because they looked expensive, I built reproducible benchmark harnesses covering the complete path from incoming audio chunks to detection results.
The benchmarks measure:
- individual inference stages;
- end-to-end server throughput;
- event-loop blocking;
- audio buffering overhead;
- model initialization time;
- memory consumption;
- wake-word decoder behavior;
- and, separately, actual wake-word detection performance.
One important lesson came immediately from the measurements:
The biggest performance problem was not Python itself. It was how the machine-learning runtime was being used.
The TFLite thread-count trap
The sound detector was originally giving TensorFlow Lite os.cpu_count() as its number of inference threads.
On the benchmark machine, that meant 20 threads.
It sounds reasonable. It was actually disastrous.
The benchmark showed:
| TFLite threads | Median inference time |
|---|---|
| 1 | 7.7 ms |
| 2 | 5.7 ms |
| 3 | 3.5 ms |
| 4 | 2.7 ms |
| 6 | 2.4 ms |
| 8 | 1.8 ms |
| 12 | 7.9 ms |
| 16 | 228.7 ms |
| 20 | 309.4 ms |
The important part is not that 4 threads happened to be the fastest value in this particular run.
It is the shape of the curve.
Once the number of XNNPACK workers became too large for the available CPU resources, performance collapsed. Twenty threads could make a single inference more than 100 times slower than a properly bounded configuration.
The release therefore introduces a bounded inference-thread selection, using the process CPU affinity and capping the default at four threads.
This one change completely changed the performance characteristics of the sound detector.
Around 100× faster sound inference
With the production configuration, sound inference went from approximately:
185 ms → 1.84 ms per window
That is roughly a 100× improvement.
More importantly, this was not achieved by changing the model, reducing the quality of the classifier, changing the detection threshold, or introducing approximate inference.
It is the same YAMNet model and the same detection pipeline, with the inference runtime configured appropriately.
The final benchmark shows that TFLite inference still accounts for approximately 98% of the time spent processing a detection window.
That is actually a useful result: it tells us that the optimization work is now focused on the right place.
The event loop no longer waits for machine learning
The second major problem was architectural.
Sound detection was being executed directly on the asyncio event loop.
A synchronous machine-learning inference therefore blocked the event loop while it was running.
With a single connection this can be difficult to notice. With several simultaneous audio streams, it becomes a serious problem.
The benchmark makes the difference very clear.
For four concurrent connections:
| Configuration | Throughput | Event-loop blocking |
|---|---|---|
| Original configuration | 6.8 windows/s | 3,542 ms |
| Optimized configuration | 368–435 windows/s | 1.7 ms median / 5.1 ms max |
The optimized server processes roughly 54–64× more detection windows per second in this benchmark.
But I consider the event-loop result even more important.
The event loop is now essentially free to continue handling network activity while the expensive sound inference runs in the executor.
This changes the behavior of the server under concurrency from "one client can make everybody wait" to something much closer to what a real-time WebSocket service should be doing.
There is a deliberate trade-off here: offloading inference introduces some overhead compared with running the same operation inline. The measured single-engine throughput cost was roughly 12–35% in controlled comparisons.
That is a worthwhile trade for a network server because the alternative is blocking every connection.
Fixing concurrency correctness, not just speed
Moving inference to worker threads exposed another issue.
The original implementation used a shared audio consumer containing a mutable audio buffer. Multiple connections could therefore write into the same buffer.
That was not merely a performance concern.
Two clients could potentially interleave their audio data and produce corrupted detection windows.
The new implementation gives every WebSocket connection its own buffering state while keeping the expensive YAMNet interpreter shared and protected.
This provides an important combination:
- per-connection audio state;
- one shared model instance;
- bounded memory usage;
- serialized access to the native interpreter.
The additional memory cost is tiny: approximately 16 KB per connection for the consumer buffer.
A 1,400× improvement in audio buffering
The benchmark also found a surprisingly expensive piece of Python code in the hot path.
An 8 KB audio chunk was previously copied byte by byte:
for byte in chunk:
buffer[index] = byte
index += 1That meant thousands of Python-level operations for every incoming audio frame.
Replacing that with a slice assignment reduced the measured buffering cost from:
0.766 ms → 0.0005 ms
or approximately 1,400× faster.
The optimization also removes an unnecessary copy before converting the completed buffer for inference.
The old behavior was kept as the semantic reference, and regression tests verify that the new buffering path produces the same windows for the tested chunk-boundary cases.
The wake-word decoder is now safe under concurrency
Wake-word processing already ran in the thread pool, but the underlying PocketSphinx decoder was shared.
That meant multiple threads could potentially access the same native decoder simultaneously.
This is a particularly unpleasant class of bug: everything can appear to work for a long time, until concurrency produces a native-state race.
The release serializes access to the PocketSphinx decoder.
This prevents concurrent mutation of the same decoder and makes the existing shared-decoder architecture safe from that particular race.
There is still an architectural limitation here: a shared decoder does not provide completely isolated utterance streams between simultaneous clients. A future bounded per-language decoder pool is the natural next step.
That is intentionally a separate architectural change because each decoder is expensive — roughly 160–200 ms to create and around 22 MB of memory.
Then came the wake-word study
Performance was only half of the work.
The other important question was much simpler:
Does Oremi actually recognize "oremi" reliably enough?
Rather than changing thresholds until the result "felt better", I built a dedicated wake-word study harness and evaluated the actual production WakewordEngine.
The study compared different pronunciation configurations using labelled positive and negative audio samples.
For French, the key question was how PocketSphinx should represent the pronunciation of oremi.
The previous pronunciation was approximately:
oo rr ei mm iiThe study identified another pronunciation that better matched the intended French pronunciation:
oo rr ai mm iiThe important thing is that the change was tested through the real production detection path, not just against an isolated phonetic representation.
French wake-word results
With the same:
- acoustic model;
kws_threshold = 1e-15;kws_delay = 10;- 16 kHz sample rate;
- decoder configuration;
- and production
WakewordEngine;
the measured results were:
| French configuration | TP | FN | TPR | FP | FPR | Youden J |
|---|---|---|---|---|---|---|
| Previous pronunciation | 39 | 273 | 12.50% | 19 | 4.09% | +0.0841 |
| New pronunciation | 47 | 265 | 15.06% | 24 | 5.17% | +0.0989 |
The French true-positive rate therefore increased by 2.56 percentage points, or about 20.5% relative.
The negative side of the trade-off is also visible: the false-positive rate increased by 1.08 percentage points.
This is important because I don't want to present the change as a magic improvement without showing its cost.
The additional false positives were concentrated in synthetic near-miss material. On the real-human French negative samples, the false-positive rate remained 1/190.
The change is also not a strict superset of the old pronunciation: some samples that matched the previous phonetic representation no longer match the new one. The net result is nevertheless an improvement on the labelled corpus.
English remained completely unchanged
One of the constraints of this experiment was that the French work must not accidentally change English recognition.
The English configuration was therefore treated as a control.
The result:
| Language | TPR before | TPR after | FPR before | FPR after |
|---|---|---|---|---|
| French | 12.50% | 15.06% | 4.09% | 5.17% |
| English | 25.16% | 25.16% | 1.70% | 1.70% |
The English verdicts were bit-identical across the benchmark samples.
No English acoustic model, threshold, pronunciation or workaround was changed.
The production path and the independent probe agree
I also wanted to make sure that the benchmark wasn't accidentally measuring a simplified test implementation.
So the pronunciation experiment was independently checked through the production engine and through the separate probe path.
For French:
- production path: 47 true positives after the change;
- independent probe: 47 true positives after the change.
For English:
- 77 true positives before and after.
The final production-engine verification compared 776 French samples and 836 English samples, with zero disagreements after the change.
That gives considerably more confidence than a benchmark based on a mocked decoder or a hand-written phoneme test.
Regression testing
The release also comes with a stronger regression safety net around these changes.
The test suite now covers:
- detector buffering equivalence;
- connection-local audio consumers;
- thread-safe detector inference;
- serialized wake-word decoder access;
- pronunciation selection;
- separation between French and English pronunciations;
- KWS configuration;
- production
process_rawbehavior.
The complete suite reached:
81 tests passed.
Static and security checks were also run against the changed code, including Ruff, Bandit and EditorConfig validation.
The only Ruff finding remaining is a pre-existing unused import in the test suite.
What did not change
A useful part of a performance release is being explicit about what was not changed.
I deliberately did not change the acoustic models.
I did not change the YAMNet detection threshold.
I did not silently alter the audio protocol.
I did not enable every pronunciation variant simultaneously just to increase recall.
I also did not change the semantics of the audio-window pipeline simply because the benchmark revealed a potential opportunity there.
There are still documented areas for future work, including:
- bounded pools of wake-word decoders for true session isolation;
- multiple YAMNet interpreters for higher aggregate throughput;
- separating detection workloads into dedicated executors;
- moving expensive engine initialization away from the event loop;
- and revisiting the handling of audio bytes that remain after a detection window is filled.
Those are deliberately left for future releases because they require either additional benchmarking or an explicit product-level decision.
WebSockets 17.1
This release also moves the server to the modern asyncio WebSocket API provided by websockets 17.1.
It is not the headline performance improvement of this release, but it puts the networking layer on the current API rather than continuing to build on the deprecated legacy implementation.
The important part is that this migration was kept separate from the ML benchmark work, so the measured performance improvements above are attributable to the actual inference and concurrency changes rather than being presented as an unexplained effect of a dependency upgrade.
What 4.0.0.b11 means in practice
Oremi Ohunerin started this cycle as a working audio server.
It comes out of the cycle as a much better measured audio server.
The most significant numbers are:
| Metric | Before | After |
|---|---|---|
| Sound inference | 185.2 ms | 1.84 ms |
| 4-connection throughput | 6.8 windows/s | 368–435 windows/s |
| Event-loop blocking | 3,542 ms | 1.7 ms median / 5.1 ms max |
| Audio buffering per 8 KB chunk | 0.766 ms | 0.0005 ms |
/openapi.json | 346 µs | 20.9 µs |
| French wake-word TPR | 12.50% | 15.06% |
| English wake-word TPR | 25.16% | 25.16% |
The most satisfying part is that these improvements did not come from one giant rewrite.
They came from understanding the system all the way down to the actual hot path:
WebSocket
↓
audio chunk
↓
per-connection buffering
↓
executor
↓
ML inference
↓
wake-word / sound detection
↓
JSON eventThen measuring each part.
That is the approach I want to keep for Oremi going forward: measure first, change deliberately, benchmark again, and keep the numbers honest.
Oremi Ohunerin 4.0.0.b11 is now available as the latest beta release of the project, with the complete API and deployment documentation published alongside it.
Installation and Usage
Ready to try Oremi Ohunerin? The complete installation instructions, configuration options, API reference, and usage examples are available in the official documentation.
See the Oremi Ohunerin Documentation for everything you need to get started.
Related articles
- Sep 25, 2026 Ṣeto: a small orchestrator for DockerDocker · Networking · Security · Hardware · Traefik · Linux · Python
- Sep 19, 2026 Taking Oremi Izwi to the GPU: What We Learned from a Deep TTS Performance InvestigationAI · GPU · Python · TTS
- May 12, 2025 Crafting Clear FastAPI Docstrings for OpenAPI: A Human-Centered ApproachPython · FastAPI · Web