Benchmark · frontier-ASR baseline
How a frontier ASR hears Bantu-accented English.
Every code-switch take is time-segmented at the Bantu→English boundary, so the English half is isolated, transcribed by a frontier ASR (Google Cloud Speech-to-Text), and scored against the gold reference. The result is a measured, reference-backed accent-robustness benchmark — run your own model on the same audio and compare to this baseline.
The systematic failure mode
The accent breaks function words before it breaks meaning.
The substantive English is mostly recovered (content recall 81%),
but the discourse marker means is mis-transcribed in
69% of takes — heard as “men's”, “main”, even
“Main Street”. A high-frequency word a frontier ASR gets wrong almost every time on this
accent. That is exactly the gap an accent benchmark is built to surface.
Listen + compare
Gold reference vs the ASR hypothesis.
Play the English segment (what the STT was given), then read the gold reference against Cloud STT's transcript.
Audio pending.
Audio pending.
Audio pending.
Audio pending.
Audio pending.
Audio pending.
Audio pending.
Audio pending.
Audio pending.
Per-language intelligibility
Measured accent difficulty.
| Language | Scored | Content recall | Deviation WER | “means” err | Not-English | Mean conf. |
|---|---|---|---|---|---|---|
| 🇸🇿 Swati | 215 | 82% | 1.14 | 72% | 7% | 0.73 |
| 🇿🇲 Bemba | 12 | 46% | 0.56 | 25% | 33% | 0.66 |
The offer
A turnkey accent-robustness eval, with a frontier baseline.
Run your ASR on the same time-segmented audio, score against the gold references, and
compare to this Cloud-STT baseline. It is verification/validation test data for
Bantu-English speech — native, consented, reference-backed, and growing. Pull it from
/api/v1/tone/asr-benchmark.