Cantonese Voice Benchmark 2026.1 · Window 2
Production voice systems evaluated on Hong Kong insurance contact-centre audio across four pillars — hearing, understanding, speaking, and conversation.
The first scored window has not closed. Reference baselines are shown below so the scale is meaningful the moment real entries land.
| A | Open-source pipeline | 76.2 | 11.1% | 3.81 | 3.44 | 3.29 | 2.11sp95 3.88 |
| B | Proprietary voice LLM | 82.7 | 8.6% | 4.14 | 3.62 | 4.01 | 0.88sp95 1.52 |
Cantonese carries six lexical tones, writes no spaces between words, and switches into English mid-sentence without marking the boundary. A model tuned on Mandarin does not degrade gracefully here — it fails in ways that read as fluent.
buy · sell
Low rising against low level. In a sales call this inverts who is doing what.
insure · report
High rising against mid level. 保險 is a policy; 報險 is filing a claim against one. Both are said constantly on the same call.
medical · two
High level against low level. A claims call is dense with both hospital words and read-aloud digits, so this one lands in the amount as easily as the diagnosis.
Hear all six tones on the recognition pillar — the contours are plotted and playable.
No files are uploaded and nothing is scored on request. We connect to a server the entrant runs, stream audio at it, and measure what comes back on our own clock.
A WebSocket server you host. We are the client, so nothing of yours has to be handed over and nothing of ours has to be trusted.
12 clips with published ground truth, downloadable once you have an account. Unlimited runs, full per-clip diffs, no effect on your score.
30 clips drawn at random from a pool of 55, stratified across difficulty tiers, seeded and recorded so the draw can be reconstructed.
Each clip resolves to a verdict and an owner: yours, ours, the network, or nobody. A run we break costs you nothing and does not consume quota.
A bootstrap interval accompanies every number, and entrants whose intervals overlap are published as tied rather than ranked against noise.
A benchmark is only worth reading if it is willing to say less than it could. These are the four cases where we withhold a number rather than produce one.
Entrants whose confidence intervals overlap share a rank band and are reported as tied.
A run that completed too few clips is published as insufficient data. Turning a broken integration into a bad score would misattribute an operations problem as model quality.
Pillars are opt-in. A team that did not enter speech synthesis is shown as not entered, never as scoring nothing.
The language-model judge runs 3 times across 2 model families and is correlated against human raters before its pillar counts at all.
Organised in Hong Kong, on Hong Kong audio, judged by Cantonese speakers. The dataset is real contact-centre recordings, consented and masked, never scraped.