Announcement · 27 Jul, 05:14 HKT
Cantonese speech evaluation is not missing. Evaluation on real Hong Kong enterprise call audio is.
There are already Cantonese speech benchmarks, and any claim that this is the first would be false. WSYue-eval covers recognition and synthesis. HKCanto-Eval, CantoNLU, MDCC and Common Voice zh-HK all exist and all predate us. We cross-anchor against WSYue-ASR-eval precisely so that a character error rate published here can be read alongside one published there.
What does not exist is an evaluation on real Hong Kong insurance contact-centre audio: consented recordings of actual calls, with the hesitation, the crosstalk, the background of an open floor, and the code-switching that a Hong Kong caller does without noticing. Read speech is a different problem and models that do well on it do not reliably do well on this.
Four pillars, because a voice system is four systems. Can it hear, can it understand, can it speak, and can it hold a conversation. Entrants may enter any subset. A team that enters two pillars is ranked on the Core Score alongside everyone else, and a pillar not entered is reported as not entered, never as a zero.
Cantonese Voice Benchmark · Published 27 Jul, 05:14 HKT.