Methodology · 27 Jul, 05:14 HKT
Prose describing how a transcript is cleaned before scoring is not enough to reproduce a number.
Every benchmark that reports character error rate applies some normalization first, and the difference between two published CER figures is often the normalization rather than the model. Full stops, spacing around Latin words, digits written 三千 or 3000, traditional and simplified variants of the same character: each of those decisions moves the number, and a paragraph of prose describing them is not enough for anyone to arrive at the same figure.
Ours is published as the function that runs, versioned, with the version stamped on every score. If we change it, scores computed under the old version keep their old stamp and the leaderboard will not compare across versions without saying so.
The rule we hold ourselves to: you should be able to take a transcript, apply our published normalization, compute CER yourself, and get the same number we did. If you cannot, that is a bug and we want to hear about it.
Cantonese Voice Benchmark · Published 27 Jul, 05:14 HKT.