Announcement · 27 Jul, 05:14 HKT
One to two pages on what was measured. The Full Score is withheld until it is in.
The rules have always required a system description. Until this week nothing collected one, which made it a sentence rather than a rule.
Entrants now write it in the portal, against the six questions a reader of the results actually needs answered: what the stack is, what was trained and on what, what runs at inference, where the humans are, what hardware it runs on, and where it is weak. It is submitted once and fixed after that, because a description that can be edited after the fact describes nothing.
The Core Score is unaffected. The Full Score is withheld until the description is in, and the leaderboard says which of the two reasons is holding it — awaiting human evaluation, or awaiting the description.
Nothing here is published unless an entrant is named and agrees to it.
Cantonese Voice Benchmark · Published 27 Jul, 05:14 HKT.