METRIC · REFERENCE · PUBLIC AUDIO
One sound,
five forms accepted
A benchmark can keep several correct transcripts for the same sound. In one 0.9-second Hindi clip, the same output scores 33.3% WER against the first written reference and 0.0% WER against the full accepted set.
WHY THIS IS USEFUL If you use WER to say what a system heard, the choice of reference is part of the answer.
VISIBLE EVIDENCE · VoiceArena / Monsoon hi-IN · 753 clips in the public split
The source row contains public voice audio. We do not play it here without a separate rights and re-identification-risk check.
view the source datasetOne output, two judgements
OUTPUT improve होता है
Reference selected: the first form in the list. The score includes one substitution.
SHORT ANSWER the reference is part of the measurement, not decoration.
What does the score measure?
Not only the difference between voice and text. WER also measures the distance between the output and the text chosen as its reference.
Only the reference state changes: from the first string to the source's accepted forms.
In our example, `improve` is an accepted form for the first word. If the sole reference is `इम्प्रूव होता है`, the different word becomes a substitution. If the evaluator can choose `improve / होता / है` from the lattice, the output matches without an error.
The lattice keeps the options
Each slot has a small list of accepted forms. The line does not say which form is “better”; it says what the benchmark allowed to be compared as the same speech.
- 01slot 1इम्प्रूवइंप्रूवimproveimprovimproov5 accepted forms
- 02slot 2होताhota2 accepted forms
- 03slot 3हैhai2 accepted forms
Two judgements, one output
The interaction changes only access to references. The operation stays a transparent, repeatable word-level edit distance.
(I + D + S) / reference words × 100
I = insertions · D = deletions · S = substitutions
| State | Reference | Errors | Result |
|---|---|---|---|
| One reference | इम्प्रूव होता है | 1 S / 3 | 33.3% |
| Lattice | improve · होता · है | 0 / 3 | 0.0% |
The demonstration uses exact tokens so it can be checked by eye. The cited OIWER scorer adds Indic normalization and dynamic alignment across variants; this page does not reproduce it as a complete implementation.
How dense is the object?
The public hi-IN split has 753 clips. To keep the object inspectable, we show the shape of a 12-row sample, not a new statistic about every speaker.
- 009Row 0: 9 variants across 3 slots, 0.9 sec
- 0118Row 1: 18 variants across 5 slots, 2.0 sec
- 0242Row 2: 42 variants across 21 slots, 9.6 sec
- 034Row 3: 4 variants across 1 slots, 0.5 sec
- 0426Row 4: 26 variants across 10 slots, 2.8 sec
- 0540Row 5: 40 variants across 16 slots, 9.0 sec
- 06104Row 6: 104 variants across 36 slots, 12.8 sec
- 0781Row 7: 81 variants across 30 slots, 14.8 sec
- 0837Row 8: 37 variants across 23 slots, 11.6 sec
- 0955Row 9: 55 variants across 24 slots, 7.0 sec
- 102Row 10: 2 variants across 1 slots, 0.9 sec
- 1113Row 11: 13 variants across 7 slots, 1.9 sec
POPULATION LIMIT No person, location, or device field is used here. The audio is included only under the dataset's stated licence.
A ranking can inherit the reference
The release article describes a comparison where the same hypotheses are scored against flattened references or a lattice. It reports that error rates move unevenly and that some model pairs can reverse order.
An error is not a verdict
A lower score in this demonstration does not prove that a model understands the speaker better.
- 01
The accepted forms come from the benchmark's review process; this page does not independently certify Hindi orthography.
- 02
The example shows one row and the shape of 12 rows. It is not a full reanalysis of all 753 clips.
- 03
We publish no contributor metadata and make no comparison by city, region, income, device, or gender.
- 04
We do not evaluate intelligence, fairness, understanding, or overall quality. We measure the effect of reference choice on a score.
Where the demonstration comes from
The row, method, and limits can be followed back to the original materials. Sources were archived locally and checked before publication.
30 AUG
2026
Credit: VoiceArena, MonsoonASR-Open-ASR-leaderboard-hi-IN — CC BY 4.0. Change made: we reproduce redacted lattice text and the demo score; we do not redistribute source audio or the original row's contributor fields.