METRIC · REFERENCE · PUBLIC AUDIO

One sound,
five forms accepted

A benchmark can keep several correct transcripts for the same sound. In one 0.9-second Hindi clip, the same output scores 33.3% WER against the first written reference and 0.0% WER against the full accepted set.

WHY THIS IS USEFUL If you use WER to say what a system heard, the choice of reference is part of the answer.

VISIBLE EVIDENCE · VoiceArena / Monsoon hi-IN · 753 clips in the public split

PUBLIC CLIP · MONSOON hi-IN0.9 sec
CLIP NOT RE-HOSTED

The source row contains public voice audio. We do not play it here without a separate rights and re-identification-risk check.

view the source dataset

One output, two judgements

OUTPUT improve होता है

one reference 33.3% 1 substitution / 3 words
accepted lattice 0.0% 0 errors / 3 words
Change the judge

Reference selected: the first form in the list. The score includes one substitution.

SHORT ANSWER the reference is part of the measurement, not decoration.

01 / THE QUESTION

What does the score measure?

Not only the difference between voice and text. WER also measures the distance between the output and the text chosen as its reference.

same output
33.3%0.0%

Only the reference state changes: from the first string to the source's accepted forms.

In our example, `improve` is an accepted form for the first word. If the sole reference is `इम्प्रूव होता है`, the different word becomes a substitution. If the evaluator can choose `improve / होता / है` from the lattice, the output matches without an error.

02 / THE OBJECT

The lattice keeps the options

Each slot has a small list of accepted forms. The line does not say which form is “better”; it says what the benchmark allowed to be compared as the same speech.

output path accepted alternative
03 / THE CALCULATION

Two judgements, one output

The interaction changes only access to references. The operation stays a transparent, repeatable word-level edit distance.

FORMULA (I + D + S) / reference words × 100

I = insertions · D = deletions · S = substitutions

Exact demonstration score
StateReferenceErrorsResult
One referenceइम्प्रूव होता है1 S / 333.3%
Latticeimprove · होता · है0 / 30.0%

The demonstration uses exact tokens so it can be checked by eye. The cited OIWER scorer adds Indic normalization and dynamic alignment across variants; this page does not reproduce it as a complete implementation.

04 / CONTEXT

How dense is the object?

The public hi-IN split has 753 clips. To keep the object inspectable, we show the shape of a 12-row sample, not a new statistic about every speaker.

753public clips
1.33declared hours
468declared speakers
  1. 00
    9Row 0: 9 variants across 3 slots, 0.9 sec
  2. 01
    18Row 1: 18 variants across 5 slots, 2.0 sec
  3. 02
    42Row 2: 42 variants across 21 slots, 9.6 sec
  4. 03
    4Row 3: 4 variants across 1 slots, 0.5 sec
  5. 04
    26Row 4: 26 variants across 10 slots, 2.8 sec
  6. 05
    40Row 5: 40 variants across 16 slots, 9.0 sec
  7. 06
    104Row 6: 104 variants across 36 slots, 12.8 sec
  8. 07
    81Row 7: 81 variants across 30 slots, 14.8 sec
  9. 08
    37Row 8: 37 variants across 23 slots, 11.6 sec
  10. 09
    55Row 9: 55 variants across 24 slots, 7.0 sec
  11. 10
    2Row 10: 2 variants across 1 slots, 0.9 sec
  12. 11
    13Row 11: 13 variants across 7 slots, 1.9 sec
Each bar is one local sample row; the number on the right is that row's total accepted variants.

POPULATION LIMIT No person, location, or device field is used here. The audio is included only under the dataset's stated licence.

05 / WHAT THE SOURCE SAYS

A ranking can inherit the reference

The release article describes a comparison where the same hypotheses are scored against flattened references or a lattice. It reports that error rates move unevenly and that some model pairs can reverse order.

06 / WHAT IT DOES NOT SAY

An error is not a verdict

A lower score in this demonstration does not prove that a model understands the speaker better.

  • 01

    The accepted forms come from the benchmark's review process; this page does not independently certify Hindi orthography.

  • 02

    The example shows one row and the shape of 12 rows. It is not a full reanalysis of all 753 clips.

  • 03

    We publish no contributor metadata and make no comparison by city, region, income, device, or gender.

  • 04

    We do not evaluate intelligence, fairness, understanding, or overall quality. We measure the effect of reference choice on a score.

07 / FOLLOW THE EVIDENCE

Where the demonstration comes from

The row, method, and limits can be followed back to the original materials. Sources were archived locally and checked before publication.

CHECKED
30 AUG
2026

Credit: VoiceArena, MonsoonASR-Open-ASR-leaderboard-hi-IN — CC BY 4.0. Change made: we reproduce redacted lattice text and the demo score; we do not redistribute source audio or the original row's contributor fields.