# One sound, five accepted forms

In one public 0.9-second Hindi clip, the same output — `improve होता है` — gets 33.3% WER when the first written string, `इम्प्रूव होता है`, is the only reference. With the accepted lattice forms, the same output gets 0.0% WER.

## What changes

WER measures the distance between an output and its reference. In the simple case: one substitution out of three words = 33.3%. In the lattice case: `improve` is accepted for the first word and the other two match = 0.0%.

## What the source contains

The public Monsoon hi-IN split has 753 clips, 1.33 declared hours, and 468 declared speakers. This project shows one row and the shape of a local 12-row sample; it does not reanalyse the full leaderboard and does not display contributor metadata.

## Limits

The release article reports that flattening variants can change error rates unevenly and reverse model pairings. We do not recompute that ranking: the public release does not contain every model's hypotheses. The demonstration does not measure intelligence, fairness, understanding, regional performance, or overall model quality.

## Sources

- [Monsoon hi-IN dataset card](https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-hi-IN) — lattice, public split, and CC BY 4.0 licence.
- [Release article](https://huggingface.co/blog/open-asr-leaderboard-global-south) — the score and ranking implication, presented here as source-reported.
- [OIWER](https://arxiv.org/abs/2603.00941) — alignment with reference variants.
- [`voi-oiwer` on PyPI](https://pypi.org/project/voi-oiwer/) — documented public scorer.

Checked 30 August 2026. Demo audio used under the CC BY 4.0 licence stated by the dataset.
