How the score is calculated. Each system translates whole documents, one paragraph per request
(never a page or a full document in one call), and the paragraph outputs are re-joined in order.
- Two directions, scored separately: English → X (the English edition is translated into the language and compared with the language edition) and X → English (the language edition is translated into English and compared with the English edition). Use the Direction selector (default: English → X). The two are never averaged together.
- chrF2 (sacreBLEU) is computed with the whole document as a single segment against the human English edition of the same document. Higher is better, and chrF is the single metric this leaderboard ranks by.
- Results are recorded per dataset (see the Datasets tab). Within a direction, a language's score is the mean of its per-dataset scores, so a new dataset never changes the numbers of the existing ones.
- Length is output words / reference words. Values far from 1 usually mean untranslated passes-through or padding, so read chrF together with it.
- A system is only scored on the languages and directions it supports; unsupported ones are listed, not scored as zero. chrF2 is chosen over BLEU because it scores a low-resource target language less harshly: it works on characters, so rich morphology and spelling variation cost far less than they do for BLEU.
Global Leaderboard
Per-Language
All Results
Datasets
Best model per language, ranked by chrF2. The dataset columns show the chrF2 of that model on each dataset.
✓ Scored
chrF by document
✗ Not scored
| Model | Direction | Reason |
|---|
0 rows
| Language | Model | Track | Dataset | chrF | Length | Docs |
|---|
Every score on this leaderboard is tied to a dataset. More datasets will be added over time; each one gets its own column and its own numbers.