Last Translation Benchmark

Current machine translation benchmarks are saturated, and evaluation metrics are either unreliable or unscalable. The Last Translation Benchmark addresses this by curating a dataset of inputs (text, image, audio, video) that demonstrably break state-of-the-art translation models.

Contribute by submitting an input that modern machine translation models get provably wrong. Provide a robust pass/fail verification rule to automatically evaluate shown and future translations. Contributors with 10 approved submissions are welcome to co-authorship on the live publication.

Invalid or expired magic link. Please request a new one from the administrators.
Paper on arXiv LTBv1 (3456 examples, 90MB) LTBv1 on 🤗 Model Leaderboard

We will be collecting submissions continuously and will release future dataset version and updated paper with new authors on a rolling basis until end of 2026. Reach out to last-translation-benchmark@vilda.net with inquiries or file bug reports on github.com/zouharvi/last-translation-benchmark. Please do not reach out about the status of your pending submissions. To speed up the review process, you can invite other speakers of your languages who can review your submissions or nominate yourself to be a reviewer.

NameAffiliationAccepted submissions
Loading…