arXiv · 2022-10-18
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
打开证据记录
研究问题
How can a retrieval benchmark cover languages with very different resource levels while using native queries and relevance judgments?
方法
An 18-language, same-language Wikipedia retrieval benchmark built with native-speaker queries and judgments; lexical, dense and hybrid baselines were compared on the released pools.
证据支持范围
Evaluate each language explicitly, document corpus and judgment construction, and keep native review central. The reported hybrid result is a historical benchmark observation.
局限与可迁移性
This is same-language information retrieval, not translation assessment or web SEO. Heuristic segmentation, candidate-pool gaps and unfinished test labels limit transferability; it says nothing about hreflang or Google rankings.
主张分类
- 有支持
- 历史性
- 仅学术
对NAVINES的影响: 多语言页面比较实验室