arXiv · 2022-10-18
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
根拠記録を開く
研究課題
How can a retrieval benchmark cover languages with very different resource levels while using native queries and relevance judgments?
方法
An 18-language, same-language Wikipedia retrieval benchmark built with native-speaker queries and judgments; lexical, dense and hybrid baselines were compared on the released pools.
根拠が支持する範囲
Evaluate each language explicitly, document corpus and judgment construction, and keep native review central. The reported hybrid result is a historical benchmark observation.
限界と転用可能性
This is same-language information retrieval, not translation assessment or web SEO. Heuristic segmentation, candidate-pool gaps and unfinished test labels limit transferability; it says nothing about hreflang or Google rankings.
主張の分類
- 支持あり
- 歴史的
- 学術限定
NAVINESへの反映: 多言語ページ比較ラボ