arXiv · 2022-10-18
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
Evidenzdatensatz öffnen
Forschungsfrage
How can a retrieval benchmark cover languages with very different resource levels while using native queries and relevance judgments?
Methode
An 18-language, same-language Wikipedia retrieval benchmark built with native-speaker queries and judgments; lexical, dense and hybrid baselines were compared on the released pools.
Was die Evidenz stützt
Evaluate each language explicitly, document corpus and judgment construction, and keep native review central. The reported hybrid result is a historical benchmark observation.
Grenzen und Übertragbarkeit
This is same-language information retrieval, not translation assessment or web SEO. Heuristic segmentation, candidate-pool gaps and unfinished test labels limit transferability; it says nothing about hreflang or Google rankings.
Klassifikation
- Gestützt
- Historisch
- Nur akademisch
NAVINES-Einfluss: Labor für mehrsprachigen Seitenvergleich