arXiv · 2022-10-18
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
근거 기록 열기
연구 질문
How can a retrieval benchmark cover languages with very different resource levels while using native queries and relevance judgments?
방법
An 18-language, same-language Wikipedia retrieval benchmark built with native-speaker queries and judgments; lexical, dense and hybrid baselines were compared on the released pools.
근거가 뒷받침하는 범위
Evaluate each language explicitly, document corpus and judgment construction, and keep native review central. The reported hybrid result is a historical benchmark observation.
한계와 전이 가능성
This is same-language information retrieval, not translation assessment or web SEO. Heuristic segmentation, candidate-pool gaps and unfinished test labels limit transferability; it says nothing about hreflang or Google rankings.
주장 분류
- 뒷받침됨
- 역사적
- 학술 한정
NAVINES 반영: 다국어 페이지 비교 연구실