LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents
LëtzCross is a novel cross-lingual page-level benchmark for multimodal retrieval over Luxembourgish PDF documents, addressing a gap in low-resource, multilingual RAG evaluation The benchmark indexes document pages as images and supports queries in four languages: English, French, German, and Luxembourgish ColPali-style page-image retrievers outperform OCR-based text-only retrievers across all query languages in system-level comparisons Fine-tuning transfers across query languages, with French si
Analysis
TL;DR
- LëtzCross is a novel cross-lingual page-level benchmark for multimodal retrieval over Luxembourgish PDF documents, addressing a gap in low-resource, multilingual RAG evaluation
- The benchmark indexes document pages as images and supports queries in four languages: English, French, German, and Luxembourgish
- ColPali-style page-image retrievers outperform OCR-based text-only retrievers across all query languages in system-level comparisons
- Fine-tuning transfers across query languages, with French single-language fine-tuning yielding the highest mean performance on Luxembourgish queries
- Multilingual fine-tuning that includes Luxembourgish produces the strongest overall results and substantially improves retrieval specifically for Luxembourgish queries
Why It Matters
This work addresses a critical gap in multimodal retrieval research by evaluating page-image retrievers in cross-lingual, low-resource settings — an area with little existing empirical knowledge. For AI practitioners building PDF-based RAG systems targeting multilingual or low-resource document collections, these findings provide actionable guidance on retriever selection and fine-tuning strategies. The benchmark also contributes to the broader effort of making multimodal retrieval accessible beyond high-resource languages like English.
Technical Details
- Benchmark Design: LëtzCross indexes Luxembourgish PDF document pages as images and provides queries in four languages (English, French, German, Luxembourgish), combining text-focused QA pairs with visually grounded QA pairs to cover both textual and visual retrieval needs in PDF-based RAG
- Retriever Comparison: System-level evaluation comparing OCR-based text-only retrievers against ColPali-style page-image retrievers, with the latter demonstrating superior performance across all query languages
- Fine-Tuning Analysis: Examines both single-language and multilingual fine-tuning approaches, revealing that fine-tuning transfers across query languages and that including Luxembourgish in multilingual training yields the strongest results
- Key Finding: Among single-language fine-tuning settings, French fine-tuning achieves the highest mean performance on Luxembourgish queries, while multilingual fine-tuning with Luxembourgish inclusion substantially improves retrieval for Luxembourgish queries specifically
Industry Insight
- Organizations deploying PDF-based RAG systems for low-resource or multilingual document collections should prioritize page-image retrievers over OCR-only approaches, as they demonstrate consistently better cross-lingual retrieval performance
- When fine-tuning retrieval models for Luxembourgish or similar low-resource languages, including the target language in multilingual training data yields the strongest results — single-language fine-tuning on a related language like French can serve as a viable fallback
- The LëtzCross benchmark highlights an underexplored area in multimodal retrieval evaluation, suggesting that researchers and practitioners should invest in cross-lingual, low-resource benchmarks to better understand and improve retrieval systems for diverse document collections
Disclaimer: The above content is generated by AI and is for reference only.