Evaluating retrieval-augmented generation (RAG) systems is difficult for low-resource languages, where conventional reference-based metrics often fail to capture quality. A new preprint on arXiv tackles this problem for Romanian, proposing an LLM-as-a-Judge approach adapted from existing frameworks like Ragas and comparative ranking.
The authors argue that standard metrics rely on reference answers that are scarce or unreliable in low-resource settings. By adapting Ragas—a suite of RAG evaluation metrics—and comparative ranking methods to Romanian, they aim to provide a more robust way to assess system outputs without needing extensive gold references.
The paper specifically investigates whether LLM-as-a-Judge can serve as a viable evaluation strategy for Romanian RAG systems. While the abstract does not report final results, the work highlights a growing need for language-specific evaluation tools as RAG deployment expands beyond high-resource languages. The study is limited to Romanian, but its methodology could inform similar efforts for other low-resource languages.