Levantine Arabic (LA) is spoken by tens of millions of people, yet speech-language technologies for it lack shared evaluation standards. A new paper introduces SHAMS, an audio-grounded pronunciation benchmark designed to fill this gap.

The benchmark specifically targets the difficulty of evaluating LA technologies given the dialect's internal diversity. By grounding evaluation in audio, SHAMS aims to provide a more reliable measure of pronunciation performance across different variants.

As a shared resource, SHAMS could help researchers and developers compare systems consistently. The paper highlights the pressing need for such benchmarks, though it does not yet detail the benchmark's construction or results.