Yorùbá is a widely spoken tonal language in which diacritics are essential for distinguishing meaning. In practice, however, much written Yorùbá omits these marks, leaving words ambiguous and complicating downstream natural language processing tasks such as translation and text analysis.

A new preprint on arXiv introduces Yo-ByT5, a model designed specifically for diacritic restoration in Yorùbá. The abstract describes the approach as both efficient and high-fidelity, suggesting it aims to restore missing marks accurately without excessive computational cost.

If the method holds up under scrutiny, it could provide a practical tool for cleaning up real-world Yorùbá text and making the language more accessible to NLP systems. As a preprint, the work has not yet undergone peer review, so further evaluation will be needed to confirm its claims.