Question difficulty is one of the most fundamental properties in AI evaluation: it determines whether a question can meaningfully discriminate between models of differing ability. Without a clear grasp of what makes a question hard, benchmark results are hard to interpret. According to a new arXiv paper, a variety of methods can now estimate or predict difficulty, but the authors argue for something more—explaining difficulty in natural language.

The paper's abstract frames this as a move beyond prediction. Rather than simply outputting a difficulty score, the proposed approach appears to generate explanations of why a question is hard, in plain language. This could help researchers understand the specific reasoning gaps or ambiguities that trip up models, and it may inform the design of more targeted evaluation sets.

The source is limited to the abstract, which is truncated mid-sentence, so the full methodology, evaluation, and results are not yet visible. Still, the stated goal is clear: to make difficulty interpretable, not just measurable. If successful, such explanations could become a useful tool for diagnosing model behavior and improving benchmark design.