An arXiv preprint introduces EnigmaForge, a benchmark that breaks with the usual question-answer format. Instead of presenting a model with a direct query, it supplies a stack of old documents and no question at all. The task is to find and solve a small logic puzzle buried somewhere in the material.
The puzzle is hidden across letters, receipts, and logbook margins, and the benchmark claims its solution is unique, with a proof of uniqueness. The abstract ends mid-sentence, so the details of that proof are not included in the available text.
Because there is only one source, there are no conflicting accounts to compare. The main limitation is that the abstract is truncated, leaving the evaluation methodology and proof strategy unspecified.