Researchers have released a new arXiv preprint introducing ROTE (RollOut Testing of Exact memorization), a benchmarking protocol for evaluating symbolic memorization in neural sequence models. The protocol is built around complexity-controlled symbolic sequences, allowing tests to probe when a model has exactly memorized a sequence rather than merely approximated it.
ROTE is intended to study two related capabilities: memorization of symbolic rules and extension of those rules to new cases. By controlling sequence complexity, the benchmark aims to separate rote recall from genuine generalization. The abstract does not report specific results, so the main contribution at this stage is the protocol itself.