A new arXiv paper introduces OpenProblemBench, a benchmark built from 82 unresolved problems in the foundational theoretical sciences. The authors argue that the next frontier for artificial general intelligence is tackling unresolved scientific problems, and that existing benchmarks mostly measure performance on established knowledge.
According to the abstract, OpenProblemBench is intended to assess progress beyond what is already known. The paper does not list the specific problems or fields in the excerpt, but describes them as spanning foundational theoretical sciences.
Because only the abstract is available in the source, the article cannot detail the benchmark's construction or evaluation methods. The key claim is that such a benchmark could help track whether AI systems are moving toward genuine scientific discovery rather than pattern matching on solved examples.