Extraction benchmarks often suffer from bias and opacity, making it hard to trust or compare model performance. Datalab's new OmniExtractBench directly addresses these issues by introducing a more transparent evaluation framework.
The benchmark uses content-based row matching, which compares extracted values based on their actual content rather than superficial row positions or formatting. Each value is then judged against one of six possible verdicts, with a dedicated null rule for cases where no extraction should occur. This granular approach gives a clearer picture of where and why models fail.
Because the evaluation logic is explicit and rule-based, OmniExtractBench is designed to be auditable by anyone. The result is a benchmark that aims to reduce hidden biases and make extraction quality easier to inspect and verify.