Researchers have introduced HakushoBench, a new benchmark for evaluating vision-language models (VLMs) on chart and table question answering in Japanese. The benchmark is constructed from governmental white papers, providing a real-world document understanding task that goes beyond typical synthetic or web-scraped datasets.

While English-language benchmarks for chart and table VQA have advanced quickly, comparable resources for other languages remain scarce. HakushoBench aims to address this imbalance by offering a Japanese-language evaluation set, allowing researchers to test how well VLMs handle structured data in a non-English context.

The paper notes that understanding chart and table images is essential for applying VLMs to practical document analysis. By grounding the benchmark in official government documents, the authors hope to make the evaluation more representative of actual use cases, though the abstract does not yet include specific results or comparisons with existing models.