An arXiv preprint (2610.04206) proposes fine-tuning vision-language models to improve their spatial intelligence, specifically their understanding of rotations in two and three dimensions. The abstract frames spatial intelligence as a foundational skill across fields such as STEM, medicine, architecture, and construction, and points to recent studies on VLMs in this area.

Because the available abstract is truncated, the paper's methods and results are not fully described here. The title and opening indicate that fine-tuning is the proposed enhancement strategy, but no quantitative claims can be drawn from the source. With only one source, there are no conflicting findings to compare.