Small vision-language models (VLMs) are capable of reading external evidence, but according to a new arXiv paper, they often have difficulty obtaining that evidence in the first place. This gap between reading and retrieval motivates the paper's proposed solution.

The authors introduce Harness Compilation (HC), an offline procedure that adjusts the division of work between a frozen small VLM and its external components. Because the procedure runs offline, it does not require retraining the model itself; instead, it reallocates decisions about which tasks the model should keep and which should be delegated.

The abstract does not report experimental results or specific benchmarks, so the effectiveness of HC remains to be evaluated. The contribution at this stage is the framing: rather than making the small model larger or training it on new data, Harness Compilation treats the model as fixed and optimizes the surrounding harness.