A new arXiv preprint highlights a key limitation in visual agentic search: standard single-step retrievers. According to the abstract, these retrievers create a severe bottleneck for LLM agents that rely on retrieval tools to access external knowledge.
The paper notes that current pipelines require the agent to issue text queries for every step. This is inefficient for visual tasks, where the relevant information may not be easily expressed as text at each stage.
To address this, the authors propose learning to route in visual space via multi-step embedding retrieval. This approach aims to let the agent retrieve embeddings directly, rather than repeatedly falling back on text queries. The preprint is the only source used here; no other papers or results were considered.