According to a new paper on arXiv, coding agents can produce changes that pass functional tests but still miss the contribution requirements of a given repository. The abstract argues that following repository-specific norms is a separate challenge from passing tests.
These norms are not always in one place. The paper notes that guidance is dispersed across repository sources, which makes it difficult for an agent to know what is expected. The research appears to focus on how to acquire and verify these norms, though the abstract does not detail the proposed method.
At a practical level, the work points to a gap in current evaluation of coding agents: passing tests is not enough if the agent does not respect the conventions and processes of the project it is contributing to.