Language models that can recognise their own output might favour it when acting as judges, and models monitoring each other could collude. A new arXiv paper tests this ability zero-shot in current commercial models. The authors had five LLMs generate solutions and then attempt to attribute code to themselves.

According to the abstract, the results indicate that surface cues—style—explain the attribution, rather than any deeper self-knowledge. This suggests that apparent self-recognition may be an artefact of stylistic consistency, not a reliable mechanism for identity tracking.

The finding matters for AI evaluation and oversight. If models rely on style rather than true authorship, then systems designed to have models judge or monitor each other could be fooled or could inadvertently collude. The paper's full results are not yet available in the abstract, but the stated conclusion points to a need for caution in deploying LLMs as self-monitors.