A new preprint on arXiv addresses a growing trend: using multi-agent systems built on large language models to generate scientific hypotheses. The authors note that while such systems are increasingly common, two aspects are hard to interpret before deployment: the effect of refinement on the outputs and the diversity of the delivered hypothesis set.
The paper proposes an auditing approach centered on pairwise equivalence judgments—comparing hypotheses to determine whether they are meaningfully distinct. It also examines self-critique effects, where the system evaluates its own outputs, and suggests ways to measure diversity across the generated set.
Because the abstract is brief, the full methodology and findings are not yet available from the source. The work appears to be a step toward making multi-agent hypothesis generation more transparent and interpretable, rather than treating the system as a black box.