A new benchmark, the Active Causal Discovery Benchmark (ACDB), aims to evaluate how well LLM agents can uncover causal structure. The environment is grounded in structural causal models (SCMs), providing a controlled setting for testing causal reasoning.

ACDB focuses on two core tasks: recovering causal graph structure from observational data and performing hard interventions under a limited budget. This setup reflects real-world constraints where experiments are costly and agents must choose interventions wisely.

The benchmark is designed to assess whether LLM agents can move beyond passive pattern matching and engage in active, goal-directed causal exploration. By pairing observations with budget-constrained interventions, ACDB challenges agents to balance information gain against intervention costs.