Large language models are increasingly used for financial document analysis, including earnings call transcripts. A new paper on arXiv highlights a shift in user expectations: rather than standalone claims, users prefer grounded analyses that pair each claim with supporting citations.
The paper introduces a citation-grounded benchmark designed to assess how well LLMs produce trustworthy analyses of earnings call transcripts. The abstract indicates the benchmark is meant to support this move toward verifiable outputs, though it does not disclose dataset size or specific evaluation metrics.
As a single source, the paper offers no competing findings to compare. Its contribution is to frame a concrete evaluation target for a growing use case—making AI-generated financial analysis more transparent and accountable.