1.Delip Rao and Chris Callison-Burch. 2026. BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation. https://doi.org/10.48550/arxiv.2604.03159
Open Access: doi.org · DOI
Large language models with web search are increasingly used in scientific publishing agents, yet they still produce BibTeX entries with pervasive field-level errors. Prior evaluations tested base models without search, which does not reflect current practice. We construct a benchmark of 931 papers across four scientific domains and three citation tiers – popular, low-citation, and recent post-cutoff – designed to disentangle parametric memory from search dependence, with version-aware ground…
1.Diletta Abbonato. 2026. CheckIfExist: Detecting Citation Hallucinations in the Era of AI-Generated Content. https://doi.org/10.48550/arxiv.2602.15871
The proliferation of large language models (LLMs) in academic workflows has introduced unprecedented challenges to bibliographic integrity, particularly through reference hallucination – the generation of plausible but non-existent citations. Recent investigations have documented the presence of AI-hallucinated citations even in papers accepted at premier machine learning conferences such as NeurIPS and ICLR, underscoring the urgency of automated verification mechanisms. This paper presents…
References
1.Delip Rao and Chris Callison-Burch. 2026. BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation. https://doi.org/10.48550/arxiv.2604.03159
1.Diletta Abbonato. 2026. CheckIfExist: Detecting Citation Hallucinations in the Era of AI-Generated Content. https://doi.org/10.48550/arxiv.2602.15871