ARXIV:2604.03159 · LLM AGENTS · SUBMITTED 06 APR · 20:12 UTC · FRESHNESS UNKNOWN

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation

Delip Rao · Chris Callison-Burch · arXiv

A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.

Ship in 2-4 weeks›Score7.0Evidence unverified

Opportunity summary

Pain A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.

Evidence 0 refs | 0 sources | 0% coverage

Blocker Evidence unverified

Open Build Read PDF Signal Canvas Track

PROBLEM

A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records. Prior evaluations tested base models without search, which does not reflect current practice.

METHOD

Full abstract

Large language models with web search are increasingly used in scientific publishing agents, yet they still produce BibTeX entries with pervasive field-level errors. Prior evaluations tested base models without search, which does not reflect current practice. We construct a benchmark of 931 papers across four scientific domains and three citation tiers -- popular, low-citation, and recent post-cutoff -- designed to disentangle parametric memory from search dependence, with version-aware ground truth accounting for multiple citable versions of the same paper. Three search-enabled frontier models (GPT-5, Claude Sonnet-4.6, Gemini-3 Flash) generate BibTeX entries scored on nine fields and a six-way error taxonomy, producing ~23,000 field-level observations. Overall accuracy is 83.6%, but only 50.9% of entries are fully correct; accuracy drops 27.7pp from popular to recent papers, revealing heavy reliance on parametric memory even when search is available. Field-error co-occurrence analysis identifies two failure modes: wholesale entry substitution (identity fields fail together) and isolated field error. We evaluate clibib, an open-source tool for deterministic BibTeX retrieval from the Zotero Translation Server with CrossRef fallback, as a mitigation mechanism. In a two-stage integration where baseline entries are revised against authoritative records, accuracy rises +8.0pp to 91.5%, fully correct entries rise from 50.9% to 78.3%, and regression rate is only 0.8%. An ablation comparing single-stage and two-stage integration shows that separating search from revision yields larger gains and lower regression (0.8% vs. 4.8%), demonstrating that integration architecture matters independently of model capability. We release the benchmark, error taxonomy, and clibib tool to support evaluation and mitigation of citation hallucinations in LLM-based scientific writing.

RESULT

ScienceToStartup currently rates this 7.0/10 on the public viability pass. An ablation comparing single-stage and two-stage integration shows that separating search from revision yields larger gains and lower regression (0.8% vs. Code availability is…

WHY NOW

LLM Agents moved forward this cycle; last verified April 2026. Public score 7.0/10. Production flags indicate code availability.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score7.0

PainA tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.

Evidence0 refs | 0 sources | 0% coverage

Blockerno shell-level blocker reported

Analysis summary

A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

Competitive landscape

A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.

Segment

LLM Agents

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "20434c6b-2656-46e0-b571-98bfe5607c28", "arxiv_id": "2604.03159", "canonical_route": "/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation", "endpoints": { "paper_pack": "/api/v1/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation/paper-pack", "build_passport": "/api/v1/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation", "normalized_query": "2604.03159", "route": "/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation", "paper_ref": "bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation#webpage", "url": "https://sciencetostartup.com/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation", "name": "BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation", "description": "A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation#scholarlyArticle", "headline": "BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation", "description": "A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.", "url": "https://sciencetostartup.com/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation", "sameAs": "https://arxiv.org/abs/2604.03159", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2604.03159" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-04-03T16:30:58.000Z", "author": [ { "@type": "Person", "name": "Delip Rao" }, { "@type": "Person", "name": "Chris Callison-Burch" } ], "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 7 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "LLM Agents" }, { "@type": "PropertyValue", "propertyID": "commercialReadiness", "value": "code" } ] }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "LLM Agents", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "BibTeX Citation Hallucinations in Scientific Publishing Agen", "item": "https://sciencetostartup.com/paper/bibtex-citation-hallucinations-in-scientific-publishing-agents-evaluation-and-mitigation" } ] } ] }

Competitive landscape

A tool to significantly improve the accuracy of BibTeX entries generated by LLM-powered scientific publishing agents by integrating authoritative citation records.

Segment

LLM Agents

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation

BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline