Checking one citation is the whole job
A fabricated citation is easy to catch. The dangerous one is real, resolvable, correctly formatted, and does not say what the sentence citing it says. Nothing about its appearance will warn you.

Do not check whether a citation exists; check whether it supports the sentence attached to it. The failure that survives casual review is a real, resolvable source paired with a claim it does not make. A human audit of four commercial generative search engines found only 74.5% of citations supported the sentence they were attached to.{{2}} Retrieval-augmented tools reduce the failure without eliminating it: the first preregistered evaluation of commercial legal research tools marketed as eliminating hallucination found that they still hallucinated between 17% and 33% of the time.{{1}} The routine that catches it is four checks on one claim — does the source exist, does it say this, is it current, and is it the right kind of source — and it takes about two minutes. Run it on the load-bearing claim in any report you intend to act on, not on all of them, and never on the summary sentence.
The failure that looks like success
There are two ways an AI citation goes wrong, and they need completely different defences. The first is invention: a paper, case, or page that does not exist. This is the famous one — it produced the sanctioned court filings that made the news — and it is genuinely easy to catch: the link 404s, the DOI resolves to nothing, the case number returns no result.
The second is misattribution: a real source, correctly cited, attached to a claim it does not support. The link works. The publisher is reputable. The date is plausible. The formatting is impeccable. And the paper says something adjacent to, weaker than, or occasionally the opposite of the sentence in front of you. A human audit of four commercial generative search engines put a number on how common this is: only 74.5% of citations actually supported the sentence they were attached to, and just 51.5% of generated sentences were fully supported by their citations.[2]
Only the first is fixed by grounding the model in retrieved documents. The second is what remains, and it remains even in systems built specifically to cite well: the ALCE benchmark, which scores end-to-end retrieve-and-cite systems on citation quality, found that on one open-ended dataset even the best models lacked complete citation support half the time.[3] An evaluation of commercial legal research tools — products explicitly marketed on the elimination of hallucination — found that retrieval reduced hallucination relative to a general-purpose chatbot while the tools still hallucinated between 17% and 33% of the time, and that the authors had to define a typology separating answers that were wrong from answers that were merely not supported by the authority cited.[1] That second category is the whole problem: unsupported is invisible unless someone opens the source.
- Invented source: caught by a click. Rare in grounded systems.
- Real source, wrong claim: survives every check except reading it.
- Real source, stale: was true; the figure has since changed.
- Real source, wrong authority: a blog post cited for a regulatory requirement.
- Overreach: the source shows a correlation; the sentence asserts a cause.
Four checks, about two minutes
Pick the claim your decision actually rests on. Not the opening summary, not a background sentence — the one that, if wrong, changes what you do. Then:
- Search the source for the specific figure rather than reading around it. If the number is not in the document, stop there.
- Watch for the citation attached to the wrong half of a sentence — the source supports the first clause, the second is the model's own.
- Treat a claim citing three sources with more suspicion, not less. It often means no single source said it.
- If a claim carries no citation in a report that cites everything else, that absence is the finding.
| Check | The question | What failure looks like |
|---|---|---|
| Exists | Does the link resolve to the cited document? | 404, a paywall stub, or a homepage rather than the page |
| Supports | Does that document state this specific claim? | The number differs, is a projection, or appears only as something the source is arguing against |
| Current | Is this the latest version of the fact? | Correct on the cited date, superseded since — common for pricing, limits, and regulation |
| Authoritative | Is this the right kind of source for this claim? | A vendor blog cited for a competitor's benchmark; a summary cited instead of the study it summarizes |
Which claims deserve the two minutes
Verifying everything is not a workable discipline, and pretending otherwise is how teams end up verifying nothing. A twelve-source brief might carry forty factual claims; four of them matter.
Check the claims that are load-bearing, surprising, or numeric. Load-bearing means the recommendation changes if it is wrong. Surprising means it contradicts what you expected — which is either the most valuable sentence in the report or the incorrect one, and you cannot tell which without looking. Numeric means a specific figure, because figures are what get quoted onward into a slide, a memo, or a decision, stripped of the citation they arrived with.
Skip the background paragraphs. Skip the definitional sentences. Skip anything you already knew. The purpose of the exercise is not to audit the model; it is to make sure the thing you are about to act on is true.
Making it survive contact with a deadline
A verification habit that depends on discipline will not survive a busy week. The ones that last are structural: they change what the tool produces, not what the reader promises to do.
Prefer tools that attach citations at claim level rather than listing sources at the end — a bibliography cannot be checked, because you cannot tell which sentence any entry was meant to support. Prefer reports that keep the retrieval date visible, so staleness is legible without a second search. Prefer a shorter run with twelve sources you can actually open over a fifteen-hundred-source report whose citations you will never test, because a report nobody verifies is a report nobody should act on.
And write down what you checked. A brief where four claims are marked verified, with who checked them and when, is a different artifact from the same brief unmarked. It is the difference between research that survives review and research that merely reads well.
- Claim-level citations, not an end-of-document source list.
- Visible retrieval dates on every source.
- A named person against each verified claim.
- A report length you will realistically audit, not the largest one available.
Frequently asked questions
How do I check if an AI citation is real?
Open it. A fabricated source fails immediately: the link does not resolve, the DOI returns nothing, the case or paper cannot be found. This check takes seconds and catches outright invention — but it is the easy half, and passing it is not evidence the citation supports the claim.
Do citations mean an AI answer is accurate?
No. Citations mean sources were retrieved and attached. A human audit of four commercial generative search engines found that only 74.5% of citations supported the sentence they were attached to, and that barely half of generated sentences were fully supported by their citations. The failure that survives review is a real, resolvable source paired with a claim it does not make. A preregistered evaluation of retrieval-grounded commercial legal research tools found them hallucinating between 17% and 33% of the time despite marketing that promised elimination.
Which claims should I verify?
The load-bearing ones — where being wrong changes your decision — plus anything surprising and anything numeric. Skip background and definitions. Four checked claims in a report you act on is far better practice than forty unchecked ones.
Is a report with more sources more reliable?
Not necessarily, and it can be worse in practice. A fifteen-hundred-source report is one nobody will audit; twelve sources is a number a person can actually open. Reliability comes from claims being checkable and checked, not from the size of the bibliography.
What is the fastest way to catch a misattributed citation?
Search the cited document for the specific figure or phrase in the claim, rather than skim-reading around it. If the number is not in the source, the citation does not support the sentence — regardless of how relevant the source otherwise looks.
Sources and evidence
Primary and authoritative sources used for factual claims. Company research and executive forecasts are labeled as such in the article.
- 1Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh et al., Stanford RegLab / HAI · 2024-05
- 2Evaluating Verifiability in Generative Search EnginesLiu, Zhang and Liang, Stanford · 2023-04
- 3Enabling Large Language Models to Generate Text with Citations (ALCE)Gao et al., Princeton NLP · 2023-05
- 4Kendr deep research: claim-level citations and dated sourcesKendr · 2026-09