How to — check an AI-generated citation before you rely on it
An AI-drafted citation can point at a real document and still misdescribe it — that is the failure mode a link check misses. Verification means opening the source and reading the sentence your citation is attached to.
This is a job you can finish in an afternoon for a normal reference list, and it is the difference between using AI as a research assistant and signing your name to something you never read.

1. Separate the two checks — one is mechanical, one is not.
There are two different questions hiding inside "is this citation good." Question one: does this document exist, with this title, these authors, at this venue, on this date? Question two: does the document actually say what my sentence claims it says? The first is a list operation you can run over every reference at once. The second is a reading job you do yourself, and it is where the damage lives.
A New Mexico criminal defense lawyer learned the difference the expensive way. In an appeal brief he filed in August 2025, he quoted testimony from witnesses who never took the stand — including fabricated police officers — and misdescribed real cases he had cited. The state Supreme Court held him in contempt, fined him $5,000, struck his briefs and referred him to a disciplinary board. His own account: he fed a transcript of the trial into ChatGPT and assumed it produced a bulletproof summary. He did not have fake case law. He had real cases described wrongly, plus quotes that did not exist.
2. Check the document exists, mechanically, before you check anything else.
Work the reference list top to bottom and confirm each entry resolves to a real record: the title, the author list, the year, and the venue all appear together on a publisher or repository page you did not get from the model. Do this on the record itself, not on the first search hit — a fabricated paper usually fails at the byline rather than the title, which is exactly why a plausible-looking title in a search result tells you nothing. Document identifiers are the fastest signal here: a working identifier points at one specific document, and an entry that has none while claiming to be a peer-reviewed paper deserves a second look.
Where an entry claims a specific numeric result, hold it. Numbers are the highest-yield thing to verify because they are the easiest thing to regenerate plausibly: a summarizer asked to describe a table will happily produce a table.
3. Read the cited source at the claim, not the abstract.
Now the unglamorous part. For each citation, open the source and find the passage your sentence depends on. Not the abstract, not the introduction's framing, not a review of the paper — the passage. Then answer three questions in writing: does this source state this? does it state it in this direction? and does it state anything that contradicts me?
Real sources get misdescribed constantly, and the misdescription is usually not about the main finding — it is about framing. A paper that reports an association becomes "the paper found a cause." A preprint becomes a peer-reviewed result. A result measured on a narrow slice of data becomes a general claim about the field. Those are the errors that survive every automated check, and they are the ones a reader will catch later with your name on them.
4. Check the two things a model is worst at: quotes and specificity.
If your citation carries quotation marks, find those words in the source. Verbatim quotation is the single easiest thing to falsify and the single easiest thing to confirm — there is no interpretation to argue about. Same for the specificity words that make a sentence feel researched: "first," "only," "largest," "state of the art," "according to the 2024 study." Each of those is a checkable claim, and most of them get quietly hedged once you look.
How long this should take is a function of what the citation is carrying, not how many there are. A background reference in a list of thirty is a fifteen-second existence check. The two citations holding up your central argument are worth the hour — you would spend the same hour if a junior colleague handed you the draft.
5. Use a tool for the roster, never for the verdict.
Services now push a reference list against records and flag entries that do not resolve, and some check whether a source is still standing — a retracted paper can cleanly pass every existence check and still be poison in your argument, which is why retraction databases belong in your loop. The research is moving toward checking citations at the passage level: one 2026 study of page-level citation verification reports roughly 93% accuracy at flagging a bad citation against the page it points to, and similar work is training small open models to recover which paper a claim actually came from.
Take the roster flag as a prompt to read, not as a verdict. A tool can tell you an identifier dead-ends. It cannot tell you that the sentence you wrote drifts from the sentence you cited.
The move not to make.
Do not ask the model that produced the citations to check them. It is the wrong reader for two reasons: it is the system that generated the error, and its failure mode is confident agreement. Ask it to confirm a fabricated quote and it will often explain, fluently, why the quote is accurate. The same goes for asking a second general-purpose model to bless the list without the documents in front of it — that is a vibes check with extra steps.
And do not treat "the link resolves" as verification. A live URL proves a document exists. It says nothing about whether the document supports your sentence.
How you'll know it worked.
Four checks. Every entry in the reference list has a record you opened yourself, not a search result you accepted. Every citation has been read at the passage level, with the specific sentence identified. Every quotation matches word for word. And every number attributed to a source has been traced to where that source states it. When you have all four, you are in the position the New Mexico court described: it does not matter whether a tool wrote the draft, the person who signs it must be able to attest to it — and now you can.
If you want the causes and failure modes behind this in more depth, AI 101 — What is an AI hallucination? covers why models fabricate confidently, and AI 101 — What is LLM-as-a-judge? explains why an automated grader is a useful filter and a bad final word. The scale of the problem in published research — including work showing that automated checks already flag thousands of errors in the literature — is in AI audited the AI literature — 99.2% of papers flagged.
Has an AI-generated citation ever slipped past you, or do you check every reference by hand? Tell us in the comments.
Sources: New Mexico Supreme Court order · The Verge · AtomCite (arXiv) · ATTRICITE (arXiv) · Hallucinations in LLMs: a lifecycle survey (arXiv)