A free, public-interest tool for checking whether citations correspond to real scholarly records.

The process · Methodology 2026.08

How CitationCheck works

Verification checks the record. Human review checks the meaning. This page covers both the process and the full methodology behind every result.

Check citations Read the methodology

Process at a glance

  1. Parse

    Your list is split into fields — authors, title, source, year, identifier.

  2. Match

    Identifiers are resolved with the matching service; the rest is searched across open scholarly indexes.

  3. Compare

    Fields are shown side by side, with every difference marked.

A recorded walkthrough will replace this panel. Until then, this is the whole process in three moves.

Step by step

Five steps, one pass.

Nothing is rewritten, corrected or scored on your behalf — each step only adds evidence you can see. No language model generates or fills in bibliographic facts anywhere in the pipeline.

  1. Parse the reference

    The text is split into title, authors, year, and venue, and scanned for identifiers — DOI, PMID, arXiv ID. Parsing is deterministic; nothing rewrites or "fixes" your citation.

    OutputStructured fields and any detected identifiers.

  2. Find the record

    A DOI is looked up at Crossref (OpenAlex as fallback), a PMID at PubMed, an arXiv ID at arXiv; references without identifiers are searched against Crossref and OpenAlex. A resolved identifier identifies a candidate record — it never by itself proves the citation is correct, because a real DOI can travel with wrong authors, a wrong year, or details from a different paper.

    OutputCandidate scholarly records.

  3. Compare the fields

    Titles and venues are normalized first — case, accents, punctuation, journal abbreviations don't count against a match. Authors are compared by family-name overlap, tolerant of initials, name order, and particle surnames. Years within ±1 still agree (±2 where a preprint is involved). Meaningful differences are kept and shown.

    OutputField-by-field agreement.

  4. Score against the rules

    Two numbers are kept deliberately separate: the agreement score (how well the available fields match) and evidence coverage (how much identifying detail your citation actually carried — missing fields lower coverage, they never raise confidence). Some conflicts are disqualifying no matter how much else agrees: an identifier pointing at a different publication, a major author mismatch, an incompatible year. Exact thresholds are in the table below.

    OutputAgreement, coverage, and any disqualifying conflicts.

  5. Classify and show the evidence

    One of five states: Verified, Partial match, Needs review, Not found, or Service unavailable. When the evidence is ambiguous, the classifier prefers Needs review over Verified — a false green is more dangerous than a conservative one. Every result card lists what matched, what differed, which sources were checked, and what the result does not mean.

    OutputThe evidence card you see.

In practice

Three worked examples.

Verified does not mean good, and Not found does not mean fake. Every card states its own honest reading.

Verified

The citation includes a DOI. It resolves to a Crossref record whose title, authors, and journal all agree; the year differs by one, consistent with online-first versus print publication.

The bibliographic information corresponds to a real record. That is all Verified means — it says nothing about quality, retraction status, or whether the paper supports the claim it was cited for.

Needs review

The DOI resolves and the title matches — but the citation lists Smith, Jones and Patel while the record lists Smith, Brown and Patel. Two of three cited authors agree.

A real identifier and title can travel with corrupted details — a known pattern in generated citations. The card shows the author comparison and asks for human review rather than issuing a false green.

Not found

The reference is well-formed but no sufficiently matching scholarly record was located in any enabled source, and the closest candidates differ substantially.

Coverage gaps, spelling, indexing delays, and non-scholarly sources (books, theses, news, webpages) can all cause a real work to go unfound — treat Not found as a prompt to check the original source, never as proof of fabrication.

The methodology

CitationCheck answers one narrow question: does this citation appear to correspond to a real scholarly record, and do its bibliographic details agree with that record? It does not judge quality, correctness, retraction status, or whether the work supports any claim. The exact rules are below; if you can't tell why you got a result, that's a defect — report it via the link on any result card.

The rules, in numbers

Classification thresholds by result state
ResultRequires
Verified Agreement ≥ 85% · evidence coverage ≥ 65% · no critical contradiction. Via an identifier, the record's title must also agree at ≥ 75%.
Partial match Agreement ≥ 55%, or an identifier that resolves to a clearly different publication.
Needs review Everything ambiguous — including an identifier with matching title but conflicting authors or year, and citations too sparse to corroborate (a bare DOI or bare title can never verify).
Not found Best agreement < 45%, with at least 2 providers having responded successfully. Weak candidates are still listed.
Service unavailable A provider failed, timed out, or rate-limited. Provider failure is never converted into Not found.
Critical contradictions Title similarity < 55% · author overlap < 25% (both sides listing ≥ 2 authors, "et al." exempt) · year gap > 3 years. Any one blocks Verified regardless of the aggregate score.
Field weights Title 45% · authors 25% · year 15% · venue 15%, weighted over the fields actually compared. Neither number is a probability that a source is real, and neither is shown without the underlying comparison.

Preprints and provider coverage

An arXiv preprint and its published journal version are related but different records: when a preprint is involved, venue mismatches are softened and year tolerance widens, and the result card shows which version you are actually citing. The enabled sources also have real gaps — regional publishers, books, theses, non-English venues, very recent works, and non-scholarly sources may not appear, and misspellings defeat search.

Is AI involved?

No. Parsing, normalization, identifier resolution, matching, and classification are deterministic code, and suggested records come only from data actually returned by providers.

How thresholds are set — and how to correct us

Thresholds live in configuration, are exercised against a fixture suite of representative citations (correct, malformed, mismatched, fabricated-looking, preprint/published pairs), and are adjusted when incorrect-result reports show systematic error. They are judgment calls made inspectable — not claims of statistical optimality. To report a wrong result, email the support address in the footer with the citation you checked and, if possible, the DOI or a link to the record you believe is correct.

Methodology version 2026.08. Last substantive update: August 8, 2026 — identifier resolution no longer auto-verifies; evidence coverage and critical-contradiction checks added.

Check your references before you submit.

Free citation checks with the evidence shown for every result. No account required.

Check citations