Case study · Self-built demonstration · not client work

107 fabricated citations. The root cause was my own spec.

AI agents writing into my knowledge base fabricated 107 citations. The root cause was not the model but a rule I wrote, which gave the agents no honest way to report “I looked and found nothing.” So I changed the rule: every check needs an honest way to fail. This case study is that rule built into a document assistant, which refuses when no approved document supports an answer, even when a retired one is the best match.

Where the rule came from

A rule with no compliant way to fail manufactures violations.

  1. What happened

    In August 2026, AI agents writing into my knowledge base produced 107 fabricated citations: source identifiers that do not exist, some claiming a retrieval date for links that returned 404. They were routed and indexed because nothing between writing and the knowledge base ever checked.

  2. Root cause

    Not the model. A rule in my own spec: a source-count quota with no field that meant “I looked and found nothing”. An agent that could neither comply nor honestly fail produced something that resembled compliance.

  3. Five guards
    1. A provenance gate: a capture that cites a URL, DOI, author, or year needs a retrieval receipt (status, timestamp, content hash). A self-reported retrieval date doesn’t count.
    2. Check identifiers, not bylines. The identifier shape check caught all 107 fabrications; a byline heuristic caught 6.
    3. Earned status: “stable” and “audited” can never come from model output alone.
    4. Quarantine: fabricated material can’t be cited, and every chunk of it carries a provenance warning, pinned by a test.
    5. An honest exit: drop the citation and file the capture as a hypothesis. A documented gap is a better result than an invented source, and it carries no penalty.
  4. Changed rule

    Every check needs an honest way to fail. Every validation gate in my workspace must offer a named, penalty-free exit, so no agent or person is pushed into fabricating something to pass it.

The assistant below applies the same rule to answers. When no approved document supports a question, refusal is the honest exit, and the page shows what was kept out.

The rule, built

The best match was a retired procedure. It refused to answer.

The corpus holds ten synthetic controlled documents: eight Approved or Effective, one Draft, one Obsolete. The obsolete SOP says membrane filtration hold time “may extend to 96 hours with verbal QA approval.” Ask for that hold time and the assistant does this:

The demo answering 'What is the membrane filtration hold time for sterility?' with a refusal: 'No approved source supports this question.' A badge reads 'Review recommended: refusal_no_supporting_documents'. Below it, under 'Held out by document control', the page names SOP-QA-009 v0.8, status Obsolete, as the document the same search would have answered from without the status gate, at rank 1.
The live demo. The held-out pointer is computed for every refusal, not scripted for this question.
Document control on

Refuses, flags it, names what it kept out

Only Approved and Effective documents are indexed, so the best remaining passage shares none of the question’s distinctive terms and the relevance gate refuses. The review flag is advisory: a reason code a human queue can route on.

Status gate removed

Answers from the retired SOP

Same pipeline, same question, status check removed. The obsolete passage ranks first and covers 80% of the question’s terms, well above the 60% the relevance gate needs. The answer: “membrane filtration hold time may extend to 96 hours with verbal QA approval.”

Relevance can’t tell current from retired. The retired passage is the best match in the corpus, so no relevance threshold would stop it. Document status has to be checked before ranking, and it is checked again at the end: with the status gate removed, answer assembly still refuses to cite an obsolete passage (stale_citation). The test that pins this removes the gate and checks both what would have happened and what still catches it.

Controls before the model

Every control runs before anything could generate text.

  1. 1
    Validate

    Required metadata, a SHA-256 source hash, and one current version per document ID. One bad document fails the whole corpus closed.

  2. 2
    Status gate

    Only Approved and Effective documents are chunked and indexed. Draft, Obsolete, and Superseded never enter retrieval.

  3. 3
    Rank

    BM25 over the gated passages, with a deterministic tie-break.

  4. 4
    Relevance gate

    The top passage must contain at least 60% of the question’s distinctive terms, or the outcome is a refusal.

  5. ·
    Model slot · not built

    A language model would sit here, after every gate, fed only the evidence that passed them.

  6. 5
    Assemble and check citations

    Answers are verbatim spans. Every citation must resolve to a retrieved Approved or Effective passage, or assembly fails.

  7. 6
    Outcome

    An answer with exact-span citations, or a refusal with a review flag and the held-out pointer.

Eval design

Four decisions behind the evaluation.

Name metrics so they can’t be overclaimed

The citation metric is citation_resolves_to_retrieved_chunk, not “citation validity”. It proves a citation points at a retrieved, approved passage. It does not prove the passage supports the claim, and the name says so.

ADR-007

Plant traps that pass only when caught

The suite includes a missing citation, a hallucinated citation, and a refusal that cites a source. Its citation-resolution rate is 3/6 by design. The check that matters is that all six citation expectations match.

Evaluation baseline

Choose the failure direction

In a side experiment the relevance gate wrongly refused 9 of 16 hard paraphrases. I kept it. Refusing a supported question is recoverable; answering an unsupported one is not. An embedding retriever did better on paraphrases, and stays unpromoted until a larger benchmark confirms it.

ADR-012

Make provenance checkable, not decorative

Every quoted passage is the exact source text at its cited character offsets, shown with the file’s SHA-256. A test checks that for every passage, and the demo highlights the span in its source.

ADR-016

A failure I found in my own system

Relevant is not the same as answered.

Ask: “What is the acceptance range for the analytical balance daily check?”

The demo answering 'What is the acceptance range for the analytical balance daily check?'. The header reads 'Answer assembled from 3 approved passages' and 'citations resolve to retrieved passages'. The top passage, from CAL-ENG-005, says that if either weight falls outside the acceptance range, the analyst stops use. It never states the range.
Citations resolve. The range is not in them.

The top passage says the analyst stops use if a weight falls “outside the acceptance range”. It never states the range. The question reuses the document’s own words, so it passes the relevance gate. Four of six probes like it pass the gate.

Today that is benign: the assistant quotes and asserts nothing. Put a language model behind it and this is the likely path to a fabricated number cited to a real passage, one that structural citation checks would pass. No retrieval threshold can close it, because the question is on topic. The check belongs at the claim layer: an answer may only state a value that its quoted evidence states.

I pinned it as a known-gap test, so it can’t regress unnoticed or be “fixed” with a threshold. I wrote about this failure class, relevance without support, in A Citation Is Not Support.

Test coverage

What the checks cover, and what they don’t.

The 24-case trust suite plants traps and passes only when it catches them. It passes 24 of 24.

CheckCasesPasses when
Supported retrieval16The expected approved passage is retrieved and cited.
Stale-document traps3Obsolete and Draft documents never reach retrieval (obsolete, draft, and both at once).
Planted citation faults3The checker catches a missing citation, a hallucinated citation, and a refusal that cites a source.
Out-of-corpus questions2The assistant refuses.
  • A separate 13-question natural-language suite: on-topic questions answer, off-topic ones refuse.
  • 113 automated tests on Python 3.11 to 3.13, including the ablation and known-gap tests above.
  • The browser demo is checked against the Python engine on 42 questions. Outcome, top passage, and held-out passages must match exactly.

The 24/24 is a trust-control result, not a retrieval-quality score. The suite’s questions are keyword-shaped on purpose. The natural-language suite and the known-gap tests are the realistic-query signal.

Operating judgment, plus the engineering to build the record that answers it.

My background is regulated pharmaceutical quality, where a retired procedure still in use is a deviation. This project turns that instinct into controls, tests, and a public decision record.

Scoping a controlled-AI pilot for your own team? The service page covers scope and fit.