Research track · exploratory

Do reviewers catch the errors in AI-assisted work?

In regulated work, the usual answer to AI risk is "a qualified person reviews it." I want to know how well that holds up, one case at a time.

The question

When a person reviews AI-assisted work in a regulated setting, do they catch the errors that matter?

Human review is the control most AI-assisted workflows lean on. It is also the control that is hardest to see working. A reviewer who signs off a clean-looking output has not shown that they would have caught a bad one.

How it works right now

  • Three synthetic pharma-quality cases: one clean, one with a material omission stated with false certainty, and one where a real citation does not support the claim it is attached to.
  • For each AI output, the reviewer chooses accept, correct, reject or escalate, points to the evidence, and writes a short reason.
  • Scoring keeps things apart that are easy to blur: errors caught, false alarms on the clean case, whether the decision rests on evidence, escalation, and review time.

What it is not

  • Not a product, a course or a certification.
  • No claim of validation, qualification or GxP compliance.
  • Not a judgement of any person or organization. The cases are invented.

How to take part

If you review AI-assisted documents, investigations or submissions in pharma, medical devices or another regulated setting, send a note with your role and the kind of review you do.

I share the case package one-to-one after a short exchange, not automatically. The cases are still changing, and some people will be asked to wait for the frozen version so they can test it fresh.

What happens to your note

I use it to decide whether to follow up about this research. Nothing you send is published. I won't name you or your employer in a write-up without asking first.

Updates

  1. Track page opened. A three-case proof build exists. Packages go out one-to-one for individual pilot runs. The cases are not frozen, and usability with someone unfamiliar with the tool has not been run yet.

Want to take part, or think the question is wrong?

Send a short note about your role and the kind of work you do. Please don't attach or paste confidential documents, batch records, client data or exports.