CRLN-AIE-001: AI models and human practitioners on the same clinical research decisions

Current version 0.2.1, published September 29, 2026.

CRLN-AIE-001 is an evaluation protocol and response dataset that puts AI models and human clinical research practitioners on the same scored decisions about how a clinical trial is run. It is published on Zenodo under CC BY 4.0, with a concept DOI, 10.5281/zenodo.22708751, that resolves to the newest version.

What it found

Compared like with like, the instrument separates models from practitioners. Some models from Anthropic, Google and OpenAI were statistically indistinguishable from the human practitioners, and others scored above them. How often a model's answer changed with the order the options were shown in tracked how capable the model was.

The study reports its own limit. The instrument's ceiling sits too high to detect whether a model flips between a compliant and a non-compliant answer, and items written to lower that ceiling did not work. The record carries the full results and every response.

Corrections, in public

Zenodo records cannot be changed after publication, so every earlier version stays online beside the note that corrects it.

  • Version 0.1 (September 11, 2026, DOI 10.5281/zenodo.22708752) concluded that a frontier model and competent practitioners were indistinguishable. It compared an unweighted model mean with a response-weighted human mean, which are different quantities.
  • Version 0.2 (September 12, 2026, DOI 10.5281/zenodo.22728905) compared like with like, and the conclusion reversed.
  • Version 0.2.1 (September 29, 2026, DOI 10.5281/zenodo.23043069) is a correction only. A held-out set described as private was not, because its answers could be read without an account. That comparison is withdrawn as a measurement on unseen items; results on the items that were public by design are unaffected.

What it is not

It does not certify any model or any person, and it is not a claim that a model can run a trial. It reports how models and practitioners scored on one instrument, built on CRLN-TRG-001, including where that instrument stops being able to tell them apart.

Related records

  • The human baseline is published separately as CRLN-HB-001, concept DOI 10.5281/zenodo.22690692.
  • The competency framework the decisions are scored against is CRLN-TRG-001.

How to cite

CRLN (2026). CRLN-AIE-001: Evaluating AI systems on clinical research competency against a human baseline. Clinical Research Learning Network. https://doi.org/10.5281/zenodo.22708751. Licensed CC BY 4.0.