Reporting guidelines like CONSORT exist because badly reported trials are hard to use and easy to misread. Checking adherence by hand is slow, which is part of why efforts to improve it have had mixed results. So we asked a narrow question: can a large language model do the checking?
We took 113 published sports medicine and exercise science clinical trials and asked GPT-4 Turbo and Llama 2 nine reporting-guideline questions about each one. GPT-4 Vision got two extra questions about the participant flow diagram. The data was split 80/20 for training and testing.
What we found
- GPT-4 Turbo answered questions about the article text with an F1 of 0.89 and 90% accuracy (95% CI 85–94). Accuracy was above 80% for every guideline item.
- Llama 2 started poorly — F1 0.63, accuracy 64% (57–71) — but fine-tuning on GPT-4’s answers pulled it up to F1 0.84 and 83% accuracy (77–88).
- GPT-4 Vision spotted every participant flow diagram (100%, 89–100), but was much weaker at noticing when details were missing from one (57%, 39–73).
That last result is the interesting one. Finding a figure is easy; knowing what should have been in it and isn’t isn’t. Absence is harder to see than presence — which is precisely the skill guideline checking demands.
The practical takeaway: this looks feasible, but the version worth building is an open-source one. A closed model that costs money per paper and can change under you is a poor foundation for research infrastructure. That the fine-tuned Llama 2 got most of the way there is the encouraging part.
Read it
Wrightson, J.G. Blazey, P. Moher, D. Khan, K.M. and Ardern, C.L. (2025), GPT for RCTs? Using AI to determine adherence to clinical trial reporting guidelines, BMJ Open 15:e088735. Open access, also on PubMed Central.
I presented this work at the METRICS International Forum at Stanford in February 2024. The recording is on YouTube.
AI disclosure
This blog post is an AI-generated summary. It was drafted by Claude Opus 5 from the published article, and reviewed by me before publication. The study it describes, and the paper reporting it, are the authors’ own work.
Anthropic. (2026). Claude Opus 5 (model version claude-opus-5) [Large language model]. https://www.anthropic.com/claude