The Six-Platform Study — What I Found When I Tested Clinical AI

https://leanpub.com/AiforDoctors

Why I Built This Study

https://leanpub.com/AiforDoctors

The Six Platforms

https://leanpub.com/AiforDoctors

The Rubric: 20 Items, 100 Points, One Standard

https://leanpub.com/AiforDoctors

The Results

https://leanpub.com/AiforDoctors

The Final Scores

https://leanpub.com/AiforDoctors

What the Clinical AI Tools Got Right

https://leanpub.com/AiforDoctors

Where the General LLMs Failed

https://leanpub.com/AiforDoctors

GPT-5 (Score: 34 / 100)

https://leanpub.com/AiforDoctors

Claude Opus 4 (Score: 35 / 100)

https://leanpub.com/AiforDoctors

Gemini 2.5 Pro (Score: 29 / 100)

https://leanpub.com/AiforDoctors

Llama 4 (Score: 21 / 100)

https://leanpub.com/AiforDoctors

What the 2.4x Gap Actually Means

https://leanpub.com/AiforDoctors

The Benchmark Paradox

https://leanpub.com/AiforDoctors

What No Rubric Captures

https://leanpub.com/AiforDoctors

Why the Gap Exists: Architecture, Not Intelligence

https://leanpub.com/AiforDoctors

What This Means for Your Practice

https://leanpub.com/AiforDoctors

Study Limitations

https://leanpub.com/AiforDoctors

Before You Turn the Page

https://leanpub.com/AiforDoctors