Skip to content
PaperFren

Concept

Human–AI collaboration and comparison

6 studies1 discoveryEvidence last moved Sep 27, 2026

Human–AI collaboration studies compare AI outputs with human experts or test AI as an assistant to people, from triage chatbots and LLM-drafted patient replies to agents guiding users in a task. The evidence is small evaluations: clinical vignettes, expert ratings of 7–10 cases, two-participant user studies and one online experiment.

Headlines say 'AI matches doctors'. These studies show what such claims rest on — vignettes, a handful of raters, AI-written text that looks longer — and that how AI assists people affects their sense of control.

Studies

6

Findings

5

8 supporting · 0 challenging · 0 qualifying citations

Open tensions

1

Latest change

Concept page published

Human–AI collaboration and comparison

Currently

What we know

  1. AI outputs can look as good as or better than clinicians' on curated cases.
  2. AI drafts are good starting points, not finished products.
  3. Assistance, not replacement — especially for numbers.
  4. Subtle guidance can keep performance and preserve autonomy.
  5. Expert review stays necessary.

Largest unresolved question

Judges disagree on AI quality: in the triage study one judge rated the AI's differentials comparable to doctors' and another rated them worse.

Common misconceptions

  • 'AI outperformed doctors' in these studies means it's safer for real patients.

    Evidence comes from GP-played vignettes and hand-picked messages rated by physicians, not real patients; patients never rated responses.

  • Higher ratings reflect better medical content.

    Raters may prefer longer, more detailed text, and in the CDS study text style may have revealed which suggestions were AI-written.

Related