Concept · artificial-intelligence
Human–AI collaboration and comparison
6 studies1 discoveryEvidence last moved Sep 27, 2026
Human–AI collaboration studies compare AI outputs with human experts or test AI as an assistant to people, from triage chatbots and LLM-drafted patient replies to agents guiding users in a task. The evidence is small evaluations: clinical vignettes, expert ratings of 7–10 cases, two-participant user studies and one online experiment.
Headlines say 'AI matches doctors'. These studies show what such claims rest on — vignettes, a handful of raters, AI-written text that looks longer — and that how AI assists people affects their sense of control.
Studies
6
Findings
5
8 supporting · 0 challenging · 0 qualifying citations
Open tensions
1
Latest change
Concept page published
Human–AI collaboration and comparison
Currently
What we know
- AI outputs can look as good as or better than clinicians' on curated cases.
- AI drafts are good starting points, not finished products.
- Assistance, not replacement — especially for numbers.
- Subtle guidance can keep performance and preserve autonomy.
- Expert review stays necessary.
Largest unresolved question
Judges disagree on AI quality: in the triage study one judge rated the AI's differentials comparable to doctors' and another rated them worse.
Common misconceptions
'AI outperformed doctors' in these studies means it's safer for real patients.
Evidence comes from GP-played vignettes and hand-picked messages rated by physicians, not real patients; patients never rated responses.
Higher ratings reflect better medical content.
Raters may prefer longer, more detailed text, and in the CDS study text style may have revealed which suggestions were AI-written.
Related
Claim ledger
What the evidence shows
Drawn from 6 studies in this library. Mix labels say which citation roles are present; they are not a strength score. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope.
AI outputs can look as good as or better than clinicians' on curated cases.
On simulated cases, AI can rate comparably to clinicians: a triage AI's triage was safe in 97.0% of vignettes vs 93.1% for doctors with similar appropriateness, and ChatGPT patient-message replies received the highest expert ratings, with only 1 of the top 20 from the actual doctor.
- Can an AI symptom checker match GPs at diagnosis and triage?
- Can fine-tuned language models draft replies to patient messages?
Study Role Design N Population Outcome Can an AI symptom checker match GPs at diagnosis and triage? Supports OtherVignette-based role-play comparison (OSCE style): GPs acting as patients consulted both locum GPs and a Bayesian-network AI triage system; outputs rated against the vignette disease and by blinded judges N=100 · 100 clinical vignettes in the main role-play; a further 30 public vignettes for a benchmark comparison with three doctors Simulated adult primary-care cases (vignettes) played by GPs; locum GP doctors versus the Babylon Triage and Diagnostic System Precision and recall of differential diagnosis, judge ratings of differentials, triage safety and appropriateness Can fine-tuned language models draft replies to patient messages? Supports Computational / modellingTwo LoRA-fine-tuned LLaMA-65B models (local data only vs local plus GPT-improved open-source data) compared with ChatGPT-3.5, ChatGPT-4 and real provider replies, rated blind by physicians on Likert scales plus BERTScore Training: nearly half a million local message-response pairs; evaluation: 70 responses to 10 rephrased patient messages, each rated by 4 primary care physicians Adult primary care patient-portal messages and provider replies from one US academic medical centre Physician ratings (1-5) of empathy, responsiveness, accuracy and usefulness; BERTScore similarity to provider replies AI drafts are good starting points, not finished products.
But AI suggestions are often rated less usable as-is: ChatGPT decision-support suggestions matched humans on understandability and relevance but scored lower on usefulness (2.7 vs 3.5) and acceptance without edits (1.8 vs 2.8), though 9 of the top 20 were ChatGPT's.
Assistance, not replacement — especially for numbers.
AI assistance can speed expert work while humans check errors: in a two-person user study, AI-assisted screening improved recall by 71.4% with 44.2% less time, but extraction of numerical results was much less accurate than study design.
Subtle guidance can keep performance and preserve autonomy.
How an agent helps shapes perceived autonomy: implicit and explicit guidance agents both improved task success over a supportive agent, but participants felt more in control with implicit guidance.
Expert review stays necessary.
Hallucinations and outdated advice appear in AI output: a hallucinated drug name in decision-support suggestions and outdated antibiotic and COVID advice in model replies were flagged.
- Can ChatGPT suggest useful fixes for hospital alerts?
- Can fine-tuned language models draft replies to patient messages?
- Can ChatGPT turn messy pathology reports into clean data?
Study Role Design N Population Outcome Can ChatGPT suggest useful fixes for hospital alerts? Supports Human experimentFive blinded CDS experts rated a randomised mix of ChatGPT and human suggestions for 7 EHR alerts on eight Likert items. No single N: 36 suggestions from ChatGPT and 29 from humans for 7 alerts, each rated by 5 experts. Physician and pharmacist clinical-decision-support experts at two US academic medical centres. Likert ratings of understanding, relevance, usefulness, acceptance, workflow, redundancy, inversion and bias, plus an overall score. Can fine-tuned language models draft replies to patient messages? Supports Computational / modellingTwo LoRA-fine-tuned LLaMA-65B models (local data only vs local plus GPT-improved open-source data) compared with ChatGPT-3.5, ChatGPT-4 and real provider replies, rated blind by physicians on Likert scales plus BERTScore Training: nearly half a million local message-response pairs; evaluation: 70 responses to 10 rephrased patient messages, each rated by 4 primary care physicians Adult primary care patient-portal messages and provider replies from one US academic medical centre Physician ratings (1-5) of empathy, responsiveness, accuracy and usefulness; BERTScore similarity to provider replies Can ChatGPT turn messy pathology reports into clean data? Supports Computational / modellingPrompt engineering on a small training set of public pathology reports, then held-out evaluation of GPT-3.5-turbo against expert-curated labels, a keyword-search baseline and a fine-tuned BERT NER model N=774 · 774 valid TCGA pathology reports in the test set; 78 CDSA reports used for prompt development; a separate 191 osteosarcoma reports for an extension De-identified lung cancer pathology reports from public archives (CDSA, TCGA) Accuracy and coverage for pathologic T stage, N stage, overall stage and histology type
Debates
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes.
Judges disagree on AI quality: in the triage study one judge rated the AI's differentials comparable to doctors' and another rated them worse.
Judges disagree on AI quality: in the triage study one judge rated the AI's differentials comparable to doctors' and another rated them worse.
Qualified studies asking the same question reach different answers. The disagreement is listed, not scored.
Timeline
How understanding moved
Study years are when the paper was published. Evidence edits are dated changes to this page's claims. Explanations are when PaperFren added a Discovery — not a claim that the science happened that day.
2026
- Dropping race from a model can make it less fair, and fixing one gap can widen another
Concept page published
Human–AI collaboration and comparison
Change log
What changed
Dated edits to this page's evidence: studies added or removed from a claim, claims added or withdrawn, and new explanations tagged here. Rewordings are not listed.
- Concept page published
Papers
6 studies in this library bear on Human–AI collaboration and comparison, ordered by citations.
- Can ChatGPT suggest useful fixes for hospital alerts?
Experts found ChatGPT's suggestions for improving clinical alerts clear and relevant but less useful and less ready to adopt than human suggestions, though several ranked among the best.
- Can ChatGPT turn messy pathology reports into clean data?
With carefully engineered prompts, ChatGPT extracted cancer stage and tumour type from pathology reports with about 89% average accuracy, beating older NLP methods, but it misapplied staging rules and invented answers for blank reports.
- Can an AI symptom checker match GPs at diagnosis and triage?
In simulated consultations, a company's Bayesian-network symptom checker matched GPs at listing the right condition and gave safer (more cautious) triage advice, but expert judges disagreed about how good its diagnosis lists were.
- Can fine-tuned language models draft replies to patient messages?
A language model fine-tuned only on doctors' real portal replies wrote drafts no better than the doctors, but adding richer, empathetic example replies to its training made it about as good as ChatGPT.
- Can an LLM pipeline speed up systematic reviews without losing quality?
Breaking systematic-review work into structured steps that a large language model performs and experts can check found far more relevant studies and extracted data more accurately than simply prompting GPT-4.
- Can an AI teammate guide you without making you feel bossed around?
An agent that hinted at the best goal through its own movements helped people pick the better target as often as an agent that told them outright, while leaving them feeling more in control.
Compare studies
Select 2–10 studies. Design and N are labels, not a ranking.
Nothing selected yet.
Questions
What is still open
Judges disagree on AI quality: in the triage study one judge rated the AI's differentials comparable to doctors' and another rated them worse.
Ask PaperFren about Human–AI collaboration and comparison
Study this conceptflashcards and short-answer questions
Critically evaluate the claim that an AI triage system is as safe as doctors.
On vignettes, the AI's triage was safe 97.0% vs 93.1% for doctors. But vignettes played by GPs are not real patients, many authors were company employees, the gold standard came mainly from one judge, and no inferential statistics were reported. Judges also disagreed about differential quality.
What does the implicit guidance agent study add beyond accuracy?
Both guidance agents improved object capture over a supportive agent, with no significant difference between them. But participants felt more autonomous with implicit guidance. Tiny mazes, a single Likert autonomy item and a mostly male online sample limit generalisation.