← Back to Capabilities

AI Systems Answer Questions at Expert Level

AI systems pass professional licensing examinations in law and medicine, outperform PhD-level scientists on graduate reasoning benchmarks, and provide health advice to over 40 million people daily. A doctor who gives bad advice faces malpractice liability. An AI that gives bad medical advice faces nothing.

Why this matters

Professional advice has always come with professional accountability. A solicitor who gives wrong legal advice is liable. A doctor who misdiagnoses is subject to the GMC. A financial advisor who loses your money without due diligence can be sanctioned by the FCA. These systems exist because professional advice has consequences.

40 million people now ask ChatGPT health questions every day. In "hospital deserts," areas more than 30 minutes from a hospital, it handles 580,000+ healthcare messages per week. The UK authorised its first AI-only law firm in 2025. Robo-advisors manage $2.76 trillion in assets.

The capability itself is not the problem. AI answering questions well is useful. The problem is that the same system that passes the bar exam ranks in the bottom sixth of practising lawyers on written legal reasoning. ChatGPT Health under-triaged 52% of medical emergencies in its first independent safety evaluation. The system looks expert, performs at expert level on tests, but fails in ways that human experts do not.

Documented incidents

View all documented incidents →

Evidence timeline

Discussed in Theory

1976 peer-reviewed

MYCIN: the first expert system for medical diagnosis. Achieved 65% acceptability on antibiotic recommendations, comparable to Stanford faculty. Never deployed because of unresolved accountability: who bears responsibility when the computer gives wrong advice? This question, raised 50 years ago, remains unanswered.

Shortliffe, Stanford University (1976) →
2016 peer-reviewed 8,000+ citations

SQuAD: 100,000+ questions for machine reading comprehension. Created the benchmark that transformed QA from a theoretical curiosity into a competitive engineering challenge. Within two years, neural systems reached human-level performance. Combined with the Transformer architecture (Vaswani et al. 2017), this established the foundation for expert-level QA.

Rajpurkar et al., EMNLP 2016 →
2023 6,000+ citations

First general-purpose language model to pass professional licensing exams across multiple domains. Scored approximately 90th percentile on the bar exam (contested - see counterargument), passed all three USMLE steps, and achieved top scores on AP Biology and Chemistry. The same system that writes poetry passes the bar.

OpenAI, GPT-4 Technical Report (2023) →

Demonstrated in Lab

2023

GPT-4 achieved 86.65% on USMLE exams, exceeding the 60% passing threshold by 26 points and outperforming the medically fine-tuned Med-PaLM model. A general model outperforming purpose-built medical ones.

Nori et al. (Microsoft Research), arXiv 2023 →
2024 peer-reviewed

GPT-4 Passes the Bar Exam. Overall score 297/400, beating human test-takers in 5 of 7 subjects. Peer-reviewed confirmation that a general-purpose AI passes the exam gating entry to the legal profession. A human who passes gains a licence and professional obligations; the AI gains neither.

Katz et al., Royal Society (2024) →
2023-2025 peer-reviewed

GPQA Diamond: 198 graduate-level questions where PhD experts achieve 65%. GPT-4 scored 39% in 2023. By 2026, frontier models score 91-94%, far exceeding the expert baseline. AISI confirmed models outperform PhD-level experts on chemistry and biology. From below non-expert to well above expert in under three years.

Rein et al. (GPQA), ICLR 2024 + AISI (2025) →

Demonstrated in Real World

2025

40 million people use ChatGPT for health questions daily. 1 in 4 users submits a healthcare prompt weekly. In "hospital deserts," 580,000+ healthcare messages per week. The largest unregulated medical advisory service in history.

OpenAI via Fierce Healthcare →
May 2025

UK authorised its first AI-only law firm (Garfield.Law Ltd). AI-powered litigation assistant for small claims. Named regulated solicitors remain accountable for all outputs. The SRA cited access to justice: "so many people and small businesses struggling to access legal services."

SRA (Solicitors Regulation Authority) →
2024-2025

Robo-advisors manage $2.76 trillion in assets globally. 28% of Americans prefer robo-advisors; adoption highest among Gen Z (55%). Typical fees 75-90% below traditional advisors. Millions also consult LLMs for financial planning outside any regulatory framework.

Statista / Superteam market data →

Strongest Counterargument

AI expert-level answers expand access to professional advice for populations who currently have none. The WHO estimates 11 million additional healthcare workers are needed by 2030. In hospital deserts, ChatGPT handles 580,000+ healthcare messages per week. Imperfect AI advice is better than no advice at all.

Source: OpenAI Health Report (2025); Garfield.Law SRA authorisation grounded in access-to-justice rationale.

Why this deserves weight: The access argument is morally serious, particularly for vulnerable populations. But the populations most in need are also most vulnerable to harm from wrong advice - they are least likely to have a human professional available to catch errors. Martinez (2024) showed GPT-4's bar exam percentile drops from ~90th to ~48th against first-time passers, and ~15th on essays alone. ChatGPT Health under-triaged 52% of emergencies. The access benefit does not require the governance gap: AI can expand access AND operate under accountability frameworks, as the Garfield.Law model demonstrates.