Hero illustration for “The Overconfidence Effect: Why Experts Are Wrong More Often Than They Think”
Skip to article
Hero plate: The Overconfidence Effect: Why Experts Are Wrong More Often Than They Think
HPC  ·  Science Deep Dive  ·  revised

The Overconfidence Effect: Why Experts Are Wrong More Often Than They Think.

The gap between how certain experts feel and how often they are right is not a personality flaw. It is a measurable, neurally encoded failure mode that appears in every high-stakes professional domain tested, and the overconfidence bias science shows it can be reduced. Here is what the science actually says, and what to do with it.

01The 30-Point Gap

Experts at 98% Confidence Are Right 68% of the Time

Ask yourself a factual question: something you feel confident about. A date, a number, a prediction about next quarter's results. Now assign a probability: how certain are you? If you said 98%, Lichtenstein, Fischhoff, and Phillips have a number for you. Across more than 15,000 probability judgments aggregated from over twenty studies, people who stated 98% confidence were correct 68% of the time.[8] That is not a rounding error. That is a 30-percentage-point calibration gap: the systematic distance between how certain humans feel and how often they are actually right.[7]

The finding did not stay in the laboratory. Svenson's canonical 1981 study asked American drivers to rate their own safety: 93% placed themselves above the median.[9] The mathematics are unyielding (half of all drivers are, by definition, below median), but the overconfidence bias does not answer to mathematics. It answers to something deeper in the architecture of human judgment.

That matters because this is not a story about amateurs misjudging trivia questions. The overconfidence effect scales with expertise, stakes, and consequence. Tetlock spent twenty years tracking 284 political and economic experts (professors, government advisors, think-tank analysts) as they made 28,000 predictions about world events. The result: they barely outperformed random chance.[4] The experts who appeared most frequently in media, the ones who spoke with the greatest certainty, performed worst of all.[5]

The history

The scale of the problem becomes clear only when you see it across domains. In medicine, physicians claiming complete certainty about a diagnosis were wrong approximately 40% of the time in autopsy-comparison studies (a figure that applies to certainty-claiming diagnoses specifically, not to clinical judgment generally, since calibration improves in domains with rapid diagnostic feedback).[22] In finance, CFOs asked to provide 90% confidence intervals for the S&P 500's annual return captured the true value only 36.3% of the time.[51] In national security, a 2025 study of approximately 1,900 NATO-cleared officials found they were correct about 58% of the time on forecasts they rated at 90% confidence, a 32-point gap that mirrors the laboratory findings from four decades earlier.[31]

These are not isolated failures. Grežo's 2021 meta-analysis of 34 studies confirmed that overconfidence exerts a statistically significant effect on financial decision-making across investment, trading, and innovation contexts.[28]

The pattern is robust enough to name. Overconfidence bias science (the study of why humans systematically overestimate their knowledge, overrate their relative standing, and set confidence intervals too narrow) is one of the most replicated research programmes in behavioural science.[1][6] The question is no longer whether the effect is real. The question is why the brain generates it, and what can be done about it.

02The Mechanism

Three Types of Overconfidence and the Neural Architecture That Produces Them

For forty years, the overconfidence literature produced apparently contradictory results. Some studies found people wildly overestimated their abilities. Others found they underestimated themselves on easy tasks. The confusion persisted until Moore and Healy published their definitive 2008 taxonomy in Psychological Review, which demonstrated that researchers had been conflating three mechanistically distinct phenomena under a single label.[6]

The first is overestimation, believing your absolute performance is better than it actually is. The second is overplacement, believing you rank higher relative to others than you do, the engine behind Svenson's 93% above-median drivers.[9] The third, and most dangerous, is overprecision: the excessive certainty that your beliefs are correct, expressed as confidence intervals that are far too narrow.[6] Moore and Healy showed that on easy tasks, people underestimate their absolute performance but overplace themselves against peers. On hard tasks, the pattern reverses. The three types respond differently to task difficulty, which is why decades of research that treated them as one thing generated decades of confusion.

The critical insight is that overprecision is the most persistent of the three forms. It appears regardless of task difficulty, population, or domain.[6][23] Sanchez and Dunning's 2023 interdisciplinary review confirmed that experts are uniformly overprecise regardless of expertise level; expert status does not reduce and may actually exacerbate excessive certainty.[23]

Confidence signal 01 reward processed Striatum 02 certainty rewarded vmPFC / rACC 03 confirming evidence up Lateral PFC 04 errors coded weakly

The reward architecture of overconfidence: the brain treats certainty as a reward signal in the striatum, while the vmPFC and rACC selectively amplify confirming evidence, leaving the lateral PFC to encode disconfirmation weakly, hardwiring a systematic gap between felt certainty and actual accuracy.

Diagram · HPC

The neural evidence explains why. Molenberghs and colleagues conducted the largest fMRI study on confidence and accuracy ever published (308 participants, exceptional for neuroimaging) and found that the brain processes certainty and correctness through structurally separate systems.[15] Higher confidence activated the striatum and hippocampus, regions associated with reward processing. Higher metacognitive accuracy (actually knowing when you are right and when you are wrong) correlated with decreased activation in the anterior medial prefrontal cortex.[15]

That matters because the implication is stark: the feeling of being certain and the reality of being correct are generated by opposing neural systems. Confidence feels like a reward. The brain treats it as one. Fleming and Dolan's review of the neural basis of metacognitive ability confirmed the dissociation: the lateral prefrontal cortex governs retrospective accuracy judgments, while medial structures govern the prospective feeling of knowing.[16]

The architecture gets worse under the influence of asymmetric belief updating. Sharot, Korn, and Dolan demonstrated that participants selectively updated their beliefs in response to better-than-expected information but failed to update appropriately after worse-than-expected news.[17] The ventromedial prefrontal cortex and rostral anterior cingulate cortex show heightened activation for confirming evidence, while the lateral PFC codes negative prediction errors weakly.[17][18] Approximately 80% of people demonstrate this optimism bias.[18]

03Evidence

The Five Strongest Studies on the Overconfidence Effect

01The claim

The single load-bearing finding

The hero study finds 28,000+ predictions.

Not all evidence carries equal weight. A twenty-year prospective longitudinal study tells you something different from a single-session lab experiment with thirty undergraduates. The five studies ranked below represent the strongest available evidence on overconfidence, chosen for design quality, measurement precision, causal clarity, and replication value. Together, they establish that the calibration gap is not an artefact of how researchers ask questions. It is a stable property of how human minds process uncertainty.

Pooled estimate

28,000+ predictions

02How we measured

Grading the calibration studies

Studies scored on design, sample, rigour, causality, replication, citations.

For overconfidence science, ecological validity is decisive: laboratory calibration gaps mean little unless they replicate in real expert populations making real consequential forecasts over years.

Rubric weights

Design/30
Sample/20
Rigour/15
Causality/15
Replication/10
Citations/10

03The spread

Heterogeneity across 5 studies

Methodological quality across the ranked studies.

The hierarchy reveals a pattern that is easy to miss when reading individual studies: the evidence gets stronger, not weaker, as you move from the laboratory to the field. Tetlock's experts were not performing contrived tasks under artificial conditions. They were making real predictions about real events (elections, economic shifts, military conflicts) in their areas of professional specialisation.[4] The TNSR 2025 study extended this to approximately 1,900 NATO-cleared national security officials making 60,000-plus assessments about geopolitical events.[31] The gap held.

Rubric spread

91 → 70 /100

Highest to lowest rubric score across the ranked studies.

04What does not hold

Negative knowledge

What the evidence base does not support.

One important nuance deserves emphasis. Gigerenzer's ecological rationality framework challenges the blanket framing of heuristics as error-prone. Simple heuristics, he argues, can outperform optimisation under specific ecological conditions.[3] This is not a refutation of overconfidence research. It is a boundary condition. Weather forecasters approach calibration because they operate in environments with rapid, unambiguous feedback.[8] The problem is not that heuristics always fail.

The studies

5 trials. One pooled answer.

Below: the anchor study in full; then the forest plot at scale; then the supporting trials in ranked order.

The Key Study Highest rubric · 91/100 · load-bearing

01Anchor

Expert Political Judgment: How Good Is It? How Can We Know?

Tetlock Superforecasting: The art and science of prediction 2005 20-yr Longitudinal · Expert Population · Pre-Registered Outcomes

Philip Tetlock spent two decades collecting predictions from 284 political and economic experts (university professors, government advisors, think-tank analysts). He verified every prediction against actual outcomes.

Rubric breakdown

Design27/30
Sample17/20
Rigour14/15
Causality13/15
Replication10/10
Citations10/10
Total 91/100

The strongest studies, ranked by methodological weight.

Each scored 0–100 against a six-criterion rubric, tagged by design and year; the anchor leads.

050100 rubric 90 01 Tetlock Cohort · 2005 91 02 Lichtenstein, Fischhoff & Phillips 1982 85 03 Moore & Healy 2008 79 04 Molenberghs, Trautwein & Böckler 2016 74 05 Mellers, Ungar & Baron 2014 70 rubric score · out of 100
Anchor (Rank 1) Supporting
Rank Authors & title Journal · Year Finding Score

02

Lichtenstein, Fischhoff & Phillips

Calibration of Probabilities: The State of the Art to 1980

1982

When subjects stated 98% confidence, accuracy was 68%, a 30-point calibration gap replicated across populations, domains, and decades. Calibration only approached accuracy in rapidly-feedback domains like experienced weather forecasting.

85/100

03

Moore & Healy

The Trouble with Overconfidence

2008

Three distinct forms of overconfidence (overestimation, overplacement, and overprecision) respond differently to task difficulty. Overprecision is the most persistent, present regardless of context.

79/100

04

Molenberghs, Trautwein & Böckler

Neural correlates of metacognitive ability and of feeling confident

2016

Higher confidence correlated with striatal/hippocampal (reward) activation; higher metacognitive accuracy correlated with decreased anterior medial prefrontal activation. The feeling of being right and the reality of being right recruit opposing neural systems.

74/100

05

Mellers, Ungar & Baron

Psychological Strategies for Winning a Geopolitical Forecasting Tournament

2014

The CHAMPS KNOW probabilistic reasoning training produced reliable accuracy gains (6–11% Brier score improvement) sustained across multiple tournament years. Comparison class use was the single most powerful component.

70/100

04Stakes

The Cost of Misplaced Certainty Across Four Domains

Overconfidence is not an abstract laboratory finding. It has a measurable cost in money, health, security, and democratic function, and the price scales with the confidence of the decision-maker.

01 System 01

Financial

Consistent with an overconfidence account, male investors traded 45% more than female investors and earned risk-adjusted returns 1.4 percentage points lower per year (overconfidence inferred from gender-linked trading patterns).[24] Overconfident CEOs were 65% more likely to make acquisitions, and markets punished them: −90 basis points at announcement versus −12 for rational CEOs.[27] The most active traders in Odean's study earned 11.4% annually versus a market average of 16.4%.[25]

In practice

portfolio churn, impulsive trades, chronic underperformance vs. index

02 System 02

Medical

Physicians claiming complete certainty were wrong approximately 40% of the time in autopsy-comparison studies (a figure specific to certainty-claiming diagnoses, not clinical judgment generally).[22] In Senegal, overconfident healthcare providers were 26% less likely to correctly manage patients and performed 18% fewer diagnostic actions.[36] High-confidence residents spent 12% less time per case with no accuracy advantage. Certainty compressed effort, not error.[37]

In practice

premature diagnostic closure, skipped differentials, unexamined confidence

03
System 03

Strategic / Geopolitical

Overconfident leaders systematically overestimate capabilities and underestimate adversaries, a pattern documented across the First World War, Vietnam, and the Cuban Missile Crisis.[29] The 2025 TNSR study confirmed the pattern in active national security professionals: the calibration gap at 90% confidence was 32 points.[31] The same bias that made junior analysts overconfident in cold reading tasks operates at the level of strategic military planning.

In practice

intelligence failures, planning overruns, escalation through misread signals

04 System 04

Information / Democratic

Three in four Americans overestimate their news discernment by 22 percentile points, and this overconfidence, not actual news literacy, predicts willingness to share false political content.[32] Low health-literacy individuals with high confidence showed 62% tobacco use versus 29% in the high-literacy group.[33] The Dunning-Kruger effect pattern means the people most confident in their media literacy are often the least equipped to exercise it.

In practice

uncritical content sharing, resistance to correction, false sense of expertise

05Protocol

A 4-Step Calibration Protocol for High-Stakes Decision-Making

Every step does the same thing at a mechanistic level: it reintroduces the feedback signal that professional environments typically remove. Overconfidence is the predictable output of a brain designed for rapid-feedback environments operating in a slow-feedback world.

The protocol, as a sequence.

Ongoing → Pre-decision → Estimation → Training

Ongoing 01 Prediction Tracking Pre-decision 02 Premortem Estimation 03 Reference ClassForecasting Training 04 Structured Debiasing
01 Step 01 · Ongoing

Prediction Tracking

Maintain a simple prediction log: write down what you expect, your confidence as a percentage, and your reasoning, then track actual outcomes.

Why

Overconfidence persists because feedback is absent in most professional domains. The only populations who achieve naturalistic calibration operate in rapid-feedback environments.[8] Tracking artificially closes the loop. Chang et al. found comparison class use combined with tracking was the most powerful component of the CHAMPS KNOW protocol.[39]

Common mistake

Tracking in vague terms ("I thought this would go well") defeats the purpose. Confidence must be quantified as a percentage before the outcome is known. Minimum 20–30 tracked predictions before calibration patterns become visible.[45]

02 Step 02 · Pre-decision

Premortem

Before finalising any major plan, assume it has already failed completely. Write down every plausible reason it failed.

Why

Veinott, Klein, and Wiggins found the premortem technique produced a 25-point reduction in plan confidence, outperforming standard critique, pro/con listing, and cons-only generation in a five-condition controlled experiment.[42]

Common mistake

Running the premortem after commitment, when social pressure makes generated concerns feel disloyal rather than informative.

03 Step 03 · Estimation

Reference Class Forecasting

When estimating costs, timelines, or success rates, find a reference class of comparable past projects and anchor to the empirical median before adjusting.

Why

Flyvbjerg validated reference class forecasting on 258 large infrastructure projects across 20 nations; it systematically reduces the planning fallacy.[43] Kahneman and Lovallo showed the "inside view" is the primary driver of planning overconfidence; the corrective is the outside view.[46]

Common mistake

Treating your current project as so unique that no reference class applies. The literature suggests this is almost always the inside view in disguise.

04 Step 04 · Training

Structured Debiasing

Complete a structured probabilistic reasoning training covering base rates, Bayesian updating, comparison classes, disconfirming evidence, and aggregated views.

Why

Morewedge's single-session debiasing training produced at least 31.94% bias reduction in the game condition, with effects persisting at two months.[40] Sellier et al. found trained participants were 19% less likely to choose the inferior solution on a real business case (the corrected figure per a published corrigendum that revised the original 29% estimate downward).[41]

31.94% Complete a structured probabilistic reasoning
Common mistake

Believing domain expertise eliminates the need for calibration training. Sanchez and Dunning confirm experts are uniformly overprecise regardless of domain.[23] A 2025 meta-analysis of 54 RCTs (N = 10,941) found small but significant debiasing effects (Hedges' g = 0.26).[49]

06Verdict

The verdict.

Bottom line

The brain will not stop generating certainty. The only question is whether you build the systems that test it before you act on it.

The overconfidence effect is not a pop-psychology curiosity. It is a precisely measured, neurally grounded, and universally replicated property of human cognition: a 30-point gap between how certain people feel and how often they are right, present in every high-stakes domain from medicine to military intelligence. The gap does not close with expertise, experience, or education. It closes only when the environment provides calibration feedback (rapid, unambiguous, and unavoidable) or when the individual builds the external structures that substitute for it. The science is clear: certainty is a feeling, not a signal, and treating it as a signal is the single most expensive cognitive error in professional life.

The reframing this evidence demands is simple but uncomfortable. Every professional who has ever said "I'm 90% sure" about a strategic judgment was probably operating with something closer to 58% accuracy, with no internal mechanism for detecting the difference.[31] That is not a failure of character. It is a design specification. The brain was built for environments where confidence was tested by consequences within hours, not years. The modern professional world removed the consequences and left the confidence.

The practical implication is equally direct. Calibration is a skill, not a trait. The evidence from Mellers' forecasting tournaments,[38] Morewedge's debiasing interventions,[40] and Sellier's field-transfer studies[41] converges on the same conclusion: structured training in probabilistic reasoning produces measurable improvement, and the improvement transfers to real-world decisions. The solution is not to feel less confident. It is to build the feedback loops, the premortems, the reference classes, and the prediction-tracking systems that translate the feeling of confidence into a testable, falsifiable commitment.

What Tetlock called epistemic humility is not a personality orientation. It is an engineering specification: a set of external structures that compensate for a known limitation in the brain's confidence architecture. The question was never whether we are overconfident. The question is what we build to correct for it.

The whole argument, one axis

Stated Certainty vs Actual Accuracy

0 25 50 75 100 percentage STATED CONFIDENCE 98% ACTUAL ACCURACY 68%
01Claim

The gap is systematic

The 30-point calibration gap is not noise, bad luck, or poor education. It is a stable architectural feature of human cognition, present in every population and domain measured across fifty years of research, from trivia questions to military intelligence.

Claim
02Consequence

Certainty suppresses correction

Overconfidence does not just produce wrong answers. It produces wrong answers that feel right and disable the error-detection systems that would catch them. Physicians order fewer tests. Investors check fewer sources. Planners skip the reference class.

Consequence
03Lever

Calibration is trainable

The IARPA forecasting tournaments, single-session debiasing interventions, and field-transfer studies converge: structured probabilistic reasoning training produces measurable, durable, and transferable accuracy gains. The correction is external scaffolding, not willpower.

Lever

Editorial confidence

High · 35 sources · 50 years of converging evidence · replicated across populations, domains, and methods · strongest support from prospective longitudinal data and randomised interventions

- 30 -

Put it to work

Where this science goes next on HPC

07Bibliography

The bibliography.

35 sources · ~5h est. corpus read · 35 visible

Meta · 1 Review · 4 Journal · 24 Book · 4 Chapter · 2
Type
Sort
  1. 01 Journal

    Judgment under uncertainty: Heuristics and biases

    doi: 10.1126/science.185.4157.1124
  2. 03 Book

    Gut feelings: The intelligence of the unconscious

  3. 04 Book

    Expert political judgment: How good is it? How can we know?

  4. 05 Book

    Superforecasting: The art and science of prediction

  5. 06 Review

    The trouble with overconfidence

    doi: 10.1037/0033-295X.115.2.502
  6. 07 Journal

    Knowing with certainty: The appropriateness of extreme confidence

    doi: 10.1037/0096-1523.3.4.552
  7. 08 Chapter

    Calibration of probabilities: The state of the art to 1980. In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under uncertainty: Heuristics and biases (pp. 306–334). Cambridge University Press

    doi: 10.1017/CBO9780511809477.019
  8. 09 Journal

    Are we all less risky and more skillful than our fellow drivers? Acta Psychologica, 47(2), 143–148

    doi: 10.1016/0001-6918(81)90005-6
  9. 15 Journal

    Neural correlates of metacognitive ability and of feeling confident: A large-scale fMRI study

    doi: 10.1093/scan/nsw093
  10. 16 Journal

    The neural basis of metacognitive ability

    doi: 10.1098/rstb.2011.0417
  11. 17 Journal

    How unrealistic optimism is maintained in the face of reality

    doi: 10.1038/nn.2949
  12. 18 Journal

    The optimism bias

    doi: 10.1016/j.cub.2011.10.030
  13. 22 Journal

    Overconfidence as a cause of diagnostic error in medicine

    doi: 10.1016/j.amjmed.2008.01.001
  14. 23 Review

    Are experts overconfident? An interdisciplinary review

  15. 24 Journal

    Boys will be boys: Gender, overconfidence, and common stock investment

    doi: 10.1162/003355301556400
  16. 25 Journal

    Volume, volatility, price, and profit when all traders are above average

    doi: 10.1111/0022-1082.00078
  17. 27 Journal

    Who makes acquisitions? CEO overconfidence and the market's reaction

    doi: 10.1016/j.jfineco.2007.07.002
  18. 28 Meta

    Overconfidence and financial decision-making: A meta-analysis

    doi: 10.1108/RBF-01-2020-0020
  19. 29 Book

    Overconfidence and war: The havoc and glory of positive illusions

    doi: 10.4159/9780674039162
  20. 31 Review

    Texas National Security Review / Good Judgment Project

    source
  21. 32 Journal

    Overconfidence in news judgments is associated with false news susceptibility

    doi: 10.1073/pnas.2019527118
  22. 33 Journal

    Overconfidence in managing health concerns: The Dunning-Kruger effect and health literacy

    doi: 10.1007/s10880-022-09895-4
  23. 36 Journal

    Overconfident health workers provide lower quality healthcare

    doi: 10.1016/j.joep.2019.102213
  24. 37 Journal

    Overconfidence, time-on-task, and medical errors: Is there a relationship? Advances in Medical Education and Practice, 15, 133–140

    doi: 10.2147/AMEP.S442689
  25. 38 Journal

    Psychological strategies for winning a geopolitical forecasting tournament

    doi: 10.1177/0956797614524255
  26. 39 Journal

    Developing expert political judgment: The impact of training and practice on judgmental accuracy in geopolitical forecasting tournaments

  27. 40 Journal

    Debiasing decisions: Improved decision making with a single training intervention

    doi: 10.1177/2372732215600886
  28. 41 Journal

    Debiasing training improves decision making in the field

    doi: 10.1177/0956797619861429
  29. 42 Journal

    Evaluating the effectiveness of the PreMortem technique on plan confidence

  30. 43 Journal

    From Nobel Prize to project management: Getting risks right

    doi: 10.1177/875697280603700302
  31. 45 Chapter

    Debiasing. In D. J. Koehler & N. Harvey (Eds.), Blackwell handbook of judgment and decision making (pp. 316–337). Blackwell

  32. 46 Journal

    Timid choices and bold forecasts: A cognitive perspective on risk taking

    doi: 10.1287/mnsc.39.1.17
  33. 47 Journal

    The effect of calibration training on the calibration of intelligence analysts' judgments

    doi: 10.1002/acp.4236
  34. 49 Journal

    Debiasing training reduces confirmation bias in national risk analysts

    doi: 10.1038/s41598-025-28794-w
  35. 51 Review

    The impact of cognitive biases on professionals' decision-making: A review of four occupational areas

    doi: 10.3389/fpsyg.2021.802439

THE HAND-OFF — Performance Scan: find your limiting factor. Three minutes, 24 questions, six systems — and your answers never leave your device.

High-Performance Insights