Who Predicts Well, and Why: The Evidence on Expert Forecasting
Essay. What separates experts who predict well from those who don't? A survey admitting peer-reviewed and academic-press sources only — September 2026.
Process note: this essay was produced with AI assistance. The AI surfaced candidate sources; the author directed which sources to admit, set the methodology, and challenged the drafts through iterative adversarial review before publication. All figures are dated and linked so readers can check them.
Source standard: peer-reviewed journals and academic presses only — no news, no social media, no vendor reports, no industry-funded work where funding bears on the claim. September 23, 2026.
An essay on what the evidence says about who predicts well, and why. It states no argument it cannot source: every factual claim is tied to a peer-reviewed or academic-press source, linked where a DOI or official page exists.
Abstract
What separates experts who predict the world well from those who don't? The evidence points to three conditions rather than a ranking of professions: how the forecaster thinks, what the task allows, and what the institution rewards. Eclectic, adaptive thinkers beat framework-bound ones, and training and teaming lift accuracy further. Weather forecasting shows what quantified feedback can do; clinical judgment loses to simple statistical rules where feedback is noisy; economists nail inflation while missing recessions — same profession, different tasks. And institutions shape outcomes: the IMF's groupthink missed 2008 while heterodox outsiders warned; prediction markets aggregate dispersed knowledge into prices; Moore's law worked because an industry coordinated around it. The throughline: forecasters do well where questions are well-defined, feedback arrives quickly, and methods adapt — and poorly where feedback is slow, questions are vague, and institutions reward confident performance over accuracy.
1. How experts think beats what experts know
Over roughly twenty years, Philip Tetlock collected 82,361 forecasts from 284 experts on political and economic outcomes (Tetlock, 2005, Princeton University Press; figures via the peer-reviewed review by Tschoegl & Armstrong, 2007). Borrowing Isaiah Berlin's categories, he sorted forecasters into hedgehogs — who view the world through one elegant overarching framework — and foxes — who draw on eclectic sources and revise. Hedgehogs performed at roughly the chance benchmark; foxes did significantly better, though neither beat simple statistical extrapolation rules. The chance benchmark itself was a design choice — uniform probability across researcher-partitioned outcome categories — and critics have contested it. Later work points the same way: the IARPA-backed Good Judgment Project found training, teaming, and habits of mind lifting accuracy (Mellers et al., 2014; Mellers et al., 2015), and The Forecasting Collaborative (2023, Nature Human Behaviour) found social scientists statistically indistinguishable from a lay crowd — credentials are no guarantee of accuracy. Technical knowledge alone did not improve judgmental forecasting accuracy either (Sanders & Ritzman, 1992, via Lawrence et al., 2006). Within forecasting, how you think predicts accuracy better than what you know.
2. The task sets the ceiling
No field shows the power of task structure better than weather forecasting. Numerical weather prediction improved not through a single breakthrough but through "steady accumulation of scientific knowledge and technological advances" — denser observations, better physical models, and ensemble methods that quantify uncertainty probabilistically (Bauer, Thorpe & Brunet, 2015, Nature). Forecast skill has risen decade after decade, and the authors judge the impact of weather prediction "among the greatest of any area of physical science." The task is ideal for learning: well-defined questions, scored daily, with errors feeding directly back into the models.
The general theory behind this pattern was stated by Kahneman & Klein (2009, American Psychologist): skilled judgment develops only in high-validity environments — where stable relationships connect cues to outcomes — combined with adequate opportunity to learn those regularities through practice with rapid, unequivocal feedback. Weather forecasting qualifies; stock picking and long-term political forecasting, they argue, are near zero-validity. Their joint verdict also explains a puzzle running through this literature: subjective confidence is no cue to accuracy — the "illusion of validity."
Where feedback is noisy and judgment is complex, experts often lose to their own data used mechanically. Dawes, Faust & Meehl (1989, Science) compared clinical judgment — combining predictors in the head — against actuarial rules combining the same predictors by formula. The formula won across domains: interpreting psychological test profiles, diagnosing progressive brain dysfunction, predicting cancer survival time. The experts' knowledge wasn't the problem; their integration of it was. Consistency beats brilliance when the signal is buried in noise.
Economics shows both sides of the task divide. Consensus forecasts missed 58 of 60 recessions a year in advance — "the record of failure to predict recessions is virtually unblemished" (Loungani, 2001, International Journal of Forecasting) — discrete turning points in a complex adaptive system. Yet professional survey forecasts of U.S. inflation beat standard statistical models out of sample (Ang, Bekaert & Wei, 2007, Journal of Monetary Economics). Same profession, different tasks, different accuracy.
Even small-scale estimation follows the task. Underestimation is more frequent than overestimation across judgment-based time predictions (Halkjelsvik & Jørgensen, 2012, Psychological Bulletin) — the planning fallacy (Buehler, Griffin & Ross, 1994). Where each task feels unique and feedback arrives too late to matter, even experts anchor on the best case.
3. Institutions decide what gets rewarded
Institutions also decide which forecasts get heard. The IMF's Independent Evaluation Office attributed the Fund's pre-2008 failure to "groupthink, intellectual capture, [and] a general mindset that a major financial crisis in large advanced economies was unlikely" (IEO, 2011). Meanwhile twelve analysts publicly warned of a housing-led recession — nearly all outside mainstream neoclassical economics (Bezemer, 2010). The accurate forecasts existed; the institutions weren't built to hear them.
Prediction markets are an institutional fix for exactly this problem: instead of polling experts, they let dispersed information aggregate into prices. A contract paying $1 if a candidate wins, trading at 53 cents, means the market believes the chance is about 53% (Arrow et al., 2008, Science). Such markets have forecasted elections, company sales, and product launch timing — working not because traders are brilliant but because the mechanism pays for being right.
Institutions can also manufacture accuracy, not just suppress it. Moore's law is the mirror image: a forecast that worked because an industry organized itself around it. The "law" was "a number of different laws that have replaced each other in succession," whose predictions "become the basis for future production goals," grounded in economic rather than technical reasoning (Mollick, 2006, IEEE Annals of the History of Computing). By 2003 more than 900 companies were working from biennial roadmaps with fifteen-year targets — "a self-fulfilling prophecy supported for three decades by interfirm cooperation and synchronized R&D" (Misa, 2019). Moore deliberately built the national and international roadmaps to "set the direction and cadence of innovation" (Lécuyer). The industry's own 2005 roadmap concedes the point from inside: it was "not so much a forecasting exercise as a way to indicate where research should focus." Moore's law did not predict the future; it coordinated it.
Extreme-ultraviolet lithography is that roadmap system made concrete. Because the 13.5 nm wavelength makes refractive lenses impossible, the industry coordinated for decades around all-reflective optics and tin-plasma sources — a development arc stretching back to the late 1980s (Wu & Kumar, 2014, Applied Physics Reviews; Tomie, 2012). A prediction made into a plan.
Conclusion: design for feedback, not prestige
Stripped of overclaim, the evidence leaves a practical lesson rather than a hierarchy. Accurate forecasting clusters where questions are well-defined, feedback arrives quickly, methods are adaptive, and institutions coordinate on targets instead of merely publishing them. Policy forecasters improve by thinking like foxes; weather services by quantifying uncertainty; clinicians by letting the formula integrate; economists by respecting what the task allows; institutions by rewarding accuracy over confidence — and by recognizing that some predictions succeed because they are made into plans. The discipline that predicts well may simply be the discipline whose world talks back fastest. The only ranking the evidence licenses is of the conditions under which anyone predicts well.
Glossary
Actuarial vs clinical judgment — combining predictors by fixed formula vs in the head; the formula wins where feedback is noisy (Dawes, Faust & Meehl, 1989).
Base rate — the background frequency of an outcome in the relevant population; neglecting it is a classic forecasting error.
Brier score — the mean squared error of probabilistic forecasts; 0 is perfect. The standard scoring rule for probability judgments.
Calibration — the match between stated probabilities and observed frequencies: of all events assigned 70%, about 70% should occur.
Ensemble forecasting — running many model versions with varied initial conditions to quantify uncertainty probabilistically, as in modern weather prediction.
Resolution (discrimination) — the ability to sort events into higher- vs lower-probability buckets; distinct from calibration.
Superforecaster — a top performer identified in the Good Judgment Project tournaments (Mellers et al.).
Validity (of an environment) — in Kahneman & Klein's sense, whether stable cue–outcome relationships exist to be learned; high-validity environments can produce skilled intuition, zero-validity ones cannot.
References
Ang, A., Bekaert, G., & Wei, M. (2007). Journal of Monetary Economics, 54(4), 1163–1212. doi.org/10.1016/j.jmoneco.2006.04.006
Arrow, K. J., et al. (2008). "The promise of prediction markets." Science, 320(5878), 877–878. doi.org/10.1126/science.1157679
Bauer, P., Thorpe, A., & Brunet, G. (2015). "The quiet revolution of numerical weather prediction." Nature, 525(7567), 47–55. doi.org/10.1038/nature14956
Bezemer, D. J. (2010). Accounting, Organizations and Society, 35(7), 676–688. doi.org/10.1016/j.aos.2010.07.002
Buehler, R., Griffin, D., & Ross, M. (1994). Journal of Personality and Social Psychology, 67(3), 366–381. doi.org/10.1037/0022-3514.67.3.366
Caplan, B. (2007). Critical Review, 19(1). (On the contestability of Tetlock's chance benchmark; no verified DOI.)
Dawes, R. M., Faust, D., & Meehl, P. E. (1989). "Clinical versus actuarial judgment." Science, 243(4899), 1668–1674. doi.org/10.1126/science.2648573
The Forecasting Collaborative (2023). Nature Human Behaviour. doi.org/10.1038/s41562-022-01517-1
Halkjelsvik, T., & Jørgensen, M. (2012). Psychological Bulletin, 138(2), 238–271. doi.org/10.1037/a0025996
IMF Independent Evaluation Office (2011). IMF Performance in the Run-Up to the Financial and Economic Crisis: IMF Surveillance in 2004–07. IEO evaluation page
Kahneman, D., & Klein, G. (2009). "Conditions for intuitive expertise: A failure to disagree." American Psychologist, 64(6), 515–526. doi.org/10.1037/a0016755
Lawrence, M., Goodwin, P., O'Connor, M., & Önkal, D. (2006). International Journal of Forecasting. doi.org/10.1016/j.ijforecast.2006.03.007 (reporting Sanders & Ritzman, 1992).
Lécuyer, C. (2022). "Driving Semiconductor Innovation: Moore's Law at Fairchild and Intel." Enterprise & Society, 23(1), 133–163. doi.org/10.1017/eso.2020.38
Loungani, P. (2001). "How accurate are private sector forecasts? Cross-country evidence from consensus forecasts of output growth." International Journal of Forecasting, 17(3), 419–432. doi.org/10.1016/S0169-2070(01)00098-X
Mellers, B., et al. (2014). Psychological Science. doi.org/10.1177/0956797614524255
Mellers, B., et al. (2015). Perspectives on Psychological Science. doi.org/10.1177/1745691615577794
Misa, T. J. (2019). HoST: Journal of the History of Science and Technology, 13(1), 106–109. doi.org/10.2478/host-2019-0005
Mollick, E. (2006). "Establishing Moore's Law." IEEE Annals of the History of Computing, 28(3), 62–75. doi.org/10.1109/MAHC.2006.45
Tetlock, P. E. (2005). Expert Political Judgment: How Good Is It? How Can We Know? Princeton University Press. (Reviewed in Tschoegl & Armstrong, 2007.)
Tomie, T. (2012). Journal of Micro/Nanolithography, MEMS, and MOEMS, 11(2), 021109. doi.org/10.1117/1.JMM.11.2.021109
Tschoegl, A. E., & Armstrong, J. S. (2007). International Journal of Forecasting, 23(2), 339–342. doi.org/10.1016/j.ijforecast.2007.02.002
Wu, B., & Kumar, A. (2014). Applied Physics Reviews, 1, 011104. doi.org/10.1063/1.4863412
Funding and conflict notes
Mellers et al. (2014, 2015) were sponsored by IARPA, the U.S. intelligence community's research arm.
Bauer, Thorpe & Brunet (2015): the authors were at ECMWF and Environment Canada, both publicly funded; no industry funding.
Wu & Kumar (2014) were Applied Materials employees; read as an industry-standpoint review.
Arrow et al. (2008) is a Science Policy Forum by academic economists; no industry funding bearing on the claims.