Section 1
The five challenges at a glance
Every operator forecasts constantly, whether they admit it or not. A hiring plan is a forecast of demand; a cash runway model is a forecast of collections; a pricing change is a forecast of client behavior. The trouble is that most founders forecast the way pre-tournament experts did in Tetlock's earlier research: with confident narratives rather than calibrated probabilities, and without ever scoring themselves afterward. Five failure modes recur. Overconfident point estimates replace probability ranges, so plans break the moment reality deviates. Base-rate neglect leads founders to plan from the inside view, ignoring how similar projects historically performed. Belief inertia keeps stale assumptions alive long after disconfirming evidence arrives. The absence of feedback loops means forecasting skill never improves because accuracy is never measured. And volatility paralysis pushes operators to either stop planning entirely or to over-plan with false precision. The Good Judgment Project's findings address each of these directly, because the project was, in essence, a multi-year controlled study of what separates accurate forecasters from confident ones. The table maps the five challenges to causes, victims, and evidence; the sections that follow take the three most consequential ones apart and rebuild them as operator practices.
Section 2
Challenge one: point estimates and the case for probabilistic thinking
The Good Judgment Project emerged from a forecasting tournament run by IARPA, the US intelligence community's research agency, beginning in 2011. Across four years, roughly 500 questions, and over a million forecasts, the project led by Philip Tetlock and Barbara Mellers at the University of Pennsylvania won decisively, outperforming rival academic teams by 35% to 72% and beating professional intelligence analysts with access to classified information by roughly 30% (Good Judgment Project; Wikipedia, Good Judgment Project). The single most transferable finding for operators concerns granularity. Superforecasters did not say 'likely' or 'unlikely'; they distinguished 63% from 68%, and that precision was not false: when forecasts were scored with Brier scores, finer-grained probability use correlated with accuracy (Tetlock and Gardner, 2015). The business translation is direct. A revenue forecast of 'we will hit 1.4 million' is unfalsifiable theater. A forecast of '70% probability we land between 1.25 and 1.5 million, 15% above, 15% below' forces explicit assumptions, enables scenario-based cash planning, and can actually be scored later. Operators who convert their three biggest planning numbers, revenue, pipeline conversion, and hiring need, into probability ranges gain two assets immediately: honest conversations about downside cases, and a paper trail that turns next quarter's results into forecasting training data.
Section 3
Challenge two: base-rate neglect and the outside view
Tetlock's research found that accurate forecasters consistently begin with the outside view: before analyzing the specifics of a situation, they ask what usually happens in situations of this class (Tetlock and Gardner, 2015). This habit, rooted in Kahneman and Tversky's reference-class forecasting work, was a reliable discriminator between superforecasters and the merely confident. Most founders do the opposite. When launching a new service line, the inside view dominates: our team is strong, the early client conversations are warm, the deck is good. The outside view asks colder questions: what fraction of new service lines at firms like ours reach break-even within a year? How long did our last three offerings take to produce repeatable revenue? The inside view is vivid; the outside view is accurate. For service businesses, building base rates is cheaper than it sounds because the firm's own history is the best reference class available. Proposal win rates by segment, average days from proposal to signature, ramp time for new hires to full utilization, client churn by cohort: a spreadsheet with twenty historical observations beats intuition on every one of these questions. The GJP evidence suggests the sequencing matters: anchor on the base rate first, then adjust for what is genuinely different this time. Operators who reverse the order, starting from the special case, systematically overestimate speed and underestimate cost.
Section 4
Challenge three: belief inertia and the discipline of updating
Among the strongest behavioral findings from the Good Judgment Project is that superforecasters update their beliefs far more frequently than average forecasters, and in smaller increments (Tetlock and Gardner, 2015). They treat each forecast as a living estimate, nudged by every new piece of evidence, rather than a position to defend. Tetlock's summary has become the project's signature line: for superforecasters, beliefs are hypotheses to be tested, not treasures to be guarded. The research also found that training helped: forecasters given even a short course in probabilistic reasoning and cognitive debiasing showed measurable accuracy improvements in tournament conditions (Mellers et al., 2014, as documented by the Good Judgment Project research). Operators face a structural obstacle the tournament forecasters did not: their forecasts are often public commitments. A founder who announced a growth target to the team feels that revising it signals weakness, so the official forecast stays frozen while reality drifts. The fix is separating the two artifacts. Keep the public goal as a motivational target, but maintain a private rolling forecast, updated monthly or whenever material evidence arrives: a lost anchor client, a hiring miss, a sudden referral surge. Record each update with one line of reasoning. The update log does double duty: it improves the current forecast, and it becomes the curriculum for the firm's calibration reviews.
Section 5
Innovative solutions
Several practices adapt tournament-grade forecasting to the scale of a small firm. The internal forecasting tournament: quarterly, each leader independently submits probability forecasts on the same ten company questions before any discussion, then the team compares and aggregates. This replicates the independence-then-aggregation structure that drove GJP team performance (Mellers et al., 2014) and doubles as a noise audit on the leadership team's assumptions. Calibration dashboards: lightweight tools and spreadsheets now make Brier scoring trivial; the firm tracks each leader's calibration curve across quarters, making overconfidence visible and improvable rather than personal. Rolling forecasts over annual budgets: replacing the once-a-year budget theater with a four-quarter rolling forecast updated monthly imports the superforecaster habit of frequent small updates into the finance rhythm. Prediction-market-style polling: for firms above a dozen people, simple internal polls on operational questions, will this project ship on time, will this client renew, surface front-line information leadership does not have; the GJP found aggregated crowd estimates were a strong baseline even before selecting top forecasters. Finally, AI-assisted base-rate retrieval: language models are now adequate at assembling reference classes from a firm's own CRM and project history, provided the operator treats the output as a starting draft for verification, not as evidence. Each practice is cheap; together they amount to a forecasting capability most competitors lack entirely.
Section 6
Solution framework
The operator's forecasting system reduces to four components, each grounded in tournament evidence. Component one: question design. Convert vague aspirations into falsifiable, time-bound questions with probabilities. 'Grow the firm' becomes '75% probability of 20%+ revenue growth by Q2 close'. Bad questions cannot produce learning regardless of method. Component two: the outside view first. For every material forecast, write down the base rate before the special case: our historical win rate, ramp time, churn cohort data. Adjust from the base rate, and cap adjustments unless the evidence for 'this time is different' is concrete (Tetlock and Gardner, 2015). Component three: update cadence. Revisit the rolling forecast monthly inside the weekly operating rhythm's monthly financial review; small frequent updates, each with one recorded line of reasoning. Component four: scoring and review. Quarterly, score the previous quarter's logged predictions, compute simple calibration by probability band, and discuss the two worst misses without blame, asking what information was available and ignored. Teams should forecast independently before discussing, then aggregate, capturing the team effect documented in the GJP research (Mellers et al., 2014). None of this requires software beyond a spreadsheet. It requires the cultural move Tetlock identified as foundational: treating beliefs as hypotheses, and accuracy as a skill the firm deliberately trains.
Section 7
Evidence-based action plan
Month one: build the forecast log. Create a shared sheet with columns for question, probability, reasoning, date, and outcome. Seed it with ten falsifiable predictions about the coming quarter spanning revenue, pipeline, delivery, and hiring. Require every leadership team member to submit probabilities independently before the numbers are discussed. Month two: assemble base rates. Pull the firm's own history: proposal win rates by service line, time-to-cash by client type, new-hire ramp times, churn by cohort. One afternoon of CRM archaeology yields reference classes that will anchor every future plan. Convert the annual budget into a four-quarter rolling forecast updated at the monthly financial review. Month three: first scoring session. Score expired predictions, chart calibration by band, and run a blameless review of the two largest misses. Expect overconfidence in the high-probability bands; the GJP training research suggests even brief calibration feedback measurably improves subsequent accuracy (Mellers et al., 2014). Quarter two onward: institutionalize. Fold the forecasting tournament into quarterly planning, track each leader's calibration trend, and start using probability ranges rather than point estimates in board and bank conversations. Within a year the firm owns something rare among small businesses: a measured, improving forecasting capability, and a leadership team that knows the difference between what it believes and what it knows. For adjacent evidence in this pillar, see [The Premortem and Decision Hygiene: Klein and Kahneman's Protocols for Small Firms](/blog/growth-premortem-decision-hygiene-small-firms) and [Experimentation as Strategy: The Evidence for Cheap Tests Before Big Bets](/blog/growth-experimentation-as-strategy-cheap-tests).