Section 1
The five challenges at a glance
Service-business founders are professional believers; conviction wins clients and rallies teams. But conviction applied to strategy under uncertainty produces big bets on untested assumptions, and the research record on untested assumptions is grim: at firms that measure rigorously, only 10-20% of expert-designed experiments produce positive results (Kohavi and Thomke, HBR 2017). Five challenges keep firms betting instead of testing. The HiPPO problem: the highest-paid person's opinion settles debates that data could settle cheaper and better. Big-bet bias: strategic moves arrive as all-or-nothing commitments, a full rebrand, a new office, a second service line built before a single paying test. The sample-size excuse: operators assume experimentation requires web-scale traffic, when the relevant discipline, falsifiable hypothesis, control condition, predefined success metric, scales down to ten prospects. Vanity-metric theater: tests get run but measured on metrics that cannot disconfirm anything. And the missing portfolio: experiments happen as one-offs rather than as a managed pipeline, so the firm never compounds what it learns. The table maps the challenges; the deep dives follow, anchored in Thomke's experimentation research and the largest startup A/B study available.
Section 2
Challenge one: the HiPPO problem and what testing giants learned
Stefan Thomke's two decades of Harvard Business School research on business experimentation reach an uncomfortable conclusion for confident leaders: expert intuition about what will work is wrong most of the time, and the only way to find out cheaply is to test. In 'The Surprising Power of Online Experiments' (HBR 2017), Thomke and Microsoft's Ron Kohavi reported that at Google and Bing, only about 10-20% of experiments generate positive results; the great majority of ideas, designed by skilled professionals, fail when measured against a control. Thomke's 2020 follow-up, 'Building a Culture of Experimentation' (HBR), profiles Booking.com, which runs roughly 25,000 tests a year and permits any employee to launch an experiment on millions of customers without management sign-off; even radical ideas from senior leaders must survive testing like everyone else's. The cultural prerequisite Thomke identifies is that data must trump opinions, which directly attacks the HiPPO dynamic, decisions defaulting to the highest-paid person's opinion. For a service firm, the lesson is not Booking.com's volume but its epistemology. When the firm's leaders debate whether productized pricing will convert better, or whether a niche repositioning will lift close rates, the honest answer is that nobody knows, the base rate says most confident guesses fail, and a two-week test on the next twenty prospects costs less than one wrong quarter.
Section 3
Challenge two: big-bet bias and the startup evidence
The strongest causal-leaning evidence that experimentation improves firm performance comes from Rembrand Koning (Harvard), Sharique Hasan (Duke), and Aaron Chatterji (Duke), published in Management Science as 'Experimentation and Start-up Performance: Evidence from A/B Testing' (2022). The team tracked more than 35,000 high-tech startups founded between 2008 and 2013, observing which adopted A/B testing tools. Adopters improved performance, measured by page views, by 30-100% after a year of use, and introduced new products at a 9-18% higher rate than non-experimenters. Strikingly, fewer than one in five firms used A/B testing at all, about a quarter in Silicon Valley versus 19% elsewhere, meaning a practice with measurable performance benefits sat unused by the large majority (Koning, Hasan, and Chatterji, 2022; Duke Fuqua, 2021). The researchers' interpretation matters for strategy: experimentation helped startups fail faster as well as scale faster, killing weak ideas before they consumed the company. That is the precise antidote to big-bet bias. A service firm considering a second service line typically faces a 100,000-dollar-plus commitment in salaries and attention. The experimental alternative: sell a pilot version to three existing clients at a defined price before hiring anyone. Either outcome of that test is cheap; either outcome of the untested bet is expensive. The startup evidence says firms that habitually choose the test outgrow the firms that habitually choose the bet.
Section 4
Challenge three: the sample-size excuse and small-firm experiment design
The most common objection from service-business operators is statistical: we have forty clients, not forty million sessions, so A/B testing cannot work here. The objection confuses the tool with the discipline. Thomke's research program covers experimentation well beyond high-traffic web testing, including field experiments in retail and service settings where samples are small and randomization is rough (Thomke, Experimentation Works, 2020). What transfers to a small firm is the experimental structure: a falsifiable hypothesis stated before the test, a comparison condition, a predefined success metric with a kill threshold, and a fixed time box. With small samples, firms trade statistical significance for decision relevance: a test on twenty prospects cannot detect a 3% lift, but it reliably detects the large effects that actually matter to strategy, whether a new offer converts at meaningfully higher rates, whether prospects will pay a premium price at all, whether a niche message books more calls. Practical small-firm designs include sequential testing (this month's pitch versus last month's baseline), split-batch outreach (two offers across two hundred cold contacts), pilot pricing (the next five proposals carry the new structure), and smoke tests (a landing page for the unbuilt service measuring booked calls). The standard is not academic publication; it is making fewer expensive mistakes than competitors who decide by debate. Against that standard, small-sample experiments clear the bar comfortably.
Section 5
Innovative solutions
The 2026 toolkit makes small-firm experimentation cheaper than at any point since the research was published. AI-accelerated test assets: landing pages, offer variants, and outreach sequences that once took a designer a week now take an afternoon, collapsing the cost side of every smoke test. Synthetic pre-tests: language models can simulate buyer objections to an offer before it reaches real prospects; the evidence base for treating these as substitutes for real tests is thin, so the disciplined use is filtering which hypotheses earn a real-world test, never replacing one. Pilot-as-product: rather than building a new service line, firms sell a fixed-scope, fixed-price pilot to three existing clients; the pilot is simultaneously revenue and an experiment with a built-in success metric. Pricing experiments through proposals: because service firms issue proposals serially, every batch of five proposals is a natural test cell for structure, anchoring, and packaging changes, an approach consistent with Thomke's argument that everyday business processes can become experimental instruments (HBR 2020). Public experiment scoreboards: posting live experiments and their kill criteria internally enforces the data-over-opinions norm and prevents quiet reinterpretation of failed tests. And reference-class kill rules: borrowing from the forecasting practices in this pillar, firms set kill thresholds using their own historical base rates, which removes the negotiation that usually rescues a sponsor's dying experiment. Together these lower the cost of disciplined testing below the cost of one unstructured strategy debate.
Section 6
Solution framework
A service firm can operationalize experimentation as strategy through four commitments. First, classify decisions by reversibility and cost before deciding how to decide. Cheap, reversible moves should simply ship; expensive, hard-to-reverse moves, rebrands, new lines, market entries, must first be decomposed into their riskiest assumptions, each phrased as a falsifiable hypothesis. This is the strategic core: big bets become sequences of small tests. Second, enforce the experiment contract. Every test gets a one-page charter before launch: hypothesis, metric, success threshold, kill threshold, sample, time box, owner. The charter is what separates experimentation from anecdote collection, and it institutionalizes Thomke's data-over-opinions norm (HBR 2020). Third, run a portfolio, not episodes. Maintain a ranked hypothesis backlog, keep two or three tests live, review them weekly inside the operating rhythm, and log every result in a learning repository. The Koning, Hasan, and Chatterji findings associate sustained adoption, not one-off use, with the 30-100% performance gains (2022). Fourth, protect the failure budget. If far more than half your tests succeed, you are testing things you already know; the Google and Bing base rate of 10-20% wins (Kohavi and Thomke, 2017) suggests genuine uncertainty produces mostly failures, and mostly-failures is what productive search looks like. Leadership's job is to make that failure rate culturally safe and financially survivable, which is exactly what small stakes accomplish.
Section 7
Evidence-based action plan
Week one: inventory your bets. List every initiative currently consuming money or senior attention, and mark which rest on untested assumptions about buyer behavior. Pick the single riskiest assumption in the portfolio. Week two: write your first experiment charter, hypothesis, metric, success and kill thresholds, time box, owner, and launch the cheapest test that could disconfirm the assumption: a split outreach batch, a pilot offer to three clients, a smoke-test landing page. Weeks three to six: run two to three charters in parallel and review status weekly in the operating rhythm. Expect failures; the measured base rate at elite testing organizations is 10-20% success (Kohavi and Thomke, HBR 2017), and your early hit rate teaching you the same lesson is the system working. Week seven: hold the first learning review, every completed test, what it killed, what it validated, what it spawned, and start the permanent learning log. Quarter two: tie experimentation to capital allocation with a graduation rule: no initiative above your materiality threshold receives full funding until its riskiest assumption has survived a charted test. That single rule converts the research, the 30-100% performance gains among testing adopters (Koning, Hasan, and Chatterji, 2022) and Thomke's cultural findings, into binding firm policy. Within two quarters the firm is no longer debating opinions; it is arbitrating evidence it generated itself. For adjacent evidence in this pillar, see [Cyber and Operational Resilience: The Rising-Risk Evidence Every Growing Firm Should Read](/blog/growth-cyber-operational-resilience-small-firms) and [AI as a Decision Partner: Evidence, Failure Modes, and the Operator's Protocol](/blog/growth-ai-decision-partner-operators-protocol).