Business Growth

Experimentation as Strategy: The Evidence for Cheap Tests Before Big Bets

When researchers tracked more than 35,000 startups, the ones that adopted A/B testing improved performance by 30-100% within a year and shipped more products, yet fewer than one in five firms used it at all (Koning, Hasan, and Chatterji, Management Science 2022). Meanwhile, the organizations that test most rigorously report the most humbling base rate in business: only 10-20% of expert-designed experiments at Google and Bing beat the control (Kohavi and Thomke, HBR 2017). Both findings point the same direction: confident strategic intuition is usually wrong, and cheap tests are the rational precursor to expensive bets. This article reviews Stefan Thomke's experimentation research and the startup evidence, then shows how a service business without web-scale traffic can still make testing its default strategic instrument.

Joshua Agonya Pi'Rwot

By Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator

Executive summary

Startups that adopted A/B testing grew page views 30-100% within a year, yet fewer than one in five used it. Here is the research case for building an experimentation culture before committing capital to big strategic bets.

Section 1

The five challenges at a glance

Service-business founders are professional believers; conviction wins clients and rallies teams. But conviction applied to strategy under uncertainty produces big bets on untested assumptions, and the research record on untested assumptions is grim: at firms that measure rigorously, only 10-20% of expert-designed experiments produce positive results (Kohavi and Thomke, HBR 2017). Five challenges keep firms betting instead of testing. The HiPPO problem: the highest-paid person's opinion settles debates that data could settle cheaper and better. Big-bet bias: strategic moves arrive as all-or-nothing commitments, a full rebrand, a new office, a second service line built before a single paying test. The sample-size excuse: operators assume experimentation requires web-scale traffic, when the relevant discipline, falsifiable hypothesis, control condition, predefined success metric, scales down to ten prospects. Vanity-metric theater: tests get run but measured on metrics that cannot disconfirm anything. And the missing portfolio: experiments happen as one-offs rather than as a managed pipeline, so the firm never compounds what it learns. The table maps the challenges; the deep dives follow, anchored in Thomke's experimentation research and the largest startup A/B study available.

Section 2

Challenge one: the HiPPO problem and what testing giants learned

Stefan Thomke's two decades of Harvard Business School research on business experimentation reach an uncomfortable conclusion for confident leaders: expert intuition about what will work is wrong most of the time, and the only way to find out cheaply is to test. In 'The Surprising Power of Online Experiments' (HBR 2017), Thomke and Microsoft's Ron Kohavi reported that at Google and Bing, only about 10-20% of experiments generate positive results; the great majority of ideas, designed by skilled professionals, fail when measured against a control. Thomke's 2020 follow-up, 'Building a Culture of Experimentation' (HBR), profiles Booking.com, which runs roughly 25,000 tests a year and permits any employee to launch an experiment on millions of customers without management sign-off; even radical ideas from senior leaders must survive testing like everyone else's. The cultural prerequisite Thomke identifies is that data must trump opinions, which directly attacks the HiPPO dynamic, decisions defaulting to the highest-paid person's opinion. For a service firm, the lesson is not Booking.com's volume but its epistemology. When the firm's leaders debate whether productized pricing will convert better, or whether a niche repositioning will lift close rates, the honest answer is that nobody knows, the base rate says most confident guesses fail, and a two-week test on the next twenty prospects costs less than one wrong quarter.

Section 3

Challenge two: big-bet bias and the startup evidence

The strongest causal-leaning evidence that experimentation improves firm performance comes from Rembrand Koning (Harvard), Sharique Hasan (Duke), and Aaron Chatterji (Duke), published in Management Science as 'Experimentation and Start-up Performance: Evidence from A/B Testing' (2022). The team tracked more than 35,000 high-tech startups founded between 2008 and 2013, observing which adopted A/B testing tools. Adopters improved performance, measured by page views, by 30-100% after a year of use, and introduced new products at a 9-18% higher rate than non-experimenters. Strikingly, fewer than one in five firms used A/B testing at all, about a quarter in Silicon Valley versus 19% elsewhere, meaning a practice with measurable performance benefits sat unused by the large majority (Koning, Hasan, and Chatterji, 2022; Duke Fuqua, 2021). The researchers' interpretation matters for strategy: experimentation helped startups fail faster as well as scale faster, killing weak ideas before they consumed the company. That is the precise antidote to big-bet bias. A service firm considering a second service line typically faces a 100,000-dollar-plus commitment in salaries and attention. The experimental alternative: sell a pilot version to three existing clients at a defined price before hiring anyone. Either outcome of that test is cheap; either outcome of the untested bet is expensive. The startup evidence says firms that habitually choose the test outgrow the firms that habitually choose the bet.

Section 4

Challenge three: the sample-size excuse and small-firm experiment design

The most common objection from service-business operators is statistical: we have forty clients, not forty million sessions, so A/B testing cannot work here. The objection confuses the tool with the discipline. Thomke's research program covers experimentation well beyond high-traffic web testing, including field experiments in retail and service settings where samples are small and randomization is rough (Thomke, Experimentation Works, 2020). What transfers to a small firm is the experimental structure: a falsifiable hypothesis stated before the test, a comparison condition, a predefined success metric with a kill threshold, and a fixed time box. With small samples, firms trade statistical significance for decision relevance: a test on twenty prospects cannot detect a 3% lift, but it reliably detects the large effects that actually matter to strategy, whether a new offer converts at meaningfully higher rates, whether prospects will pay a premium price at all, whether a niche message books more calls. Practical small-firm designs include sequential testing (this month's pitch versus last month's baseline), split-batch outreach (two offers across two hundred cold contacts), pilot pricing (the next five proposals carry the new structure), and smoke tests (a landing page for the unbuilt service measuring booked calls). The standard is not academic publication; it is making fewer expensive mistakes than competitors who decide by debate. Against that standard, small-sample experiments clear the bar comfortably.

Section 5

Innovative solutions

The 2026 toolkit makes small-firm experimentation cheaper than at any point since the research was published. AI-accelerated test assets: landing pages, offer variants, and outreach sequences that once took a designer a week now take an afternoon, collapsing the cost side of every smoke test. Synthetic pre-tests: language models can simulate buyer objections to an offer before it reaches real prospects; the evidence base for treating these as substitutes for real tests is thin, so the disciplined use is filtering which hypotheses earn a real-world test, never replacing one. Pilot-as-product: rather than building a new service line, firms sell a fixed-scope, fixed-price pilot to three existing clients; the pilot is simultaneously revenue and an experiment with a built-in success metric. Pricing experiments through proposals: because service firms issue proposals serially, every batch of five proposals is a natural test cell for structure, anchoring, and packaging changes, an approach consistent with Thomke's argument that everyday business processes can become experimental instruments (HBR 2020). Public experiment scoreboards: posting live experiments and their kill criteria internally enforces the data-over-opinions norm and prevents quiet reinterpretation of failed tests. And reference-class kill rules: borrowing from the forecasting practices in this pillar, firms set kill thresholds using their own historical base rates, which removes the negotiation that usually rescues a sponsor's dying experiment. Together these lower the cost of disciplined testing below the cost of one unstructured strategy debate.

Section 6

Solution framework

A service firm can operationalize experimentation as strategy through four commitments. First, classify decisions by reversibility and cost before deciding how to decide. Cheap, reversible moves should simply ship; expensive, hard-to-reverse moves, rebrands, new lines, market entries, must first be decomposed into their riskiest assumptions, each phrased as a falsifiable hypothesis. This is the strategic core: big bets become sequences of small tests. Second, enforce the experiment contract. Every test gets a one-page charter before launch: hypothesis, metric, success threshold, kill threshold, sample, time box, owner. The charter is what separates experimentation from anecdote collection, and it institutionalizes Thomke's data-over-opinions norm (HBR 2020). Third, run a portfolio, not episodes. Maintain a ranked hypothesis backlog, keep two or three tests live, review them weekly inside the operating rhythm, and log every result in a learning repository. The Koning, Hasan, and Chatterji findings associate sustained adoption, not one-off use, with the 30-100% performance gains (2022). Fourth, protect the failure budget. If far more than half your tests succeed, you are testing things you already know; the Google and Bing base rate of 10-20% wins (Kohavi and Thomke, 2017) suggests genuine uncertainty produces mostly failures, and mostly-failures is what productive search looks like. Leadership's job is to make that failure rate culturally safe and financially survivable, which is exactly what small stakes accomplish.

Section 7

Evidence-based action plan

Week one: inventory your bets. List every initiative currently consuming money or senior attention, and mark which rest on untested assumptions about buyer behavior. Pick the single riskiest assumption in the portfolio. Week two: write your first experiment charter, hypothesis, metric, success and kill thresholds, time box, owner, and launch the cheapest test that could disconfirm the assumption: a split outreach batch, a pilot offer to three clients, a smoke-test landing page. Weeks three to six: run two to three charters in parallel and review status weekly in the operating rhythm. Expect failures; the measured base rate at elite testing organizations is 10-20% success (Kohavi and Thomke, HBR 2017), and your early hit rate teaching you the same lesson is the system working. Week seven: hold the first learning review, every completed test, what it killed, what it validated, what it spawned, and start the permanent learning log. Quarter two: tie experimentation to capital allocation with a graduation rule: no initiative above your materiality threshold receives full funding until its riskiest assumption has survived a charted test. That single rule converts the research, the 30-100% performance gains among testing adopters (Koning, Hasan, and Chatterji, 2022) and Thomke's cultural findings, into binding firm policy. Within two quarters the firm is no longer debating opinions; it is arbitrating evidence it generated itself. For adjacent evidence in this pillar, see [Cyber and Operational Resilience: The Rising-Risk Evidence Every Growing Firm Should Read](/blog/growth-cyber-operational-resilience-small-firms) and [AI as a Decision Partner: Evidence, Failure Modes, and the Operator's Protocol](/blog/growth-ai-decision-partner-operators-protocol).

FAQ

Direct answers for operators.

Can a service business with a small client base really run experiments?

Yes, by keeping the discipline and adapting the instrument. A falsifiable hypothesis, a comparison condition, a predefined success metric, and a time box all work at small scale: split outreach batches, pilot offers to a handful of clients, sequential month-over-month pitch tests, and smoke-test landing pages. Small samples cannot detect tiny effects, but strategy decisions hinge on large effects, which small tests detect reliably.

What did the 35,000-startup A/B testing study actually find?

Koning, Hasan, and Chatterji (Management Science, 2022) tracked high-tech startups founded 2008-2013 and found A/B testing adopters improved performance, measured by page views, 30-100% after a year, and introduced new products at a 9-18% higher rate. Fewer than 20% of firms used testing at all. The authors note experimentation helped firms both fail faster on weak ideas and scale winners.

Why do most experiments fail, and is that a problem?

Kohavi and Thomke (HBR 2017) report only 10-20% of experiments at Google and Bing produce positive results, despite being designed by skilled professionals. That failure rate is not waste; it is the measured cost of searching an uncertain landscape, and it is exactly why testing beats betting: each failure costs a test's budget instead of a big bet's budget. A high success rate usually means you are testing the obvious.

How do we stop the boss's opinion from overriding test results?

Pre-commitment is the mechanism. Thomke's culture research (HBR 2020) shows experimentation works only when data trumps opinions, so write the success and kill thresholds into a one-page charter before launch, with the sponsor's signature. At Booking.com even senior leaders' ideas must survive testing. In a small firm, the equivalent norm is that charters are arbitrated by their pre-agreed metrics, never renegotiated after results arrive.

Joshua Agonya Pi'Rwot

Written by

Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator · Country Director, AVODA Group Uganda · EMBA

Joshua helps service-business operators turn scattered marketing into a clear path from first attention to booked call. He is Founder of Business Growth Accelerator and Country Director of AVODA Group Uganda.