Section 1
The five challenges at a glance
The first ten hires are the highest-variance decisions a founder makes, and the failure research explains why. Each early hire is 10-20% of total capacity, sets cultural defaults that outlast them, and is usually made under time pressure with no hiring infrastructure. The evidence identifies five recurring failure patterns. Founders hire for comfort, friends, family, people like themselves, which Wasserman's data flags as a stability risk that surfaces as co-founder and early-team conflict (Wasserman, 2012). They hire ahead of revenue, building payroll on projections; CB Insights post-mortems repeatedly pair team failure with running out of cash (CB Insights, 2021). They hire capacity when they need judgment, a mistake AI has made far more expensive, since field experiments show AI now performs much of the junior task work that used to justify early junior hires (Noy & Zhang, 2023; Brynjolfsson et al., 2023). They leave roles and equity ambiguous, the single most documented source of splintering teams in Wasserman's founder dataset. And they skip structured assessment entirely, despite mis-hire costs that SHRM benchmarks at $5,475 on average before accounting for the outsized disruption in a five-person firm (SHRM, 2025). The table maps each pattern; the next three sections analyze the most consequential.
Section 2
Challenge 1: The 65% claim, what the evidence actually supports
Citation integrity matters here, because this statistic drives real decisions. The 65% figure's lineage runs through Noam Wasserman, the Harvard Business School professor whose book The Founder's Dilemmas drew on data from nearly ten thousand founders. Wasserman has repeatedly cited research on VC-backed startup failures, including an oft-referenced 1989 study of venture portfolios, finding roughly 65% of failures among high-potential startups attributable to people problems rather than product, market, or technical causes (Wasserman, 2012; Entrepreneur, 2021). So the figure is real, but note its boundaries: it describes VC-backed, high-potential companies, the underlying portfolio study predates the modern startup era, and it is regularly misattributed to CB Insights. CB Insights' own analysis of 400-plus startup post-mortems found 'not the right team' cited in 23% of failures, third behind product-market fit and cash (CB Insights, 2021). The honest synthesis: across methodologies and decades, team and people factors consistently rank among the top three failure causes, somewhere between a quarter and two-thirds of cases depending on definition. For a service-business founder the precise number matters less than the asymmetry it reveals, founders systematically over-invest diligence in offers and pricing, and under-invest it in the ten decisions the failure literature says are most lethal. Treat each early hire with at least the rigor you give a major client contract.
Section 3
Challenge 2: Composition, complementarity beats likeness
Wasserman's central finding on team composition is uncomfortable: the choices that feel safest are statistically the riskiest. Founding and early teams built from friends, family, and homogeneous backgrounds feel stable but carry elevated conflict risk, because social ties suppress exactly the hard conversations, equity, roles, underperformance, that keep teams intact (Wasserman, 2012). The research-backed alternative is complementarity: early members whose skills, networks, and dispositions cover different territory, with explicit role boundaries. Jim Collins' large-sample research on enduring companies reached the same conclusion from the other direction, the ultimate constraint on growth is the ability to get and keep enough of the right people, prioritized even before strategy (Collins, 2001). What does complementarity mean concretely for a service firm's first ten? A workable composition pattern from the evidence: hires one through three should close the founder's biggest skill gap (typically delivery depth or sales, whichever the founder is not), not duplicate the founder's strength. Hires four through six professionalize the engine, operations, client management, quality. Hires seven through ten add leverage and optionality, which in 2026 increasingly means AI-workflow skill across every role rather than a dedicated 'AI person.' The Microsoft/LinkedIn finding that 71% of leaders prefer AI-skilled juniors over experienced non-users applies doubly in small teams, where each member's tooling habits propagate culturally (Microsoft/LinkedIn, 2024).
Section 4
Challenge 3: AI rewrites the hiring sequence itself
The classic first-ten playbook assumed a pyramid: a senior person or two, then juniors to absorb task volume. The field-experiment evidence undermines that template. Brynjolfsson, Li, and Raymond found AI assistance lifted novice worker productivity 34% precisely because the tool encodes and distributes the judgment of experienced workers (Brynjolfsson et al., 2023); Noy and Zhang found it compresses performance gaps between weaker and stronger professionals on writing tasks (Noy & Zhang, 2023). Read together for team design: AI is a substitute for junior capacity but a complement to senior judgment. The pyramid inverts. A 2026-era service firm's first ten should be judgment-dense, fewer people, each able to direct, verify, and correct AI-assisted work, rather than capacity-dense. This also changes what failure looks like. The old failure mode was overhiring juniors and drowning in management overhead; the new one is hiring judgment-light teams whose AI-assisted output looks plentiful but degrades silently without review. Practitioner data on AI-native startups, directional, not peer-reviewed, shows founding teams of two reaching meaningful revenue before hire five, with team sizes around twenty at valuations that previously required hundreds (Owyang, 2025). You need not match that; you should ask its question. Before each of your first ten hires: is this role buying judgment AI cannot supply, or capacity it increasingly can?
Section 5
Innovative solutions
Founders applying the research are converging on several practices. First, the pre-hire charter: a one-page document per early hire defining role boundaries, decision rights, success metrics at 90 days, and, for equity-bearing hires, vesting and exit terms, written before the search starts. This directly targets the ambiguity Wasserman's data identifies as the top splintering cause (Wasserman, 2012). Second, the judgment-or-capacity test from the previous section, run before opening any role. Third, dual work-sample screening, the same with-and-without-AI assessment Gartner predicts will reach 75% of hiring processes by 2027, scaled down to founder-friendly form: one paid sample, one live reasoning conversation (Gartner, 2025). Fourth, hiring fractionally before permanently: trying senior capability part-time for a quarter before committing a full-time seat, which converts the highest-stakes hires into reversible experiments (HBR, 2024). Fifth, reference checks focused on conflict behavior rather than competence, asking former colleagues how the person handled disagreement, feedback, and ambiguity, since those are the documented failure channels. None of this slows hiring much; a charter takes an afternoon and a work sample a week. Against a failure literature where people problems claim somewhere between 23% and 65% of companies, the hours are the cheapest insurance available to a growth-stage founder.
Section 6
Solution framework: the 10-seat map
Plan all ten seats before filling seat one. Draw a grid of ten positions in three tranches. Tranche one (hires 1-3, the complement tranche): close the founder's largest gap; every candidate scored on complementarity to existing strengths, not similarity. Tranche two (hires 4-6, the engine tranche): operations, client management, quality, roles that convert founder-dependent delivery into a system. Tranche three (hires 7-10, the leverage tranche): added only after revenue per FTE justifies them, with AI-workflow proficiency mandatory. Three rules govern the map. Rule one: every seat passes the judgment test, if AI plus an existing team member can cover 70% of the work, the seat is deferred and the budget moves to tooling or training, consistent with the substitution evidence (Brynjolfsson et al., 2023; Noy & Zhang, 2023). Rule two: every seat gets a charter before a candidate, neutralizing the ambiguity risk in the failure data (Wasserman, 2012). Rule three: no tranche is skipped, founders who jump to leverage hires before the engine exists create the judgment-light failure mode, and founders who over-build the engine before revenue recreate the cash-plus-team death spiral in the CB Insights post-mortems (CB Insights, 2021). Revisit the map quarterly; in the AI era seat definitions depreciate like skills do, the WEF estimates 39% of core skills change by 2030 (WEF, 2025).
Section 7
Evidence-based action plan
This quarter, run four steps. Step one (week 1): audit your existing team against the failure patterns. Any undocumented roles, unvested equity, unresolved role overlaps, or simmering conflicts are live instances of the 65%/23% risk, schedule the awkward conversations now, since Wasserman's data shows delay compounds them (Wasserman, 2012). Step two (week 2): build your 10-seat map, including seats already filled. Mark each as complement, engine, or leverage; many founders discover their first hires duplicated their own strengths, which tells you what the next seat must correct. Step three (weeks 3-4): for the next open seat, write the charter, run the judgment-or-capacity test, and design a paid work sample with both AI-assisted and unassisted components, per the screening evidence (Gartner, 2025). Step four (ongoing): instrument the team. Track revenue per FTE monthly and run a short quarterly health check on the documented conflict channels, role clarity, decision rights, fairness perceptions. Two honest caveats as you execute. The 65% figure describes VC-backed companies and may overstate people-failure rates for bootstrapped service firms, though no better service-business dataset exists. And the AI-substitution evidence covers specific task classes, writing, support, analysis, so apply the judgment test per task, not per job title. The companion articles on AI-proficiency screening and fractional leadership extend steps three and four. For adjacent evidence in this pillar, see [Reskilling the Existing Team: AI Training ROI and the Internal-Academy Approach](/blog/growth-reskilling-team-internal-academy) and [Fractional Everything: The Research on Part-Time Executives and the Borrowed C-Suite](/blog/growth-fractional-executives-model).