Section 1
Three different predictions, often confused for one
Sales teams say prediction and mean three separate jobs. The first is deal level: will this specific opportunity close, and when. The second is aggregate: how much will the period land at, across every deal in the funnel. The third is account level: which customers are drifting toward churn or ready to expand. They need different data and tolerate different error. Aggregate forecasts survive shaky individual calls, because errors cancel. Deal level scores do not: one wrong answer changes how a rep spends a week. Decide which of the three you are buying before anyone demos a product. Adjacent systems face the same split, which [How FinTech Is Leveraging AI Automation](/blog/how-fintech-is-leveraging-ai-automation) works through in a regulated context.
Section 2
The data problem arrives before the model does
A model learns from labelled history. Most CRMs do not have any. Stages are defined loosely, closed-lost reasons sit in a free text field that half the team skips, and the activity record is whatever the email plugin happened to capture. There is also a volume floor vendors rarely raise. A company closing forty deals a year does not have a training set. It has anecdotes. Those firms can still get value from aggregate forecasting, but a per deal probability learned from forty examples mostly reflects noise. Before any modelling, fix three things: one written definition per stage, a short mandatory closed-lost picklist, and automatic activity capture so the data does not depend on discipline.
Section 3
Where scores go to die
The common failure is not an inaccurate score. It is an accurate score that changes nothing. It appears in a dashboard, reps glance at it, and behaviour stays identical. A score is only operational once you have written down what a person does differently at 0.3 versus 0.7. Low probability plus high value goes to a manager for a rescue plan. Low probability plus low value gets closed out so the pipeline stops lying. High probability that has not moved in three weeks triggers a check on the champion. Put those rules inside the CRM, where the work happens. The forecast also has to be explained upward, which is where [Storytelling in the Age of AI and Automation](/blog/storytelling-in-the-age-of-ai-and-automation) is more relevant than it sounds.
Section 4
Run it in shadow mode first
Do not put a new model in front of a commission conversation on day one. Run it silently alongside the existing process for a full sales cycle, recording both predictions. At the end of the period, compare what the reps called, what the model called, and what actually happened. You will usually find the model beats reps on aggregate and loses to them on named accounts where the rep knows something the CRM does not. That result tells you how to deploy it: as a correction to the roll-up, not a replacement for judgement. Only after the comparison holds for two cycles should the score influence territory or deal review agendas.
Section 5
When a score touches someone's income
Sales prediction is unusual among automation projects because the subjects of the model are also its users, and they are paid on the outcome. That creates two risks most governance checklists miss. The first is gaming. If a score improves when an opportunity has more logged calls, calls will be logged. Any input a rep controls will drift, so prefer inputs from buyer behaviour rather than seller activity. The second is quiet disqualification. A lead scored low and routed to a slow queue never gets the chance to prove the score wrong, and the model then learns from its own decision. Keep a random holdout that bypasses scoring, tell the team how the scores work, and keep a named owner for anything that changes compensation.
Section 6
Calibration matters more than accuracy
Accuracy is the wrong headline metric. A model that calls everything lost will be accurate in a business with a low win rate and useless to everyone. Track calibration instead: of the deals scored around 70 percent, roughly seven in ten should close. Track period forecast error in both directions, because consistent overcalling and consistent undercalling have different fixes. Track how often the rep and the model disagreed and who was right. One review question keeps this honest. Did anyone change a decision because of the score this month, and did that decision turn out better than the one they would have made anyway.