Section 1
Ten questions that separate vendors quickly
Ask these early, in writing, and note how fast the answers come back. What model sits underneath, and what happens when the provider updates it. Is your data used to train their models. How long are prompts and outputs retained, and where. Who are the sub-processors, and do you get notice before that list changes. In which region does processing happen. What is the uptime record, with incident history rather than a marketing number. Who owns the prompts and configurations you create. How does pricing behave as volume grows. What can you export on exit, in what format. Which of your systems does it write to, with what permissions. A vendor with production customers answers these quickly. Hesitation on retention and export is the most informative signal in the set.
Section 2
Design the pilot to produce a decision
Most pilots are structured so that no outcome can lose. Fix that first. Define the workflow, not the technology. Agree a pass mark and a measurement method in writing before the pilot begins. Use your own data, including the difficult cases, not their sample set. Set a fixed end date. Nominate who decides and on what evidence. Run it against a baseline. Without a measured before, the comparison becomes a conversation about impressions, and impressions favour whoever has been in the room most. Pay for the pilot where you can. A free pilot costs you priority when something breaks and removes your standing to demand a fix.
Section 3
Where the vendor makes money, and how it shapes their advice
This is not cynicism. It is a reading instruction. Per seat pricing produces advice about rolling out widely. Per task or per token pricing produces advice about automating more steps. Implementation services attached to the licence produce advice about a longer build. A vendor whose renewal depends on measured outcomes will push you toward a narrow first use case, because they need it to work. Ask directly how they are compensated and what renewal depends on. The answer tells you which recommendations to weight and which to discount, and whether their incentive is aligned with your outcome or with your usage. Ask it of anyone advising you, consultants included.
Section 4
Reference calls, asked properly
Vendor-supplied references are selected, so ask questions their selection cannot control. What broke, and how long did it take to fix. What did you have to build yourself that you expected to be included. How long from signing to production use. Who inside your company dislikes it and why. What will your renewal decision hinge on. Ask for a reference at your size and stage, not their flagship enterprise logo. The experience of a twenty person company is not the experience of an account with a dedicated success manager. And ask about churned customers. A vendor who will discuss why one left, without defensiveness, is usually worth working with.
Section 5
Contract terms that matter more than the price
NIST frames AI risk management around trustworthiness, design, evaluation, and use, and a contract is where those become obligations rather than intentions. Paper the following: no training on your data unless explicitly agreed, a defined retention period with deletion on termination, a sub-processor list with notice before changes, processing region, incident notification within a stated window, export of your data and configuration in a usable format, and a term that does not auto-renew for a long period without a window to leave. Watch the pricing structure as carefully as the number. Minimum commitments, per-seat floors, and usage cliffs turn a reasonable quote into an unreasonable bill in year two. Negotiate the second-year price before signing the first, because that is the only moment in the relationship when your position is strong.
Section 6
Deciding, and knowing when to walk
Score candidates on outcome fit, integration depth with what you already run, data terms, exit cost, and total cost at projected volume rather than current volume. Weight integration heavily. A product that does not reach your systems does not reach your business. Track after signing: time to first production use, the share of the promised workflow actually running, support response against what was promised, cost per unit of work, and internal engineering time consumed against the estimate. You are ready to sign if a paid pilot beat a measured baseline on a metric you set in advance, the data terms are in writing, and you know what leaving costs. You are not ready if the case rests on a demo, a case study from a company unlike yours, and a discount expiring this quarter. See also [Startup Success: How AI Automation Transformed Our Business](/blog/startup-success-how-ai-automation-transformed-our-business), [Using AI and Data Analytics to Enhance Storytelling](/blog/using-ai-and-data-analytics-to-enhance-storytelling), and [Top AI Automation Tools for Startups in 2026](/blog/top-ai-automation-tools-for-startups-in-2026).