Our AI pilot or product is not delivering. Fix, scale or drop it?

Find out why it stopped short before you decide on the product, the vendor or the team. AI efforts usually stall on an undefined measure, unreliable data, no business owner or no path to production. Fix those first, then decide on evidence from one well-run use case.

Draft for Tamir's review. Not published.

Tamir Khason · Updated · Decision guides

Find out where it stopped

For each pilot, model or tool, write who asked for it, who would use it, and where it stopped and why. Ask whether a success measure was agreed before it started, and get the results from both sides.

If users stopped trusting the answers, collect the answers that broke their trust and classify each failure. Retrieval errors, stale documents and bad prompts each need a different fix. If the source documents were never good enough to answer from, a rebuild will not help.

Measure output rather than usage

High usage measures curiosity. Pick three teams with measurable output, such as tickets closed or time to respond, and measure before and after with the tools used in a defined way. Keep the licenses where the output moved.

A pilot that worked in one team often had a motivated team, clean data and close attention. A rollout gets none of those by default. Recreate the pilot's conditions in two teams and measure again before spreading further.

Depending on your seat

If you're on the board, require each pilot to name the decision it is waiting to resolve and an owner who can adopt the result, or close it. If you learn the product relied on manual work behind its AI claims, establish the facts first and decide on disclosure with counsel.

If you run the product or the data team, define one use case with a business owner, a deployment path and a measure. Decide on the team after it. If a pilot succeeded and there is no budget, state its result in hours or cases, and look for funding in the department that gains.

What to check before you decide

  • For each pilot or model, write who asked for it, who would use it and where it stopped.
  • Ask whether a success measure was agreed before the pilot, and compare your results with the customer's or the users'.
  • Classify the wrong answers that broke trust: retrieval, stale documents, prompts or data.
  • Measure one output per team before and after, and report that to the board instead of usage.
  • Check what deployment needs beyond the license: integration, training and support.
  • Name a business owner with authority to adopt the result, or close the pilot.
  • Add each failure type you find to the evaluation set and the release checklist.

Questions people ask

We found out the AI product relied on manual work behind the claims in the pitch, what should the board do?

Establish the facts first: what the product does automatically, what is manual, what was claimed and to whom, including customers. Then decide on disclosure, on the plan to close the gap and on the founders, with counsel. It depends on how far the claims went, whether customers were told the same and whether the gap can be closed.

Should the health board close innovation pilots that never reach a decision?

Require each pilot to resolve a named decision or close. Continuation depends on a live decision owner, remaining uncertainty, and evidence that another activity can change the answer.

Our enterprise pilot failed on accuracy and the team blames the customer's data, how do I know what really happened?

Get the pilot's measure and results from both sides, because pilots fail on undefined success as often as on accuracy. If the customer's data was the problem, the pilot should have found that in week one. Fix the pilot process first: agreed measure, data check at the start, weekly results. It depends on whether a success measure was agreed before the pilot and on what the results actually showed.

We gave everyone AI tools, usage is high, but I can't see any business result, keep paying or cut?

Pick three teams and measure one output each, before and after, with the tools used in a defined way; usage alone measures curiosity. Keep the licenses where the output moved and change how the tools are used where it did not. It depends on which work the tools actually touch and on whether anyone changed a process around them.

Our AI pilot succeeded with clinicians but there's no budget to deploy, what are my options?

Turn the pilot's result into a number the budget holders understand, such as hours saved per week or cases handled, and look for funding in the department that gains rather than in IT. Also ask the vendor for a staged deployment that pays as it spreads. It depends on what the pilot measured and on whether the gain can be shown in money or capacity.

A customer posted our AI product's embarrassing wrong answer and left, what do we change and what do we tell prospects?

Find the failure type, fix that class rather than the one answer, and add the safeguard that would have caught it, such as topic limits, review for sensitive outputs or confidence thresholds. Then tell prospects what changed, in specifics. It depends on whether the failure was a known weakness or a new one.

Our AI pilot worked in one team and the rollout to everyone showed nothing, why and what now?

Pilots succeed with a motivated team, clean data and attention, and rollouts get none of those by default. Find which of the three the other teams lacked and fix that in two teams before any further rollout. It depends on whether the pilot's conditions were documented and on how different the other teams' work and data are.

Should we keep expanding automated outreach when sales conversations get worse?

Judge outreach by useful conversations and adverse effects. Reply volume alone says little. Expansion depends on targeting, message accuracy, escalation, and whether automation is worsening the relationship with prospects.

Our internal RAG assistant failed because users stopped trusting it, should we rebuild or drop it?

Find out which answers were wrong and why before deciding, because retrieval failures, stale documents and bad prompts each need different fixes. Rebuild only if the failing cause is fixable with your data and ownership. It depends on whether the source documents were ever good enough to answer from.

Our data science team built models that never reached production, what should change before we decide on the team?

Find out why each model stopped: no business owner, no engineering path to production, data that was never reliable, or a problem nobody needed solved. The fix is usually in ownership and engineering, and the team's skills are the smaller issue. It depends on whether business owners exist for the use cases and on whether engineering can deploy and run a model.

Should we renew an appointment bot that cannot hand requests to staff?

Judge renewal using completed service outcomes and recoverable handoffs. Suitability depends on whether staff receive context, ownership is clear, and failed transfers remain visible.

How I can help with this decision

Ask or talk (Free)
I give my view on why efforts like yours usually stop short, and the per-pilot or per-model review that shows whether yours is fixable.
Review (Pay if it was worth it)
I write an independent post-mortem on the pilot, product or tools and their conditions for scaling. I recommend to fix, narrow, scale or stop, and what to tell the board.
Retain (When it makes sense)
I stay close through the next pilot or the first production use case to review the measure, the data checks and the results.