Draft for Tamir's review. Not published.
Find out what kind of error it is
Take a month of errors, alerts or workarounds and group them by cause. Separate design gaps from bad master data, untrained users and expected events such as a defrost cycle. A few causes usually drive most of the manual work.
Trace a complete transaction or order from the customer's promise to the final record. Where two systems disagree, decide which one is authoritative for each business event and who owns corrections. Workarounds that still exist months after go-live are usually design gaps.
Cap, fix and set the rule for resuming
Count what the manual process can handle correctly per month and cap growth there. Ask the vendor whether the fix is configuration, a change request or a new module, with a date. Each has a different timeline.
Check every manual output issued so far before customers or the regulator find the errors. Tell waiting customers what they will get and when. Agree a measure, such as alert quality or exceptions you can explain, before the rollout resumes.
Depending on your seat
If you're the CEO, the choice is rarely all or nothing. Compare restricted operation, a controlled restart and replacement. Have the customer, finance and legal owners agree how disputed accounts are treated before the next run.
If you run the rollout, let the next site walk through its most critical processes on the current design and record where it fails. Check what the contract says about defects after go-live and who pays for the fix. Set a rule that new offers are checked against the systems before launch.
What to check before you decide
- Classify a month of errors, alerts or workarounds by cause.
- Separate design gaps from data, training and expected-event causes.
- Count what the manual process can handle correctly and cap growth at that level.
- Ask the vendor whether the fix is configuration, a change request or a new module, and by what date.
- Check every manual output issued so far for errors.
- Define which system is authoritative for each business event and who owns corrections.
- Agree the measure that must hold before the rollout resumes.
Questions people ask
We launched a new energy offer and our billing system can't handle it, pause sales or keep going while IT fixes it?
Cap sales at what the manual process can bill correctly, and give IT a dated plan for the tariff in the system; a wrong bill costs more than a delayed sale. Find out whether the fix is configuration, a vendor change or a new module, because each has a different timeline. It depends on how many accounts the manual process can carry and on what the billing vendor says the change takes.
Should we expand online food sales while our systems promise unavailable stock?
Fix the customer promise boundary before pursuing more demand. The remedy depends on available-to-sell rules, stock ownership, and when the business can confirm fulfilment with certainty.
Should we fund more insurance portal features while staff still reenter customer information?
Trace the complete transaction before approving more portal features. The remedy depends on decision ownership, information compatibility, and whether existing controls require duplicate entry or merely inherited practice.
Should we expand finance automation after errors require manual investigation?
Expand only when exceptions remain understandable and manageable. The decision depends on error causes, correction authority, and whether routine savings exceed the operating burden created.
Should we keep automated billing after errors reach a large number of customers?
Decide using the established cause, correction evidence, and controls for future billing. It depends on whether the failure is contained, affected accounts can be reconciled, and responsible owners approve the remaining risk.
ERP rollout stuck after first plant go live with workarounds, should we push plant two or fix the template first?
Workarounds that still exist months after go-live are usually design gaps, and copying them to a second plant doubles them. Count the workarounds, group them by cause and fix the few that drive most of the manual work. It depends on whether the gaps are in the template or in the first plant's data and training.
Cold chain monitoring rollout halfway and operations muted the alerts because most are false, what do we do?
Stop adding sensors until the alerts mean something, because a muted system is worse than none in an audit. Look at a month of alerts, group them by cause and change thresholds, sensor placement or escalation for each cause. It depends on whether the false alerts come from thresholds, from door and defrost cycles or from sensor placement.
Should we pause the factory rollout when ERP and production records disagree?
Pause the affected acceptance decision until the systems agree on business meaning and correction ownership. It depends on whether differences concern production, quality release, financial posting, or recoverable reporting delays.
How I can help with this decision
- Ask or talk (Free)
- I give my view on whether this is a design problem or a bad start, and how to size the cap. I tell you which grouping exercise or vendor question sets the real timeline.
- Review (Pay if it was worth it)
- I write an independent assessment of the errors, their causes, the vendor's position and your exposure. I recommend whether to proceed, cap, fix first, change scope or stop, and the measure for resuming.
- Retain (When it makes sense)
- I stay close through the fix to review readiness at each new site and the vendor's delivery before the cap is lifted.