Draft for Tamir's review. Not published.
Set the boundary before the tool
List the decisions the AI would make and the information it would only present. A note an agent edits is different from an entry in a medical record. Advice to an operator is different from triggering an action.
Your compliance, legal, clinical or process safety owners confirm the permitted use for your sector. Keep the safety systems and the human decision unchanged until they have. Confirm a manual path that can take over if the AI is paused.
Evidence on your own cases
Test on your own history, including the difficult and adverse cases. An evaluation that excludes failure cases without a reason tells you little. Count omissions, inventions and wrong answers separately.
Make review real. If staff approve outputs without reading them, pause what goes out and change the approval step. Show the reviewer what the system changed and how confident it is. Log every recommendation and every edit.
Depending on your seat
If you're on the board, first ask for an inventory of every model that influences decisions, with its purpose, owner, approval and monitoring. Start with the ones that decide about customers, patients or employees. Then set a routine: approval before use and a periodic report to you.
If you run the systems and a vendor supplies the model, check your contract rights to its inputs, outputs and versions. Then ask the vendor to explain ten decided cases, including any complaint, in terms the committee can read.
What to check before you decide
- Write down which decisions the AI makes and which information it only presents.
- Have the accountable compliance, legal, clinical or safety owners confirm the permitted use.
- Test on your own historical cases, including the difficult and adverse ones, and record any exclusions with reasons.
- Check your contract rights to the model's documentation, inputs, outputs and version history.
- Make a named person the author of record, and log every recommendation and edit.
- Confirm a manual path that can take over if the AI is paused.
- Keep an inventory of every model that influences decisions, with its owner and monitoring.
Questions people ask
Risk committee discovered AI models influence customer decisions with no inventory or monitoring, what should we require?
Require an inventory of every model in use with its purpose, owner, approval and monitoring, then a governance routine: approval before use, performance monitoring, periodic review and a report to the committee. Start with the models that decide about customers. It depends on your regulator's expectations for model risk and on how many models are vendor-supplied.
Management wants AI advising control room operators, what must the board require before AI touches safety related operations?
Require that AI stays advisory with the operator's decision and the safety systems unchanged, that the tool is tested on the plant's own history, that operators are trained on its limits, and that every recommendation is logged. Require a safety review by the plant's process safety function before use. It depends on which operations the tool advises on and on the plant's safety management system.
Should we approve hospital automation when adverse-event evidence is excluded?
Do not rely on an evaluation that omits material failure cases without justification. Approval depends on intended use, representative adverse conditions, and assessment by accountable clinical and safety specialists.
What should directors require before AI telematics influences driver disciplinary decisions?
Separate operational indicators from evidence sufficient for a consequential decision. Use depends on context, error handling, and review by appropriate safety, HR, and legal owners.
Our AI document pilot sent errors to citizens because staff approved without reading, stop or fix?
Pause the sending, keep the drafting, and fix the approval step: approval must require reading, and the system must show what it changed and flag uncertainty. Measure the error rate on a sample before any sending resumes. It depends on the error types and on whether the staff can review at the volume the pilot produces.
Should our insurance business let AI make personalised sales recommendations?
Define the allowed sales role before buying the capability. Viability depends on evidence, escalation, customer understanding, and applicable requirements confirmed by compliance and legal owners.
Vendor AI underwriting model in production, how do I prove to the risk committee that decisions can be explained?
Explanation means you can show, for a given decision, which inputs drove it and that the same inputs give the same answer. Ask the vendor for per-decision explanations and test them on complaint cases. It depends on whether your contract gives you access to the model's inputs, outputs and versions.
Call centre wants AI call summaries written into member records, what do we need to check before a pilot?
Decide first what the summary is allowed to do: a note the agent edits is different from an entry in the medical record. Test on real calls with agents checking every summary, and measure omissions and inventions separately. It depends on your privacy approvals for recording and processing calls and on who signs the entry.
What would justify moving refinery AI from recommendations to operational control?
Widen authority only after the intended operating scope and failure handling have suitable specialist assurance. It depends on the consequences of incorrect actions, operator oversight, and independent protective functions.
How I can help with this decision
- Ask or talk (Free)
- I give my view on which use is safe to pilot and which conditions matter most. I also name the contract right to check before you ask the vendor for anything.
- Review (Pay if it was worth it)
- I write an independent assessment of the proposed use, its evidence on your cases and its governance. I recommend conditions for approval, a narrower scope, or declining the use.
- Retain (When it makes sense)
- I stay available to review the results, the vendor's model changes and any incidents against the agreed scope.