Draft for Tamir's review. Not published.
What the number does and doesn't tell you
Accuracy is one number across many kinds of decision. It hides which claims the model gets wrong, and those may be the ones that cost you most. A model can do well on routine claims and badly on complex ones.
Ask how the figure was measured: on whose data, from which period, and against whose labels. If the vendor can't answer clearly, the number can't support a decision yet.
How I would test it
Take a sample of your own past claims where the right outcome is already known. My illustrative example uses 500 cases and a two-week test; the right size depends on your volume and claim types.
Agree the pass criteria with the vendor before the test starts. Include errors by claim type, what the model does when it is unsure, and how a handler overrides it.
When the answer depends
If the AI only sorts claims for a handler who still decides, a lower accuracy can still be useful. If it pays or rejects claims on its own, the bar is much higher, and you should check what your regulator expects.
Asking me is free. You get my initial view and the question I think matters most. A full review runs under a short agreement, and the first one includes a no-value, no-fee clause.
What to check before you decide
- Ask on which data the 95% was measured, from which period, and who labelled the correct answers.
- Ask for accuracy on a sample of your own past claims, with pass criteria agreed before the test.
- Break the results down by claim type and size, and look hardest at the expensive errors.
- Ask what the model does when it is unsure, and how a handler sees and overrides its decision.
- Check what the contract says if accuracy on your claims drops after go-live.
- Confirm where your claim data goes during the test and who may use it.