The errors we'd rather make

Share
The errors we'd rather make

Most of us have had the text from the bank asking whether a transaction was really us. It is irritating, particularly when it arrives shortly after the card was declined in front of a queue. But I would be considerably more than irritated if a large fraudulent payment went through and no text arrived at all.

So in this case the answer is easy. I would rather the bank stopped a hundred of my genuine transactions than approved one fraudulent one. Yet, having decided that, I realised nobody had asked me. The bank set the line somewhere, and it has never shown anyone the working.

Two names for two errors

The bank would of course prefer to block every fraudulent payment and approve every genuine one. So would I. That system does not exist. However sophisticated the detection becomes, there are people whose full-time occupation is staying ahead of it, and the bank is left drawing a line through a distribution it cannot see clearly. Draw it tighter and more genuine transactions get stopped. Draw it looser and more fraud gets through. This is the part that took me a while to see properly: moving the line does not remove the trade-off. Tighten it and you will usually prevent more fraud while stopping more genuine transactions; loosen it and the reverse happens. Reducing both at once requires a better system, and a better system takes years and money, while the line can be moved on a Tuesday afternoon.

The two errors have names. A false positive is flagging something that turns out to be fine: the declined card, the healthy patient recalled for further tests. A false negative is missing something that was not fine: the fraudulent payment approved, the illness the screening did not catch. Any useful system that sorts things under uncertainty risks producing both, and the person who sets the threshold is deciding, whether or not they put it this way, which of the two they would rather have more of.

Where there is no right answer

Once you see the shape of it, it turns up everywhere, including in places too small to feel like decisions at all. If you have a traditional boiler, you choose how long to leave the hot water on. Would it be worse to run out when you need it, or to waste energy and add to the bill? I would ideally heat exactly as much as my family needs, but the amount varies day to day in ways I cannot predict, so I am setting a threshold on a timer and accepting that I will be wrong in one direction or the other most days.

With the bank I did not have to think about it. With the boiler I have thought about it and I still could not tell you that my answer is the correct one, because I do not think there is a correct one. Somebody who cared more about the energy bill than I do would set the timer differently and would not be making a mistake. Whether a false positive is worse than a false negative is often not a finding but a preference. It is not the sort of thing that gets settled by collecting more data, which is inconvenient, because it means that two people looking at identical evidence can reach opposite conclusions and both be reasoning correctly.

In What's your risk appetite for other people? I wrote that I would rather be open and kind and be exploited by tens of people than close myself off and miss the one sincere person. That is the same structure. I was stating a tolerance for false positives in order to protect against false negatives, and I meant it. I can also see the case for the opposite policy, held by people who have been exploited more often than I have and have priced the risk accordingly. Neither of us is misreading the evidence. We are weighting the errors differently, which is why that kind of disagreement is so difficult to argue anyone out of.

The numbers underneath

The stakes rise quickly from here. In a mental health setting, would it be worse to detain someone who did not need detaining, or to leave someone at risk without care? In social care, would it be worse to separate a child from safe and loving parents, or to leave a child with dangerous ones? In safeguarding, is it worse to put tens of people through a process that finds nothing, or to miss the one case where the concern was real?

It is tempting to leave those as rhetorical questions. They deserve better than that, because there are numbers underneath them and the numbers are not comfortable. A 2016 meta-analysis of thirty-seven longitudinal studies looked at how well psychiatric patients could be sorted into higher and lower risk of suicide. Over an average follow-up of a little over five years, 5.5 per cent of those categorised as high risk died by suicide, against 0.9 per cent of those categorised lower risk. That is a real and substantial difference. It also means that roughly nineteen out of every twenty people identified as high risk did not die by suicide, and that just under half of those who did had been placed in the lower-risk group. The authors also found evidence of publication bias favouring the studies that made risk categorisation look strongest, so the true picture is probably no better than this and may be worse. NICE guidance published in 2022 now instructs clinicians in England not to use risk assessment tools or low, medium and high stratification to predict suicide or repeat self-harm, and not to use them to decide who is offered treatment or who is discharged.

The reason is arithmetic rather than incompetence. When the outcome you are looking for is rare, even a genuinely informative test produces far more false positives than true ones. That is true of suicide prediction, of safeguarding referrals, of screening for uncommon cancers, and of watchlists for rare crimes. It is not a flaw in the clinicians or the social workers. It is a property of low base rates that no amount of professional diligence removes, and it means the threshold question does not disappear. It only stops being able to disguise itself as a technical one.

The errors you never find out about

There is a further difficulty, which is that most systems never learn which errors they made. The bank finds out fairly quickly, because the customer telephones or the fraud gets reported. Almost nothing else works like that. If someone is detained and comes to no harm, nobody can say whether the detention prevented something or whether there was nothing to prevent. If a child is removed and grows up well, the version where they stayed is not available for inspection.

Even the cases we treat as settled are not always settled. The Cleveland removals of 1987, in which 121 children were taken into care on the strength of a diagnostic test that was subsequently discredited and most of whom were returned home, entered English social work as the standard example of intervening too readily. Later archival work has argued that a substantial proportion of the original findings may have been sound after all. Nearly forty years on, the number of false positives is still in dispute.

You cannot calibrate against outcomes you never observe. Which means that in most of these systems the line is not being tuned by feedback at all, because the feedback does not exist. Something else must be doing the tuning.

Which error gets a face

The two errors are not equally visible. A false positive usually produces a named, frustrated person who can describe what happened to them: the family wrongly investigated, the detained patient, the innocent man on a watchlist. A false negative splits into two extremes. Most of the time it is invisible, a missed opportunity or a counterfactual nobody can point to. But occasionally one materialises catastrophically — a child left in a dangerous home who later dies, a person assessed as low risk who does not survive the month — and when that happens the line moves hard, towards whichever tragedy is most recent. After the death of Peter Connelly became public in 2008, referrals and care applications in England rose sharply and stayed high for years. After a wave of intervention is judged to have gone too far, the bar goes back up. The threshold is being set by the last thing that made the news rather than by anyone's considered view of the relative costs.

I suspect this accounts for a good deal of political disagreement that presents itself as factual. Two people argue about whether a system is too harsh or too lax, marshalling evidence, each convinced the other is either callous or naive. Often they do not disagree about the evidence at all. They disagree about which error they would rather live with, and that disagreement is invisible to both of them because neither has said it out loud.

What this does not settle

I want to be careful about how much I claim here. Saying this clearly does not settle anything. If the costs on the two sides are genuinely incommensurable, such as someone's liberty against someone else's life, then knowing that you are choosing a threshold tells you precisely nothing about where to put it. I have found the framing useful for understanding why certain arguments cannot be won, not for winning them. And I notice that I have written all of this from the position of the person setting the line, which is the comfortable end of it.

We will not get it right. That much is settled before we start. What is not settled is which of our inevitable errors we would rather make, and who ends up carrying them.

So: in the decisions you actually make — at work, with your children, about the people you let close — which way are you already erring, and when did you last check whether you meant to?


Notes and sources

Suicide risk categorisation. Large, M., Kaneson, M., Myles, N., Myles, H., Gunaratne, P. and Ryan, C. (2016), 'Meta-Analysis of Longitudinal Cohort Studies of Suicide Risk Assessment among Psychiatric Patients: Heterogeneity in Results and Lack of Improvement over Time', PLoS ONE 11(6): e0156322. doi:10.1371/journal.pone.0156322. Thirty-seven studies, fifty-three samples, 315,309 people. Mean follow-up 63 months. Positive predictive value of a high-risk categorisation 5.5 per cent; rate among lower-risk patients 0.9 per cent; pooled sensitivity 56 per cent, specificity 79 per cent. Egger's test indicated publication bias favouring stronger associations.

The false-positive problem in suicide prediction. Nielssen, O., Wallace, D. and Large, M. (2017), 'Pokorny's complaint: the insoluble problem of the overwhelming number of false positives generated by suicide risk assessment', BJPsych Bulletin 41(1): 18–20. doi:10.1192/pb.bp.115.053017. See also Pokorny, A. D. (1983), 'Prediction of suicide in psychiatric patients: report of a prospective study', Archives of General Psychiatry 40(3): 249–257.

NICE guidance. National Institute for Health and Care Excellence (2022), Self-harm: assessment, management and preventing recurrence, NICE guideline NG225, recommendations 1.6.1–1.6.6. https://www.nice.org.uk/guidance/ng225

Care applications after 2008. Cafcass publishes monthly and annual care application figures at https://www.cafcass.gov.uk.

Macleod, S., Hart, R., Jeffes, J. and Wilkin, A. (2010), The Impact of the Baby Peter Case on Applications for Care Orders, NFER/LGA. The National Audit Office reported in 2010 that local authorities became more cautious and more likely to apply for care orders following the death of Peter Connelly.

Cleveland, 1987. Butler-Sloss, E. (1988), Report of the Inquiry into Child Abuse in Cleveland 1987, London: HMSO.

Own work. 'What's your risk appetite for other people?' https://www.mujsyed.com/whats-your-risk-appetite-for-other-people/