When the Algorithm Is Wrong, You Own the Decision
A 93% confidence score reads to most executives as near-certainty. It is a probability estimate, produced by a model trained on data nobody in your organisation reviewed, validated against benchmarks nobody there set, and applied in conditions the vendor did not fully anticipate. The number carries authority it has not earned, and it is not a defence.
A Florida lawsuit reported by The Guardian makes the point concrete. Robert Dillon was arrested at home, 300 miles from where the crime occurred, because a facial recognition system told Jacksonville Beach police there was a 93% probability he was the man on the security camera. The charges were dropped later. The arrest still happened, and it is the organisation that acted on the output, not the vendor that produced it, now facing the consequences.
Australian executives and general counsel should read that carefully — not because facial recognition is arriving in your sector, but because the accountability structure underneath it already sits inside every consequential algorithmic decision your organisation makes.
The Vendor Score Is Not the Decision
Most AI procurement conversations collapse a distinction worth keeping: a system produces an output, and an organisation chooses to act on it. Vendors sell the first. The second is yours.
A 93% match is information. Arresting someone on that information alone is a governance choice, and one that skipped a review which would have taken about thirty seconds to establish that the suspect lived hundreds of kilometres away. Nothing about that is a technology failure. It is a process failure, and processes belong to whoever designed them.
The reason this matters here is pace. Australian organisations under board pressure to show an “AI strategy” are deploying algorithmic tools faster than their governance can follow, and the accuracy figure in the sales deck is doing the work of due diligence. An accuracy metric describes how a model performed on a test dataset under controlled conditions. It says nothing about how it behaves on your population, in your context, on decisions carrying legal or regulatory weight.
The questions that do not get asked in those procurement meetings are the ones that create liability: Who reviews the output before consequential action is taken? What does the reviewer actually know about the model’s limitations? What is the documented basis for a decision to override or accept the recommendation? What happens when the model is wrong?
Where Australian Law Is Heading
The Australian Privacy Act reform process and the Attorney-General’s current review of automated decision-making are not abstract regulatory exercises. They are the legislative framework inside which your AI governance decisions will eventually be judged.
The direction is clear. The reform agenda is moving toward requiring that organisations using automated decision-making for consequential outcomes — outcomes that significantly affect individuals’ rights, opportunities, or access to services — provide explainability and ensure meaningful human review is in place. “The vendor’s algorithm recommended it” will not be a satisfactory explanation to the Office of the Australian Information Commissioner, and it will not be a satisfactory explanation to a court.
Organisations in financial services are already working within APRA CPS 234’s requirements around information security governance, and the principle it encodes — that boards and senior management are accountable for governance frameworks, not just for adopting technology — extends directly to AI risk. If you are in a regulated sector and you are deploying automated decision tools that affect customers, members, or employees, the regulatory expectation is that your governance framework is commensurate with the risk. A vendor accuracy metric in a procurement document is not a governance framework.
The sectors with the most exposure are not necessarily the obvious ones. Financial services and insurance have regulatory pressure that is at least forcing the conversation. Healthcare, education, recruitment, tenancy, and local government are deploying algorithmic tools in contexts that affect individuals significantly, with governance structures that in many cases have not caught up at all.
The Second-Order Problem Boards Are Missing
The immediate risk is legal liability when an algorithmic system produces a wrong output and harm follows. The second-order risk is subtler and more damaging: organisational cultures that learn to launder decisions through algorithms precisely because it diffuses accountability.
This happens. When a system produces a recommendation, the human reviewer is under implicit pressure to accept it. Overriding the algorithm requires justification. Accepting it does not. Over time, the human in the loop stops being a genuine check and becomes a formality — a box ticked to satisfy governance requirements while real decision authority has migrated to the model. The governance documentation shows human review. The operational reality is that the algorithm decides.
This is the failure mode boards need to be asking about, not whether the AI is accurate enough. The question is whether the human oversight your organisation has documented actually functions as oversight, or whether it is theatre designed to satisfy a compliance requirement while leaving consequential decisions effectively unreviewed.
The Florida case is instructive here too. A basic cross-check — does the suspect actually live near the crime scene — would have caught the error immediately. That check either did not happen or did not register against the weight of a 93% algorithmic confidence score. The number displaced judgement.
What Due Diligence Requires Now
Buying one of these tools responsibly takes a different set of questions than most procurement processes currently ask. Accuracy is the starting point rather than the answer: accurate under which conditions, on which population, against which benchmark?
From there, the documented failure modes. What does this model handle badly, and does that group overlap with the people our decisions affect?
Then the decision boundary, settled before deployment rather than after an incident. At what confidence threshold does automated action require a human? Who performs that review, and are they interrogating the output or initialling it?
Legal and compliance belong in this before go-live, not summoned afterwards. And at the end of it, someone with authority — not a vendor compliance pack — needs to be able to explain to a regulator or a court why the organisation was satisfied the system suited this use case, which controls were in place, and how those controls behaved in practice.
The Accountability Lands Where the Decision Was Made
The vendor will say the tool performed within its documented parameters, and they may well be right. It may have done precisely what it was built to do. None of that moves the liability for how it was deployed, what decisions were constructed around it, and which safeguards were absent.
Executives treating vendor accuracy claims as the end of their due diligence are mistaken about what they bought. What you have purchased is an input to a decision. The decision itself, and everything downstream of it, stayed with you.
Which makes the governance gap in most AI deployments a very ordinary one: no clear, documented, genuinely functioning chain of human accountability between an algorithmic output and a consequential action. That chain is considerably cheaper to build before deployment than to reconstruct while defending a decision in court.