Beyond the Black Box: why trusting AI isn't enough for auditors
Internal logs are not an audit trail. Learn the separation-of-powers architecture - an independent, cryptographically immutable Forensic Witness - that makes automated decisions truly auditable.
About the author+
Jamil Luketic
Executive Director at Decision Keep
Former Data & Tech Leader at Oracle, Mastercard, Coles, Optus, and Reece.
Connect on LinkedInIn the last eighteen months, the conversation around AI has shifted from "What can this model do?" to "How can we prove what this model did?" McKinsey's state of AI research shows organisations moving from experimentation to value, while Gartner predicts audit and compliance will be transformed by generative AI and EY argues integrity is the defining issue for the audit profession in the age of AI.
As AI agents move from experimental tools into the core of enterprise decision-making - approving loans, triaging claims, automating compliance - they create a new class of systemic risk. We have built powerful AI models, but we have failed to build the audit infrastructure to support them.
Most organisations are operating on a fundamental misunderstanding: they believe their internal logs constitute an audit trail for automated decisions.
The fallacy of internal governance
In traditional IT, logs are sufficient. In AI-driven decision-making, they are a liability.
Most enterprises rely on internal governance platforms that monitor model performance, drift, and operational metrics. These tools are valuable for optimisation, but they are architecturally flawed for auditing automated decisions. Because they reside inside the same production environment as the AI, they are subject to the same systemic pressures, overrides, and administrative access.
If an auditor asks you to prove a decision was compliant, and your evidence is stored inside the same system that generated the decision, you do not have evidence - you have a self-reported story. To a forensic auditor, that is not an audit trail; it is a point of failure.
The separation of powers
To achieve true enterprise-grade accountability for automated decisions, look to the same principles that govern financial systems: the separation of powers.
The entity that executes the decision cannot be the same entity that acts as the Forensic Witness. Move toward an architecture defined by two distinct layers:
- The Operational Layer - where the AI lives, breathes, and makes decisions. Managed by your internal governance tools.
- The Forensic Witness Layer - a cryptographically immutable, external record that remains untouched by your internal systems.
For the full legal and regulatory context, see The Forensic Witness: why automated decisions need an independent record.
This "Forensic Witness" acts as an independent auditor's ledger. It does not interfere with the AI's speed or logic, but it provides the cryptographic proof that an automated decision occurred, when it occurred, and exactly what the inputs were. Once that record is signed and chained, it becomes permanent - incapable of being manipulated, patched, or overwritten by any administrator within your production stack.
Why this matters for auditability
For the Chief Risk Officer and the Lead Auditor, this architectural separation changes everything:
- Zero-trust proof - you no longer need to trust the integrity of your production logs. You verify the signature against a published key.
- Offline verification - auditors no longer need access to your live AI stack. They export the record chain and verify it independently, reducing security friction and privacy risk.
- Tamper-proof integrity - because the Witness layer is decoupled, your AI system effectively loses the ability to "edit its own history."
The next frontier: accountable AI
We are exiting the era of "move fast and break things" and entering the era of "move fast and be accountable." The black box is no longer an acceptable excuse for regulators, and "trust us" is no longer a valid response to an audit.
The future of AI in the enterprise belongs to those who treat automated decision-making with the same cryptographic rigor as our most sensitive financial records. It is time to move beyond the internal log and establish a Forensic Witness for every automated decision your AI makes.
Building for the future
The good news: this is not a research project. It is a discipline of producing, for every automated decision, a signed, chained, time-anchored, verifiable record your organisation controls.
That is the whole idea behind Decision Keep: the Forensic Witness for Automated Decisions – the independent, verifiable record for every automated decision - sovereign and portable, so the evidence never leaves infrastructure you control.
FAQ
Questions auditors, risk and legal actually ask
Why aren't internal logs enough for an audit?+
What is the 'separation of powers' for automated decisions?+
How does an immutable Forensic Witness help auditors?+
Does this slow the AI down?+
Sources
References & further reading
Independent analysis and standards cited in this article.
- The state of AI in 2025: Agents, innovation, and transformation
McKinsey & Company · 2025
- AI Regulations to Drive Responsible AI Initiatives
Gartner · 2024
- AI assessments: enhancing confidence in AI
EY · 2025
Move beyond the internal log
Decision Keep is the Forensic Witness for Automated Decisions – the independent, immutable witness layer for automated decisions - sovereign and portable, so the evidence never leaves infrastructure you control. See it on your stack.
Keep reading
Authorization decision logging: how fast AI risk scoring stops an authorization decline spike
An authorization decline spike is the worst kind of payment event: revenue hemorrhages, customers complain, and every denied transaction is now a potential d…
What tool can automatically identify at-risk accounts before they cancel?
A churn model can flag an account likely to cancel in milliseconds. But the flag itself the automated decision to treat this customer differently is the thin…
AI-native vs AI-enhanced risk decisioning: where the evidence gap widens
Most organisations can tell you whether an AI system is "enhanced" or "native." Far fewer can prove what each decision was , and that nobody changed the reco…
Documentation