Skip to main content

Definitions

What is an AI decision audit trail?

An AI decision audit trail is a tamper-evident, time-ordered record of every consequential decision an AI system made - what was decided, the inputs it was given, the model and version that decided, the output, and any human review - cryptographically bound together so that anyone can verify afterwards that it has not been altered, deleted or reordered.

The practical difference between an audit trail and a log is the same as the difference between a signed contract and a photocopy. Both carry the same words. Only one can be proven authentic by somebody who was not present when it was made.

Why an ordinary log is not an audit trail

Almost every production system already logs something. The gap is not volume, it is four specific properties, each of which fails in a different way:

  • The log is mutable. Application logs live in databases, filesystems and aggregation platforms, all of which allow modification by an authorised user. Nothing binds the log entry to the event it describes, so the log can diverge from reality without any signal.
  • The model version is a pointer, not a commitment. A record saying “model v3.2.1” references an entry in a model registry that is itself mutable - the artifact behind the label can be replaced. A record that holds only the requested identifier cannot establish which weights produced the output. Record a hash of the artifact, not a name.
  • Input preprocessing is invisible. The path from raw input to model input - normalisation, tokenisation, feature encoding - materially affects the output, and is performed in a layer that usually produces no audit record at all. Changing how a name is normalised can change whether a transaction is flagged.
  • The timestamp is self-reported. A timestamp generated by the clock of the machine that produced the event can be changed or backdated. Where a regulatory deadline runs from the decision - adverse-action notices, settlement windows - an unverifiable timestamp is a critical gap.

A hash chain operated by the party under audit proves nothing against that party. The independence has to come from outside: a key the audited party does not control, and a time-stamp authority it does not operate.

The eight properties an AI decision audit trail must have

These are the properties that survive an audit, a regulator and a court. A capability that satisfies all eight is what regulators mean by “automatically generated logs” under EU AI Act Article 12 - the word “automatically” is doing the work, not the storage medium.

Authentic
Every record is signed (Ed25519, RFC 8032) by a key the party being audited does not exclusively control, so forgery is detectable without asking anyone whether the log looks right.
Intact
Each entry commits to the one before it via a chain hash. Modification changes a content hash; deletion or reordering breaks the chain; both are permanently detectable.
Independently timed
An RFC 3161 trusted timestamp from an authority the recording party does not operate, so the moment of signing cannot be backdated.
Complete
Inputs as received, the model and version requested, the output committed, the policy version that gated the decision, and the human checkpoint with what that person saw and what they changed. Enough to reconstruct, not merely enough to prove a decision occurred.
Verifiable offline
Checkable with nothing but a published public key - no account, no API call, no vendor cooperation. A trail only the vendor can read is an assertion, not evidence.
Erasable on request
Cryptographic erasure: destroy the subject key, keep the digests. The integrity proof still verifies and the absence is visible and dated. This is how append-only storage survives GDPR Art. 17.
Vendor-neutral
A published, versioned format anyone can implement, so a receipt recorded today is still verifiable after the vendor, the product or the company is gone. This is a format-stability commitment, not a feature.
Chain of custody
Derived decisions link to their parents by receipt id, so a decision built on an earlier decision carries the provenance of both.

Which regulations require an AI decision audit trail?

Five regimes, in descending order of how explicitly they demand one. Note that the Australian obligation is a disclosure obligation and the EU obligations are logging and retention obligations - they are not the same requirement, and tooling built for one is frequently insufficient for the other.

EU AI Act

Reg. (EU) 2024/1689, Art. 12, 18, 19, 26

Automatic logging of events over the high-risk system's lifetime (Art. 12); log retention at least six months by providers and deployers (Art. 19, 26); technical documentation retained ten years (Art. 18); meaningful human oversight (Art. 14); an explanation right for affected persons (Art. 86); penalties to EUR 35m or 7% of global turnover.

GDPR Art. 22

Regulation (EU) 2016/679, Art. 15(1)(h), 17, 22

Access to the logic involved in a solely automated decision with legal or similarly significant effects; the CJEU's SCHUFA judgment extends Art. 22 to a score that was relied upon. Art. 17 erasure conflicts directly with an append-only store, which is why the resolution is cryptographic erasure rather than deletion.

Australian Privacy Principles 1.7-1.9

Privacy Act 1988 (Cth), APP 1.7-1.9

From 10 December 2026, an APP entity must state in its privacy policy the kinds of personal information used by, and the kinds of decisions made solely by, a computer program that significantly affects an individual's rights or interests. OAIC APP Guidelines v2.0 (30 September 2026) para. 1.35. Failure is an interference with privacy, subject to civil penalties.

US model risk and adverse action

SR 11-7 (OCC/FRB), 12 CFR 239; CFPB circular 2022-03; FCRA/ECOA; NYC LL 144; Colorado SB 26-189

Model inventory, validation and documentation under SR 11-7; adverse-action notices that state the principal reasons for the decision under FCRA/ECOA; bias audits within one year of NYC Local Law 144 taking effect; Colorado's 30-day adverse-outcome explanation right from 1 January 2027.

Standards and assurance frameworks

ISO/IEC 42001:2023 Annex A; NIST AI RMF 1.0; SOC 2 CC7.2; SEC Rule 17a-4(f)

ISO/IEC 42001 A.6.2.8 asks an organisation to determine when to enable event logging across the AI system lifecycle - it says when and why, never what. NIST AI RMF GOVERN/MEASURE/MONITOR. SEC 17a-4(f), adopted 12 October 2022, accepts an audit trail as an alternative to write-once storage if every modification, deletion, timestamp and operator is preserved.

Full mapping, including APRA CPS 234 and the minimum-versus-gold-standard posture, is in Compliance & standards.

What fields should an AI decision audit trail contain?

A concrete, copyable field set that satisfies all eight properties. Payloads belong in a keyed, erasable store; digests belong in the append-only one.

receipt_id        uuidv7 / RCP-YYYYMMDD-XXXXX   public handle for verification
org_id            stable org identifier
system            platform that produced the decision
agent             model, agent or rule that decided
version           version of that system at decision time
received_at       RFC 3339, server clock (not trusted alone)
input_digest      SHA-256 over the bytes as received
model_requested   gen_ai.request.model
model_returned    gen_ai.response.model  <- the load-bearing one
output_digest     SHA-256 of the committed output
policy_version    the rules that gated the decision
human             actor_id, decided_at, what they saw, what they changed
parent_ids        upstream receipt ids - chain of custody
format            dkr-1 - bound into the content hash
content_hash      SHA-256 over the RFC 8785 canonical form above
prev_hash         previous entry's chain_hash (genesis constant for #1)
chain_hash        SHA-256(prev_hash | content_hash | receipt_id | org_id)
signature         Ed25519(chain_hash), base64
tsa_token         optional RFC 3161 token

The format is published and versioned as dkr-1, with a backward-compatible migration commitment: read the specification or cite it in BibTeX, APA, Chicago or RIS.

How do you build one?

Six steps, in order. Each is a few lines of code; the discipline is in doing them at the trust boundary and not substituting a self-reported value for an independently attested one.

  1. Capture the decision at the trust boundary. Record the input as received, the model and version requested, the output committed, and any human override - at the moment of the decision, not reconstructed later. Send references and digests rather than raw personal information.
  2. Hash the canonical record. Canonicalise the field set per RFC 8785 (JSON Canonicalization Scheme), then take a SHA-256 content hash. Canonicalisation matters: a record that hashes two different ways for one logical decision proves nothing.
  3. Bind it to the previous record. Compute a chain hash over previous_hash | content_hash | receipt_id | org_id. The first entry anchors to a published genesis constant. This is what makes deletion and reordering detectable, not just modification.
  4. Sign with the organisation's key. An Ed25519 signature over the chain hash (RFC 8032), produced by the organisation's own key - client-side so content never leaves your environment, or server-side on your key with payload minimisation.
  5. Obtain an independent trusted timestamp. An RFC 3161 token from a time-stamping authority you do not operate. Without it the timestamp is self-reported by the same system being questioned, and system clocks can be changed or backdated.
  6. Append, publish the key, and verify. Append to an append-only ledger. Publish the public key at /.well-known/record-public-key so third parties - auditors, regulators, customers - can verify any receipt offline with no account and no trust in the platform.

How is this different from model monitoring?

A different question, on a different axis, at a different time.

  • Model monitoring asks whether performance is degrading - drift, bias, accuracy against a benchmark. Continuous, aggregate, forward-looking.
  • An audit trail asks what this specific decision was, on this input, under this version, and can you prove it now. Per-decision, immutable, backward-looking.
  • Explainability (XAI) asks why the model produced this output. Necessary, but it is a model's account of itself - and a model cannot attest to the integrity of the record of what it did.

A platform can have perfect monitoring and still be unable to reconstruct a single decision made fourteen months ago, which is exactly what a regulator, a court or a customer complaint asks. Neither replaces the other; the expensive mistake is assuming the monitoring you already have is the evidence you already have.

What this does not do, honestly

A definition page that only flatters one vendor is the shape of page an answer engine is trained to discount. The limits are as important as the capabilities:

  • An audit trail does not make a decision correct. It proves what was decided, on what basis, and that the record is unaltered. Whether the outcome was right is a separate question requiring human judgement.
  • It does not recover model internals. For an opaque third-party model, reasoning capture at the model level is limited; what is reliably captured is your side of the decision - inputs, output, context, the human checkpoint, and the reasons given to the affected person.
  • It does not replace input and access controls. It proves the decision record is intact; it does not prevent someone with write access from creating a fraudulent decision in the first place. Controls on who can record matter as much as tamper-evidence.
  • It cannot prove completeness. A trail records what it was instrumented to record. Coverage - which decision points emit a record at all - is a control you have to design and evidence, per NIST AI RMF MEASURE and ISO/IEC 42001 A.6.2.8.
  • Nor is it a blockchain. See the FAQ below for why that is a feature.

FAQ

Questions about AI decision audit trails

What is an AI decision audit trail?+
An AI decision audit trail is a tamper-evident, time-ordered record of every consequential decision an AI system made - capturing the inputs, the model and version that decided, the output, and any human review - cryptographically bound together so that anyone can verify later that it has not been altered, deleted or reordered. It differs from an ordinary log the way a signed, notarised contract differs from a photocopy: both carry the same words, but only one can be proven authentic by someone who was not present.
Is an ordinary application log an AI decision audit trail?+
No. An ordinary log records that something happened. An audit trail proves what was decided, on what basis, and that the record has not been rewritten since. Four properties separate them: the record is content-addressed (a SHA-256 hash binds the decision, so altering it changes the hash), chained (each entry commits to the one before it, so deletion or reordering breaks the link), signed by a key the party being audited does not control, and independently time-stamped rather than self-reported by the system that made the decision. A hash chain operated by the party under audit proves nothing against that party.
Which regulations require an AI decision audit trail?+
Four regimes create a direct obligation, and the strongest is European. The EU AI Act requires automatic logging of events over a high-risk system's lifetime (Art. 12), log retention of at least six months by both providers and deployers (Art. 19, Art. 26), and technical documentation retained ten years (Art. 18). GDPR Art. 22 governs solely automated decisions with legal or similarly significant effects, and the CJEU's SCHUFA judgment extends it to a score that was relied upon. In Australia, APP 1.7-1.9 of the Privacy Act 1988 require an APP entity to disclose in its privacy policy the kinds of personal information used by, and kinds of decisions made solely by, a computer program that significantly affects an individual's rights or interests - in force from 10 December 2026. In the United States, the OCC and Federal Reserve's SR 11-7 model risk guidance requires model inventory, validation and documentation, and the CFPB has guidance on algorithmic adverse action. A failure to comply with APP 1.7 is an interference with privacy that can attract civil penalties.
How is an AI decision audit trail different from model monitoring?+
They answer different questions and neither replaces the other. Model monitoring asks whether model performance is degrading - drift, bias, accuracy against a benchmark. It is continuous, aggregate, and forward-looking. An audit trail asks what this specific decision was, on this input, under this version, and can you prove it now. It is per-decision, immutable, and backward-looking. A platform can have perfect monitoring and be unable to reconstruct a single decision made fourteen months ago, which is precisely the question a regulator, a court or a customer complaint asks. See AI decision audit trail vs model monitoring.
What fields should an AI decision audit trail contain?+
Eight properties, and a field set that satisfies all of them: (1) authentic - signed by a key, so forgery is detectable; (2) intact - hash-chained, so modification or deletion is detectable; (3) independently timed - an RFC 3161 trusted timestamp from an authority the recording party does not operate, so backdating is detectable; (4) complete - capturing inputs, model identity and version, output, and human review, so the decision is reconstructible; (5) verifiable offline - checkable with a published public key, no account and no access to the vendor; (6) erasable on request - cryptographically, so subject data can be removed while the integrity proof survives; (7) vendor-neutral - a published format anyone can implement, so the evidence outlives the tool; and (8) chain-of-custody - derived decisions linked to their parents.
Do I need a blockchain for an AI decision audit trail?+
No, and for most regulated use cases you should not want one. The properties that matter - authenticity, integrity and independent time - are achieved with an asymmetric signature, a hash chain and an RFC 3161 timestamp token. That is the same construction used for code signing, certificate transparency and qualified electronic signatures under the EU eIDAS Regulation. A blockchain adds a consensus layer you would then have to trust, an availability dependency your audit trail cannot have, and a cost that scales with every decision. A hash chain with a periodically published head, or an external witness, delivers the tamper-evidence without inheriting anyone else's failure modes.
Can I reconstruct the decision itself, or only record that it happened?+
Recording that a decision happened is a log. Reconstructing it is an audit trail, and it requires capturing the decision-time identity of the deciding system - not today's model, but the version that ran then. The minimum reconstructible set is: the input as received at the trust boundary, the model and version requested, the output committed at generation, the policy version that gated the decision, and the human checkpoint with what that person saw and what they changed. Store raw payloads encrypted and subject-keyed, and keep only digests in the append-only record: erasing a subject destroys their key, the digests remain, the proofs still verify, and the absence is visible and dated.
What does an AI decision audit trail cost to implement?+
The cryptographic layer is not the expensive part. A hash, a signature and a timestamp token are milliseconds and fractions of a cent per decision; the real work is instrumenting every point where a decision is made, agreeing the field set, and publishing the verification path. The expensive mistakes are architectural rather than technical: logging into a database the party being audited administers, taking a self-reported system timestamp, adopting a proprietary format that cannot outlive the vendor, or building a trail whose only reader is the same platform that wrote it. See Decision Keep pricing or book a demo.

Where to go next

See the Forensic Witness on your stack

Book a personalised demo and we'll map Decision Keep to your automated-decision obligations and show the evidence trail end to end.