iterthink
Can You Prove a Human Reviewed What Your AI Changed?
A practical guide to human oversight for AI-edited documents: what's actually required, and how to evidence it
Published
AI now drafts, rewrites, and "cleans up" documents in seconds. A paragraph gets tightened, a number gets reformatted, a clause gets rephrased. The file moves on toward sign-off. The edit was effortless to make. The problem is that it's nearly impossible to audit after the fact.
For teams that produce documents other people rely on (plans, reports, tenders, specifications, contracts), that's a new and uncomfortable kind of risk. Not the risk that AI makes a mistake. The risk that a change happened, nobody noticed, and you can't show who approved it.
This guide is for the people who own that risk: quality leads, project directors, BIM and document managers, and compliance officers. It covers what "human oversight" actually means in practice, why an "Approved" tag is not the same as evidence, and how to close the gap, whether or not a specific regulation applies to you yet.
It is not legal advice. Where the EU AI Act is referenced, treat it as orientation, and confirm specifics with your own counsel.
The gap nobody planned for
Most document workflows were designed for a world where humans made the changes. You could reconstruct intent from email threads, comments, and the memory of whoever did the work. Slow and imperfect, but traceable.
AI-assisted editing breaks that model in three ways:
Changes are now cheap and invisible. A human reviewer leaves a trail: a comment, a tracked change, a message. An AI edit can rewrite half a page with no explanation of what it touched or why. The formatting often looks cleaner afterward, which makes errors harder to spot, not easier.
Approval and evidence have drifted apart. Teams still sign things off. But the sign-off increasingly records a conclusion ("this is approved") without recording what the approver actually saw, what changed since the last version, or how long they spent. When someone later asks "did a human really review this?", the honest answer is often "we think so."
Knowledge leaves with people. When the person who understood why a change was made moves on, the reasoning goes with them, unless it was captured somewhere durable at the moment the decision was made.
None of this requires a regulation to be a problem. It's already costing teams rework, client trust, and the occasional very expensive surprise. But regulation is now pointing in the same direction, which is worth understanding.
What the EU AI Act actually says about human oversight
The EU AI Act (Regulation 2024/1689) entered into force on 1 August 2024 and phases in over several years. Its obligations scale with risk: minimal-risk systems carry almost none, while "high-risk" systems carry the heaviest requirements.
For high-risk systems, Article 14 (Human Oversight) is the relevant principle. In plain terms, it requires that such systems be designed so a competent person can:
- understand the system's capabilities and limitations,
- monitor its operation and detect anomalies or unexpected behaviour,
- stay alert to automation bias, the tendency to over-trust an AI's output,
- correctly interpret what the system produced, and
- decide not to use the output, or to override or stop the system.
Article 26 extends this to deployers (the organizations using the system): assign oversight to trained, competent people, and retain the system's automated logs. Article 86 introduces a right for affected individuals to an explanation of the role AI played in a decision.
The throughline across all of these is a single idea: oversight has to be real and demonstrable, not asserted. A policy document saying "a human reviews everything" is not evidence. Built-in mechanisms, retained logs, and a record of what the human actually saw and decided: that's evidence.
An honest note on timing and scope
Two caveats matter, because getting them wrong damages credibility:
The high-risk enforcement date is in flux. The obligations for high-risk systems were originally set for 2 August 2026. A "Digital Omnibus" simplification process has since proposed deferring them, and as of mid-2026 a provisional political agreement points to 2 December 2027 for high-risk Annex III obligations, with formal adoption still pending. The exact date may keep moving. The requirement, that high-risk AI be subject to meaningful, demonstrable human oversight, is not what's being debated.
Most document work is not "high-risk" under the Act. The Act's high-risk categories cover specific domains (biometrics, critical infrastructure, employment decisions, and similar). A general-purpose tool that helps you compare document versions is not, by itself, a high-risk AI system, and nobody should tell you the Act forces you to buy one.
So why does this matter to a document team? Because the Act codifies a direction of travel that extends well past its strict scope. Clients, auditors, public-sector procurement, and your own risk function are all converging on the same expectation: if AI touched the work, you should be able to show what it did and that a person stood behind the result. Building that capability now is good practice regardless of which of your systems are formally in scope, and regardless of the final date.
"Approved" is not evidence
Here's the failure mode that catches teams off guard. The workflow produces an approval: a status, a checkbox, a signature. What it doesn't produce is the context behind that approval.
Picture the question an auditor, a client, or your own legal team eventually asks: "Show me that a qualified person reviewed the changes the AI made to this document before it went out."
A defensible answer needs four things, and most workflows capture only the last one:
- What changed. The specific differences between versions, down to the word, including the ones the AI introduced.
- Who reviewed it. A named, competent person, not just a service account or "the system."
- What they saw and decided. That the reviewer was actually shown the changes, and accepted or rejected them.
- A record that can't be quietly rewritten. An immutable trail from draft to final.
An "Approved" tag with none of the first three behind it fails as evidence. It tells you a conclusion was reached; it can't tell you the review actually happened. Reconstructing that context months later, from filenames, scattered comments, and chat threads, is expensive, often incomplete, and sometimes impossible.
A checklist: seven things you should be able to show
Use this to pressure-test your own workflow. For any document that AI helped produce, can you produce each of these on demand?
- A precise diff of every change between versions, human and machine, at the word level. Not just "a new file appeared."
- Attribution for each change: human edit vs. AI-generated, and which person or tool was responsible.
- Proof the reviewer saw the changes, not just that a final status was set.
- The reviewer's identity and competence. A named person with the authority and knowledge to sign off.
- An immutable history from first draft to approval that can't be edited after the fact without leaving a trace.
- Retained logs for an appropriate period (the Act references at least six months for in-scope systems; many sectors keep far longer).
- Recoverable reasoning. Enough context that someone who wasn't in the room can understand why a change was accepted, even after the original people have left.
If you can produce all seven today, you're in a strong position. If you can't, that gap is the thing to close, well before anyone asks.
How to close the gap
Three approaches, roughly in order of robustness:
Process discipline (necessary, not sufficient). Write down who reviews what, require sign-off before anything is final, and standardize version naming. This helps, but it relies on people remembering to do it and doesn't capture what the reviewer actually saw. It produces conclusions, not evidence.
General tooling (partial). Track-changes in your word processor, PDF comparison, or a cloud suite's version history will capture some of the picture. The common weak spots: changes made by AI often aren't distinguished from human ones; the audit trail can be edited; and your data may live somewhere you don't control. That last point is a real problem in regulated, public-sector, or privacy-sensitive work.
A dedicated review layer (most robust). A purpose-built layer sits between drafting and approval and captures all seven items above as a matter of course: word-level diffs across any version, attribution of human vs. AI changes, an explicit review-and-approve step, and an immutable trail, ideally running locally so sensitive documents never leave your control.
This last category is the gap iterthink was built to fill. It produces word-level diffs across documents and plans, distinguishes and surfaces AI-introduced changes alongside human ones, adds an explicit review and approval step, and records an immutable audit trail from first draft to sign-off. Because it's local-first and source-available, your documents and your evidence stay under your control rather than on someone else's cloud. That matters most in the regulated and public-sector settings where oversight questions get asked. It is not a compliance product and won't make you "AI-Act compliant" on its own; it's a practical way to operationalize human accountability for AI-assisted document work.
The bottom line
Strip away the regulatory detail and the situation is simple. AI has made document changes effortless to produce and difficult to audit. The expectation, from regulators, clients, and your own risk function, is moving toward demonstrable human accountability for those changes. The teams that will handle this calmly are the ones that can already answer one question without scrambling:
"Show me what the AI changed, and prove a person reviewed it before it became final."
You don't need a deadline to make that worth solving. You need it to be true the next time someone asks.
This guide is for general information and is not legal advice. Regulatory requirements depend on your specific systems, jurisdiction, and use case; confirm details with qualified counsel. Timelines referenced reflect the position as of mid-2026 and are subject to change.
Want to see what a defensible review trail looks like in practice? Start reviewing for free with iterthink.