AI Governance needs an exception ledger, not another approval committee
Image: Depositphotos, Google’s Nano Banana AI Image Generator
The summer of rogue AI is sending a useful warning to enterprise leaders. Recent incidents and tests involving frontier AI systems are forcing companies to rethink governance around access, identity, and agency. The underlying concern is operational, not speculative. OpenAI disclosed in July that models in a cyber evaluation found an unintended path out of a constrained environment and reached Hugging Face infrastructure while pursuing the benchmark objective. For companies, the lesson is not to surround every AI action with another approval committee. That recreates the friction automation was supposed to remove. The better response is to make exceptions visible, reviewable, and owned.
Dr. Tom Davenport wrote in his Substack and on Cognitive World that the enterprise AI conversation is moving from model capability to organizational deployment. That shift matters because the hardest failures often live in the space between technology and the workflow around it: unclear permissions, ambiguous handoffs, missing context, and people who do not know when an automated action needs to stop.
Every agentic workflow should therefore have an exception ledger before it scales. (1) The ledger is a lightweight operating record of the moments when an AI system does something outside the expected path or requires a person to intervene. It is not a log of every prompt and action. It is a management tool for finding the few exceptions that reveal where a workflow is brittle.
The first practical question is what counts as an exception. A useful definition is any event in which the workflow cannot safely complete its intended path without material human judgment, correction, escalation, or recovery. That includes a person overriding the AI, repairing an output before it can be used, stopping an automated action, resolving conflicting evidence, handling a case the system cannot classify, or reopening downstream work because an earlier AI-supported step was wrong. Routine human review that confirms a correct result does not need its own ledger entry. The goal is to capture deviations that expose a weakness in the design.
A useful exception ledger records six things. First, what the system was trying to accomplish. Second, what triggered the exception. Third, what the agent did next. Fourth, what a human had to correct, approve, reverse, or investigate. Fifth, the downstream consequence. Sixth, the person who owns the redesign.
Take an HR agent that helps employees answer policy questions and initiate routine requests. Most interactions may be harmless. But suppose the agent encounters a leave request involving a policy that varies by jurisdiction, gives a confident answer based on the wrong policy version, and routes the employee to the wrong process. The exception ledger should capture the policy ambiguity, the incorrect routing, the human correction, and whether the fix belongs in retrieval, workflow logic, permissions, or escalation rules.
The same discipline applies to finance. An agent may be allowed to prepare a payment but not release it. If it repeatedly creates payments with incomplete vendor context, the problem is not solved by asking a manager to approve more carefully forever. The recurring exception tells the team that the workflow needs a stronger precondition before the agent can reach the approval step.
This is why an exception ledger differs from an approval queue. An approval queue asks, “May the system proceed this time?” An exception ledger asks, “Why did this workflow require human rescue, and what should change so the same rescue is less likely next time?” One controls an individual action. The other improves the system.
The ledger should be reviewed weekly by the people who can redesign the work: the process owner, an operational user, the technical owner, and the relevant risk or compliance partner. They should look for clusters rather than anecdotes. Ten different exceptions caused by the same missing customer field are one design problem. Five cases where people override an agent because the policy is unclear may be a policy problem rather than a model problem.
That distinction is crucial. Cognitive World has also emphasized the importance of redesigning end-to-end processes and making the knowledge required at decision points explicit. An exception ledger gives leaders evidence about where that redesign is failing in practice. It shows whether the trouble comes from the model, the data, the workflow, the policy, the permissions, or the human role.
It also protects human capability. If people only click “approve” or “reject,” they may become passive monitors. When they must classify why an exception occurred and decide what should change, they stay engaged with the underlying work. The review becomes a learning loop rather than a rubber stamp.
Leaders should track three outcomes from the ledger. The first is repeat-exception rate: are the same failures recurring? The second is recovery effort: how much human time is spent fixing exceptions after they happen? The third is redesign closure: how many recurring exceptions result in a workflow, policy, data, or permission change? Those measures reveal whether governance is improving operations.
The goal is not zero exceptions. Novel work will always produce them, and a system that never raises an exception may simply be hiding risk. The goal is to make exceptions informative. A mature organization should become faster at detecting patterns, assigning ownership, and changing the workflow before a small failure becomes a costly one.
Getting started can be deliberately modest. Pick one AI-supported workflow that already has a named business owner and enough volume to reveal patterns. For one week, ask the people doing the work to record only material exceptions using the six fields above. At the end of the week, group the entries by cause and select the most frequent or costly recurring pattern. Assign one owner to change the relevant prompt, data source, policy, permission, handoff, or escalation rule. Then run the workflow for another week and see whether the same exception declines.
That two-week cycle is enough to test whether the ledger creates useful management information without building new bureaucracy. If it does, add a second workflow. If it produces noise, tighten the exception definition. If staff record nothing while still spending substantial time fixing AI output, that is evidence that the reporting rule is too narrow or the culture is discouraging disclosure.
Enterprise AI governance will fail if it becomes a layer of meetings sitting above unchanged work. It succeeds when it changes what happens inside the workflow. An exception ledger is modest enough to start with one workflow tomorrow and rigorous enough to show leaders where human judgment is still carrying the system. That is exactly the information they need before they give agents more authority.
(1) Adapted from: The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026). https://disasteravoidanceexperts.com/aibook
About the Author
Gleb Tsipursky, Ph.D.
Dr. Gleb Tsipursky, a behavioral scientist called the “Office Whisperer” by The New York Times, helps tech-forward leaders stop overpaying for AI while boosting engagement and innovation. He serves as the CEO of the AI consultancy Disaster Avoidance Experts, and wrote eight books, including The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026).