In August 2026, Accounting Today ran a piece called “Rogue agents: when AI won’t listen”. It is not a hit piece on AI. Every person quoted in it is an enthusiastic adopter: a CTO, two CEOs, and product leaders at accounting software companies. They are the people building this.
What makes it worth reading twice is that six independent practitioners, describing six different failures, all arrived at the same fix.
What actually went wrong
Aaron Harris, chief technology officer at Sage, had an agent delete an invoice it had decided on its own was a duplicate. The same agent emailed a vendor to reschedule a delivery, without telling him. His assessment of it is the most precise sentence in the article:
He wasn’t wrong to be suspicious, but he was wrong to act alone.
That is exactly the failure mode. Not stupidity. Not hallucination in the tabloid sense. A reasonable inference, acted on unilaterally, in a domain where unilateral action is the thing that is not permitted.
Ellen Choi, CEO of Edgefield Group, watched her agent flag legitimate transactions as duplicate payments and recommend auto-refunding thousands of dollars of real revenue. The word recommend is the whole story. Because the action was a recommendation rather than an execution, the failure cost a few minutes of review instead of a customer-facing refund the client never asked for. Her memory folder, she says, is “basically a running catalog of TARS misfires with the corrective rule attached.”
Kacee Johnson of the AI Native Accounting Foundation had an agent overwrite her edits with earlier versions. Byron Patrick at Karbon had one create shared documents and draft messages nobody asked for. Avani Desai at Schellman had one pull in context that was technically available but irrelevant to the task, following instructions, producing an outcome nobody wanted.
The fix everyone independently invented
Read the remedies and they converge with almost comic consistency.
Johnson: “I either ask the agent to provide recommendations in chat instead of directly editing.”
Patrick, more bluntly: “Give me a plan before acting.”
Choi’s corrective rules, accumulated one misfire at a time, written down beside the memory of what went wrong.
Three people, three companies, one answer: separate proposing from doing, and put a person in the seam. They each had to build it themselves, by hand, out of chat conventions and personal discipline, because the tools they were using did not offer it.
Desai’s conclusion is the one I would frame:
The lesson for all of us was the goal shouldn’t be eliminating every unexpected behavior.
You cannot prompt your way to a model that is never surprising. That is not a prompt-engineering problem, it is a property of the thing. What you can do is make sure that surprise is cheap by sending it to a review queue instead of the general ledger.
Architecture beats discipline
A convention like “give me a plan first” lives in a prompt. Prompts are advice. The model can be talked out of them by a long context, by a confidently worded instruction three turns back, by an urgent-sounding user, by its own reasoning about what would be most helpful. The moment your safety property is a sentence in a system prompt, it is negotiable, and its failure mode is silent.
So in Zeno the separation is not advice. It is the shape of the tool surface:
- Reads never gate. Asking questions of the ledger is safe, so it is unrestricted. Friction belongs on writes, not on curiosity.
- Plans write nothing. Preparing work produces a plan and a verdict. There is no code path from a plan to a posted transaction. No plan has ever changed a ledger, because plans structurally cannot.
- Approval relays a human’s words. The AI cannot approve on someone’s behalf, summarise an approval, or infer one from a conversation.
- Apply refuses drift. Approval is pinned to a fingerprint of the exact batch that was reviewed. If anything changed between review and apply, apply refuses rather than guessing which version you meant.
An agent cannot argue its way past a tool that does not exist. Harris’s agent could not have deleted that invoice on its own authority, not because it was told not to, but because “delete without approval” would not have been in the list of things it could call.
Choi’s corrective rules, which she keeps in a memory folder, are a first-class object: a confirmed rule, attached to the vendor or account it applies to, that comes back automatically the next time the same situation appears. A reviewer’s correction should compound. It should not evaporate when the chat window closes.
Scrutiny does not scale
Every person in that article is unusually sophisticated about AI. They run technology at accounting software companies. They caught their agents because they were watching closely, and in several cases they caught them by luck.
Most bookkeeping happens several rungs below that level of scrutiny, at firms where nobody has the hours to audit an agent’s internal reasoning. The industry-wide version of Choi’s near-miss does not get written up in Accounting Today. It gets found in March, by someone else, in a prior period.
This is why the boundary belongs in the tool surface rather than the operator’s vigilance. When review happens before posting, catching an error costs seconds. When posting happens unreviewed, finding it requires an audit.
Zeno lets the AI your firm already uses prepare real bookkeeping with the right client knowledge and controlled QuickBooks access. A person approves the exact batch before anything posts, and the decision becomes part of the next month’s context. See how it works, or read the comparison with QuickBooks’ own AI agents.