BACK TO INSIGHTS

BACK TO INSIGHTS

Agentic AI

8 mins

Agents Need an Undo Button

Why better models won’t stop agents from making costly mistakes, and how to design agentic products where every mistake is cheap to take back.

By Srijan, founder of WBOS · October 2026 · ~8 minute read

The short version

  • Agents act rather than answer, so their mistakes leave the screen and land in inboxes, databases and bank accounts.
  • Better models won’t fix this on their own, because small per-step error rates compound across long chains of actions.
  • The real question is how cheaply you can recover when the agent is wrong, and the answer is to design reversibility in from day one.

A wrong answer costs a sentence, but a wrong action costs a customer

For the last couple of years, most of the AI we shipped was the kind that answered questions. If the model got something wrong, a person would read it, frown a little, and move on, and the damage more or less stopped at the screen.

Agents are a different story, because they don’t just answer, they act. They send the email, update the price, delete the duplicate record and deploy the page, so when an agent gets something wrong, the mistake doesn’t stay on the screen. It lands in the real world, in someone’s inbox, in your database, or in a bank account.

Most teams building agents put almost all their energy into making the agent right more often, and that does matter. But the question we think matters more is what happens when it is wrong anyway. Our answer is that every agent needs an undo button, and that you have to design it in from day one rather than bolting it on after the first incident.

Why better models won’t save you

The usual response we hear is that the next model will be smarter, so it will make fewer mistakes. That’s true, but it misses the point in a fairly important way.

Agents rarely take one action; they take long chains of them. Say each step in a task is right 99% of the time, which is honestly better than what most teams see when they actually measure it. Over a 50-step task, the odds that every single step goes right are 0.9950, which works out to roughly 60%, so about four runs in ten will contain at least one mistake somewhere.

Even if you push per-step accuracy all the way to 99.9%, a 50-step task still goes wrong about 5% of the time, and if that agent runs a thousand times a day, you’re looking at around fifty bad outcomes every single day. Models will keep getting better, but we’ll also keep handing agents longer and more ambitious tasks, so the error rate per task doesn’t really disappear, it just moves around.

There’s a second, sneakier problem too. The mistakes that slip past a strong model tend to be the confident ones. A weak model usually fails loudly and obviously, whereas a strong model fails plausibly, in a way that sails right past a quick human glance, and those are exactly the errors you need to be able to catch and fix after the fact.

So the real question was never how to make the agent perfect:

It’s not “how do we make the agent never wrong?” but rather “how cheaply can we recover when it is?”

Reversibility is a design property

The first step is boring but essential. You list every action your agent is able to take, and then you sort each one by how hard it would be to take back. We like to use four tiers for this.

Tier What it means Examples Default policy
0. Read-only Changes nothing at all Searching, fetching a record, reading a file Let it run freely
1. Reversible Undo restores the exact prior state Editing a draft, changing a setting, moving a file Let it run, but log the prior state
2. Compensable Can’t be undone directly, but a follow-up action fixes it Charging a card (refund it), publishing a page (unpublish it), creating an order (cancel it) Let it run, with the compensating action recorded upfront
3. Irreversible Once it’s done, it’s out in the world Sending an email or SMS, hard-deleting data, paying a vendor, posting publicly Put it behind a human or a hold window

Two things usually fall out of this exercise. The first is that most teams discover their agent has far more tier 3 actions than they had assumed, because something like “send a notification” sounds completely harmless right up until the agent sends 4,000 of them to the wrong segment.

The second is that a lot of tier 3 actions can be moved down a tier with a bit of engineering. A hard delete can become a soft delete with a 30-day purge, an email can sit in an outbox that only flushes after ten minutes, and publishing can go to a staging URL before it goes live. A big part of your job as an architect is pushing as many actions as you can down towards tier 0 and tier 1.

Five patterns that make agents undoable

1. Plan first, then act

Have the agent produce a structured plan before it touches anything, covering which actions it will take, on which records, and in what order. A plan is cheap to inspect, diff and reject, and it also gives you a natural place to check policy, for example “this plan touches 3,000 rows, which is more than we allow for an unattended run.”

2. Dry runs and shadow writes

Let the agent execute against a copy of your data, or in a mode where writes are recorded but never actually committed, and then show the diff. Developers already trust this pattern because it’s exactly how git and terraform plan work, and there’s no good reason agents should work any differently.

3. An action log that keeps the prior state

Every write the agent makes should be logged along with whatever was there before it, so that undo becomes a mechanical replay in reverse instead of a forensic investigation. This needs to be a proper event log and not a chat transcript, because “the agent said it updated the price” is useless to you at 2 a.m., when what you actually need is the old price and the new one.

4. Compensating actions, recorded at write time

For tier 2 actions, the agent should record the reverse action at the very moment it acts, whether that’s the refund that would reverse a charge or the unpublish call that would pull a page. If you wait until something has gone wrong to work out how to reverse it, you’ll be working it out under pressure, and that rarely goes well.

5. Checkpoints and human gates

Long tasks should have checkpoints, so that a failure at step 38 rolls you back to step 30 and not all the way to zero. Tier 3 actions should go through a person, or through a hold window where, say, the email waits in an outbox for ten minutes where anyone can see it and cancel it. A good gate shows the human what will actually happen rather than what the agent intends, and a one-line diff will beat a paragraph of the agent’s reasoning every time.

None of this is particularly new. Databases have had transactions for decades, finance has reversing entries, and ops teams have runbooks and rollbacks. What is new is that the thing making the changes is probabilistic, which moves these patterns from being nice to have to being genuinely load-bearing.

What it looks like in a real build

Take an agent that edits a merchant’s live online store, where the merchant types something like: “Run a 20% Diwali sale on all ethnic wear and put a banner on the homepage.”

A naive agent goes straight to work. It updates the prices on every product it thinks counts as ethnic wear, rewrites the homepage, and cheerfully reports back “Done!” If it has misread the catalogue and discounted the entire store, the merchant only finds out when the first underpriced orders start coming in.

An undoable version does exactly the same work, just in a different shape:

  1. Plan. It first produces a change set listing 142 products with their old and new prices, along with a diff of the homepage, and nothing gets written yet.
  2. Policy check. Because the change set touches prices, it gets flagged for review, whereas a banner change on its own would have gone straight through.
  3. Preview. The merchant sees the store on a staging URL along with a short summary like “142 products discounted, lowest new price ₹480,” and can approve or edit it with a single tap.
  4. Commit with a log. Once approved, every write is logged with its previous value and grouped together under one change ID.
  5. One-click revert. If something looks off the next morning, “Undo Diwali sale” puts all 142 prices and the old homepage back in a single action, and since ending the sale is the same operation, it doubles up as a scheduled rollback too.

The model behind both versions is identical. The second one is simply safe enough to hand to a non-technical merchant, and in our experience that’s what decides whether a feature like this ships at all.

A checklist before you ship an agent

If you’re a founder about to put an agent in front of real users, these are the questions worth asking your team. If the answer to any of them is “we’ll add that later,” then later usually turns out to mean right after the first incident.

  • Do we have a list of every action the agent can take, sorted by how reversible it is?
  • Can we show what the agent is about to do before it does it, in a form a user can actually read?
  • Is every write logged along with the state it replaced?
  • For every action that can’t be undone, do we have a recorded way to compensate for it?
  • Do irreversible actions go through a person or a hold window?
  • Are there hard limits on how much a single run can do, such as rows touched, emails sent or money moved?
  • Can a user undo the agent’s last task in one action, without having to call support?
  • Have we tested the undo path as seriously as we’ve tested the happy path?

Trust is built on recovery, not perfection

We don’t trust banks because they never make mistakes, we trust them because a wrong charge can be reversed. We trust version control for the same reason, because nothing is ever really lost. People will come to trust agents in much the same way, not when they become flawless, but when their mistakes become cheap to fix.

The teams that win with agents probably won’t be the ones with the cleverest prompts. They’ll be the ones whose users feel safe pressing “go,” because they know they can always press “undo.”

If you’re building an agentic product and would like a second pair of eyes on its architecture, this is the kind of problem WBOS works on. Book a 30-minute build session.


Srijan is the founder of WBOS, an AI and product engineering studio. The storefront walkthrough is an illustrative example, and the error-rate figures are simple compounding arithmetic rather than measured benchmarks.

Have Ideas?

We will me more than excited to take your idea to production for you

Partner. Build. Scale

Partner. Build. Scale

READY TO START ?

Let's discuss how we can engineer it into reality. We are ready to partner with you.

All rights reserved. WBOS.
POWERED BYOPMIZ