top of page

My AI Made 44 Changes in a Client's Ad Console Last Week. Here's the Supervision System.

  • Jul 17
  • 4 min read

Last week an AI made 44 changes inside a client's live ad console. Bid adjustments, budget shifts, pauses, negative keywords — executed directly in the platform, not suggested in a doc for someone to action later. That sentence makes some people lean in and others reach for the fire alarm. Both reactions miss the point.


The novelty is not the story. The supervision system is the story. Autonomous execution without guardrails is reckless whether a human or a model is doing it. With the right guardrails, it is simply faster operating. Here is the system that makes it safe.


Guardrail 1: Every change writes a log


Nothing happens silently. Each change produces a record: what changed, the before state, the after state, and the reason. Forty-four changes means forty-four rows I can read in two minutes. If I cannot explain a change from its log line, the change should not have happened. The log is not paperwork — it is the difference between an operator and a black box.


Guardrail 2: Screenshot verification


Executing a change and confirming a change are two different things. The system captures the console state after each action, so 'paused' means a screenshot shows it paused, not that an instruction was sent hopefully into the void. This closes the gap between what a tool intended and what the platform actually did — a gap that quietly wrecks more accounts than bad strategy does.


Guardrail 3: A keep bar, not a recency reaction


Every candidate for pausing gets scored against a defined keep bar — a threshold tied to the ad's best sustained performance window, not just its worst recent days. This stops the single most common operating error in paid media: panicking at a bad week and killing something that was carrying the account. The model does not get to pause on vibes. It has to clear the bar.


Guardrail 4: Kill criteria defined in advance


Before anything runs, the conditions for stopping are written down. Spend thresholds, performance floors, integrity checks. If a data feed looks wrong, execution halts rather than acting confidently on bad numbers. Kill criteria set beforehand are worth ten times more than judgment applied after the money is gone.


The keep bar is not just theory — it catches real mistakes, including human ones. On one account I ran the model back over a week of pause decisions a person had made. Eleven of eighteen pauses were justified. Several were not — including one genuine winner that got paused twice. The pattern was textbook recency bias: a strong ad hit a rough few days, the buyer panicked, and off it went. The model, judging each ad against its best sustained window instead of its worst recent one, kept the winner alive. This is not an argument that humans are bad at the job. It is an argument that a keep bar protects everyone from the same predictable error.


Guardrail 5: A scheduled scorecard


Every Monday the system runs a check against the account: what moved, what the changes produced, where reality diverged from the plan. It is the standing review that turns a pile of individual actions into a story I can defend to a client. Automation without a scheduled review is just faster drift.


What I don't let it do


Autonomy has a boundary, and the boundary is where a mistake stops being reversible. The AI adjusts bids, shifts budgets within limits, adds negatives, and pauses against the keep bar — all things I can undo in minutes. It does not launch brand-new campaigns unsupervised, it does not touch pricing, and it does not make irreversible or account-structural changes without a human in the loop. The rule is simple: the more permanent the consequence, the more human judgment sits in front of it. Speed is for the reversible decisions. People are for the one-way doors.


The most valuable thing the system does is not even execution — it is pushback. More than once it has refused to build the chart I asked for, because the underlying data would have been misleading, or flagged a distortion I had glossed over in my own first draft. An operator that only says yes is a liability. One that occasionally says 'this number doesn't mean what you think it means' is worth keeping around.


The real lesson


People fixate on the 44 changes. What actually made the week safe was boring: a log, a screenshot, a threshold, a stopping rule, and a weekly review. Those five things would make a human media buyer more accountable too. The AI just makes following them non-negotiable.


Autonomous execution is not risky because a machine is doing it. It is risky when nobody wrote down how it is allowed to fail.

If you are experimenting with AI in your ad accounts, do not start with autonomy. Start with the change log. Add screenshot verification. Define your keep bar and your kill criteria. Schedule the scorecard. Autonomy is what you earn once the guardrails hold — not the feature you switch on first.


Comments


bottom of page