Playbook

How to build a marketing agent for production: put the controls in the code that performs the write

An antique fountain pen with a small padlock locked through its clip, resting on a blank sheet of paper on a worn wooden workbench, in grainy black and white

One disclosure before the argument. This publication is produced with support from Docket, which sells software in this category. Nothing below names or scores a product, and every design choice in it can be built on any stack.

Our argument is that a marketing agent is ready for production only when the controls it depends on are enforced by ordinary code at the point where the write happens, checked at the moment of the write, on every path that can change what a customer sees. When the controls live in the prompt, in a policy document, or in a reviewer's memory, that is where they fail.

We have described the states and gates this publication runs on, and the tests a system should pass before it earns the word agent. This piece is about the layer underneath both: where each control physically lives in the build, and how it gets bypassed.

Split the system into a proposer and an executor

The first architectural decision we would make is to separate the part that thinks from the part that acts.

The model proposes. It researches, drafts, classifies claims, and emits a structured request: publish this package, send this email, update this record.

A separate executor decides whether the request runs. It is deterministic code with no language model inside it, and it is the only component that holds the API credentials.

The OWASP GenAI Security Project's guidance on what it calls excessive agency points the same way. Its mitigation, complete mediation, reads: "Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not."

Anthropic's guidance on building agents notes that "the autonomous nature of agents means higher costs, and the potential for compounding errors." A rule written into a prompt is one more input the model weighs. A rule written into the executor is a condition the request must satisfy before anything leaves the building.

A practical test: could the model, given a persuasive enough input, publish something your rules forbid? If the answer depends on the model declining, the rule lives in the wrong layer.

Check the approval at the moment of the write

Security engineers have a name for the most common approval bug. MITRE catalogues it as CWE-367, time-of-check time-of-use: "The product checks the state of a resource before using that resource, but the resource's state can change between the check and the use in a way that invalidates the results of the check."

In a marketing agent, the check is a human approving a draft on Tuesday. The use is the publish on Thursday. Anything edited in between, a fixed typo, a swapped statistic, a regenerated paragraph, rides through on Tuesday's approval unless the executor looks again.

So the executor looks again. This publication's runtime, described here as method rather than offered as proof, has a publish command that re-runs every validation check, reads a pause flag, and confirms that the exact-version approval still matches the package's current hash and has not expired. That command is built to do all of this inside the same step that builds the CMS request from that same package, so the window between check and write shrinks to a single command.

That command is not yet our live path. Today our pages reach the CMS through a connector that the drafting agent's session calls directly, and the hash check runs afterwards, when the publication receipt is recorded. On that path the check comes after the write, and the separation between proposer and executor does not yet hold. By the standard in our argument, that path is not yet production grade.

A control that misses one path is a suggestion

The most instructive failure in our own runtime came from a gap in coverage. The gate itself worked.

In July 2026 the piece on the agent ownership gap was one of the two July corrections described in our piece on owning the loop. The correction was made directly in the CMS editor, outside the runtime, where the approval gate was simply not on the path.

The corrected page stayed live for 401 hours with no approval on record for the content readers were seeing. What surfaced it was reconciliation, the receipt-time hash match described above: the runtime refuses to issue a publication receipt for live content unless it can match that content, by hash, to a recorded approval that predates publication. When a receipt was finally requested, that refusal surfaced the gap. The page was reconstructed into the runtime verbatim, approved at its exact hash, and logged as a retroactive approval exception rather than receipted as if nothing had happened.

There is a real reason the fast path exists. A correction to a live error has to be quick, and forcing every emergency edit through a two-day approval cycle would keep wrong information in front of readers longer. Reconciliation catches the bypass without preventing it, and it only catches it when it runs. That is why we now recommend running it on a schedule rather than waiting for someone to ask for a receipt.

The general rule: inventory every way content can reach the public, including the admin screen, and either route it through the executor or, as a stopgap that falls short of that standard, reconcile it against the executor's record on a timer.

Give each credential only the writes its job needs

The model never holds a credential in this design, but the executor does, and its scope decides the blast radius.

OWASP's guidance is to "limit the permissions that LLM extensions are granted to other systems to the minimum necessary." OpenAI's practical guide to building agents makes it operational: rate each tool low, medium or high "based on factors like read-only vs. write access, reversibility, required account permissions, and financial impact," and use the rating to pause for checks or escalate to a human before high-risk calls.

In four to six recent conversations with marketing and revenue operations leaders whose teams had tried building their own agents, the failure they described was a bad write: leads routed to the wrong queue, existing contact owners overwritten in the CRM, one system's field filled with another system's data. That is our observation from those conversations, and it is not a measured rate.

An agent that drafts nurture emails needs read access to segments and write access to drafts. It does not need write access to the owner field, the lifecycle stage, or the suppression list, and a credential that can reach them eventually will.

Make retries return the first result

The failure that looks least like an AI problem is the retry.

An executor calls the email provider, the connection times out, and the executor cannot tell whether the send happened. Retrying blindly risks a second send to the whole list. Stopping blindly risks a send that never went.

Payment APIs handle this with idempotency. Stripe's design lets a client retry "without accidentally performing the same operation twice": the server saves the result of the first request for a given key and returns that same result to any retry. For a marketing executor, a natural key is the package identifier plus its approved version hash, so one approved version can produce at most one send.

Where a provider offers no idempotency, record partial state instead of guessing. When a live publish fails partway, this publication's publish command is built to write down whether the page verified, marks the package as published or blocked accordingly, and opens an exception for a human rather than retrying.

Put the kill switch in the executor's path

The stop conditions matter less than where they are read. In the design above, pausing is a flag in a control file that the executor reads before every write it performs, and the model has no path to change it. Publishes through our CMS connector do not read it yet.

Where each control lives

This is our framework, assembled from the sources above.

ControlEnforced byChecked whenFails open if
Content rules and claim checksExecutor validationInside the publish commandThey live only in the prompt
ApprovalHash comparison in the executorAt the moment of writeIt is checked only when granted
CoverageReconciliation against the live surfaceOn a timer, after any bypassA human edits the CMS directly
PermissionsCredential scope per toolAt provisioningOne broad key serves every tool
RetriesIdempotency key or recorded partial stateOn every retryThe executor retries blindly
StopPause flag read by the executorBefore every writeThe model is asked to stop itself

Anthropic recommends "extensive testing in sandboxed environments, along with the appropriate guardrails." The table doubles as a test plan: trigger each row's failure deliberately in staging and confirm the executor refuses.

The cost is speed of change

Every row in that table makes the agent slower to change. A new tool needs a risk rating and a scoped credential. A new publishing path needs routing or reconciliation. A copy fix needs a fresh approval.

That friction is the price of letting the interesting part, the model, stay flexible upstream. Our prediction, offered as inference: the first serious incident in most marketing agent deployments will not be a hallucinated sentence. It will be a write that went through a path nobody gated.