Playbook

How to use AI in marketing: start with the jobs where a mistake is cheap to catch

A vintage jeweler's loupe resting on a blank sheet of paper on a worn wooden workbench, in grainy black and white

One disclosure before the argument. This publication is produced with support from Docket, which sells software in this category. Nothing below names or scores a product, and the method works the same whichever vendor, or no vendor, you use.

The first jobs to give a marketing agent are the ones where a wrong answer stays inside your team and a person can catch it faster than they could have done the work. That rules out most of the jobs current AI budgets favor.

Budgets follow the jobs that are easiest to show

MIT's Project NANDA asked executives to allocate a hypothetical $100 of generative AI spend across functions. In The GenAI Divide: State of AI in Business 2025 (Challapally and colleagues, July 2025), sales and marketing took the largest share, while back-office automation often delivered the better return. Fortune's coverage reports the same split.

The report is loose with its own headline number. The takeaway for that section says 50% of budgets go to sales and marketing; the paragraph beneath it says roughly 70%. The authors describe the functional split as consistent across interviews and warn that only the use-case breakdowns beneath it are "directional at best." Treat the direction as the finding.

Sales and marketing dominate, the authors write, "not only because of visibility, but because outcomes can be measured easily."

Our inference from that: the property that makes a marketing job attractive to fund is often the property that makes it a bad first job. An output that moves a board metric is an output that reaches the market, and a mistake in it reaches the market too.

Plausible and wrong is the failure to design around

The best public evidence on where AI help turns harmful comes from a preregistered field experiment with 758 Boston Consulting Group consultants, published in Organization Science in March 2026.

Both halves of the experiment were go-to-market work. On 18 tasks the researchers judged inside the model's capability, among them segmenting a footwear market and drafting press release copy, consultants with GPT-4 completed 12.2% more tasks, 25.1% faster, at significantly higher quality.

The task chosen to sit outside that capability was a brand analysis: recommend which of a company's brands held the most potential for growth, using spreadsheet data that looked complete and interview notes that changed the answer. The control group got it right about 84.5% of the time. The two AI groups scored 60% and 70.6%, and spent less time on the task.

Our reading of the design: the harm tracked whether the right answer depended on a detail a fluent summary would skip, and a single task can only suggest that.

The same paper measured how convincing the answers were. Graders who were not told the correct solution scored each recommendation for clarity, persuasiveness, and coherence. The AI groups outscored the control group on that measure, and they did so among the participants who got the answer wrong as well as those who got it right.

The experiment ran on a 2023 model, and capability has moved since. The finding that should carry over is the second one: the wrong answers read better. We infer that a reviewer skimming for quality would have approved them.

So the check is the control, and a first job has to be one whose check is cheap.

That cost is the one marketing teams piloting agents raised with us this quarter. A lifecycle marketing manager at a B2B software company described spending several hours previewing and testing a single agent-built campaign before it could launch.

The gates a first job has to clear

A job qualifies as a first job only if it clears all four gates. Failing one is enough to wait.

The mistake stays inside the team. The output reaches a colleague before it reaches a buyer, a list, or a budget. We have argued the ordering principle separately: the least correctable action sits behind the most gates. A first job should have no uncorrectable action in it at all.

Checking is faster than doing. The success criteria exist before the job runs, so review means comparing the output against something already written down.

It recurs at least weekly. Enough runs to see an error rate within a month, and enough volume that the setup pays back.

The inputs already exist where the agent can read them. If the job first needs someone to assemble a spreadsheet by hand, the bottleneck is the spreadsheet.

Running a demand gen calendar through the gates

We listed the recurring jobs on a typical B2B demand gen calendar and scored each one as the job with the agent taking the action it names. The inventory and the scores are our editorial judgment, and yours will differ.

JobStays insideFaster to checkWeeklyInputs on handFirst job?
Pre-launch campaign QAYesYesYesYesYes
Database hygiene proposalsYesYesYesYesYes
Weekly performance readoutYesYes, if scopedYesYesYes
Competitor change logYesYesYesYesYes
Webinar and form comment taggingYesYesYes, if pooledYesYes
Nurture email sendsNoNoYesYesNo
Social publishingNoYesYesYesNo
Paid budget reallocationNoNoYesYesNo
Lead score model changesNoNoNoYesNo
Landing page publishingNoYesNoYesNo
Positioning and messagingYesNoNoNoNo
Quarterly plan and forecastYesNoNoNoNo

Five of twelve pass. That count comes out of the gates, and a calendar heavier on events or partner work would produce a different one.

Pre-launch campaign QA. The agent receives the campaign brief and the built assets a day before launch and returns every mismatch between them: a UTM that does not match the naming convention, a link that resolves to the wrong page, a form missing a required field, a suppression list that is not attached. Each finding cites the line of the brief it violates. It reports and fixes nothing. This inverts the lifecycle manager's afternoon: the agent checks the person's build against a written spec, instead of the person checking the agent's.

Database hygiene proposals. Duplicate candidates, job titles mapped to your persona list, records missing the fields your reports depend on. The agent writes to a review queue, never to the record. A person approves a batch after sampling twenty rows from it.

Weekly performance readout. This is the one that complicates the method, because it clears every gate and still sits closest to the brand-analysis trap. A narrative about why pipeline moved is exactly the persuasive, checkable-looking output the experiment above caught out. So scope it to what moved: every figure linked to the report it came from, every week-over-week change past a set threshold flagged. The why stays with a person.

Competitor change log. A weekly diff of competitors' pricing pages, product pages, release notes, and job postings, each entry carrying the URL and the date seen. The inputs are public and the check is one click.

Webinar and form comment tagging. Webinar questions and free-text form comments, pooled weekly and tagged against a fixed theme list, with a count and three raw examples per theme. The fixed list is what makes it checkable; letting the agent invent themes moves the job back outside the gates.

The rejected rows fail for one of two reasons. Either the mistake leaves the building, or nobody can say in advance what a correct answer looks like. Lead scoring fails both. A growth marketing leader at a B2B software company told us sellers stopped trusting a model score once they found cases where it was wrong and nobody could explain why.

The obvious objection is to keep the agent on nurture email but have it draft for approval, so the mistake stays inside. It does, and the job moves to the second gate, where it fails. A single social post is one artifact against one brief. A nurture program is several steps across several segments with no written answer to compare against, so the reviewer rereads every variant. That is where hours like the lifecycle manager's go. Drafting for review becomes a reasonable second job once the review has criteria of its own.

What the rejected column costs

NANDA's exhibit lists six sales and marketing uses. Two of them, competitor analysis and social sentiment analysis, map onto jobs that pass the gates. The other four, AI-generated outbound emails, personalized content for campaigns, follow-up automation, and lead scoring, fail as first jobs, because each puts output in front of buyers or sellers before anyone has written down what correct looks like. The use-case list is directional by the authors' own account, so read that as an illustration of the gates sorting jobs within a single budget category.

A team that starts there will spend its pilot checking output. Its pilot results will then measure how much review the team could stand, not what the agent can do, and a result like that tends to decide the next budget cycle either way.