An agency running a dozen accounts produces the same five artefacts over and over: research, briefs, drafts, QA notes and reports. AI can take real hours out of every one of them, and it can also put a fabricated statistic in front of a client with total confidence. The difference between those outcomes isn't the tool or the prompt. It is whether each output gets checked against something real before anyone acts on it.
Why the same model helps and fabricates
A language model writes the most plausible continuation of whatever you give it. Both halves of its behaviour follow from that one mechanism. Paste in the top-ranking pages for a keyword and ask what they cover, and the plausible continuation is a faithful summary, because the answer sits in the input. Ask what share of searchers click the first result, and it has nothing to measure, so it produces a figure shaped like the figures it has seen. It isn't malfunctioning. Fabrication is the same completion process running without source material.
Three fabrications show up constantly in agency work, because they read exactly like the real thing:
- Statistics and studies. A specific figure with a credible-sounding source is the model's default move when a claim needs support.
- Quotes. Ask what an expert or a document says and you can get a clean, confident sentence nobody ever wrote.
- Internal links. Ask a draft to link related pages on the client's site and it will invent URLs that match the site's pattern and don't exist. They look right, and they 404.
That gives you the test for any step you're thinking of handing over. Did you supply the source material, and can the output be checked against something concrete? Two yeses: give the step to AI and check the result. Either no: keep a human on it, or redesign the step until both answers are yes.
Research and briefs: strong, because you feed it
Research is the step AI compresses hardest, and the safest, because you control the input. Paste in the ranking pages and ask what they all cover and what none of them answer. Paste an export of a client's queries and ask for clusters by intent. Paste a client call transcript and ask for the questions customers keep raising. Each answer is checkable in minutes, because everything it drew on is in front of you.
Briefs sit one step up. AI turns a keyword, an intent and an audience into a structured brief quickly, and the structure is usually sound. What it can't know is strategy: which cluster this client should chase this quarter, what the page must promise, what the client won't say in public. A workable split is that AI drafts the skeleton and a human writes the three lines that matter, the goal, the angle, and the one thing the page must get right.
Drafts and QA: fast on structure, blind on truth
A model produces a competent first draft far faster than a person, and the competence is real: coverage, headings, readability. What doesn't arrive is anything only your client knows, plus facts you can trust. So the human pass on a draft isn't a vibe read. It has named checks: every statistic traces to a source you can name or it comes out, every quote is verified or cut, every internal link gets clicked, and someone who knows the client reads every claim the client would have to defend.
No automated score replaces that pass. Scoring engines measure form: length, structure, headings, keyword placement. A draft with an invented figure in every section can still score well, because correctness isn't something a scorer can see.
QA is the reversal, and it's where AI quietly earns the most. Checking a draft against a brief is a rubric task: whether every section is present, whether the intro answers the query, whether the tone has drifted, whether anything is asserted without support. You supply the draft and the rubric, so the output is checkable by construction, and it catches the boring misses a tired reviewer skips.
Reporting: it narrates whatever you hand it
Turning a table of movements into a plain-English paragraph is a genuine AI strength, and clients read the paragraph, not the table. Two cautions keep it honest. The model narrates wrong numbers as fluently as right ones, so the figures come from your tracking, confirmed by a human, before any prose is written around them. And it will happily supply causes, "rankings rose because of the new content", that nothing in front of it supports. Let AI write the description of what moved. The sentence claiming why it moved is yours, and Proving ROI honestly covers how to earn it.
What to take away
- AI is safe on steps where you supplied the input and the output can be checked, and dangerous where you ask it for facts it has no way to hold.
- Statistics, quotes and internal links are the three fabrications to hunt in every draft, because they read as real and fail as facts.
- Research, brief skeletons and rubric-based QA are the high-return uses; final facts, strategy and causal claims stay human.
- No scoring engine can detect a made-up fact, so a human fact pass stays in the pipeline at any volume.
Next
AI gets far more useful when it can measure instead of guess. Claude as your SEO analyst connects an assistant to live ranking data.