Skip Navigation or Skip to Content
An editor's desk with a long-form article draft on screen covered in red review annotations beside printed research papers

Table of Contents

08 Sept. 2026

The 23-Check Citation Standard: How We Decide What Gets Published

What is a citation quality standard?

A citation quality standard is a fixed set of checks an article passes before it goes live, covering whether the page deserves to exist at all, whether every claim traces back to a primary source, whether the page is built to be quoted, and whether the live version matches what was written. Ours has 23 checks arranged in four gates.

The objection that prompted this article shows up in most first conversations. A CMO looks at a publishing operation running on AI assistance and asks the obvious question: isn't this just spam at scale? Fair question. The answer turns entirely on whether anything gets checked. Google's own spam policy draws the line in exactly that place. Scaled content abuse is defined as generating "many pages for the primary purpose of manipulating search rankings and not helping users", and the policy names one specific behaviour: "using generative AI tools or other similar tools to generate many pages without adding value for users."

Read that again. The policy does not care who typed the words. It cares whether value was added. That moves the argument off human versus machine and onto three questions: did a check happen, what did it test, and what happened to the drafts that failed it. Most operations cannot answer any of the three. This article answers all three and hands you the checklist.

45%

AI answers with a significant issue

EBU/BBC, 3,000+ responses, Oct 2025

31%

Responses with sourcing problems

Missing, misleading or wrong attribution

68.01%

US Google searches ending without a click

SparkToro and Similarweb, Jan to Apr 2026

2,022

Court decisions logged over AI-fabricated citations

Charlotin database, 5 Sept 2026

What you'll learn:

  • What Google's spam policy actually says about AI-assisted content, in its own words
  • The measured failure rate of unchecked AI output, and where the failures cluster
  • All 23 checks, grouped into the four gates they sit in, with the failure mode for each gate
  • What happened when we ran the standard against our own 313 published pages
  • How to order the checks so the standard costs hours rather than days

Key Takeaway

What separates a publishing operation from a spam operation is whether a documented check runs before publication, and whether anything ever fails it. A standard nothing fails is not a standard.

Does Google penalize AI content?

No. There is no AI penalty in Google's documentation, and no published Google system that detects authorship and demotes it. What exists is a spam policy aimed at volume without value, and it applies identically to a person writing 400 thin pages by hand and a model generating them overnight.

The same policy page also defines site reputation abuse, which catches third-party content published on a host site mainly to borrow that host's earned ranking signals. Both policies describe a pattern of behaviour rather than a production method. If you want the mechanics of how engines choose sources in the first place, we cover that separately in how LLMs decide what to cite.

The practical risk sits somewhere else entirely: publishing a page that carries a number nobody traced and a claim nobody could defend in a sales call. That page gets quoted back to you eventually, and by then it is on your domain with your name in the byline.

What actually goes wrong when nobody checks?

The largest study of the problem to date came from the European Broadcasting Union and the BBC. Twenty-two public service media organisations across 18 countries evaluated more than 3,000 AI assistant responses in 14 languages. Forty-five percent of answers carried at least one significant issue, and the single largest category was sourcing: 31% of responses had missing, misleading or incorrect attribution. Accuracy problems, including invented details, appeared in 20%.

Note which number is bigger. The models were wrong about facts less often than they were wrong about where the facts came from. Attribution is the weak joint, and attribution is precisely what an editorial standard can fix without touching the model.

A hand with a red pen circling a statistic on a printed page and drawing an arrow to the source line beneath it

The legal profession has the most complete public record of what happens next, because courts write it down. Researcher Damien Charlotin maintains a running database of court decisions involving AI-hallucinated content, mostly fabricated case citations. As of 5 September 2026 it listed roughly 2,022 decisions. These are not people who set out to deceive. They are professionals who assumed a citation that looked correct was correct and never opened it.

Marketing has no equivalent register, so the failures stay invisible. A fabricated statistic in a blog post produces no sanction and no correction. It circulates, gets quoted by the next article, and hardens into something everyone knows.

That hardening happened inside the production of this article. The research step returned 19 sources. Exactly one was the organisation that had published the underlying policy. Six were flagged by our screening step as likely AI-authored. The rest were marketing blogs of unclear provenance, several repeating each other's numbers with no traceable origin.

Source classReturnedUsed in this article
Primary source (the body that published the data)1Yes
Secondary marketing blog, provenance unclear12No
Flagged as likely AI-authored by the screening step6No
Total191

Source: peppereffect production run for this article, 8 September 2026. Screening verdicts produced by the research pipeline's source classifier.

One discarded page carried a tidy claim that a March 2024 Google change produced a "45% reduction in low-quality content", alongside a correlation coefficient attributed to a third-party tool. Neither traced to anything. Both are the kind of number that survives because it sounds specific. Every figure in this article was opened, read and dated by hand, which is why the list at the bottom is short.

Three B2B marketing professionals reviewing a printed content quality checklist pinned to a glass office wall

What are the 23 checks?

Four gates, in order. A draft that fails any check inside a gate stops there and does not move to the next one. The order is deliberate: the cheap checks that kill the most work come first, and the expensive checks that require a finished page come last.

1

Does this page deserve to exist?

Five checks, all runnable before a word is written. This is the gate almost every content operation skips, and skipping it is how a blog ends up with four pages competing for one query.

2

Is every claim traceable?

Seven checks against the evidence itself. This is where drafts die, and it is the gate that separates an editorial operation from a content mill.

3

Is it built to be cited, and to earn the click?

Six checks on structure. Getting quoted by an answer engine and getting a human to click through are two different jobs, and both have to be designed in.

4

Is it live, and is what is live what we wrote?

Five checks against the published page. A draft passing is not the same thing as a page passing.

Gate 1: does this page deserve to exist?

#CheckWhat failing looks like
01Live-inventory duplicate checkA second URL on a topic that already has a page
02Cannibalisation check in Search ConsoleTwo of your own pages splitting impressions on one query
03Click-defensibilityEverything on the page can be restated by an AI answer in three sentences
04Query evidenceInvented FAQ questions nobody has ever typed
05Topic fitA high-volume topic that has nothing to do with what you sell

Source: peppereffect editorial standard v1.0, September 2026.

Check 03 fails most often. A topic with volume and no defensible asset feels like an opportunity right up to the moment an AI Overview answers it completely. The five things that survive are an interactive comparison, a calculator, a diagnostic, a template, and a proprietary verdict. If a planned page carries none of them, it does not get written, whatever the search volume says. That rule is why we publish two or three pieces a week instead of ten.

Gate 2: is every claim traceable?

#CheckWhat failing looks like
06Named source on every number"Studies show" with no study named
07Specific page URL, never a root domainA statistic linked to a company homepage
08Date on every figureA 2021 benchmark presented as current
09Primary source, not a repeatThree blogs citing each other and no study underneath
10Blocked-numbers list checkedReusing a figure the team already investigated and refused
11Multiple always paired with the absolute"Up 83x" with no base, so growth reads as scale
12Correlation described as correlationA causal verb sitting on an observational finding

Source: peppereffect editorial standard v1.0, September 2026.

Checks 07 and 09 fail together. A number gets attributed to a well-known company that never published it, and the link points at that company's homepage, so nobody can check. We keep a standing blocked-numbers list for exactly this, and the entries are specific: one widely repeated zero-click figure sits on it because the company it is attributed to never published it. The real source was a different firm entirely, and the number had drifted in the retelling.

Check 11 exists because of one uncomfortable review. A client result of 83 times more AI-sourced visits is true and verifiable, and standing alone it invites a reader to imagine a flood. Paired with its absolute, 250 visits in a quarter, it tells the truth: a real multiple on a small base, in a channel that is still early. The multiple never travels alone now.

Gate 3: is it built to be cited, and to earn the click?

#CheckWhat failing looks like
13Answer first, in 40 to 60 wordsThree paragraphs of throat-clearing before the answer
14Question-shaped headings from real queriesClever headings nobody searches for
15Extractable units, source line under every tableA table an engine can quote with no attribution attached
16FAQ built from query data, one FAQ headingTwo FAQ sections and five invented questions
17Valid schema with the correct author and publisherSomeone else's name in your publisher field
18Link map in both directionsA new page with no inbound links, orphaned on publication

Source: peppereffect editorial standard v1.0, September 2026.

Structure earns the citation. Pew Research tracked 68,879 Google searches from 900 US adults during March 2025 and found that 88% of AI summaries cited three or more sources, with a median summary length of 67 words. Three slots, 67 words. A page that buries its answer under context does not get into that window, which is the whole argument for check 13. We go deeper on the structural side in our guide to schema markup for AI citation and the answer engine optimization hub.

Check 17 fails quietly. When we rebuilt our own AEO hub in September 2026, the page had been carrying JSON-LD naming a different agency as both author and publisher, inherited from a template months earlier. Every engine reading that page had been told, in machine-readable form, that someone else wrote it. Nothing on the visible page was wrong. The check that caught it was a schema read, not a proofread.

The full 23-check standard as an interactive checklist you can work through or print, with the failure mode called out for each gate.

Open the 23-check template

Gate 4: is it live, and is what is live what we wrote?

#CheckWhat failing looks like
19Body verified twice, on draft and on liveA placeholder published to the live URL
20Images on your own CDN with descriptive alt textHotlinked images and alt text that reads "IMG_4021"
21Metadata completeA meta description auto-cut at 300 characters
22Live render check in a browserComponents that render in the editor and collapse live
23Voice checkSentences that could sit unchanged in a competitor's article

Source: peppereffect editorial standard v1.0, September 2026.

Check 19 is the one people find absurd until it saves them. Content platforms keep separate draft and live versions, and the sync between them fails silently more often than any vendor admits. The failure mode is a live URL serving a placeholder while the CMS shows a finished article. Verifying the published body against the draft costs one API call and has caught real placeholder publications in our own operation.

A reviewer comparing a marketing article against the original research PDF on two screens to verify a cited statistic

What happened when we ran this against our own blog?

In August 2026 we applied the standard retroactively to every published URL on this domain. Three hundred and thirteen pages went in. Seventy-seven came out.

The pages that failed were not badly written. Most were competent articles on subjects that had nothing to do with what we sell, published during a period when the operating theory was that volume compounds. Gate 1 killed them on check 05 before anything else got evaluated. Thirty-nine more were merged into stronger pages that already existed, which is check 01 doing its work three years late.

Bar chart of 313 peppereffect URLs reviewed in August 2026 showing 75 kept, 2 rebuilt, 39 merged, 196 redirected and 1 unpublished

Source: peppereffect prune decision list, 30 August 2026 (n=313). Full decision record held internally, one row per URL with its redirect target.

Key Takeaway

We deleted or redirected 236 of our own 313 pages to pass our own gate. A standard you have never failed is a standard you have never enforced. If you want the diagnostic version of this exercise, start with why organic traffic drops and the impressions-up, clicks-down pattern.

How do you run 23 checks without stalling the calendar?

By ordering them by cost. Gate 1 runs before anything expensive exists, which is why it removes the most work: in the retroactive review of our own blog, all 236 removals were Gate 1 decisions, taken on duplication or topic fit before anyone opened a draft. Gate 4 runs last because it needs a finished, published page to test.

A hand pausing over a laptop trackpad before publishing, with a single red warning indicator visible among green status blocks

The common mistake runs the order backwards. Teams write the whole article, then fact-check it, then discover at the end that the topic was never defensible. Now there is a finished draft nobody wants to throw away, and it gets published anyway. Sunk cost does the deciding, not the standard.

Most of these checks are also mechanical. Duplicate detection, schema validation, link mapping, metadata completeness, live-body comparison and the voice check all run as scripted steps. Two stay human: check 03 and check 09, deciding whether a page has a real reason to exist and whether a source is genuinely the origin of a number. That is where the judgement sits.

The failure mode to watch

A standard with no recorded failures is decoration. If your process has never rejected a draft, never removed a statistic, and never killed a planned article, it is not a quality gate. Track the rejection count as a metric in its own right. On this article alone, Gate 2 threw out 18 of the 19 sources the research step proposed.

None of this makes a page rank on its own. The point of the standard is narrower and more useful: it makes every claim on the page defensible when a buyer, a competitor or an answer engine tests it. That is the asset. In a market where 68.01% of US Google searches ended without a click between January and April 2026, being the source that survives scrutiny is worth more than being the page that ranked. We argue that case in full in the AI traffic share objection, and show the category data behind it in our AI visibility benchmarks.

Frequently Asked Questions

Does Google penalize AI content?

No. Google has never published an AI-detection ranking system, and its documentation contains no AI-specific penalty. What it does publish is a scaled content abuse policy covering pages generated in volume to manipulate rankings without adding value, and that policy applies whether a person or a model produced them. The practical implication is that authorship is the wrong thing to worry about. Verification, originality and whether the page gives a reader something they cannot get from the answer itself are the things that decide outcomes.

Is AI content bad for SEO?

Unchecked AI content is bad for SEO in the same way unchecked human content is, and for the same reasons: unverifiable claims, nothing original on the page, and topics picked for volume rather than fit. The EBU and BBC study of more than 3,000 AI assistant responses found sourcing problems in 31% of them, which is the specific weakness an editorial standard removes. AI-assisted content that passes a real verification gate behaves like any other well-made content. Our guide to optimising for AI search covers the structural side.

Can Google detect AI-generated content?

There is no public Google tool or documented ranking system that identifies AI authorship, and Google's own policy language avoids the question entirely by describing behaviour instead. Third-party detectors exist and are unreliable in both directions, producing false positives on careful human writing and false negatives on lightly edited model output. Designing a process to evade detection optimises for the wrong thing. Traceable sources and material nobody else has produce pages that hold up whatever score a detector puts on them.

What counts as low-quality AI content?

Practically: a page that restates what is already available, cites nothing traceable, and gives a reader no reason to open it rather than reading the summary. Google's own framing is pages produced in volume with no added value. Our operational test is check 03, click-defensibility. If the entire page can be restated by an AI answer in three sentences, it is low quality by definition, no matter how well written it is. Between that test and topic fit, 196 of our own published pages were redirected in a single review.

Does AI-generated content affect SEO rankings?

Production method is not itself a ranking factor in anything Google has documented. What affects rankings is the same set of things it always did, plus a newer consideration: whether the page still earns a visit when an answer engine has already summarised the topic. Ahrefs measured a 34.5% lower click-through rate for top-ranking pages on informational queries with an AI Overview present, across 300,000 keywords in April 2025. Ranking and traffic have partly decoupled, which is the subject of traffic collapse in AI search.

Is SEO still relevant for generative AI search?

Yes, with a changed job description. Retrieval still has to find the page, so indexing, structure and authority still matter. What changed is the second half: appearing in the answer is now a separate outcome from appearing in the results, and it is measurable on its own. We track it as share of model, meaning how often a brand is named in AI answers to its category's buying questions. Teams that measure only rankings and sessions are now blind to most of what their buyers actually see.

Find out what the engines currently say about you

We measure how often your brand is named in AI answers to your category's buying questions, which competitors appear instead, and which pages the engines are citing. You get the measurement and the evidence behind it, before any conversation about work.

Request your AI visibility measurement

Or start with how to get cited by ChatGPT →

Resources

Related blog

Answer Engine Optimization dashboard showing AI search platforms citing brand content with citation indicators and visibility metrics
04
Sept.

Answer Engine Optimization: How B2B Brands Get Cited by AI Search

A marketing executive studying a traffic dashboard where AI referrals are a thin slice, beside a screen listing vendor names in an AI answer
02
Sept.

AI Is Only 2% of My Traffic. Why Your Buyer Disagrees.

A marketing executive studying AI visibility benchmark charts by category on a large office display at dawn
31
Aug.

AI Visibility Benchmarks by Category: What 3,011 B2B Domains Reveal

THE NEXT STEP

Stop Renting Leverage. Install It.

Together we can achieve great things. Send us your request. We will get back to you within 24 hours.

Group 1000005311-1