← All reports

October 4, 2026

Grounding checks for AI ad copy: what actually holds up

NAD rejected AI-generated test evidence for an ad claim, and new research shows LLM judges can be consistent and still wrong. What that means for AI copy.

Contents

What changed

Three things landed in the last few weeks. Separately they look like unrelated news. Together they describe a gap that any brand generating ad copy with a model is now sitting inside.

A self-regulatory body rejected AI-generated evidence for an ad claim

On September 17, 2026, BBB National Programs' National Advertising Division announced its decision in a case brought by Sonova USA against Starkey Laboratories over the Omega AI hearing aid. NAD recommended Starkey modify or discontinue claims covering speech intelligibility, spatial awareness, competitive performance, and "better hearing all around," and modify "the leader in hearing healthcare." The claims had run in Starkey's technical product handbook, professional pamphlets, starkeypro.com, a press release, and social posts.

The part that matters outside the hearing-aid category is how NAD treated the evidence. Starkey supported comparative claims — including "up to a 6.5 dB advantage (70% improvement) in speech intelligibility" compared to all major brands — with laboratory testing that used Whisper Automatic Speech Recognition, an AI speech-to-text system, against flagship products from six major brands. Starkey had added the word "predictive" in front of "speech intelligibility." NAD held that a reasonable consumer would still read the claim as real-world performance, that "predictive" was not enough to qualify it, and that the record did not establish that Whisper ASR had been validated to support a claim about measurable outcomes for actual hearing aid wearers. NAD's framing: AI testing is an emerging area with potential benefits, but the testing must be properly validated for the specific claims being supported, particularly when those claims convey real-world consumer benefits.

The same decision contains a useful reminder about preference data. Starkey's study found 94% preferred Omega AI for directionality and 96% for speech clarity — among respondents who expressed a preference. But 52% and 42% respectively expressed no preference. Under NAD precedent, unqualified preference claims should not be made when 20% or more of respondents express no preference. NAD recommended Starkey discontinue the claim or clearly disclose both the "expressed a preference" limitation and that the comparison was to Starkey's own prior technology.

Two papers put numbers on how unreliable an LLM judge can be

An arXiv preprint posted September 12, 2026 evaluated two local open-weight judges against human ratings. Nine human annotators scored 300 responses across 100 questions in five categories; each judge scored the same responses three times with the same rubric. LLaMA-3-8B reached a Pearson correlation of 0.275 with average human scores and a mean absolute error of 27.71. Qwen2.5-7B reached 0.340 and 18.64. Yet both were highly self-consistent: exact consistency rates of 97.3% and 92.3%. A judge can return the same score every time and still not agree with people.

A second preprint studied where judge error comes from across four benchmarks, six models, and 36 judge-examinee pairs. It reports that a judge's own task accuracy predicts its judging accuracy (Pearson r of at least 0.90), that judges show a directional bias tied to capability, and that judges grade more capable models more leniently. Its proposed fix is aggregation: a weighted majority vote across multiple judges that tracked an oracle selection within roughly half a percentage point.

Hallucination rates are a property of the task, not the model

Vectara's public hallucination leaderboard, last updated September 22, 2026, measures one narrow thing: whether a model introduces unsupported content when summarizing a provided document. On that grounded-summarization task the best models sit under 2% — and answer rates matter, because a model that declines to answer more often has fewer chances to invent. Rates reported for other, less constrained tasks run far higher. The practical read: a vendor's headline hallucination number tells you almost nothing about your ad copy unless the benchmark task looks like your task.

The rulebook did not change, which is the point

The FTC's Policy Statement Regarding Advertising Substantiation still governs. Advertisers and their agencies must have a reasonable basis for objective claims before dissemination; the representation that a reasonable basis exists is itself material to consumers; lacking one is an unfair or deceptive practice under Section 5. Where an ad says "tests prove," the advertiser must have at least the level of proof claimed. In its Operation AI Comply sweep the FTC put it plainly — then-Chair Lina Khan said there is "no AI exemption from the laws on the books" — and in the DoNotPay matter the complaint alleged the company had not tested whether its AI output matched a human lawyer's work; the proposed order carried a $193,000 payment.

On the disclosure side, Bird & Bird notes the European Commission published its Final Code of Practice on the Transparency of AI-Generated Content on June 10, 2026 and its Final Guidelines on the AI Act's Article 50 transparency obligations on July 20, 2026 — a labeling duty that sits alongside, not inside, substantiation law.

Why this matters to a brand spending $50K+/month

At that spend you are not writing ten headlines a month. You are writing hundreds, and increasingly the platform is writing some of them for you. New Frame Digital, documenting its own account, reports Meta offering AI variants immediately after creative upload, creative enhancements on by default for most new ad sets, settings that do not reliably survive duplication, and no single off switch; it cites Meta reporting more than nine million small businesses using at least one AI creative tool, and REI saying Meta auto-enrolled it in the feature behind a widely shared ad showing a bike with two sets of handlebars.

So the volume of copy nobody read before it shipped is going up, and the entity that answers for it has not moved. It is still the advertiser. Three specific exposures follow:

  • Invented specifics. A generated line that adds a percentage, a timeframe, a guarantee, or a comparative superlative that is not in your product data is an objective claim you now have to substantiate before it runs.
  • Using a model as your measuring instrument. This is the Starkey lesson. If your substantiation for "better" is an AI system's output — an ASR score, a synthetic-panel preference, an LLM judge's rating — you need validation evidence that the instrument measures the thing the claim promises, for the claim you are actually making.
  • Preference and testimonial framing. The 20%-no-preference rule and the "compared to our own prior version" disclosure are not AI-specific, but AI copy tools will happily strip the qualifiers your legal review added.

What to do this month

1. Split your copy checks into claims and style, and only automate the first with retrieval

Build one check that asks a narrow question: is every factual assertion in this line supported by the attached source record? Feed the judge the source — SKU spec sheet, approved claims doc, test report — not just the copy. MLflow's built-in RetrievalGroundedness judge is a reasonable shape to copy: it returns yes/no plus a rationale that names which statements are supported, which lack support, and the specific quotes from context that support or contradict them. That rationale is the audit artifact; the score alone is not. Keep tone, brand voice, and hook quality in a separate scorer so a style pass can never overwrite a factual fail.

2. Validate the judge against human labels before it gates anything

Pull 100–200 real generated lines, have a human who can sign off on claims label each as supported or unsupported, then measure judge agreement on that set. Arize's guidance for teams building judges is to validate against a golden dataset and aim for 75–90% agreement with human labels before scaling — treat that as Arize's recommended band, not a regulatory threshold. Arize also flags the pitfalls worth stealing: skipping validation, feeding the judge the wrong inputs (a hallucination check with no retrieved context), and reading a 100% pass rate as success rather than as evidence the eval is too easy. Given the 0.275 and 0.340 correlations reported for small local judges, also record self-consistency separately from agreement, so a stable-but-wrong judge cannot pass as a good one.

3. Write a one-page claim policy and wire it into the generation step

List the claim types that require substantiation before any variant ships: numbers, timeframes, health or income outcomes, "#1"/"best"/"more than," anything comparative. For each, name the document that substantiates it and the person who approves it. Then put that list in the generation prompt as a prohibition and in the judge rubric as a trigger. Add a standing instruction that comparatives to your own previous version must say so. On Meta specifically, re-open the Advantage+ enhancement panel on every edit and duplication rather than trusting an account-level default, and keep screenshots of what shipped.

What we would watch next

  • Whether NAD's validation language gets applied to a case where the AI system was used to write the claim rather than to test it. The reasoning transfers; no decision we found has done it yet.
  • Multi-judge aggregation moving from preprint to product. A calibrated weighted vote across several judges is the most direct answer to "how do I show my instrument is validated," and it is cheap enough at ad-copy volumes.
  • Whether platforms begin surfacing which lines their systems generated, in a form you can export. Article 50 labeling pressure in the EU pushes that way; advertiser-facing logging is a different feature and does not exist yet.
  • Benchmarks built on marketing copy with a source document, rather than on summarization or QA. Nothing we found measures groundedness on ad claims specifically.

What we could not verify

  • The full text of NAD's Starkey decision. Our reading is from the September 17, 2026 press release; NAD case decisions are published to subscribers, so the underlying reasoning may contain qualifications we cannot see.
  • Whether Starkey has appealed or complied. The press release notes advertisers state whether they will comply; we did not find a follow-up.
  • Both judge papers are arXiv preprints and, as far as we can tell, not yet peer reviewed. The 0.275 and 0.340 correlations come from small open-weight judges scoring a 124M-parameter model's output, which is not a marketing-copy setting.
  • Vectara's leaderboard is maintained by a vendor that sells grounded-generation products. The task definition is public and narrow, which is why we cite it, but we did not independently reproduce any figure.
  • The 75–90% human-agreement band is Arize's published guidance, not an industry standard or a regulatory requirement.
  • New Frame Digital's Meta observations are from its own account and from its reporting of REI's and Business Insider's accounts; we did not reproduce the default-on or duplication behavior ourselves.
  • We found no FTC enforcement action to date that turns specifically on AI-generated ad copy lacking substantiation, as distinct from deceptive claims about AI products.

Sources