What changed
Three things moved this year, and together they break the creative-testing habits most $50K+/month accounts still run on.
One: Meta has said, in its own engineering writing, that the number of creatives in its system is about to grow a lot. Meta's retrieval stage — the first step of ad selection — narrows "tens of millions of ad candidates into a few thousand relevant ad candidates" before ranking models ever see them. Meta writes that Advantage+ automation "further increases the number of eligible ads," and that with generative tools for producing ad content, "the number of ads creatives in Meta's recommendation systems is expected to grow significantly" (Meta Engineering, Andromeda). Read that as an operating instruction: the auction is being rebuilt to sort through far more creative than a human buyer can read. Your job shifts from picking the winner to supplying options that are genuinely different from each other, and reading the early signal correctly.
Two: the measurement stack split into tiers. Serious creative reporting now separates a hook signal (3-second views divided by impressions), an engagement signal (hold rate — full or long watches divided by 3-second views, plus average watch time), a conversion signal (CPA, CVR, ROAS), and a program-health layer: refresh rate, library diversity, and "fatigue half-life," defined as the number of days until an ad's CPA inflates 20% past its own baseline. Superscale's framing is blunt: teams that only look at the conversion tier "have performance reporting, not creative analytics" (Superscale). The same source puts the hook window at the first 1.5 seconds, not three.
Three: hook-rate benchmarks stopped being one number. Thumbstop rate is now published by placement and by traffic temperature. For DTC cold prospecting, in-feed video typically runs 18–28% with a median near 23%; Reels prospecting runs 24–36%, median ~30%; Stories 22–32%. Warm retargeting runs 30–45%, median ~36% — and the same source warns that retargeting thumbstop "is not a creative-quality signal," because it inflates the picture (Adsights thumbstop glossary, citing Motion's 2026 creative benchmarks). Other 2026 guides publish a single band — above 35% strong, 25–35% solid, below 25% a creative problem (Adextract) — which is the same threshold Superscale uses. Those two views disagree. We think the split-by-placement version is the more useful one for cold traffic, and we flag the conflict below rather than pretend it away.
Why this matters to a brand at this spend level
At $50K+/month you are not testing creative; you are running a creative supply chain, and every bad kill decision is expensive twice — once in wasted spend and once in the winner you never found.
The first problem is that CTR is doing work it was never qualified for. Adextract calls CTR "a diagnostic metric, not a success metric," and cites a figure that CTR influences only 4% of ROI, against a Nielsen number — cited second-hand through Meta — that creative quality drives 56% of a campaign's ROI (Adextract). We could not trace either figure to a primary document; see the verification list. But the direction is consistent with what the tiered frameworks assume: a creative can pass the attention gate and fail completely at the business gate. Adextract's own illustration is 500,000 impressions and 15,000 likes producing three conversions at $847 CPA.
The second problem is that the number on your dashboard is not the number your CFO is asking about. Platform-reported ROAS and incrementality-measured ROAS are different measurements, and incrementality-measured ROAS is "consistently lower than platform-reported ROAS, often significantly so" (Blip on holdouts and conversion lift). If you are optimising a creative library against a metric that systematically flatters the platform, you will over-invest in whatever best harvests existing demand.
The third problem is volume economics. One 2026 write-up claims teams have moved from 10–20 creatives per month in 2024 to 100–300 variants per month, with cost per creative falling from $50–200 to $2–10, and cites "internal Meta data from 2025" showing advertisers running 10+ creative variants per ad set saw 25–35% lower cost-per-acquisition than those running fewer than five (Distk). That is an unlinked claim from a vendor-adjacent source and we are not treating it as fact. What it does tell you is the direction of the arms race you are bidding against.
What to do this month
1. Replace "kill on CTR" with a two-gate rule
Gate one is attention, judged against the right band for the placement and the traffic temperature — not against your account average, which is contaminated by retargeting. For cold in-feed video, 18–28% thumbstop is the honest working range; Reels needs more, 24–36%, because people swipe faster there (Adsights). A creative that lands below the band is a creative problem, and Adextract's advice is the correct one: do not spend more budget hoping it improves.
Gate two is retention, before conversion. Hold rate — the share of 3-second viewers who reach 15 seconds — averages 40–50%, with above 60% strong and below 30% meaning the hook lands but the body loses people (Adextract). That single split tells you whether to rewrite the first 1.5 seconds or the middle of the ad, which is a different production brief and a different team.
Only then does conversion cost get a vote. And note what each gate is for: attention and retention diagnose the creative, conversion metrics judge the buy.
2. Put a written retirement rule in the account, with an impression floor
Fatigue is three signals moving together — link CTR falling while frequency climbs and CPM rises — not one metric wobbling. Link CTR dropping 10–15% week over week is the early tell, usually appearing three to five days before frequency crosses its threshold; 20%+ over two weeks is confirmed. Frequency of 2.0–2.5 is a warning in prospecting, 3.0–4.0 is danger. CPM up 15–20% is a lagging confirmation. The usable operating rule: pause or swap when the 3-day link CTR is down 30%+ against the 14-day average, with a floor of at least 1,000 impressions so you are not reacting to noise (Klipio).
The same source gives the typical arc for one creative-audience pair: strong days 1–5, link CTR softening days 5–10, cost per result creeping days 10–20, a real drop past day 20, with most winners peaking somewhere in days 7–21. Audiences under 500K can burn out in one to two weeks; larger ones take three to four. Pair that with a fatigue half-life measure — days until CPA runs 20% above the ad's own baseline — and retirement becomes a schedule instead of an argument (Superscale). Fix in the order angle, then hook, then format, then volume; Klipio rates a new angle as the highest-impact lever and raw volume as the only one that never runs dry.
3. Book one holdout test, and size it before you book it
A conversion lift study randomly splits the audience into a test group that is eligible to see ads and a control group that is not, and measures the gap. The sizing is not a detail: most platforms recommend at least 200,000 users per group for statistical power, and tests should run two weeks minimum, preferably four, to capture delayed conversions — longer if your consideration cycle is long (Blip). A neutral result, where test and control perform the same, is a signal to change creative or offer, not proof the campaign failed.
Run this once a quarter at minimum. It is the only number in your account that tells you whether the creative library is generating demand or harvesting it, and it is the number that should govern budget, while hook rate and hold rate govern production.
What we'd watch next
- Whether hook-rate bands compress. If AI-assisted production really is pushing accounts to 100–300 variants a month, the median thumbstop in a category should move, and today's bands go stale fast. Re-baseline against your own cold-prospecting median quarterly rather than against a published table.
- Whether the "10+ variants per ad set" claim gets independent replication. It is currently circulating without a linkable primary source. If someone publishes a controlled test, that changes how many variants are worth producing.
- Diagnostics under consolidation. Meta's retrieval design assumes very large creative pools; the open question for advertisers is how much per-asset diagnostic detail survives Advantage+ consolidation. Keep your own pattern-level labelling (hook type, format, angle) outside the platform so you can still answer "which pattern works" if platform breakdowns thin out.
- Hold rate as the durable metric. Hook rate is the easiest number to game with a loud first second. Hold rate and cost-per-result are harder to fake, and we expect serious teams to weight them more heavily through 2027.
What we could not verify
- "Creative quality drives 56% of a campaign's ROI" (Nielsen). Cited by Adextract as "Nielsen research cited by Meta." We did not find the underlying Nielsen document in this run.
- "CTR influences only 4% of ROI." Attributed by Adextract to "Martech" with no linked study. Treat as directional.
- "Internal Meta data from 2025: 10+ variants per ad set, 25–35% lower CPA." Distk provides no link or methodology. Unverified.
- The learning-phase threshold. The commonly repeated figure of roughly 50 optimisation events per ad set per week comes from Meta's own help documentation; Meta's help page returned no readable content to us in this run, so we have left the number out of the body rather than cite it second-hand.
- Conflicting hook-rate bands. Adsights puts DTC cold in-feed prospecting at 18–28%, while Adextract and Superscale both treat below 25% as a creative failure across the board. Part of the gap is warm-versus-cold traffic mix. We used the placement-specific version for cold prospecting and treat 35%+ as a stretch target, not a pass mark.
- Vendor case studies. Superscale's Taxfix example (+45% CTR, ~20% CPA reduction across 200+ ads) is the vendor's own client result, self-reported.