Does Google Ads Ad Strength Matter? What to Change When the Score and Results Disagree

Ad Strength can expose gaps in a responsive search ad, but it is not an auction factor or a profit metric. This guide shows how to choose the smallest defensible asset change when the score, message requirements, and business results disagree.

By
Hookin Team, Performance Editorial
Published
September 10, 2026
Reading time
17 min read
Views
48 views
On this page
  1. Ad Strength is not Ad Rank, Quality Score, or performance
  2. Why score and performance studies appear to disagree
  3. Turn each recommendation into a message decision
  4. Three score-versus-results decisions
  5. Test the edit, not the label
  6. Sources

Imagine two responsive search ads. One is rated Poor, yet its mature lead cohort is producing qualified leads at $100 against a $120 ceiling. The other is rated Excellent, but its apparently healthy $75 cost per conversion combines add-to-cart events with purchases; the true purchase CPA is $150 and the campaign is losing contribution after media.

Which ad should you change first?

The answer does not come from the label. Ad Strength matters as a creation-stage diagnostic for the assets you supply. It is not an auction score, a business outcome, or proof that an edit will improve performance. At the same time, the edits behind the rating—adding a useful message, repeating a keyword, removing a pin, or expanding the number of combinations—can change what people see and how the ad performs.

The practical rule is therefore neither “chase Excellent” nor “ignore Ad Strength.” Translate the recommendation into a specific proposed change, protect facts that must remain visible, verify that the account is measuring the right outcome, and test the change only when the decision is worth the cost and uncertainty.

Ad Strength is not Ad Rank, Quality Score, or performance

Google defines Ad Strength as an overall rating plus specific action items for responsive search ads. The current rating can be Incomplete, Poor, Average, Good, or Excellent. Its inputs include the relationship between your text assets and ad-group keywords, the number and diversity of headlines and descriptions, and available sitelinks. It recalculates when assets, keywords, or account structure change, but it does not measure the quality of the landing page or offer.

Google is also explicit about the boundary: the rating does not determine whether the ad is eligible to serve and is not used to calculate Ad Rank, Quality Score, or auction wins. That distinction matters because several Google Ads numbers can appear to answer the same question while measuring different things.

Signal What it is The decision it can support What it cannot establish
Ad Strength An ad-creation diagnostic for RSA asset coverage, relevance, diversity, sitelinks, and combination freedom Whether the supplied message set appears incomplete, repetitive, or constrained Auction position, incremental conversions, CPA, revenue, or profit
Ad Rank Auction-time values used to determine eligibility and position, based on factors including bid, ad and landing-page quality, thresholds, competition, search context, and expected asset impact Why an ad may or may not compete in a particular auction Whether a higher Ad Strength label caused the outcome
Quality Score A 1–10 keyword-level diagnostic based on expected CTR, ad relevance, and landing-page experience Where keyword, ad, or landing-page relevance may need attention A KPI or a direct auction input
Asset reporting Impressions and performance statistics associated with assets shown in responsive ads Directional patterns and assets that receive little or no exposure The isolated causal value of one headline or description
Campaign outcomes The conversions, qualified leads, purchases, value, contribution, and cost that the business has defined and measured Whether the campaign is meeting its operational and economic goal Whether every underlying creative recommendation is already optimal

The Ad Rank definition and Quality Score documentation make those auction and diagnostic boundaries clear. Google’s asset-report documentation adds another useful warning: asset-level impressions and costs are non-summable when several assets appear together, while asset-level CTR, CPA, and ROAS should be treated as directional because combinations influence the ratios.

This is why a headline row is not a private A/B test. If one served ad contains two headlines and one description, all three assets participated in the same impression. Their associated statistics do not represent three separate customers, and deleting one asset does not guarantee that its displayed CPA disappears.

Google currently reports that advertisers who improve responsive search ad and sitelink Ad Strength from Poor to Excellent see 15% more conversions on average. The cited source is Google internal data collected from August 15–20, 2025. The Help page does not publish the sample, conversion definitions, cost effect, selection process, or causal design, so this is a first-party directional claim—not a forecast that your next score increase will add 15% more sales or profit.

Why score and performance studies appear to disagree

The public evidence does not support a simple monotonic rule in either direction. That is not necessarily a contradiction: the studies use different populations, comparison units, campaign types, outcomes, and interventions.

Evidence Reported observation What survives scrutiny
Google internal data, August 2025 Poor-to-Excellent improvement was associated with 15% more conversions on average Google sees a positive aggregate relationship, but the public page does not reveal enough method detail to treat the rating change as the cause or translate it into profit
Optmyzr, April 2026 Across roughly 20,000 active accounts, Average ads had a $12.43 CPA and 12.65% CVR, while Excellent ads had a $28.68 CPA and 4.97% CVR; Poor ads had the highest reported ROAS The rating was not a reliable performance ordering in this observational aggregate. Different advertisers, offers, conversion definitions, budgets, and account choices remain confounded
Doctor Ads, August–September 2026 In same-ad-group comparisons, Good beat Average on median CTR in 71.1% of 807 groups, while Excellent beat Good in only 48.3% of 1,613 groups Holding the ad group constant produces a positive CTR signal across some rating gaps, but not the final step. These are win rates, not CTR-lift magnitudes, conversions, or profit; the portfolio was not a pure Search-only population
Big Flare replication, February–April 2024 The trial recorded 129.03 conversions versus 96.33 for the base: a calculated 33.95% increase, with similar spend The displayed 80% interval ran from −8% to +76%, crossing zero. The trial also reduced assets from 15 headlines/4 descriptions to 8/2 and added another RSA per ad group, so it did not isolate asset count or Ad Strength
PPC Hero/iProspect, July 2021 The same RSA text was fully pinned in the Good control and unpinned in the Excellent variant; the article reported lower CTR, CVR, and Quality Score for Excellent This is relevant counterevidence to a universal unpinning benefit, but the readable report does not provide the experiment size, period, or effect sizes and excluded groups with more than a 10% impression difference

The Optmyzr study calls its own data observational and directional. The Doctor Ads analysis improves comparability by examining ads in the same ad group and reporting its denominators, but its main outcome is CTR. The Big Flare article describes several tests; the published replication screenshot shows a promising point estimate but a wide interval. The older PPC Hero experiment is useful mainly as a boundary case.

The defensible conclusion is narrower than any universal headline: a rating-performance association does not show that chasing the rating caused an improvement, while one account where a lower rating won does not make the diagnostic useless. “Ad Strength” is a label over several possible edits. Adding a truthful qualifier, adding five vague synonyms, unpinning a required disclaimer, and creating a second RSA are not the same treatment just because each can change the label.

Turn each recommendation into a message decision

Before editing, rewrite the platform recommendation in plain language: What exact customer question is missing, repeated, constrained, or mismatched? Then decide whether the proposed asset improves the message without weakening the offer.

Recommendation A useful change A reason to defer or refuse it
Add more headlines or descriptions Add a genuinely new role: offer scope, audience qualifier, process, concrete benefit, evidence, risk reducer, or next step The new asset merely paraphrases an existing claim, creates awkward combinations, or invents a benefit that the landing page cannot support
Include popular keywords Express the real search concept naturally in one asset and align it with the offer and landing page The phrase is irrelevant, misleading, grammatically broken, or displaces the differentiated message
Make assets more unique Replace duplicate ideas, not just duplicate words “Unique” wording still communicates the same generic promise or makes the set less coherent
Reduce pinning Unpin optional wording when more combination freedom is valuable Keep a pin when the category, price condition, legal language, service boundary, or other required fact must appear in a position
Add sitelinks Add relevant destinations that help the searcher complete a real task Do not create empty or duplicative destinations merely to reach the score’s six-sitelink input

Use the character limit to edit the idea, not distort it

Google allows up to 15 headlines and four descriptions in an RSA, with a 30-character headline limit and 90-character description limit. Its guidance says a keyword must fit wholly inside one headline to contribute there; a longer phrase can go in a description.

For example, the fictional keyword warehouse storage installation services is 39 characters. Forcing it into a headline is impossible. A more useful pair is:

  • Headline: Warehouse Storage Install — 25 characters
  • Description: Explore warehouse storage installation services for your business. — 66 characters

The shorter headline preserves the category; the description carries the complete phrase naturally. The important test is not whether every keyword appears everywhere, but whether the ad remains accurate, recognizable, and useful for the search intent.

Treat pinning as a message-control tradeoff

Google recommends using pinning sparingly because it reduces combination freedom, and suggests several distinct alternatives in the same position when pinning is necessary. That is sound input guidance, not proof of a universal performance penalty.

Ask one operational question: Must this fact appear whenever the ad serves? If yes, pin an approved version to Headline 1, Headline 2, or Description 1. Headline 3 and Description 2 are not guaranteed to appear. If several phrasings are allowed, each pinned alternative must independently satisfy the requirement. If the fact is optional, leaving it unpinned may create more useful combinations. Changing both wording and pin status is a copy-plus-delivery package, not a clean wording test.

Also audit more than the editor’s visible list. Google’s current RSA documentation says unused headlines and descriptions can appear in link-based placements, including assets borrowed from another active RSA in the same ad group. The Ad Strength documentation says text customization can add Google-generated assets that also enter the rating. Review the active asset set, pins, generated text, neighboring RSAs, and destination alignment before concluding that one suggestion is harmless.

Use the combinations report to inspect common served combinations, not to recreate a static ad and assume it will behave the same. Google explicitly says common combinations are not guaranteed to perform identically as static versions.

Three score-versus-results decisions

The following cases are fictional teaching examples. They are not Hookin accounts, observed campaign results, forecasts, or claims that the proposed edits will improve a score or outcome.

1. Poor score, acceptable business results: protect the baseline and test the gap

A commercial warehouse-racking installer has a Poor RSA with seven headlines, three descriptions, Headline 1 pinned to Warehouse Storage Install, and Description 1 pinned to Commercial projects only. We do not offer residential installation. The score asks for more diverse messages.

The mature cohort has 30,000 impressions, 1,500 clicks, $6,000 spend, 60 deduplicated qualified leads, and 12 linked closed sales. At $900 contribution per sale before media, the qualified-lead CPA is $100 and contribution after media is:

12 × $900 − $6,000 = $4,800

The company’s CPA ceiling is $120. The score is weak; the measured business result is acceptable.

The message audit still finds one concrete gap. Trusted Installation Team and Dependable Installation perform the same generic trust role, while the real ability to phase installation around an operating warehouse is absent. Replace only:

Dependable InstallationPhased Warehouse Installation

The new headline is 29 characters. Keep the category and commercial-only pins. Do not turn “phased” into unsupported promises such as “zero downtime,” same-day installation, guaranteed savings, or a price claim.

Decision: do not rebuild or unpin the ad because it says Poor. Preserve the working baseline; treat the single replacement as a message-coverage hypothesis. If phased installation is not a real, landing-page-supported service, reject the candidate even if it would raise the score.

A hypothetical $6,000, 28-day, 50/50 test at a stable $4 CPC would produce about 750 clicks per arm. At the current 4% qualified-lead rate, that is roughly 30 qualified leads per arm—not evidence that the test can reliably distinguish a 15% CPA difference. When the likely information is too small for the risk, keeping the baseline and declining the test is a valid decision.

If the test is funded, name it after the edit, keep the protected assets and outcome definition fixed, stop new spend at the time or budget ceiling, and wait for the full conversion window plus import lag. A score increase is not an adoption criterion. A wide result is inconclusive, not permission to keep spending until a winner appears.

2. Excellent score, poor economics: repair the outcome definition first

A made-to-order desk retailer has an Excellent RSA with 15 headlines, four descriptions, and no initial pins. It reports 160 primary actions on $12,000 spend, apparently a $75 CPA.

But the 160 actions are 80 add-to-carts plus 80 verified purchases. Net sales from the purchases are $24,000, and contribution margin before media is 30%.

Calculation Result
Reported blended action CPA $12,000 ÷ 160 = $75
Purchase CPA $12,000 ÷ 80 = $150
Average order value $24,000 ÷ 80 = $300
Contribution per order $300 × 30% = $90
Contribution after media $24,000 × 30% − $12,000 = −$4,800

The platform label is strong, but the economics are not. Google’s conversion-goal documentation says primary actions enter the Conversions column and bidding when their standard goal is used; secondary actions normally remain observational, with a custom-goal exception. The first action is therefore to separate purchase from add-to-cart, verify deduplication and values, and decide which event should guide bidding and evaluation.

Only then consider a message hypothesis. The landing page says desks are made to order and require 15 business days before dispatch, but the ad does not surface that expectation. Replace the generic description:

Find a desk that combines thoughtful design with a comfortable workspace.

with:

Made to order. Allow 15 business days before dispatch. See delivery details.

The candidate is 76 characters. It does not promise delivery in 15 days; transit time remains separate. If the dispatch fact is optional, this can be a single description replacement. If it must appear on every relevant impression, first create an approved pinned Description 1 baseline. That is a copy-plus-pin change. A later wording test can compare two compliant pinned versions, but it must not put the required fact in only one arm.

Decision: Excellent does not excuse a $150 purchase CPA against $90 contribution per order, and the new description is not guaranteed to fix the loss. Checkout friction, price, traffic quality, or the offer may be the dominant problem. Even a reduction from $150 to $130 would remain above variable-contribution break-even. Limit loss first; test the expectation-setting message only after the measurement and operating baseline are clean.

3. Poor score, immature data: fix facts without inventing a winner

A packaging-design service has a Poor RSA with three headlines, two descriptions, and a pinned description stating that printing and manufacturing are not included. The landing page also says design projects start at $1,500, but the headline set does not.

The first ten days show 4,000 impressions, 120 clicks, $300 spend, four forms, two qualified leads, and two unresolved forms. The 30-day conversion window is still open. Current qualified CPA is $150.

Without changing the ad, the two pending forms can produce three different views of the same $300 cohort:

Newly qualified pending forms Final qualified leads CPA on the same spend
0 2 $150
1 3 $100
2 4 $75

These are conditional arithmetic branches, not probabilities. Google notes that conversion delay can temporarily make CPA look higher and ROAS lower. More events could also arrive before the window closes.

A truthful coverage fix is still available: add Projects Start at $1,500 as a 24-character fourth headline. Retain the pinned design-only boundary. Do not infer a new Ad Strength rating, a lift, or acceptable lead economics from the starting fee; close rate and contribution are still unknown.

Decision: check the form-to-CRM import and wait for maturity before judging performance. Add the price qualifier now only if it is approved, current, and genuinely missing from the active message set. If the information is already sufficiently visible elsewhere, deferring the edit is also defensible. The same Poor label permits both decisions because the decision turns on message coverage and evidence, not the grade.

Use the interactive decision worksheet and three completed scenarios to apply the framework.

Test the edit, not the label

When a change is worth testing, write a decision record before opening the experiment.

Field What to record
Current evidence Rating snapshot, asset revision, date range, conversion window, lag, spend, and business outcome
Exact intervention The headline, description, pin, sitelink, or multi-change package—named literally
Protected meaning Facts, scope, prices, disclaimers, and destination alignment that neither arm may violate
Outcome Purchase, qualified lead, contribution, or another defined business event—not a convenient blended action
Success boundary The minimum useful improvement plus any volume, margin, or quality guardrail
Risk boundary Maximum spend, duration, and conditions for an immediate operational rollback
Maturity rule When the last click, attribution window, offline import, returns, or sales qualification is complete
Decision states Adopt, keep baseline, repair and restart, or inconclusive

Google Ad variations can update, add, remove, or pin RSA headline and description assets. That makes them a possible implementation route for a controlled text change. Simply running two RSAs together and reading their totals is not automatically an equal-exposure experiment.

Google’s variation reporting documentation displays a flexible confidence interval, with 80% as the default and a blue asterisk for at least 95% statistical significance. Read the interval, not just the point estimate. In the Big Flare replication, +33.95% conversions and an 80% interval from −8% to +76% mean the data were compatible with both a decline and a large increase at that reporting threshold. That is inconclusive, not “the versions are equal” and not conclusive proof of a win.

Finally, do not let one rate hide the business decision. Consider a fictional expansion where contribution per conversion is $60:

Metric Baseline Expanded coverage
Impressions 10,000 20,000
Clicks 500 700
Conversions 50 56
Spend $2,000 $2,240
CTR 5% 3.5%
CVR 10% 8%
CPA $40 $40
Contribution after ads $1,000 $1,120

CTR and CVR fell, yet the expansion added $120 in contribution after media: 6 × $60 − $240. That does not prove broader message coverage caused incremental profit; it shows why rates, totals, and economics answer different questions. If the expanded spend were $2,600 instead, contribution after ads would fall to $760. The message decision still needs a cost and margin boundary.

Ad Strength is useful when it exposes an actual input problem: thin coverage, repetition, an absent search concept, insufficient sitelinks, or constrained combinations. It becomes dangerous when the label replaces the diagnosis. Preserve required meaning, measure a mature business outcome, and judge the smallest defensible edit. The score can start the investigation. It should not finish the decision.

Sources

Back to blog

Keep reading

Turn the idea into a playable

Build and test an interactive ad in Hookin. No code required.

Start free