An ad’s reported ROAS rises from 2.8x to 5.0x. The obvious line in a weekly report is: “Creative A improved. Scale it and make more versions.”
That conclusion can be wrong even when the total is correct. In the worked report below, desktop ROAS falls from 10x to 6x and mobile ROAS falls from 2x to 1x. The aggregate rises because desktop’s share of spend moves from 10% to 80%. The report has observed a better total, but it has not shown that the creative became more persuasive.
A useful creative performance report therefore keeps three layers separate:
- Observation: What changed in the measured data or in the creative configuration?
- Possible explanation: What could account for the change, and what evidence is still missing?
- Next decision: What should the team make, retain, investigate, or decline to judge?
The seven-row campaign example and product are fictional teaching materials. The arithmetic is complete, but the results are not Hookin or client data, and no edit is presented as causal evidence.
Start With a Decision, Not a Winner List
A dashboard can rank ads by spend, CTR, CPA, or ROAS. A report must preserve enough context to make the next decision defensible.
A public 2022 campaign evaluation from Arizona State University and SciStarter shows why metric labels matter. It defined visits as click-through and view-through site visitors, while an action was a click on a project-kit link—not a purchase, completed kit, or retained participant. Its evaluators reported unclear reasons to click and small, low-contrast text in some ads. These are source-reported findings, not a test showing that a later redesign improved results. See the campaign evaluation, printed pages 6 and 8–9.
That distinction is the job of your report. It may say:
| Layer | Defensible entry | Entry to avoid |
|---|---|---|
| Observation | “A-v1 ROAS increased from 2.8x to 5.0x while its device mix changed.” | “The new creative drove more revenue.” |
| Possible explanation | “More spend was delivered to desktop, which had higher ROAS in both periods.” | “The dimensions message won.” |
| Next decision | “Retain A-v1 provisionally and investigate the delivery shift before scaling.” | “Make ten more versions of the winner.” |
The next decision does not always require a new asset. “Keep this ad,” “do not rank this version yet,” and “investigate tracking before briefing production” are all completed report outcomes.
Lock the Measurement Contract Before Comparing Creative
Before comparing ads, record the contract that produced the numbers.
| Contract area | Fields the report should retain | Mistake it prevents |
|---|---|---|
| Scope and identity | Platform, account, campaign, ad group, ad ID, creative version, row grain, included networks and filters | Combining different advertising objects under one filename or label |
| Dates and maturity | Period start and end, account time zone, extraction timestamp, report date basis, conversion maturity, version effective dates | Treating equal calendar periods as equal follow-up or equal exposure |
| Outcome and credit | Objective, conversion action, primary/secondary status, counting, deduplication, attribution model, creditable channels, conversion windows | Calling unlike events “purchases” or counting attributed credit as unique customers |
| Delivery | Device, placement or network, geography, query or audience, budget, bidding, rotation, status and schedule changes | Treating selected delivery as randomized creative exposure |
| Business and page | Offer, landing-page version, price/value basis, tax, shipping, returns, inventory and tracking incidents | Assigning an offer, checkout, or stock change to the ad itself |
| Evidence and action | Raw-row IDs, observation, alternatives, contradictory evidence, next action, owner and status | Turning an unsupported story into a production request |
Define the unit called “creative”
An ad_id, creative_version, asset_id, and the content rendered to a user are not interchangeable. Retain edit dates and actual serving intervals separately; a configuration record alone does not establish what appeared in every impression.
Responsive search ads add another boundary: Google can assemble different headline and description combinations, and Description 2 is not guaranteed in every impression. Text that must appear every time should be placed and pinned in an eligible position such as Description 1, according to the current responsive search ads documentation. An ad-level result is therefore a result for an eligible configuration and its delivered combinations—not an isolated result for one line of copy.
Co-served asset rows are not mutually exclusive ad rows, so adding their impressions can overcount the ad total. Google’s ad-level RSA asset-report documentation explicitly warns that some asset metrics are non-additive and directional rather than isolated single-asset effects.
Define the outcome, date basis, and denominator
For the Search example in this article, the formulas are:
CTR = sum(clicks) / sum(impressions)CVR = sum(selected conversions) / sum(eligible interactions)CPA = sum(eligible cost) / sum(selected conversions)ROAS = sum(selected conversion value) / sum(eligible cost)
Do not average row-level CTRs or CVRs unless the rows have equal denominators and that weighting is intentional. CPA is undefined when conversions are zero; it is not $0. Conversion value is not automatically profit, and ROAS is not incremental return.
Name the selected conversion. Google’s primary “Conversions” column can reflect primary actions, counting rules, modeled conversions, and fractional attribution credit; its rate uses eligible interactions. See Understand your conversion tracking data. An attribution model allocates credit among eligible interactions; it does not establish incremental causal lift. Google documents last click and data-driven attribution in About attribution models.
Date basis is a separate setting. In the standard primary conversion columns described for Google Ads Search, a conversion is reported back to the click date; separate “by conversion time” fields report it on the date the conversion occurred. The Google Ads API maps these to distinct metrics, including metrics.conversions and metrics.conversions_by_conversion_date; see Conversion reporting. The account time zone controls time-based reporting, according to Google Ads’ time-zone documentation.
Record the conversion window and extraction time. The window is the period after an eligible interaction during which an event can be recorded. Changing a window later does not necessarily recreate previously excluded history; Google’s own March example shows a conversion that is not retroactively restored after the window changes. See About conversion windows. “Exported today” is not the same as “every cohort is mature.”
Worked Example: Seven Rows That Reconcile
The fictional report uses this measurement contract:
- Platform and grain: Google Ads Search, full RSA ad row × period × device; asset rows excluded
- Market and currency: United States, USD; Search only, with partners, Display, and video excluded
- Objective and event: Purchase action
PUR01, counted “Every,” with transaction-ID deduplication assumed in the fixture - Attribution: Last click, seven-day click-through window, Google paid scope; no view-through or engaged-view conversions
- Date basis and time zone: Click date,
America/New_York - Periods: P1 = August 3–9, 2026; P2 = August 10–16, 2026
- Value: $100 attributed product value per purchase; tax, shipping, returns, and margin excluded
- Extraction: September 8, 2026; both cohorts are complete by construction
- Delivery: Manual CPC with Optimize rotation
Three configured versions appear in the report. A-v1 uses general feature-led copy. A-v2 adds only an 11-inch width claim; depth and height remain absent. B-v1 uses a “complete setup” theme but does not state how many units are included. In P2, the mobile bid adjustment changes from 0% to -40%, A-v2 is introduced, and B-v1 runs only on August 15–16 in a different campaign and query scope. Offer and landing-page settings are assumed unchanged for the fixture, while search terms, inventory, tag firing, and delivered RSA combinations remain unknown.
Raw rows
| Row | Version | Period | Device | Impressions | Clicks | Purchases | Spend | Value | ROAS |
|---|---|---|---|---|---|---|---|---|---|
| R01 | A-v1 | P1 | Desktop | 1,000 | 100 | 10 | $100 | $1,000 | 10.0x |
| R02 | A-v1 | P1 | Mobile | 9,000 | 450 | 18 | $900 | $1,800 | 2.0x |
| R03 | A-v1 | P2 | Desktop | 8,000 | 640 | 48 | $800 | $4,800 | 6.0x |
| R04 | A-v1 | P2 | Mobile | 2,000 | 80 | 2 | $200 | $200 | 1.0x |
| R05 | A-v2 | P2 | Desktop | 300 | 24 | 2 | $30 | $200 | 6.67x |
| R06 | A-v2 | P2 | Mobile | 200 | 8 | 0 | $20 | $0 | 0.0x |
| R07 | B-v1 | P2 | Mobile | 500 | 10 | 1 | $20 | $100 | 5.0x |
The P2 rows reconcile to 11,000 impressions, 762 clicks, 53 purchases, $1,070 spend, and $5,300 value. That produces a 6.93% CTR, 6.96% CVR, $20.19 CPA, and 4.95x ROAS. The device totals reconcile to the same numbers: desktop contributes 8,300 impressions, 664 clicks, 50 purchases, $830 spend, and $5,000 value; mobile contributes 2,700 impressions, 98 clicks, three purchases, $240 spend, and $300 value.
Reconciliation verifies that creative and device rollups describe the same universe. A many-to-many join with tags, themes, or assets can duplicate spend and conversions while still looking plausible. Preserve a unique raw-row key, aggregate performance first, and then join one-to-one metadata—or test row counts and totals immediately after enrichment.
The 2.8x-to-5.0x Trap: The Aggregate Hides the Mix
Now isolate A-v1, the only version present in both periods.
| Segment | Spend share P1 → P2 | CTR P1 → P2 | CVR P1 → P2 | ROAS P1 → P2 |
|---|---|---|---|---|
| Desktop | 10% → 80% | 10.0% → 8.0% | 10.0% → 7.5% | 10.0x → 6.0x |
| Mobile | 90% → 20% | 5.0% → 4.0% | 4.0% → 2.5% | 2.0x → 1.0x |
| Total | — | 5.5% → 7.2% | 5.09% → 6.94% | 2.8x → 5.0x |
Every within-device rate declines, yet every aggregate rate shown improves. This is not a mathematical contradiction. The mix moved toward the segment that was stronger in both periods.
The denominator check makes the error visible. P1’s correct total CTR is 550 / 10,000 = 5.5%. Averaging the desktop and mobile CTRs gives (10% + 5%) / 2 = 7.5%, which incorrectly gives a 1,000-impression row the same weight as a 9,000-impression row. P2’s unweighted average similarly says 6.0%, even though the correct aggregate is 720 / 10,000 = 7.2%.
You also need the correct weight for each metric. Holding P1’s mix constant, P2’s standardized results are:
- CTR:
10% of impressions × 8% + 90% × 4% = 4.4% - CVR:
100/550 of clicks × 7.5% + 450/550 × 2.5% = 3.41% - ROAS:
10% of spend × 6x + 90% × 1x = 1.5x
Impression weights belong to CTR, click weights to CVR, and spend weights to ROAS. One universal “segment weight” would create another error.
The report can state that device mix changed, a -40% mobile bid adjustment was recorded, and desktop received more spend. It cannot say that adjustment alone caused allocation or that the creative improved. Google’s Optimize rotation uses signals such as search terms, devices, and locations to favor expected performers, and even “Do not optimize” does not guarantee equal delivered impressions. See Use ad rotation.
The defensible decision is D01: retain A-v1 provisionally and investigate delivery. Review device and query mix, auction conditions, offer and landing-page continuity, inventory, tracking, and delivered combinations before increasing production around an alleged winner.
Know When the Data Cannot Support a Creative Conclusion
B-v1 carries a 5.0x ROAS label, but it has ten clicks, one purchase, $20 spend, two days of delivery, and a different campaign/query scope. With one fewer purchase, its ROAS would be 0x; with one more, it would be 10x. That is a sensitivity check, not a confidence interval or a reason to outrank comparable ads.
The decision is D03: defer ranking and do not scale. “Insufficient comparable exposure” is a finished analytical conclusion, not an empty report cell.
Time can create a second false story. Consider a separate fictional case with two click cohorts, each with 1,000 impressions, 100 clicks, and $100 cost. The older cohort eventually generates ten purchases worth $1,000. The newer cohort has generated only two purchases worth $200 when the report is extracted on August 17; eight more purchases occur on August 18, still within the configured window.
| Report view | Older cohort | Newer cohort | Defensible reading |
|---|---|---|---|
| Click-date, extracted Aug. 17 | 10.0x | 2.0x | Newer cohort is immature; no creative conclusion |
| Click-date, frozen Sept. 8 | 10.0x | 10.0x | Completed cohort outcomes are equal in this fixture |
| Conversion-time value divided by same-calendar click cost | 2.0x | 10.0x | Late value from the older click cohort has been mixed with newer-period cost |
On August 17, the analyst cannot know eight more purchases will arrive. Label the cohort immature and refuse a stop-or-refresh brief. A later reconciliation can explain what happened in this teaching fixture without pretending the first report had future knowledge.
Use distinct statuses for zero, missing, not_run, immature, and not_comparable. They lead to different actions. A missing row may not even mean zero: the Google Ads API conversion-reporting guide notes that rows with all-zero metrics may not be returned.
A declining curve is not proof of creative fatigue. Fatigue is one hypothesis alongside delivery, audience or query mix, placement, offer, seasonality, landing page, inventory, tracking, and delayed credit.
Turn Findings Into Three Concrete Next-Creative Briefs
The fictional product is one organizer insert measuring 11 × 7 × 2 inches; the desk, drawer, and accessories are excluded. These are fixture inputs, not real-product claims.
Each brief names the input, exact change, held constants, hypothesis, counterevidence, and rejection check.
Brief 1 — A-v3: Complete the fit proof
Observation: A-v2 mentions only an 11-inch width. It received 32 clicks and two purchases, including zero mobile purchases. That performance is too limited to validate the message, but a desk review shows that the dimensions are incomplete.
Exact change: Replace the width-only Description 1 with: “Fits spaces at least 11 x 7 x 2 in. Check your drawer dimensions.” Keep the remaining copy, pinning, offer, landing page, and intended query/device scope unchanged.
Hypothesis: Putting all three dimensions in one line may reduce fit uncertainty; a paid-conversion benefit has not been shown.
Challenge and decision rule: A reviewer must derive the minimum interior space and match it to the product record. If tested comparably, evaluate purchase CPA and ROAS with query/device mix; CTR alone is insufficient.
Owners: Copywriter, product owner, and analyst.
Brief 2 — B-v2: Clarify what the offer includes
Observation: B-v1’s first description does not state the unit count. Its one purchase and 5.0x ROAS do not prove that customers understood the scope.
Exact change: Use: “Includes one organizer insert. Desk, drawer, and accessories are not included.” Keep the price, quantity, other copy, Description 1 pinning, offer, and landing page unchanged.
Hypothesis: Explicit scope may reduce an all-inclusive interpretation. No misunderstanding rate was measured, so this is a content-risk hypothesis, not a diagnosed conversion problem.
Challenge and decision rule: A reviewer must identify what and how many items are included. If tested, monitor purchase CPA and scope-related feedback. Lower CTR alone is not failure if the edit filters a wrong expectation, but define a commercial guardrail first.
Owners: Copywriter and product owner; analyst preserves the no-scale status until exposure is comparable.
Brief 3 — A-v4: Keep essential information in a renderable position
Observation: A-v1 places the accessory-exclusion statement in Description 2. Under the documented RSA behavior, Description 2 is not guaranteed to show in every impression. This is an information-placement issue, not evidence that the existing ad underperformed because users missed the statement.
Exact change: Branch from A-v1—not from A-v3—and move the statement to Description 1: “One organizer insert. Accessories not included.” Pin it to Description 1. Shorten repeated generic benefit language elsewhere without changing the product, offer, landing page, or headline claims.
Hypothesis: An eligible pinned position should make the configured information requirement more reliable. This is a QA hypothesis, not a higher-ROAS promise.
Challenge and decision rule: Confirm the line is assigned and pinned correctly. If it is absent or wrong in a supported rendered view, QA fails. Do not remove accurate scope information because another version reports a better CPA.
Owners: Copywriter, ad QA, and product owner.
When the question truly is causal—“Did this one change improve purchases?”—use a prospective experiment rather than ordinary optimized rotation. Google’s experiment reporting can show uncertainty and a “no clear winner” state; see Monitor your experiments. That documentation does not supply a universal conversion count or duration that guarantees a decisive result; define the minimum useful effect, guardrails, eligible population, and analysis plan for the actual campaign.
Copy the Reusable Creative Performance Report Template
A reusable report can stay compact when every layer has one job.
| Report layer | Required fields | Completed output |
|---|---|---|
| Measurement header | Scope, grain, filters, dates, time zone, extraction time, maturity, conversion action, counting, attribution, windows, value basis | A comparison contract that another analyst can reproduce |
| Version manifest | Version and ad IDs, configured copy/assets, launch and pause dates, offer and landing-page IDs | A record of what was configured |
| Change log | Timestamp, old value, new value, owner, affected campaigns/ad groups, known isolation limits | A list of campaign and creative changes without causal language |
| Raw performance rows | Unique row ID, period, version, device/placement/audience or query, impressions, interactions, conversions, spend, value | Recalculable counts before ratios |
| Reconciliation | Creative, period, and segment totals; denominator checks; join row counts | Confirmation that every view describes the same data universe |
| Interpretation record | Observation, possible explanations, contradictory evidence, missing inputs | A bounded account of what is known and unknown |
| Decision record | Retain, produce, investigate, defer, or stop; owner; status; source row and audit IDs | A traceable operational choice |
| Production brief | Baseline, exact diff, held constants, hypothesis, acceptance check, falsification signal, commercial metric | A buildable next step rather than “make more winners” |
The companion files include the seven-row CSV, clean schema, version/change log, no-conclusion case, and three briefs. The critical tables also appear here, so the method does not depend on a download. Download the filled seven-row report, reconciled totals, blank CSV schema, reusable report guide, version and change log, decision register, three completed briefs, cohort-maturity case.
A strong creative performance report shows what changed, protects comparisons from bad denominators and mixed cohorts, and makes the next decision traceable to its inputs. Sometimes that decision is a new proof, offer, or layout variation. Sometimes it is to keep the current ad. And sometimes the most useful answer is: the data do not yet support a creative conclusion.
Sources
- Libraries as Community Hubs for Citizen Science Ad Campaign Evaluation Report (IMLS) — Arizona State University, University Office of Evaluation and Educational Effectiveness; SciStarter host, June 27, 2022.
- Understand your conversion tracking data — Google Ads Help.
- About conversion windows — Google Ads Help.
- About attribution models — Google Ads Help.
- About your Google Ads time zone setting — Google Ads Help.
- Conversion reporting — Google for Developers; page last updated August 19, 2026 when checked.
- Use ad rotation — Google Ads Help.
- About responsive search ads — Google Ads Help.
- About the ad-level asset report for responsive search ads — Google Ads Help; full performance statistics documented for dates on or after June 5, 2025.
- Monitor your experiments — Google Ads Help.




