A playable can win the IPM column and still lose the business decision. That does not prove the ad attracted “worse players,” and it does not prove creative–product mismatch caused the loss. It means the install row is only the first clue.
The useful question is narrower: did the creative make a promise the product could not carry into the first session? Treat that as a testable hypothesis, then protect activation, retention, and mature revenue with definitions and guardrails agreed before launch.
Start with the claim you can actually test
“Fake gameplay” is often used as if it described one tactic. It does not. The disputed territory includes an invented mechanic that is absent from the product; a materially altered mechanic with different agency, win conditions, or rewards; a real but marginal minigame presented as the core experience; cinematic footage that reads like controllable play; and deceptive interface behavior, such as controls that appear functional but are not.
Truthful compression sits outside that fake-gameplay taxonomy. It is the acceptable boundary this article recommends: a real loop is made faster or clearer while its objective, agency, consequence, and product prominence remain honest. Whether a particular execution crosses that boundary still depends on its overall impression and evidence from the current product.
Those cases should not be collapsed into one performance story. A mixed-methods study by Moradzadeh and Kou found that ad–game mismatch and the user complaints around it are multifaceted. Research by Egliston, Hardwick, and Carter documents deceptive and coercive patterns in playable ads. Neither study estimates a universal effect on IPM, activation, D7 retention, LTV, or ROAS.
A creative can win on IPM while losing on activation, retention, or mature ROAS. When its promise materially differs from the product, incongruence is one hypothesis to test—not a conclusion the IPM row can prove.
Targeting, placement mix, auction pressure, store-page changes, app versions, onboarding incidents, and attribution changes can produce similar symptoms. “The ad caused disappointment” is a reasonable editorial interpretation in some cases; it is not a measured result until the design isolates it.
IPM is a rate, not a verdict
IPM is attributed installs ÷ recorded impressions × 1,000. The word “attributed” matters. In the AppsFlyer glossary, an install is recorded at first app launch, while impression definitions can differ by media source. Other measurement arrangements can use different conversion rules. A download, an attributed install, and a first launch should not be treated as interchangeable unless the measurement contract says they are.
IPM answers one precise question: for every thousand impressions counted by this source, how many installs were credited under this attribution setup? It contains no activation event, return event, revenue, cost, incrementality, or satisfaction term. It cannot tell you why the rate changed.
It is also an outcome of a system: creative, audience, placement, auction, frequency, store journey, product eligibility, and attribution all contribute. Comparing two creative IDs inside optimized delivery does not automatically create randomized cohorts. IPM remains useful; it simply should not be asked to certify player quality on its own.
Use the Hookin congruence review
Creative–product congruence is a Hookin editorial framework, not a standardized industry metric or validated predictive scale. Its job is to make a vague argument inspectable. Review the playable and the current product side by side, capture evidence for each row, and resist turning the result into a magic score.
| Review dimension | Question to answer | Evidence to keep |
|---|---|---|
| Goal and prominence | Is the advertised objective real, meaningful, and similarly prominent? | Playable capture, product path, feature frequency |
| Input and agency | Do gestures, choices, and degree of control map to the app? | Annotated input map and side-by-side recording |
| Success and failure | Are the win conditions and consequences genuine? | State diagram and outcome captures |
| Reward and progression | Does the promised reward connect to the actual economy or loop? | Reward table, unlock path, time to feature |
| Pacing and difficulty | What was compressed, accelerated, or made easier? | Timed sequence and declared compression |
| Aesthetic and character promise | Are the characters, setting, UI, and production-value expectations representative? | Rights-cleared asset map and current product captures |
| First-session handoff | Does the app deliver or honestly bridge to the advertised promise? | First-session path and time to promised value |
Useful review labels are directly representative, truthfully compressed, adjacent but real, materially altered, invented, and indeterminate. These are production labels, not legal conclusions. A real minigame can still create an unrepresentative overall impression if it is presented as the main product experience.
Read the synthetic funnel without fooling yourself
Synthetic test-cell contract
The numbers below use one explicit fictional measurement setup. It is an illustration, not a claim that every network or MMP records these events this way.
| Contract field | Definition fixed before the test |
|---|---|
| Impression source | The same fictional ad server’s impression_served event, recorded once when an eligible creative render is accepted. Both cells use the same source and rule; no cross-network rows are mixed. |
| Assignment and delivery | Eligible users are randomly assigned once, 50/50, to mutually exclusive A or B cells before creative render. Geography, OS, audience, placement eligibility, optimization goal, bid strategy, store page, app version, and run window are held constant; repeat users remain in their assigned cell. |
| Attribution and install | The same fictional MMP credits a unique first app launch after an ad click as the attributed install. The click window is seven days, view-through attribution is disabled, and duplicate first-launch events are deduplicated to one attributed install per user and cell. |
| Activation | A unique attributed installer fires tutorial_complete within 24 hours of the credited first launch. The denominator is every eligible unique attributed install in that cell; duplicate activation events, test devices, and invalid traffic are excluded. |
| D7 return | A unique attributed installer fires app_open on UTC calendar cohort day 7, with the install date as day 0. The denominator is the same eligible attributed-install cohort, not activated users; this is not a rolling 168-hour definition. |
| Cost, revenue, and maturity | Fictional delivery stops accruing eligible installs on 2026-06-28. Cost and D30 attributed revenue are frozen as of 2026-07-29 UTC, when every included cohort has matured through day 30. Spend is final USD media cost; revenue is USD IAP plus ad revenue, net of platform fees and refunds, under the same MMP attribution window. |
Illustrative synthetic example—not observed campaign data. Values are constructed to demonstrate metric relationships. They are not Hookin, client, platform, network, or industry benchmarks. Recorded impressions, attributed installs, rates, CPM, and revenue are assumed inputs; spend, IPM, activated/returning counts, and ROAS are calculated with the formulas below. No significance or causal conclusion is claimed.
| Field | Creative A: compressed truth | Creative B: altered mechanic |
|---|---|---|
| Recorded impressions (assumed input) | 1,000,000 | 1,000,000 |
| CPM in USD (assumed input) | $8.00 | $10.00 |
| Attributed media spend (calculated) | $8,000 | $10,000 |
| Attributed installs: unique MMP-credited first launches (assumed input) | 8,000 | 12,000 |
| IPM (calculated) | 8.0 | 12.0 |
24-hour tutorial_complete activation rate among attributed installs (assumed input) |
45% | 24% |
| Unique activated installers (calculated) | 3,600 | 2,880 |
| UTC calendar-day-7 return rate among attributed installers (assumed input) | 18% | 8% |
| Unique D7 returning installers (calculated) | 1,440 | 960 |
| Mature D30 attributed revenue in USD (assumed input) | $7,360 | $6,100 |
| D30 attributed ROAS (calculated) | 92% | 61% |
The arithmetic is reproducible: spend = impressions ÷ 1,000 × CPM; IPM = installs ÷ impressions × 1,000; activated installers = installs × activation rate; D7 returning installers = installs × D7 return rate; and ROAS = attributed revenue ÷ attributed spend × 100.
In this invented example, equal impressions do not mean equal spend because CPM differs. D7 return is calculated from the full attributed-install cohort, not from activated users, so it is not the next step in a sequential funnel. Revenue is assumed to be USD IAP plus ad revenue net of refunds and platform fees, joined to complete media cost through day 30. Change that scope and the ROAS comparison changes too.
The metric contract comes before launch
The table only becomes an experiment when each field has an operational definition. Write the contract before delivery starts:
- Impression and install: name the serving source, attribution source, install event, attribution window, and any deduplication.
- Activation: name the event—such as
tutorial_complete—the eligible attributed-install cohort, uniqueness rule, and window. There is no universal activation event; the GA4 recommended-event catalog supplies event names, not your product’s value definition. - D7 retention: name the exact eligible attributed-install cohort used as the denominator and unique returning users as the numerator; specify the return event, collapse duplicate opens to one returning user, and state whether “day 7” is a calendar cohort day or a rolling interval. The AppsFlyer cohort documentation, for example, uses its own cohort and denominator semantics.
- ROAS: name the cost and revenue sources, gross or net revenue, included revenue types, currency, refunds, attribution window, and data-through date.
- Maturity: freeze the point at which retention and revenue cohorts are considered decision-ready. Recent rows may be partial.
Privacy-preserving attribution can delay postbacks or withhold fields. Apple documents multiple conversion windows and delayed, privacy-preserving postbacks. Do not present immature, aggregated data as complete user-level observation.
A fair test needs more than two creative IDs
Start with a falsifiable hypothesis: “Preserving the real resource decision while compressing the opening will improve IPM without reducing 24-hour activation by more than two percentage points or D7 return by more than one point.” The margins are business choices, not universal benchmarks.
Then record the unit of assignment, eligible population, primary outcome, guardrails, non-inferiority margins, minimum detectable effect, power target, sample-size assumption, analysis method, and stop rule. A safety or policy stop may be immediate; an ordinary performance stop should not appear only after a flattering early swing. Google Ads experiment guidance emphasizes a clear hypothesis and metric, while Unity’s creative-testing documentation is explicit that “fair opportunity” does not guarantee identical impressions.
Keep audience, geography, optimization goal, bid strategy, placement eligibility, store page, app version, and run window comparable where the platform permits. If ordinary optimized delivery chooses who sees each creative, treat the cohorts as observational unless the platform documents mutually exclusive or randomized assignment.
Finally, attributed ROAS is not incremental ROAS. It describes credited revenue relative to credited cost under an attribution model. Incrementality needs a design that estimates the counterfactual, such as an appropriate holdout or randomized lift design. For a broader testing workflow, use Hookin’s A/B testing guide alongside this metric contract.
Policy risk starts with the overall impression
Performance is only half the review. In the UK, the ASA’s mobile-game guidance says gameplay should be representative and cautions that qualifications may not correct a misleading overall impression. Its Evony ruling shows that a promoted minigame can exist in the product and still be presented misleadingly when it is not representative of the core experience. The Playrix ruling similarly shows that a disclaimer is not a general cure.
In the United States, the FTC’s advertising guidance requires truthful, non-deceptive advertising and substantiation for objective claims. Google Ads’ misrepresentation policy addresses misleading design and inconsistency between ad and destination; its app-ad requirements also cover functionality and accidental interaction.
Store review is a separate layer. Apple requires accurate metadata and promotion in its App Review Guidelines. Google Play’s Deceptive Behavior policy and store-listing guidance require listing claims and assets to reflect the app. These sources apply in their own jurisdictions and products; they are not global legal advice and do not replace network review or counsel.
Write compressed truth into the brief
“Make it exciting but accurate” is not a buildable instruction. Give the team a promise boundary. For a survival builder, that might read:
- Product truth: players combine scarce resources and choose which shelter risk to address first.
- Playable opening: begin with the storm approaching, two repairs available, and the first meaningful choice one gesture away.
- Allowed compression: shorten crafting time, clarify damage, and bring the first consequence forward.
- Promise to preserve: resource decisions shape survival.
- Do not invent: aiming a weapon, solving an unrelated pin puzzle, or granting rewards disconnected from progression.
- Evidence: attach the product capture and a first-session path showing where the advertised loop appears.
The storm can arrive faster. Feedback can be louder. The first reward can become legible in seconds. That is compression. Replacing the reason to play is substitution. For opening design, see The First 3 Seconds; for AI-assisted production, encode the mechanic, stakes, feedback, and exclusions using Hookin’s prompt guide.
Make the handoff earn the next session
A sound decision rule is deliberately unglamorous. If IPM improves and mature guardrails remain within their predefined margins, the creative is a candidate winner—after checking assignment, cost, and cohort maturity. If IPM improves while activation or D7 return deteriorates beyond the margin, call it a top-funnel win with a quality tradeoff, not a victory. If ROAS is immature, make no ROAS claim yet.
If the pattern survives those checks, return to the congruence evidence. The goal is not to prove that every aggressive playable is deceptive or that literal demos always perform better. It is to build a compelling, compressed promise that the store page and first product session can continue.
Keep IPM. Just stop letting one rate speak for the whole relationship.




