At 10:06, the UA team puts twelve playable thumbnails on a review deck. “We made twelve creatives,” someone says. Then the first serious question lands: “What did we learn?”
The room goes quiet. Three builds changed the opening copy and the difficulty. Two use a different store page. One was rebuilt for another network profile. Several were exported over the same filename. The dashboard can rank ads, but nobody can reconstruct which mechanic, package, destination, or audience produced a row.
This is a composite editorial scene, not a Hookin or client result. It captures a production failure that is easy to mistake for creative velocity: many files, weak identity, and no clean causal question. The remedy is not fewer ideas. It is a frozen mechanic, declared factors, and evidence that survives the export.
Twelve is a useful production target. It is not an evidence-based optimum, a sample-size recommendation, or proof that an experiment exists.
Freeze the mechanic before multiplying the creative
A playable mechanic is more than its theme. Its causal skeleton is the path from player input to state change, success or failure, feedback, reward, approximate session length, and store transition. Freeze that skeleton as version M01. The variants may wear different openings, pacing, art emphasis, or end-card language, but the same action should still produce the same meaningful consequence.
This separation matters because “change the hook” can quietly become “change the game.” A rescue opening that also lowers difficulty, shortens input timing, adds a reward, and switches the CTA gives the team four plausible explanations for one observed difference. The build may be good; the inference is not.
Mechanic choice belongs upstream. Hookin’s playable game-type guide can help select a loop, while the prompt-writing guide can make its rules explicit. This article begins at the point where a loop has been chosen and the team needs to explore it without losing experimental identity.
Write a frozen-mechanic contract that a reviewer can replay
The contract below is a Hookin operating artifact, not an industry standard. Fill it before the first variant brief. A reviewer should be able to watch any build and decide whether the mechanic stayed inside the boundary without relying on the creator’s memory.
| Contract field | What must remain frozen | Evidence |
|---|---|---|
| Player input | The gesture, timing window, and number of meaningful actions | Input-state diagram and replay capture |
| State transition | What changes after a valid, invalid, or missed action | Named states and transition table |
| Success and failure | The conditions that resolve the attempt | Test cases for both paths |
| Core feedback | The information that tells the player what happened | Frame capture, copy, sound-off fallback |
| Reward connection | How success connects to progress or value | Reward-state map |
| Session envelope | Approximate playable length and number of loops | Timed run across target devices |
| Store transition | When a CTA may appear and which actions can open it | CTA event log and destination record |
M01. If a row changes, create a new mechanic version instead of hiding the change inside a creative label.“Frozen” does not mean visually identical. It means the team has decided which causal structure is not under test. If a pacing treatment requires a materially different success window, record a new mechanic version. That discipline may feel slower during briefing; it is much faster than explaining an uninterpretable winner after the campaign.
Name the factor, profile, destination, and cell
A filename such as rescue_final_v7_REAL.zip is a warning, not an identity system. Use a grammar that exposes the planned factors and the delivery context. For example:
M01-HR-PC-EC-CONT__NP-UA3__SD-DEFAULT__C07__R01
M01: frozen mechanic version.HR: rescue hook.PC: compressed pace.EC-CONT: “continue” end-card treatment.NP-UA3: declared network or host profile.SD-DEFAULT: declared store destination.C07: test cell;R01: artifact revision.
The code is only a readable key. The build ledger remains the source of truth for the full SHA-256 checksum, source revision, package constraints, QA evidence, assignment, and timestamps. Once exported, a build ID is immutable. Change one byte and the revision and checksum must change too.
Do not treat the network profile as an invisible export setting. A host adapter can affect lifecycle behavior, CTA handling, audio, orientation, packaging, and rendering. If the profile changes, the resulting observation cannot be described honestly as a creative-only comparison. The same rule applies when the destination moves from a default listing to a campaign-specific store page.
Use a blank twelve-cell map, not twelve mystery files
Here is one planning map: three hook frames × two pacing treatments × two end cards. It contains twelve production cells because the arithmetic is useful for coverage—not because twelve has been proven optimal. The assignment and result columns are deliberately blank. Nothing below is campaign data or a performance benchmark.
| Cell | Hook | Pace | End card | Example build key | Assignment | Result |
|---|---|---|---|---|---|---|
| C01 | Rescue | Compressed | Continue | M01-HR-PC-EC-CONT__C01 | — | — |
| C02 | Rescue | Compressed | Install to unlock | M01-HR-PC-EC-UNLK__C02 | — | — |
| C03 | Rescue | Standard | Continue | M01-HR-PS-EC-CONT__C03 | — | — |
| C04 | Rescue | Standard | Install to unlock | M01-HR-PS-EC-UNLK__C04 | — | — |
| C05 | Mastery | Compressed | Continue | M01-HM-PC-EC-CONT__C05 | — | — |
| C06 | Mastery | Compressed | Install to unlock | M01-HM-PC-EC-UNLK__C06 | — | — |
| C07 | Mastery | Standard | Continue | M01-HM-PS-EC-CONT__C07 | — | — |
| C08 | Mastery | Standard | Install to unlock | M01-HM-PS-EC-UNLK__C08 | — | — |
| C09 | Collection | Compressed | Continue | M01-HC-PC-EC-CONT__C09 | — | — |
| C10 | Collection | Compressed | Install to unlock | M01-HC-PC-EC-UNLK__C10 | — | — |
| C11 | Collection | Standard | Continue | M01-HC-PS-EC-CONT__C11 | — | — |
| C12 | Collection | Standard | Install to unlock | M01-HC-PS-EC-UNLK__C12 | — | — |
The map forces a useful conversation. Is “install to unlock” truthful for every destination? Does compressed pacing still represent the product? Are the character and environment assets cleared for paid advertising? A cell that fails those questions should be rejected before traffic, not preserved for the symmetry of the grid.
Choose the experiment mode before the platform chooses it for you
There are two legitimate modes. In an isolated-factor sequence, compare one declared axis at a time, keep a control, and advance only a build that protects its guardrails. This is easy to explain and useful when traffic is limited, but it is slow and can miss interactions. Google’s experiment guidance starts with a clear hypothesis and selected metric, while TikTok’s current split-test variable guidance describes one-variable comparisons inside its product. That platform recommendation does not make one-factor testing the only valid statistical design.
In a designed factorial screen, define factors and levels in advance, preserve cell identity and assignment, and estimate the intended main effects or interactions. Full and fractional designs answer different questions. The NIST/SEMATECH handbook explains why fractional factorial designs can screen factors efficiently—and why design resolution determines which effects are confounded. If nobody can state what is aliased with what, the team is not ready to interpret a fractional design.
Twelve ads rotating under ordinary delivery optimization are not automatically a factorial experiment. A platform can assign different audiences, placements, auctions, and volumes to each creative. TikTok documents mutually exclusive audience groups for its supported split-test workflow. Unity’s Creative Testing guidance calls for a control, explicit hypothesis, sample-size planning, consistent packs, and stable parameters—and warns that material pack differences can affect impression distribution. Product-specific testing tools can help, but their setup and assignment rules still define what can be claimed.
Give every exported build one immutable ledger row
The map says what the team intends to make. The Variant Build Ledger says what actually shipped. Create one append-only row for every artifact:
| Ledger field | Required record |
|---|---|
| Creative identity | Build ID, parent ID, mechanic version, and artifact revision |
| Source identity | Repository commit or archive ID, build command, and dependency lock |
| Artifact identity | SHA-256 checksum, byte size, file count, orientation, and package name |
| Declared change | Exact factor levels and every difference from the parent |
| Delivery identity | Network profile, host/MRAID assumptions, campaign, and test cell |
| Destination identity | Store page ID or URL and deep-link or fallback behavior |
| QA evidence | Validator output, target-device captures, telemetry check, and reviewer |
| Experiment identity | Assignment unit, eligible population, launch/stop UTC, and metric-contract reference |
| Decision | Advance, hold, reject, or inconclusive—with reason and sign-off |
A checksum is not administrative decoration. It answers whether the reviewed, uploaded, and measured files are actually the same bytes. If a network-specific repair changes the package, append a new row. If the store destination changes mid-test, append the event and flag the affected observations. The ledger should make ambiguity visible instead of smoothing it into the word “variant.”
Apply the do-not-compare gate before reading a winner
The gate protects the phrase “creative-only effect.” It does not require a perfect laboratory; it requires the team to record material differences and bound its language. Mark do not compare when any of these conditions fails without an approved analysis plan:
- The core mechanic version, network profile, package behavior, and store destination match the comparison contract.
- Audience, geography, placement eligibility, optimization goal, bid strategy, budget constraints, and run window are comparable.
- The assignment unit is documented. Optimized delivery is not called randomized unless the platform’s supported setup justifies that label.
- The app version, onboarding, offer, and telemetry definitions remain stable—or their changes are modeled explicitly.
- Primary outcomes, guardrails, exclusions, stop rule, and multiplicity handling were set before the results were inspected.
- Cohorts are mature enough for the named metric. Google’s experiment FAQ warns within its product context about conversion delay and sufficient volume; Apple’s AdAttributionKit documentation describes delayed, privacy-preserving postbacks. Neither supports a universal “seven days is enough” rule.
Store continuity is both a guardrail and a factor. Apple supports campaign-specific custom product pages with unique URLs, but that capability is not independent proof of lift. If one variant lands on a different page, the destination belongs in the hypothesis and ID. Do not hide it under “end-card copy.”
Representativeness and rights are release gates even when they are not test factors. In one UK ruling, the ASA found an Evony ad misleading although the promoted puzzle content existed in the product; the ruling is fact- and jurisdiction-specific, but it shows why mere existence is not the same as representative prominence. Keep a rights-cleared asset map too. Google Play’s IP policy applies within its store scope and is not a global ownership decision, yet it reinforces the operational need to know where characters, music, UI, and generated assets came from.
Let traffic and decision cost determine what twelve can earn
Before assigning a cell, write the minimum detectable effect, baseline assumption, power target, eligible traffic, primary metric, and decision cost. Twelve underpowered cells do not become informative because the grid looks complete. If the campaign cannot support the planned design, reduce factors, run an isolated sequence, pool only where the model permits, or treat the wave as exploratory and say so.
Multiple comparisons widen the opportunity to find a flattering result by chance. Predeclare how multiplicity will be handled and which metric can name a winner. Let activation, mature retention, revenue, complaints, crashes, and policy signals act as guardrails where they matter. Keep attribution source and data-through timestamp beside the verdict. A top-funnel winner with immature downstream cohorts is a candidate, not a settled answer.
For the measurement contract, sample planning, and interpretation layer, continue with Hookin’s A/B testing guide. The production system here does a different job: it ensures that when the analysis asks “what changed?”, the team can point to one mechanic contract, one declared factor set, one assignment, one destination, and one immutable artifact.
Return to that 10:06 review. The team may still have twelve thumbnails. This time, each one has a reason to exist—and the thirteenth will not be built until the first twelve have earned a specific next question.




