PLAYABLE ZERIND

A/B testing playable ads: hooks, difficulty, end cards and variants

Concept tests before iteration tests, one primary metric, the order we vary things in, and how to build a playable so new variants are a config change rather than a rebuild.

A playable is software, so almost everything in it is a variable: the opening frame, the difficulty, the length, the fail state, the end card, the button text, the language. That is its great advantage over video, and also its trap. Teams that test everything at once on too little traffic end up with a spreadsheet of coin flips.

This guide covers how we approach A/B testing playable ads: the two kinds of test, the one metric to pick, the order to vary things in, how to set up a clean test and how to build the playable so new variants are cheap.

Concept tests and iteration tests

There are two different questions you can ask, and they need different tests.

  • Concept tests ask which idea works. Two or three playables with different mechanics or angles for the same game, run against each other. The differences are large, so results show up quickly.
  • Iteration tests ask how to improve the winner. One variable at a time: a new opening frame, an easier first level, a different end card. The differences are smaller, so they need more volume to read.

Run them in that order. Polishing a weak concept through ten iterations rarely catches up with a stronger concept you never tested. For a parking-jam game like our Park Out concept, a concept test might compare "clear the whole lot", "free one car before the timer runs out" and "find the one car that can escape". Only once one of those wins does it make sense to test its hook.

Pick one primary metric

Decide before launch which number decides the test, and write it down.

  • Primary: IPM, installs per thousand impressions, for install campaigns. It reflects both whether people engage and whether they convert, and it is comparable across variants on the same network.
  • Guardrails: cost per install, day 1 and day 7 retention, and return on ad spend. A variant that wins on IPM but brings in players who leave immediately is not a winner. Brands can use first orders or sign-ups instead of retention.
  • Diagnostics: in-ad events such as first interaction, completion, end card shown and CTA tapped. These explain why a variant won or lost; they do not decide the test.

Avoid deciding on click-through rate alone. A confusing playable can earn accidental taps, and networks count clicks differently.

What to test, in order

Every game is different, but this is the order we usually work in, from the changes that tend to move results most to those that tend to move them least.

1. The opening frame

What the player sees before touching anything. A packed lot or a nearly empty one in Park Out. Two crossed arms or five in Hand in Hand. A full plate or a half-eaten one. Small changes here decide whether the first touch happens at all.

2. Difficulty and the fail state

Can the player lose, and how quickly does the challenge arrive? In Pop Pop Pan, the popcorn can burn if the player holds the heat too long. Testing a version where it cannot burn tells you whether the risk is driving engagement or driving players away. For a slingshot mechanic like Counting Sheep, one shot versus three changes both difficulty and length.

3. Length and CTA timing

End after the first success, or after a second, harder step? Show a small install button during play, or only on the end card? Note that some networks decide this for you: Mintegral asks for a CTA button that stays visible for the entire playable.

4. The end card

The value line, the art and the button verb. Gameplay screenshot or character art. "Play free" or "Install now". For brand playables, the offer itself: a delivery offer against a percentage discount is often a bigger difference than any visual change.

5. Language

A localised playable against the English version in a non-English market. The answer is often obvious, but it is worth confirming before you pay for every language.

6. Theme and season

The same mechanic in a new skin. A winter theme like Snow Buddy, where a snowball grows as it rolls and three stack into a snowman, is a natural seasonal refresh. Themes rarely beat a better hook, but they help fight fatigue on a proven winner.

Setting up a clean test

  • Change one thing. If the new variant has a different hook and a different end card, a win tells you nothing about either.
  • Keep everything else equal. Same network, campaign, placement type, geography, bid strategy and dates. Run the variants side by side rather than one after the other.
  • Use the network's creative testing tools where they exist, or put the variants in the same campaign or ad set so they compete for the same traffic. Check how each network splits delivery: many optimise towards the early leader, which can starve the other variant.
  • Decide the sample size in advance. Use a standard significance calculator with your current IPM and the smallest improvement you care about. Then wait for that volume rather than stopping when one line looks good.
  • Run full weeks. Weekday and weekend traffic behave differently. Ending a test on a Tuesday after starting on a Friday mixes the two unevenly.
  • Let the learning phase pass. New creatives often behave differently in their first days on a network. Do not judge them on that period alone.

Instrument the playable

Install data tells you which variant won. In-ad events tell you why. A useful minimum set:

  1. Ad loaded and visible.
  2. First interaction.
  3. Tutorial or first step completed.
  4. Outcome reached, win or loss.
  5. End card shown.
  6. CTA tapped.

If most players never make the first interaction, test the opening frame. If they start but do not finish, look at difficulty and length. If they reach the end card but do not tap, test the end card. Unity's ironSource team suggests tracking similar events, including the moment the end card is shown.

Check each network's rules before adding tracking. Most forbid outside requests during play; Unity says analytics calls may be permitted if they contain no personal data, and Mintegral offers its own event hook. Where a network gives you no in-ad events, a test link with the same build, played by a small group, still shows where people get stuck.

Build variants as configuration

The cheapest variant is one that needs no new code. When a playable is built with its variables in one config object, a new hook or end card is a few changed values and a rebuild of the network files.

{
  "variant": "hookB_ecA_de",
  "opening": "packed",
  "difficulty": 2,
  "failState": true,
  "maxSeconds": 25,
  "persistentCta": false,
  "endCard": { "art": "character", "line": "end_line_2", "button": "play_free" },
  "language": "de"
}

Two habits make this work in practice:

  • Name files after the config. A file name like parkout_hookB_ecA_de_applovin.html makes reports readable months later.
  • Keep text in string tables. Every line of copy lives in one file per language, so localisation and copy tests never touch the game code.

Our guide to making a playable ad shows where the config fits in the build.

Reading results and deciding

  • Clear win on IPM, guardrails hold: the variant becomes the new control. Test the next variable against it.
  • Clear win on IPM, guardrails slip: the variant is attracting the wrong players. Often a sign the playable is promising something the game does not deliver.
  • No meaningful difference: the variable does not matter much for this game. Move to a bigger change rather than slicing the same variable thinner.
  • Inconclusive after the planned volume: treat it as no difference. Extending the test until something wins is how false winners are made.

Creative fatigue and refresh

Even a strong playable wears out as the same audiences see it again. Watch IPM per creative over time, not just in aggregate. When a winner starts to decline, refresh in this order: the opening frame first, then the end card, then the theme, and only then the concept. Keeping two or three tested variants in reserve means you can swap without waiting for a new build.

A sample four-week plan

WeekTestVariants
1ConceptTwo or three playables with different mechanics or angles
2Opening frame of the winnerControl plus two new openings
3Difficulty or lengthControl plus one easier or shorter version
4End card, then languagesControl plus one end card; roll the winner out in your main languages

Adjust the pace to your volume: a smaller budget may need two weeks per test to reach a readable sample. For the design principles behind each variable, see our playable ad best practices, and for how playables and video compare in the same account, read playable ads vs video ads. Our templates and variants service covers new hooks, end cards and languages for testing.

Start a playable with us

Have a playable that works and needs a testing plan? Send us your store link and current results. We build new hooks, end cards, difficulty settings and languages as variants of the same build. Write to info@zerind.com or use the form on our homepage.

Sources

Checked on 8 October 2026. Networks update their specs; always confirm against the current page before you upload.

  1. Unity (ironSource content team), 5 missed opportunities that can destroy your playable ad: https://unity.com/blog/5-missed-opportunities-that-can-destroy-your-playable-ad
  2. Unity Ads, Playable asset specifications: https://docs.unity.com/en-us/grow/acquire/creatives/playable/specifications
  3. Mintegral Help Center, Playable test guide: https://helpcenter.mintegral.com/en/docs/playable-ad-guide