Uncategorized

Why app creative testing breaks down, and why it isn’t a data literacy problem


Cole Dennis is a creative strategist at Phiture, a full stack app marketing agency with hubs in New York, Berlin, London, and Barcelona. He has spent 13 years as a creative digital marketer, studied biochemistry and math before switching to writing, and holds a US patent for a system that analyzes UA creative performance and iterates on creatives using machine learning.

At Business of Apps London 2026, he laid out how growth teams should be feeding information to their creative teams, and why the answers they are missing usually have nothing to do with statistics.

Key takeaways

  • ASO test results are not statistically difficult; the hard part of creative testing sits earlier in the process, in how the test was designed
  • Very few companies bother to teach their creative teams how ad networks, DSPs, and the app store ecosystem actually work
  • A test only produces an actionable answer when one independent variable is isolated: design or messaging, not both
  • Report creative performance as a delta over the campaign average, and include business metrics like cost per subscription, not just raw CTR tables
  • Creative teams advance in three stages: logical, strategic, and driver

“Math comes in orders of magnitude of difficulty”

The question Cole gets most often from growth leaders is how to improve data literacy on the creative team. His answer is that the premise is usually wrong.

He offers a hierarchy. “Math comes in orders of magnitude of difficulty. So most difficult, chaos theory. Second most difficult is UA campaigns and campaign data. Then it’s dividing the bill at dinner, then it’s ASO test results.”

The point is that ASO results are near the bottom. Once a test reaches statistical significance, the platforms hand back a red bar and a green bar, and a green bar taller than the red one generally means the variant beat the control.

– Advertisement –

He is aware of the objection. Someone raised it with him the same day: you cannot trust the platforms, they lie to you. That answer is more nuanced – look back across your tests, adjust for noise, make sure you have enough volume over enough time. But for an individual test, the reading is simple.

“If you trust Big Daddy Apple or Big Daddy Google, then you can just trust the red and green bars.”

If the data isn’t complicated, the difficulty lies somewhere else in the process.

Train creatives as if they were growth marketers

Cole’s view is that the missing piece is almost always logistical, and that it needs to be in place long before a creative goes live. Creatives need to understand the ecosystem their work is being placed into: the networks, the placements, the constraints, the way spend gets allocated.

He has been consistently surprised by how rarely this happens, even as the industry has matured and the tooling has become more accessible.

“A year into me working there, they would go, can I ask you, what is a DSP?”

He had been referencing specific ad networks and DSPs in briefs and reviews for a year. Nobody had ever taken the time to explain what those things were.

“I really, truly think that you should train them up as if they were growth marketers, so that they can understand where their creatives are being seen.”

Doing this covers about half the work of explaining results later. ASO is simple, but UA and ad networks come with millions of caveats, and it is rare for anyone to have a straightforward answer to “what’s our best ad?” A creative team that understands the gist of how each network behaves can hear a caveated answer without losing the thread.

The algorithm itself is worth teaching explicitly. At organizations advertising mostly on one or two platforms, Cole walks creative teams through ad groups and how the platform optimizes across a set of creatives – largely because designers get despondent when one asset takes all the scale and the rest go unseen. If half of them never get impressions, what is the point of making them?

“The algorithm is going to pick a champion for you. And you kind of have to deal with that unless you’re willing to force spend.”

Map creative tactics to the metrics they move

The next most valuable exercise is connecting parts of the funnel to the creative levers that move them. Cole is upfront that the mapping is loose, but it gives a team somewhere to start when a specific metric is underperforming.

  • Views and awareness: design. It is the first thing people respond to in any visual.
  • Clicks and CTR: messaging. This is where interest gets captured.
  • Installs and CTI: the content of the message. Cole separates this from messaging deliberately – what you are talking about is a different variable from how you are talking about it.
  • Conversion: harder. This is where you have to consider the integrity of the funnel, the creative experience, and everything else the user is seeing on the way through.

The scientific method, applied to a movie poster

From there Cole moves to something concrete: the scientific method, which he teaches to most of the creative teams he works with.

“I really think that understanding logic is more important than understanding statistics when you’re trying to get creatives to understand how data works.”

For ASO in particular, the data is simple. The weight falls on the logic used to design the creative and frame the hypothesis. Without that, the results do not hold water no matter how clean the numbers look.

His worked example is a pair of movie posters: the original Pete’s Dragon and the more recent remake. Same content, entirely different design. That makes it a design test. The messaging has not changed, the value propositions have not changed, and it is not aimed at a wildly different audience. It is one attempt to capture attention in a different way.

His messaging example is a set of Campbell’s soup ads he produced years ago while building the machine learning analysis platform. The theory was that you can run as many heartwarming family scenes as you like, but not everybody likes their family, so a family cannot be used to sell to every individual person. Same content, same value proposition – you will want soup because of your home life – with the messaging changed around it.

Neither example requires statistical analysis. What they require is that the creative team, before anything goes live, has done everything it can to make the eventual result interpretable.

“All we’re trying to do when we run a creative test is decide what we’re going to do next.”

Performant creative is obviously the goal, but the compounding value is institutional learning, and that only accumulates when each test was prescriptive enough to teach you something.

Don’t engineer the expertise out of the test

The counterweight matters just as much. Cole regularly sees clients who want rigid, laboratory safe tests in the App Store or in UA, and the effect is that the thinking gets smaller and smaller.

Teams narrow the variable until it is tiny, so that when they see half a percent of movement in eCPM they can point at a cause. What they lose is any test worth running.

Take a broader view of the creative process and you can compare assets that are wildly different and still get an actionable answer. Cole’s example is two posters for Evil Dead 2 – the US release and the Japanese one. The design, the tone, and the way the main characters are portrayed have almost nothing in common. You could in theory translate each one and A/B test them against each other, but that is not how movie posters work.

Read at a higher level, the two posters are a test of something else entirely: do you use local culture experts in your marketing funnel? That is a test about where the expertise in your creative production sits, not about a specific execution.

The most direct evidence he has for this came from his first agency job, where every new client opened with the same question: what color should our button be? It drove him to run a large test – ten ads, three variations each. One used a red button, because red had won in previous button color tests. One used a button the designer built specifically for that ad, with no guidelines beyond “make a button that fits this ad.” The third ran with no button at all.

The custom button won every single time.

“Creative expertise is not something that you should engineer out of your creative tests.”

Report a delta, not a number

UA reporting is where things get harder than ASO. Making sure every creative gets comparable scale is difficult, and where it is possible it is expensive. So how do you build a report that shows a creative team what is actually working?

Cole’s answer is to stop handing over straight metric tables. Instead of comparing raw CTR, CTI, or eCPM, he frames almost everything as a delta over the average performance of that campaign. That creates a dynamic benchmark for whatever time period and whatever network you are looking at, and from there you set a threshold above which a creative counts as a success.

The second half is business metrics, and this is where reporting often fails creative teams outright.

He describes a children’s education app whose creative team was split roughly half 2D designers and half video artists – which produced, predictably, roughly half 2D ads and half video ads. The video artists were getting demoralized, because on every report their work sat below the 2D ads on installs and subscriptions.

So he started showing them the metrics the growth team had been keeping to itself: cost per paying install and cost per subscription, by platform. Across the two asset types, those were comparable.

Creative teams are built out of people with specific production expertise, and they are rarely brought in on the numbers that would tell them whether that expertise is paying off. With the data, they can prioritize production and reallocate budget going to outside vendors. Without it, a lot of that happens inefficiently.

“Inefficiency can kill creative teams.”

“You really want your creative team to be given every weapon they can to prove that they are not only useful, but efficient and performant.”

Reverse engineering the benchmark

Cole showed a straightforward IPM report – the kind you can automate and send to the creative team on a schedule – with the campaign average marked and a threshold drawn at 150% of it.

The 150% is not arbitrary. It is reverse engineered from two things the team already knows.

The first is creative capacity: how often the team can ship. Say a new video every two weeks. The second is creative decay: on your main networks, say you lose 30% of creative performance every two weeks. If the goal is to always have something running that beats the average ad in your ecosystem, you can work backward to the number a new asset has to hit on launch. In this case, 150% of average.

The value is that the creative team now knows how often it should be producing, and what “good” looks like on arrival, from a number rather than from feel or best guess. It gives them something specific to work toward.

It also matters for morale. A design team gets to operate the way every other function in marketing does – with targets, and with results on paper they can point to.

“Like treats like”

Cole’s warning about what happens without any of this comes, characteristically, from medical history. Pliny the Elder and his contemporaries practiced “like treats like”: crushed whole bees for stinging pains, onions for colds, runny noses, and irritated eyes.

“They really, truly believed that if something caused a symptom that you were experiencing, that it was the key to solving your problem.”

A creative pipeline that isn’t using data to inform itself falls into the same trap, and the symptom is ads that get more and more prescriptive. Cole’s illustration: “download game, buy coin, feel good about it.” Without a creative team making its own decisions, ingesting data, and acting on it, the work goes flat – and narrowing hypotheses in pursuit of clean data is exactly what produces that outcome.

Three stages of creative team advancement

Cole closed with a progression for teams looking to move in the other direction.

Stage one: logical. The creative team participates in hypothesis formulation, sits in on larger strategy conversations, and helps perform research. In his experience creatives are curious and eager to do this and simply are not asked. One practical note: if the creative team doesn’t have an AppTweak license and the rest of the team does, give them one. It is invaluable for creative research.

Stage two: strategic. The creative testing roadmap becomes its own artifact, separate from the ASO or UA roadmap that tracks product beats, launches, and feature changes. Starting with evergreen creative keeps it from colliding with anything already scheduled, and it starts taking testing and roadmap work off the rest of the team.

Stage three: driver. In addition to briefing themselves and owning their roadmap, the creative team manages its own testing campaigns.

“Is that scary to anyone?”

He is serious about it. Give the creative team a small budget to run their own tests. This does not mean handing over login credentials and the keys to the kingdom, but the best creative teams he has been part of were behind the wheel on at least some amount of creative spend.

“I promise we can handle it.”

There is a practical argument alongside the developmental one. When the creative team is actively involved in launching, far less has to be relayed through someone else – and, as he put it, nobody likes infinite meetings and nobody likes infinite documentation.

Source link

Visited 1 times, 1 visit(s) today

Leave a Reply

Your email address will not be published. Required fields are marked *