All posts
Content & Creative Production · 9 min read

Same product, five hooks: which one stops the scroll?

You've written five different openings for the same serum. One leads with the problem, one with the result, one with price, one with an objection, one straight with the product. Putting all five live and scaling whatever sticks is the most common decision beauty brands make — and it's the right one.

What isn't right starts afterwards: checking the dashboard two days later and crowning whichever has the highest click-through rate. At that step, most brands aren't making a decision about creative. They're making a decision about noise.

This piece covers why the hook is so decisive, the statistical cost of testing five variants at once, and when a winner is actually worth believing — with sources.

Why is the hook so decisive?

Because attention isn't one event but two stages: gaining it and holding it. Advertising research measures these separately, and together describes them as a critical bottleneck of social media advertising effectiveness (Bruns et al., 2025). The hook is the whole of stage one: however good the rest of the creative is, if stage one fails, stage two never starts.

To see how small that scale is, look at the industry's own measurement standard. Under the Media Rating Council's mobile viewability guidelines, a mobile display ad counts as a viewable impression when at least 50% of its pixels are on screen for at least one continuous second — and the guidelines state explicitly that this requirement applies equally in News Feed environments. For video the threshold is slightly longer: two continuous seconds of playback meeting the same pixel requirement (Media Rating Council, 2016).

So two seconds is all it takes for your ad to be counted as seen. The hook's job isn't to win those two seconds — it's to turn them into a third. That's why hook testing isn't an aesthetic preference; it's a test of the first door in the funnel.

The one thing to change when writing five hooks

  • Change the opening line and what appears in the first frame.
  • Keep the product, the offer, the body copy, the CTA and the closing frame identical.
  • Keep editing pace and music identical — those move performance on their own.
  • Make sure every variant runs in the same ad set, to the same audience, on the same budget.

If one variant changed the opening, the music and the CTA all at once, then even when you find a winner you won't know what won — and you'll have produced no learning you can carry into the next creative.

Is it right to test five hooks at once?

As an idea, yes: without creative variation there's no learning. The cost is statistical. Comparing two variants is one comparison; putting five in play raises the number of comparisons, and with it your chance of finding a difference that only looks meaningful.

This is a known problem in experimental design. In their KDD 2022 paper, Kohavi, Deng and Vermeer write that factors like multiple variants, iterating on ideas several times, and flexibility in data processing increase the false positive rate due to multiple hypothesis testing (Kohavi, Deng & Vermeer, 2022).

The practical upshot: you don't need to abandon five-hook tests, but a winner from a five-variant test is weaker evidence than a winner from a two-variant test. To reach the same confidence you need either more data or a replication round.

How do you pick the winner wrong?

Three common mistakes, all fed from the same source: reading the result early and selectively.

1. Watching constantly and stopping the moment it looks good

The same paper describes how a commercial A/B testing system showing near-real-time results led users to peek at the data and stop when it was statistically significant — and notes that this type of multiple testing significantly inflates type-I error rates (Kohavi, Deng & Vermeer, 2022). Checking your ad dashboard three times a day and declaring a winner is exactly that behaviour.

2. Testing with too little power

Statistical power is the probability of detecting a real difference when one exists. The same paper states at section-heading level that experiments with low statistical power are not trustworthy, and calculates the pre-experiment power in its worked example at roughly 3% — severely underpowered. A five-variant test running two days on a small daily budget usually sits on that side of the line.

3. Celebrating a striking result

The more striking the result, the more suspicion it deserves. The authors invoke Twyman's law: any figure that looks interesting or different is usually wrong. Faced with a conversion lift of over 300% in their example, they note they have been involved in tens of thousands of A/B tests at Airbnb, Booking, Amazon and Microsoft and have never seen an improvement anywhere near that size — and reject the result (Kohavi, Deng & Vermeer, 2022). When one hook doubles another, your first reflex should be to replicate, not to celebrate.

What if four out of five hooks don't win?

Nothing. That's the expected outcome. The same paper collects published experiment success rates across companies: 33% at Microsoft, 15% at Bing, 10% at Booking.com, Google Ads and Netflix, and 8% at Airbnb Search (Kohavi, Deng & Vermeer, 2022). Even at the companies with the most mature experimentation cultures in the world, the large majority of ideas fail to beat the status quo.

Also, other factors like multiple variants, iterating on ideas several times, and flexibility in data processing increase the FPR due to multiple hypothesis testing.

Kohavi, Deng & Vermeer, KDD '22

There's a second reading of that table that concerns you directly: as the success rate falls, the risk that a statistically significant result is actually a false positive rises. In the authors' calculation that risk is 5.9% at a 33% success rate, climbing to 26.4% at an 8% success rate. A new brand's hook pool has no obligation to succeed more often than these companies do — so the scepticism a "significant" result deserves is greater, not smaller.

There's something reassuring rather than discouraging in this: four hooks failing isn't a verdict on your creative ability. Eliminating four of them is the entire point of testing. The only mistake is scaling without eliminating.

Does the winning test actually sell more?

There's one more layer here. The number on your dashboard shows what the platform attributes to your ad, not the difference the ad made. Using Facebook's own data, Gordon, Zettelmeyer, Bhargava and Chapsky contrasted results from 15 US advertising experiments — 500 million user-experiment observations and 1.6 billion impressions — with those from multiple observational models, and found the observational methods often failed to produce the same effects as the randomized experiments, even after conditioning on extensive demographic and behavioural variables (Gordon et al., 2019).

Hook testing falls into this trap less often, because a platform's A/B test tool splits the audience randomly and so builds a real experiment. But if the metric you read comes from your measurement setup and that setup is broken, the test breaks with it — missing purchase events or double-counted conversions can systematically mislabel the winner. We covered where measurement quietly breaks in our piece on Meta Pixel measurement errors.

How should a beauty brand run this in practice?

  • One variable: change only the first 2-3 seconds, keep body and CTA fixed.
  • Write down your decision metric first: click-through rate is an intermediate metric; decide on add-to-cart or cost per purchase.
  • Set the duration up front and don't intervene before it's up. Stopping early is the fastest way to pick a winner out of noise.
  • If you start with five variants, treat round one as elimination: the goal isn't to find the winner but to remove the three or four that clearly don't work.
  • If you have a winner, replicate it. If it wins again, scale. A single-round win is a hypothesis, not a result.
  • Treat striking results with suspicion: if one hook doubles another, look for setup errors first (audience overlap, uneven budget distribution, different delivery times).
  • Carry the winning hook forward as a rule for the next round — a learning like "openings that lead with the result win" is worth more than any single video.

One closing note: a hook working doesn't mean the rest of the content works. Gaining and holding attention are different jobs, and a video that wins the first three seconds but loses the sixth can produce a good click-through rate in a test and never convert. For the credibility of the body, see why UGC outsells studio content; for the overall ad structure, our starter guide to Instagram advertising for beauty brands is a good next step.

Frequently asked questions

How many hooks should I start testing with?

As few as your budget can carry. Multiple variants increase the false positive rate through multiple hypothesis testing (Kohavi, Deng & Vermeer, 2022), so a winner out of five is weaker evidence than a winner out of two. A practical route on a small budget: race five in an elimination round, then compare the surviving two in a separate, longer round.

When should I stop the test?

At the duration you set up front — not the moment results look good. It's documented that users of a commercial testing tool peeking at data and stopping at significance significantly inflates type-I error rates (Kohavi, Deng & Vermeer, 2022). Looking at the dashboard is fine; deciding based on the look isn't.

Is click-through rate enough to judge a hook?

Not on its own. CTR measures the gaining-attention stage, while advertising effectiveness works through two stages, gaining and holding attention (Bruns et al., 2025). A hook that produces clicks but no sales opened the first door and failed at the second. Your decision metric should sit closer to the bottom of the funnel.

One hook performed twice as well as another — should I scale it now?

Replicate first. The authors invoke Twyman's law for striking results — any figure that looks interesting is usually wrong — and reject a lift of over 300% in their own example on the grounds that they've never seen an improvement of that size across tens of thousands of tests (Kohavi, Deng & Vermeer, 2022). If it wins the second round too, scale it.

What if none of the five beats my current creative?

That's an expected outcome, not a bad sign. Published experiment success rates run at 33% for Microsoft, 15% for Bing, 10% for Booking.com, Google Ads and Netflix, and 8% for Airbnb Search (Kohavi, Deng & Vermeer, 2022). The move is to change the angle: if all five variations are different sentences for the same idea, you've really only tested one hook.

Sources

  1. Kohavi, R., Deng, A., & Vermeer, L. (2022). A/B Testing Intuition Busters: Common Misunderstandings in Online Controlled Experiments. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '22).ACM SIGKDD
  2. Bruns, D., Kopka, J. F., Borgmann, L., Prior, S., & Langner, T. (2025). Measuring Gaining and Holding Attention to Social Media Ads with Viewport Logging: A Validation Study Using Mobile Eye-Tracking. Journal of Advertising, 54(5), 655-672.Journal of Advertising
  3. Media Rating Council. (2016). MRC Mobile Viewable Ad Impression Measurement Guidelines, Final Version, June 28, 2016.Media Rating Council
  4. Gordon, B. R., Zettelmeyer, F., Bhargava, N., & Chapsky, D. (2019). A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Marketing Science, 38(2), 193-225.INFORMS Marketing Science

Ready to grow your brand?

It takes about as long as a coffee. Fill out the form, let us listen to your brand and build a plan made just for you.