Most of your website tests should actually fail

Here is a stat that completely changed how we look at website testing: even at mature experimentation companies running thousands of tests a year with dedicated data teams, only 10% to 20% of tests produce a significant positive result.
So, when an internal dashboard shows a 70% "win rate," it usually means you might be accidentally grading your own homework.
Most ideas fail. And that is a good thing. The best experimentation cultures are built on the simple fact that our intuition is often unreliable. If teams with every possible advantage see most of their tests fail, it is totally normal for a DTC brand to see similar results. It isn't a sign of bad ideas; it's just the reality of testing.
When win rates look too good to be true, it often comes down to a few common measurement traps. Here are three of the most common ones, and how you can gently course-correct:
1. Calling the test as soon as the chart turns green
If you watch a flat test for long enough, it will eventually cross the "statistically significant" line just by pure chance. If you stop the test right then and declare a winner, you might just be capturing a lucky moment rather than a true lift in conversion.
It's generally much safer to set a fixed timeline and sample size before the test launches, and stick to it, even if the chart looks exciting early on.
2. Getting caught by the Novelty Effect
When you change a highly visible element on your site (like a button colour or a banner), returning visitors often click it simply because it is new. This can cause a massive spike in engagement for a week or two. But that lift often decays rapidly as the "newness" wears off.
If you only measure that short-term spike, it can give you a false sense of security.
3. Testing tweaks that are too small to measure
If a store only has the traffic volume to detect double-digit changes in conversion, running a test on a slightly different font size can be problematic. It becomes incredibly difficult to separate the actual result from normal background noise.
It often pays to test changes that are bold enough for your specific traffic volume to actually register.
The Takeaway
A low win rate isn't a reason to play it safe, it's actually the opposite. Because most tests lose anyway, testing tiny, safe tweaks just buys you undetectable results. To see real movement, you have to take bigger swings: testing completely different offer structures, radically different product page layouts, or entirely new pricing presentations.
It might be worth shifting the conversation internally this week. Instead of asking your team for a high win rate, try aiming for a high learning rate.
A trustworthy loss tells you exactly what your customer doesn't care about, which prevents you from rolling out a change that could hurt revenue for years. A program winning 15% of the time with clean data compounds your growth nicely, whereas chasing a 70% win rate can sometimes just compound false hope.
This week's deep dive was brought to you by Dayo Samuels, Rainy City Agency.
Deep Dive LDN: Annual Planning
November 12th, 2026 | London | 9AM - 3PM
Step away from the daily grind, silence the notifications, and let's figure out what 2027 actually looks like for your brand. We are heading to London on 12th November for a highly curated, full-day intensive focused on annual planning.
This is an intimate room where real founders and operators sit down together over a brilliant breakfast and a premium three-course lunch to exchange ideas, solve pressing problems, and genuinely help each other.
Together, we'll unpack the critical shifts, hidden bottlenecks, and big-picture strategies you need to consider before you lock in your roadmap for the year ahead.
It is the perfect environment to swap notes, make vital connections with like-minded peers who are navigating the same hurdles as you, and leave with clarity on where to focus your energy in Q1.
It is completely free for eCommerce brands to attend, but because we keep the room small, seats are vetted and move incredibly fast.
Podcast: Why Radley Stopped Giving Away Free Margin
Are you accidentally paying your best customers to check out?
In our latest episode, we sit down with Max Wright, Head of Ecom at Radley, and Dan Bond, VP of Marketing at RevLifter, to discuss how the iconic British heritage brand completely overhauled its promotional strategy.
They break down:
- How Radley stopped training their customers to wait for sales,
- moved to highly targeted "intelligent offers,"
- drove a massive 15% lift in conversion while reducing their promotional costs.
If your discounting strategy is eating into your profits, you need to hear this playbook.
In Other News…
Anthropic Adds Invisible Watermarks to Claude AI: To comply with the EU's AI Act, Anthropic is rolling out invisible, machine-readable watermarks to all text generated by its Claude models globally. The mark survives copying and pasting, making it easier to detect if a piece of content was AI-processed.
What this means for you: As detection tools become public, relying purely on raw AI output risks breaking consumer trust and potentially taking a hit in search engine rankings.
Target Appoints First-Ever Chief AI Officer: Retail giant Target has named Chandhu Nair as its new Chief AI Officer, signalling a massive push to integrate artificial intelligence directly into its C-suite and frontline operations.
What this means for you: It might be time to start exploring how AI can solve your heavy operational and inventory bottlenecks.

DTC Live Insider
Join the waitlist to receive the UK's leading eCom insights publication!

