In 2009, a Google engineer wanted to know which shade of blue got more clicks. So the team tested 41 shades—not two, forty-one—on live search traffic, comparing click-through rates down to the pixel. The winning blue reportedly added around $200 million a year in ad revenue. Marissa Mayer was running the experiment, and it became one of the most cited stories in UX design, mostly because it sounds absurd until you realize it worked.
That’s the thing about A/B testing. It sounds like the most boring tool in a designer’s kit—two versions, split traffic, wait for data—and yet it’s responsible for some of the biggest design decisions on the internet. Booking.com reportedly runs well over a thousand experiments a year, on everything from button copy to how urgency messages are worded, and the company has built an entire internal culture around treating almost every design opinion as a testable claim rather than a settled fact. Most of those experiments fail. The ones that work pay for all the ones that didn’t, and the failures themselves become a kind of archive—a running record of ideas that felt obviously right and turned out to be wrong once real users touched them.
Designers tend to have a complicated relationship with this. Testing feels like it should validate good instincts, and sometimes it does. More often it humbles them, which is uncomfortable in a field built on taste and a sharp eye. The upside is that a humbled designer with eight months of test data is far more dangerous, in the good sense, than a confident one working purely on instinct.
What A/B Testing Actually Is
Split testing, as some designers still call it, means showing two versions of a screen, page, or flow to different segments of your audience and measuring which one performs better against a specific goal. Version A might be your current landing page. Version B swaps the headline, moves the CTA button, or changes a single word in the subheader. Whichever version drives more signups, clicks, or purchases wins, and the losing version gets retired.
The metric you choose matters more than people admit. A redesigned pricing page might increase clicks on a CTA button while actual conversions drop, because more people click through to a page that ultimately confuses them. Good A/B testing setups track a primary goal metric alongside two or three guardrail metrics, so a “win” on clicks doesn’t mask a loss somewhere further down the funnel.
The Surprising Part: It’s a Conflict-Resolution Tool
Everyone frames A/B testing as a data collection method for UX design. Fair enough, but the more useful function, especially for design studios working with clients, is political. Every studio has lived through the meeting where the client wants the big red button and the design lead wants the subtle outlined one, and both sides have made their case three times already. Nobody’s mind changes in that room.
A/B testing ends the argument without anyone losing face. Run the button colors against real users for two weeks, let the data pick, and the fight evaporates. The design team gets to stop defending taste as if it were strategy, and the client gets to stop overriding UX design decisions with a hunch about what their nephew likes. This is arguably the single most underrated use of split testing in agency work, and it rarely gets mentioned because “resolves office politics” doesn’t sound as rigorous as “optimizes conversion rate.”

What’s Actually Worth Testing
Every pixel does not deserve its own experiment. Designers tend to reach for A/B testing on:
- CTA button copy, size, and placement
- Headlines and subheaders on landing pages
- Hero images and product photography
- Form length and field order on signup pages
- Pricing page layout and plan ordering
- Video versus static image on a landing page
- Navigation structure in mobile app design
A famous Bing case study found that changing how search result links were underlined—a change so small the engineer who proposed it almost didn’t bother submitting it—increased revenue by tens of millions of dollars annually once rolled out. Small, testable changes routinely beat sweeping redesigns in measurable impact, mostly because sweeping redesigns introduce too many variables at once to know which one actually moved the number.
The Process, Stripped of Jargon
- Find the leak first. Before building anything, pull your analytics. Look for pages with unusually low conversion rate, high bounce, or a CTA button nobody clicks. That’s your testing priority list, not a spreadsheet of ideas someone had in a brainstorm.
- Pick one metric to win. Signups, time on page, add-to-cart rate—whatever it is, write it down before you build anything. If you can’t name the metric in one sentence, the test is not ready to launch.
- Write down why you think B will win. This forces you to articulate a real hypothesis instead of just trying a font because you like it. “Users are abandoning the form because it asks for a phone number too early” is a hypothesis. “Let’s make it pop” is not.
- Build both versions to production quality. This is where most first-time testers get sloppy. A half-finished B version with placeholder copy will lose to a polished A version every time, and the result will tell you nothing true about the actual idea.
- Run it long enough to matter. A common mistake in mobile app design testing is calling a winner after three days because the numbers “looked good.” Give the test enough traffic and enough time to cover a full weekly cycle—weekday and weekend behavior often differ enough to skew early results.
- Read the loss, not just the win. When B loses, ask why before moving on. Sometimes the losing version reveals something about user behavior that’s more valuable than the win itself.

Where It Actually Helps
A/B testing measures what users do, not what they say they’ll do in a focus group, which turns out to matter enormously—people are famously bad at predicting their own clicks. It’s also cheap. You don’t need a research budget or a recruiting agency, just two live versions and a tool to split the traffic; Optimizely, VWO, and Google Optimize’s open-source successors all handle this without much technical lift. And because it isolates one variable at a time, it catches things user interviews miss entirely—a button three pixels too small, a headline that reads fine out loud but flops on a phone screen.
There’s also a compounding effect worth mentioning. Each test result feeds the next hypothesis. A studio that’s run twenty tests on a client’s site has twenty small, verified facts about that specific audience—facts a competitor starting from scratch simply doesn’t have.
Where It Falls Apart
Split testing only works on a finished product. You can test two versions of a CTA button, but you can’t meaningfully test a button in isolation from the copy around it, the page load speed, or the trust signals nearby—the result will be contaminated by everything you didn’t control for. It also can’t explain motive. A test tells you B outperformed A by 12%. It never tells you why, which means teams that treat testing as a substitute for actual user research end up with a pile of winning variants and no real understanding of their audience.
And testing can’t catch what’s broken if nobody thought to test it. If your checkout flow has a confusing step nobody flagged, A/B testing won’t surface it unless someone happens to design a version that removes it. For that kind of discovery work, you still need real user research—interviews, session recordings, the unglamorous stuff.
The Actual Takeaway
Treat A/B testing as a habit, not a project. Studios that run one test a quarter learn less than studios that run one a month, even with smaller sample sizes, because the compounding knowledge about a specific audience is worth more than any single big win. Start small—one button, one headline, one week—and build the muscle before chasing the Google-blue-sized experiment. The forty-one shades came after years of smaller tests, not before, and the team behind it had already built the infrastructure, the patience, and the internal trust to run something that ambitious without anyone panicking halfway through.
The next time a design review turns into a standoff over which version is right, resist the urge to win the argument in the room. Build both, ship both to a slice of real traffic, and let the users settle it. It’s a slower way to be right, but it’s a much harder way to be wrong.
Recommended Reading
Liked this? Good. We’ve got more where that came from—case studies, hot takes, and the occasional design opinion nobody asked for but everyone needed:
Precious Errors: Testing iOS Mobile Applications
Tests Go First. Usability Testing in Design
Case Study: Adam Braun. Creating Personal Website for Entrepreneur
Case Study: ID Scanning Product Icons. Graphic Design Process