← Back to list

Why So Many Tests End Up Changing Nothing

We often plan a test, and only when it is done, do we realize it had no impact.

Irit Amelchenko in Fiverr Engineering · 2026-03-24 10:32 · 394 claps · 4.3 min read
#ab-testing #data-analysis #fiverr #product-management #analytics
Open on Medium ↗
Wiki topics: BIZ · Business Strategy GRW · Growth & Analytics 📋 · Product Management

Why So Many Tests End Up Changing Nothing

We often plan a test, and only when it is done, do we realize it had no impact.

The experiment ran smoothly. We had clean allocation, complete data, and no technical issues. And yet, when the results came in, our reaction was: “That told us nothing.”

The main metrics didn’t meaningfully change. There was no clear decision we felt confident making. After 14 days, we were left with a well-executed test — and no real insight.

This isn’t a rare outcome. Many product and growth teams run experiments that are technically correct but strategically useless.

After running enough “boring” tests, I’ve come to a simple conclusion: most low-impact experiments fail long before they’re launched, during planning. In this article, I’ll explain how I reached that conclusion, highlight common planning mistakes that lead to flat results, and show how small changes to the hypothesis, population, or framing can dramatically increase impact.

What to look for before you launch a test

Your product manager comes to you with an idea: a test that will change the world, make an impact, and increase revenue. Take this as an opportunity — as a data analyst — to be critical, set the tone, and adjust the plan so your team can maximize impact.

Think about the following:

  • Decide on the impact you are looking to achieve in this test Are you looking to move the needle and impact a main metric, or are you looking for learnings that will help you make future decisions and adjust the next A/B test?
  • Clarify the population Not just the size of the population, but also its impact. For example, in e-commerce, some buyers are very experienced on the platform and have their own habits, making them harder to influence.
  • Understand what the expected outcome is What metrics will increase? What are the trade-offs? Who benefits from this test?
  • Before the test, understand the next step of each optional outcome If the test fails, what do we do? If the test shows mixed signals, what do we do? This will help you think of options and opportunities you might not come up with after the test is closed and you already have the results.

An example of a doomed flat test (and what we’d change)

Imagine a team planned the following test: adding a CTA to search results that encourages buyers to message the seller before purchasing, assuming it will lead to better conversions and higher satisfaction.

It feels like a safe and reasonable test. Search has a lot of traffic and plenty of diversity in buyer types, meaning we can learn from it, but we also expect real impact. The change is easy to roll back if we lose, and conversion is a metric that moves the needle.

The test goes live, and the results come back flat. Why?

At this point, we need to go back to the steps we established and ask ourselves what went wrong.

When thinking deeper, a few things come to mind:

  • Search includes mixed buyer types with mixed intent Some buyers arrive with a clear vision and are ready to purchase now. Others are exploring, and some have complex requirements that require more context. That mix can produce mixed results. If we’d focused on a specific population, we might have seen a clearer outcome — and potentially more impact.
  • In some marketplace areas, conversations are unnecessary The service is simple and immediate, and encouraging a conversation might create friction. Focusing on a specific area and planning ahead would allow us to achieve better results.
  • Experienced buyers know exactly when a conversation is needed They make up the majority of the buyer base. But it is the less experienced buyers who need more guidance and encouragement, and the larger group absorbs this effect.
  • Placement is key Users usually need more context before starting a conversation. Showing a CTA in search might be a bit early in the road.
  • “Success” in such tests may be defined differently from conversion Conversations are meant to improve decision quality, not necessarily speed. In a test like this, satisfaction, cancellations, revisions, disputes, or repeat behavior may be better success criteria than immediate conversion.

This wasn’t a hypothetical example. It was a real test we ran as a team at Fiverr, and it actually helped surface something useful. Once the results came in, the gaps were pretty clear. You could see where better coordination between sides could have made the test more focused, and where we could have been more precise in the way we framed the questions. A few sharper questions upfront probably would have saved us time. The test itself didn’t teach us anything new about our users or the marketplace, but it did push us to think more carefully about how we plan these tests. If we slow down a bit at the beginning and make sure we’re aligned, we’ll run more accurate tests and get better answers out of them.

The questions we should have asked

Looking back, there were a couple of questions we could have answered during planning that would have made the risk of a flat result much clearer.

  1. Whose behavior do we actually expect to change — and who will likely be unaffected? This forces you to think beyond traffic size and ask whether the population you’re testing on is sensitive to the change at all. In this case, experienced buyers and buyers with simple needs were unlikely to benefit from a conversation.
  2. How variable is this behavior today across users, segments, or contexts?Before testing, we can look at existing data to understand whether the behavior already varies meaningfully across populations. For example, are there segments, categories, price points, or entry paths where this behavior is already more common or more successful? Or is it relatively flat everywhere? This can point directly to where a test is most likely to produce impact.

None of these insights required new experiments or complex analysis. They were available to us through questions we could have asked and data we already had.

As analysts, our impact doesn’t come from running more experiments — it comes from shaping better questions upfront. The difference between a forgettable test and a meaningful one is often a few deliberate decisions made early.

If this article helps you pause before the next test and rethink how it’s planned, it has already done its job.


메타데이터
post_id
71b3f6b63a38
slug
why-so-many-tests-end-up-changing-nothing-71b3f6b63a38
url
https://medium.com/fiverr-engineering/why-so-many-tests-end-up-changing-nothing-71b3f6b63a38
canonical_url
https://medium.com/fiverr-engineering/why-so-many-tests-end-up-changing-nothing-71b3f6b63a38
author_url
https://medium.com/@irit.amelchenko
status
ok
fetched_at
2026-06-14 11:28:49