Insights · Experimentation

Why Most Startup Experiments Fail (and What to Do Instead)

Sam Gil·Sep 2025·6 min

The Illusion of Winning Tests

Your growth team is running experiments. You hit stat sig. You’ve got charts showing +8% here, +5% there. Maybe 10 “winners” in a quarter.

And yet… company growth is flat.

This is the most common failure mode we see in startups.

Not because your team is lazy. Not because your math is wrong.

But because the entire approach to experimentation is borrowed from companies at a totally different scale.

Spotify or Uber can afford to throw dozens of tests at the wall. At startup scale, with limited engineers and data science bandwidth, this approach is toxic.

The Real Problem

Most experiments are framed around outcomes, not behavior.

  • Did revenue per user increase?

  • Did ATC rise?

  • Was conversion rate better?

That’s useful… but shallow. You don’t know why.
And without the “why,” you can’t trust the result — or scale the learning.

The outcome? 10 “winning” tests that don’t compound. Leadership wastes time chasing anecdotes. The company stays stuck.

The Framework Shift

At a leading fashion brand, we rebuilt the experimentation process from the ground up.

Every test had to answer one question up front:

If we change [X], for [user group], we expect [metric] to move - because [behavioral reason].

That last piece - because - is everything.
It drags stakeholders out of half-written hypotheses and forces alignment on user intent.
It pulls designers’ thinking into the light.
And it sets you up to measure behavioral input metrics, not just outputs.

Case 1: The PDP Buy Box

Hypothesis:

If we redesign the buy box

For mobile users,

We will increase conversion rate

Because the gallery and selector being in one place will enable users to browse more variants to easily and confidently find their favorite

Result:

  • Topline: ATC looked up, but cart completion fell → net flat. No stat sig.

  • Behavioral slice: Fewer users explored >1 variant. This drove the loss.

  • Among users who did explore variants, ATC and CVR soared.

Action: A 1-day engineering tweak encouraged variant browsing.

Impact: +12% ATC, +8% CVR overall; +15% CVR and +10% ASP on high-variant products.

Learning loop: “# of variants viewed” became a standard input metric in PDP reporting.

Case 2: Free Shipping Threshold

Hypothesis:

If we raise the free shipping threshold from $0 to $1,500,

For all users,

We will increase net profit

Because users will add items to hit the threshold, offsetting any increase in price sensitivity.

Result:

  • Shipping revenue: +$2.5M annualized.

  • AOV: +9% overall, driven by +15% lift in $1,500–$2,000 carts.

  • Conversion: -3% overall, but -20% in sub-$1,500 carts, flat above $2,000.

  • EBITDA margin: +4 pts, even with topline flat.

Impact: KPI tree (shipping revenue, AOV, conversion) became standard for future monetization tests.

The Takeaway for Startups

If your experiments don’t start with because, they don’t end with growth.

Force behavioral hypotheses.

Define the input metrics that prove intent.

Slice results by behavior, not just averages.

Build KPI trees that show trade-offs across revenue, margin, and conversion.

This isn’t a “nice to have.” It’s survival.

Because at startup scale, you don’t have 1,000 tests to hide behind.

You only have the ones that actually change user behavior — and the systems to see it clearly.


SG
Sam Gil
Principal, Growth & Analytics · Meridian Growth
About the team

Keep reading

ExperimentationSep 2025 · 6 min
Why Most Startup Experiments Fail (and What to Do Instead)
Read
Retention & LifecycleSep 2025 · 5 min
Retention in Marketplaces: Why Liquidity Beats Variety Alone
Read

Get in touch.

Tell us what you are trying to figure out. If we can help, we will say how. If we are not the right fit, we will say that too, and point you somewhere better.

01
Free work first
We start by doing a real piece of work on your data, free, so the first thing you judge is output, not a pitch.
02
First build, fixed scope
A trusted core metric layer in weeks, not quarters, at a fixed price. Small enough to be safe, real enough to matter.
03
Scale when it earns it
Month to month from there, from a few thousand a month. You own every model, dashboard and definition.