Back to Blog
2026-09-25
Comparisons
Greta.sh Editorial Team

Best A/B Testing Tools in 2026 (And When You Do Not Have the Traffic)

Statsig, Eppo, GrowthBook, PostHog, VWO, AB Tasty and Optimizely compared — and the sample-size sum that tells most teams to run a holdout instead.

Best A/B Testing Tools in 2026 (And When You Do Not Have the Traffic)
ByHiteshi Soni· Marketer, Questera

Statsig and Eppo are the modern experimentation platforms, built around correct statistics and warehouse-native analysis. GrowthBook is the open-source option you can self-host. PostHog bundles experiments with analytics and session replay. VWO and AB Tasty serve marketing-led website testing. Optimizely is the enterprise incumbent. And the honest answer for a great many teams is that you do not have the traffic for a conclusive test. That is not solved by a holdout — an unbalanced split needs more total traffic, not less — so what is left is a guarded rollout with detection limits you state openly. Work out your required sample size before you shop; it changes which of these you need, and sometimes whether you need one.

Most A/B testing advice assumes you will reach significance. Do the arithmetic first — for a lot of products, the tool is not the constraint.

Do the sample-size sum before anything else

An honest rule of thumb: detecting a relative improvement of 10% on a 5% baseline conversion rate takes roughly tens of thousands of users per variant. If your signup flow sees two hundred people a week, a two-variant test needs months, during which the product, the traffic mix and the season all change — and the result is no longer about the variant. No tool fixes this. Knowing it changes what you buy and what you do instead.

At a glance

ToolShapeSelf-hostBest for
StatsigExperimentation + feature flagsNoProduct teams wanting correct stats by default
EppoWarehouse-native experimentationNoTeams with a warehouse and a data function
GrowthBookOpen-source experimentationYesTeams who want to own the stack
PostHogAnalytics + experiments + replayYesConsolidating three tools into one
VWO / AB TastyVisual website testingNoMarketing-led page testing
OptimizelyEnterprise experimentationNoLarge programmes with governance needs
GretaFlags and guarded rolloutYour deploymentMonitoring, not a substitute for power

Pricing across this category is mostly quoted rather than published as of September 2026 — verify live.

The tools, honestly

Statsig and Eppo are where serious product experimentation has moved. Both take the statistics seriously — sequential testing, variance reduction, guardrail metrics — which matters because the most common experimentation failure is not a bad tool but a team calling a result early. GrowthBook offers the same shape as open source, and is the right answer if self-hosting is a requirement rather than a preference.

PostHog is the consolidation play: analytics, flags, replay and experiments in one product, with the usual trade of breadth against depth. VWO and AB Tasty are built for marketers testing pages visually rather than engineers testing product changes, and they are good at that job specifically. Optimizely is the incumbent and prices like one.

Greta.sh

Got an idea? Build it now!

Just start with a simple prompt. No coding required — Greta.sh turns your idea into a working app in minutes.

What to do when you cannot reach significance (and what does not help)

Three options that are better than running an underpowered test and believing the result.

First, correct a claim you will see made often, including in an earlier version of this page. A 90/10 holdout does not need less traffic than a 50/50 test. For the same effect size and error rates it needs substantially more — under equal variances, the sample demand of a 90/10 split is roughly 2.8 times that of a 50/50 split, because the small arm is where the uncertainty lives. Asking "did anything get worse" rather than "which is better" does not change that by itself. What does change it is a larger threshold: you can detect a 20% regression far more cheaply than a 2% improvement, and that is the honest reason to run one.

A guarded rollout with a stated detection limit. Release to 10%, then 50%, then everyone, watching a small set of metrics agreed in advance. Before you start, write down the size of regression this setup could actually detect at your traffic, and treat anything smaller as invisible. This is monitoring, not an experiment, and "no harm detected" means "no harm above our threshold" — a sentence worth putting in the report, because the alternative is reading an underpowered null as reassurance.

A holdout, sized honestly. Keep an unexposed group if you want a comparison rather than just monitoring, and size it against the regression you care about rather than defaulting to 10%. If the arithmetic says the holdout needs to be larger than you are comfortable with, that is the finding.

Decide on judgement and say so. Sometimes the right call is to ship the better-reasoned option and record that it was a judgement, not a measurement. That is more honest than a test with a p-value nobody should trust, and it is faster.

If you built the app yourself, the flag and assignment machinery is a small amount of work rather than a subscription — describing it to Greta gets you variant assignment and a held-out group. It does not get you the power analysis, which is the part that decides whether the result means anything. Free tier is 30 credits a month (5 a day); $20/month after ($5 with the current welcome offer). The pricing experiment specification covers the design either way.

The boundary, plainly: correct sequential statistics, variance reduction, automatic guardrail monitoring and a results interface your team will trust are what the platforms above have built. Rolling your own assignment is easy; rolling your own statistics is how people convince themselves of things that are not true. If you will run experiments continuously, buy one.

FAQ

What is the best A/B testing tool? Statsig or Eppo for product experimentation, GrowthBook if self-hosting matters, PostHog to consolidate, VWO or AB Tasty for marketing page tests, Optimizely at enterprise scale.

How much traffic do I need? Enough that the sum works out — run a sample-size calculation against your own baseline rate and the smallest effect worth detecting. Do it before shopping; it is often the answer to the whole question.

Can I A/B test with low traffic? Not conclusively, and a holdout is not the workaround it is often presented as — an unbalanced split needs more total traffic than a balanced one, not less. What low traffic does permit is detecting large regressions cheaply. Run a guarded rollout, state the size of regression you could actually detect, and report "no harm above our threshold" rather than "no harm".

Is PostHog good enough for experiments? For many teams yes, particularly if the alternative is three subscriptions. Teams running experiments as a core practice usually end up wanting the statistical depth of a specialist.

End of Log Entry
↑ Return to Top

Build Something Real

If you can describe it, you can build it.