What A/B Testing Is
Show variant A to some users, B to others.
Measure which performs better.
Statistical test to determine confidence.
Foundation of data-driven decisions.
When to A/B Test
Change with hypothesis about impact.
Enough traffic to reach significance.
Reversible change.
Metric can be measured clearly.
When Not To
Low traffic: takes too long.
Reputational risk of showing bad version.
Contract commitments (uniform experience required).
Legal/compliance changes.
Complete redesigns (people prefer familiar).
Sample Size
Depends on baseline conversion, expected lift, confidence level.
Calculator: online A/B test calculators.
Rule of thumb: 1000+ conversions per variant.
Bigger changes need less traffic.
Statistical Significance
95% confidence typical target.
90% acceptable for lower-stakes.
Not: winner after first day.
Wait for sample size hit.
Primary Metric
One primary metric per test.
Don’t cherry-pick winning metric.
Directly tied to hypothesis.
Aligned with business goal.
Guardrail Metrics
Things that shouldn’t get worse.
Retention.
Revenue.
Support tickets.
Even if primary wins, kill if guardrails degrade.
Hypothesis Formulation
“If we change X, then Y will happen because Z.”
Testable prediction.
Ties to user or business outcome.
Not: “let’s try this and see”.
Test Duration
Long enough for statistical significance.
Cover weekly cycle if applicable.
1-2 weeks typical minimum.
Not longer than needed.
Common Mistakes
Peeking: stopping test early when leading.
Multiple testing without correction.
Correlation vs causation.
Simpson’s paradox in aggregated data.
Not enough traffic (underpowered).
Segment Analysis
Overall winner may lose in some segments.
Analyze by user type, plan, geography.
But: separate hypothesis for each segment.
Don’t cherry-pick winning segment.
Practical Setup
Randomize by user ID (persistent).
Not by session (users see different versions).
Consistent hashing.
Version tracking in logs.
Tools
LaunchDarkly: feature flags + experimentation.
Split.io: dedicated experimentation.
Statsig: modern, growing.
Google Optimize: was free, discontinued.
Custom: SQL + statistical tests.
Testing Culture
Not just launching tests.
Documenting hypotheses.
Sharing results (wins AND losses).
Building organizational learning.
Iterating on insights.
Long-Term Learning
Track: what tests won, what didn’t.
Patterns emerge.
Fewer tests over time.
Broader intuition builds.
Micro-tests
Button color: rarely material.
Headlines: often material.
Prices: material and legally complex.
Feature availability: often huge impact.
Focus on high-leverage tests.
Based on Real Projects
This guide is based on our work with:
Further Reading
If this guide helped you, you might also want to read our comprehensive guide on Custom SaaS Development.
רוצים לדבר על הפרויקט שלכם?
שיחת ייעוץ חינם, ללא התחייבות - הרעיון שלכם + הניסיון שלנו
רוצים לדבר על הפרויקט שלכם?
אנחנו מתמחים בפיתוח SaaS, פתרונות AI, עיצוב UX/UI ובניית אתרים. ספרו לנו מה אתם צריכים.
דברו איתנו ←