Skip to main content
Free Calculator

Meta-Analysis Calculator

Combine results from multiple A/B tests to estimate the true effect size with greater precision. Uses inverse-variance weighted fixed-effects meta-analysis.

Only pool genuinely comparable effects.

Every row must use the same metric and effect scale and include its actual standard error. Lift plus sample size is not enough. This fixed-effect result assumes one common true effect and should not be treated as universal across different brands, pages, devices, or audiences.

Results

Pooled Effect

4.74%

weighted average lift

95% CI

[2.35, 7.13]

confidence interval

p-value

0.0001

Significant

Heterogeneity

I² Statistic

0.0%

Low heterogeneity

Cochran's Q

0.57

Study Breakdown

StudyEffect (%)WeightContribution
Test 15.20%33.7%
Test 23.80%45.9%
Test 36.10%20.4%
Pooled4.74%100%

You just ran the numbers. Where will the result live?

A calculator answers one question and forgets it. GrowthLayer keeps each test your team saves — the hypothesis, the numbers, and the decision — so future planning can start from a reviewable record.

  • Structured history of saved experiments
  • Import results from CSV in one step
  • Program summaries across saved tests
  • Free to start — no card required

Methodology

This calculator uses a fixed-effects meta-analysis with inverse-variance weighting, which is the standard approach for combining A/B test results. The core formulas are: Weight for study i: w_i = 1 / SE_i² Pooled effect: θ = Σ(w_i × θ_i) / Σ(w_i) Pooled standard error: SE(θ) = √(1 / Σ(w_i)) Where: - θ_i is the observed effect (lift %) for study i - SE_i is the standard error for study i - w_i is the inverse-variance weight The 95% confidence interval is: θ ± 1.96 × SE(θ) Each study must include its actual standard error on the same effect scale as the reported effect. If you have a two-sided 95% confidence interval, derive SE as (upper - lower) / (2 × 1.96). Sample size and lift alone are not enough to reconstruct a standard error. This fixed-effect model assumes every included study estimates one common true effect. It is only appropriate when the outcome definition, effect scale, treatment contrast, analysis unit, and target population are comparable. A low I² value does not prove comparability, especially with only a few studies. If effects plausibly vary by brand, page, device, or audience, do not interpret this pooled estimate as universal; use a justified random-effects or hierarchical analysis outside this calculator. Heterogeneity is assessed using Cochran's Q statistic: Q = Σ w_i × (θ_i - θ)² Q follows a chi-squared distribution with k-1 degrees of freedom under the null hypothesis of homogeneity. The I² statistic quantifies the proportion of total variation due to true heterogeneity rather than sampling error: I² = max(0, (Q - df) / Q × 100) I² interpretation: 0–25% low, 25–75% moderate, 75%+ high heterogeneity.

Frequently Asked Questions

What is meta-analysis in A/B testing?
Meta-analysis is a statistical technique for combining results from multiple independent experiments that test similar hypotheses. In A/B testing, it lets you pool results from several tests — for example, running the same CTA change across different pages or markets — to get a more precise estimate of the true effect size. This is especially useful when individual tests are underpowered.
When should I combine A/B test results?
Combine results only when the tests estimate the same outcome, effect scale, treatment contrast, and target population closely enough for one common effect to be meaningful. Similar titles or pattern labels are not enough. Do not pool winner labels, p-values, or relative lifts calculated from incompatible metrics.
What is heterogeneity and why does it matter?
Heterogeneity measures how much the effect sizes vary across your studies beyond what random chance would explain. High heterogeneity (I² > 75%) suggests the true effect differs meaningfully between tests, which means a single pooled estimate may not accurately represent any individual context. In such cases, a random-effects model or investigating the sources of variation may be more appropriate.
What is the difference between fixed and random effects models?
A fixed-effects model assumes all studies share one true effect size, and differences are due to sampling error only. A random-effects model assumes the true effect varies between studies and accounts for this extra variability. This calculator uses a fixed-effects model, which is appropriate when your studies are similar in design and context. If heterogeneity is high, consider a random-effects approach.
How many studies do I need for a meta-analysis?
The calculator accepts two or more studies, but a small number does not automatically create a robust result. With only 2–3 studies, heterogeneity is hard to estimate and each study can dominate the pooled value. Treat a small fixed-effect synthesis as exploratory and inspect every study rather than relying on the pooled p-value.

Related Calculators

Choosing a method or double-checking another tool? Compare 11 public A/B test calculators and download the evidence matrix.

Updated for 2026. Built by GrowthLayer — built for evidence-aware experimentation teams.