Skip to content
Agency Review Guide

conversion optimization agency

Conversion Optimization Agencies: Testing Claims Honestly

Assess a conversion optimization agency by its sample-size planning, experiment discipline and willingness to explain what its results cannot prove.

By Thomas Christianson, Editor, Agency Review GuidePublished October 9, 2026Last researched October 9, 2026

Drafted with editorial AI from the Agency Review Guide calendar brief. Every figure links to its named public source; nothing here is an estimate.

An analyst compares a laptop and smartphone at a desk beside a closed notebook.

A persuasive CRO case study can conceal a weak experiment. An agency might highlight a winning variation without explaining how visitors were assigned, when the test stopped or whether additional purchases survived cancellations and refunds.

When comparing conversion optimization agencies, evaluate the evidence behind their claims before admiring the headline. The useful question is not whether an agency has ever produced a lift. It is whether its process can distinguish an improvement from noise—and explain when the evidence remains inconclusive.

Sample size starts with a business decision

There is no universal traffic threshold that makes every conversion experiment reliable. Ask the agency to prepare a feasibility assessment using your actual funnel, rather than prescribing a standard testing calendar.

The assessment should identify:

  • **The baseline outcome:** the existing purchase, qualified-lead or activation rate, with a clear denominator.
  • **The minimum detectable effect:** the change the experiment is designed to detect, expressed in commercially meaningful terms.
  • **Statistical assumptions:** the chosen false-positive tolerance, statistical power and analysis method.
  • **Eligible traffic:** visitors who can actually enter the experiment, not total website traffic.
  • **Assignment and measurement units:** whether assignment happens by person, account or another unit, and how repeat visits are handled.

Smaller effects generally require more evidence to distinguish from random variation. Repeated sessions from the same customer should not automatically be treated as independent observations.

Ask for an expected duration based on eligible traffic, plus an explanation of conversion delays and recurring business cycles. If the proposed test cannot realistically answer the question, the honest recommendation may be to change the scope, investigate usability problems or improve instrumentation first.

Confidence settings are not proof of business value

Confidence intervals express uncertainty under a statistical model. They do not certify that tracking is correct, randomization worked or an improvement will persist.

For a normal-based, two-sided interval, the standard critical values illustrate how confidence settings change the uncertainty allowance:

These are reference values, not sample-size recommendations or a prescription for every CRO platform. Ask which method the agency uses and why it suits your outcome and experiment design.

Require a written plan before launch

An experiment plan makes it harder to redefine success after seeing the results. Request it before implementation, not as an appendix written after a winner appears.

The plan should state:

  • The observed customer problem and proposed explanation.
  • The primary outcome that will determine the decision.
  • Guardrail outcomes, such as errors, refunds or lead quality.
  • Eligibility rules, exclusions and allocation logic.
  • The sample-size approach, expected duration and stopping rule.
  • How multiple variants, outcomes and subgroup analyses will be handled.
  • Who checks implementation and authorizes launch or rollback.

Use established methods, such as those documented in the NIST/SEMATECH e-Handbook of Statistical Methods, rather than accepting proprietary labels whose assumptions the agency cannot explain.

Monitoring is different from declaring a winner

Teams should monitor experiments for broken pages, payment failures and missing events. That does not justify repeatedly checking an ordinary fixed-horizon significance test and stopping as soon as the result looks favorable.

An agency proposing early stopping should explain its sequential method and the rules it follows. A conventional fixed-horizon approach should follow its planned analysis point. Neither approach excuses changing the primary outcome after inspecting the data.

Also ask how the team investigates unexpected allocation imbalances, inconsistent exposure and tracking differences between variants. More traffic does not repair an invalid experiment.

Recognize how uplift gets overstated

The American Statistical Association’s statement on p-values explains that statistical significance does not measure an effect’s size or practical importance. Buyers should therefore demand more than a significant result badge.

Watch for these presentation problems:

  • **Relative lift without the baseline.** Request the underlying outcome rates and absolute difference so the change has business context.
  • **Only successful tests shown.** Ask how the published example fits into the wider experiment history, including negative and inconclusive outcomes.
  • **Subgroups selected after analysis.** A promising device, channel or customer segment discovered afterward should usually become a follow-up hypothesis, not an unqualified success claim.
  • **A before-and-after comparison presented as a controlled test.** Promotions, traffic mix, inventory and seasonality may offer alternative explanations.
  • **Annualized revenue presented as realized revenue.** A projection requires assumptions about persistence, exposure and commercial conditions; it is not booked income.
  • **A proxy outcome presented as profit.** More form submissions may mean more low-quality enquiries. More orders may accompany lower margins or higher returns.

Request the estimated effect and its uncertainty interval. Then ask what decisions would be sensible at the less favorable end of that interval. This shifts the conversation from celebrating a point estimate to managing commercial risk.

Ask for evidence without demanding confidential data

A credible agency can demonstrate its process without exposing a client’s customer records. Ask for a redacted experiment plan, implementation checklist and results report from the same project.

The report should connect the hypothesis to the actual decision: shipped, rejected, repeated or left unresolved. It should also explain departures from the original plan and distinguish exploratory findings from planned analyses.

If a public case study omits sample size, runtime or uncertainty, record **“not published”** and identify the agency case-study URL you checked. Do not reverse-engineer missing inputs from a headline or treat missing public detail as proof of misconduct. Instead, ask whether the agency can provide supporting documentation privately.

Read our agency evaluation methodology alongside your shortlist criteria, but evaluate experiment evidence separately from general agency credentials.

Put testing responsibilities into the scope

CRO can fail operationally even when the statistical plan is sound. Establish who owns analytics repair, design, development, quality assurance and deployment.

If research points to navigation or comprehension problems, compare the proposed research scope with that of UX and UI agencies. If implementation is the bottleneck, clarify whether the CRO provider supplies engineering or coordinates with web development agencies.

For retail businesses, our guide to choosing an ecommerce marketing agency can help frame the broader brief. Keep the CRO remit explicit: which commercial outcomes it owns, which dependencies it manages and which decisions remain with your team.

Prefer commitments to sound research, reliable execution and transparent reporting over guaranteed winning tests. An agency that can explain why a result should not be trusted is demonstrating a capability buyers need—not admitting failure.

Infographic

Confidence settings change the uncertainty allowance

Figures in Standard normal critical value

Standard critical values for normal-based, two-sided confidence intervals; these are not sample-size targets or a universal CRO testing method. Source: OpenStax, Rice University — Introductory Statistics.

Frequently asked questions

How much traffic does a conversion optimization agency need?
There is no universal minimum. An agency should assess eligible traffic, baseline conversion rate, the effect worth detecting, statistical power and the analysis method before recommending an experiment.
Should a CRO agency guarantee a conversion uplift?
Treat guaranteed uplift cautiously. An agency can commit to research quality, implementation standards and transparent reporting, but a properly run experiment may produce a negative or inconclusive result.
What should a CRO case study disclose?
Look for the baseline, outcome definition, assignment method, sample size, duration, estimated effect and uncertainty. Missing public details should be recorded as not published, with the case-study source identified, and requested during due diligence.
Is stopping an A/B test early always wrong?
No. Stopping for safety or technical failures can be appropriate, and valid sequential methods can support early decisions. Repeatedly checking a conventional fixed-horizon test and stopping when it becomes significant undermines its planned error control.
Does statistical significance mean an experiment was profitable?
No. Statistical significance does not establish commercial value. Assess the effect’s size and uncertainty alongside margins, implementation costs, lead quality, refunds and other downstream outcomes.

Keep reading and check the sources

Further reading

Spotted something out of date? Send a correction.