A persuasive CRO case study can conceal a weak experiment. An agency might highlight a winning variation without explaining how visitors were assigned, when the test stopped or whether additional purchases survived cancellations and refunds.
When comparing conversion optimization agencies, evaluate the evidence behind their claims before admiring the headline. The useful question is not whether an agency has ever produced a lift. It is whether its process can distinguish an improvement from noise—and explain when the evidence remains inconclusive.
Sample size starts with a business decision
There is no universal traffic threshold that makes every conversion experiment reliable. Ask the agency to prepare a feasibility assessment using your actual funnel, rather than prescribing a standard testing calendar.
The assessment should identify:
- **The baseline outcome:** the existing purchase, qualified-lead or activation rate, with a clear denominator.
- **The minimum detectable effect:** the change the experiment is designed to detect, expressed in commercially meaningful terms.
- **Statistical assumptions:** the chosen false-positive tolerance, statistical power and analysis method.
- **Eligible traffic:** visitors who can actually enter the experiment, not total website traffic.
- **Assignment and measurement units:** whether assignment happens by person, account or another unit, and how repeat visits are handled.
Smaller effects generally require more evidence to distinguish from random variation. Repeated sessions from the same customer should not automatically be treated as independent observations.
Ask for an expected duration based on eligible traffic, plus an explanation of conversion delays and recurring business cycles. If the proposed test cannot realistically answer the question, the honest recommendation may be to change the scope, investigate usability problems or improve instrumentation first.
Confidence settings are not proof of business value
Confidence intervals express uncertainty under a statistical model. They do not certify that tracking is correct, randomization worked or an improvement will persist.
For a normal-based, two-sided interval, the standard critical values illustrate how confidence settings change the uncertainty allowance:
- A **90% confidence level** uses a critical value of **1.645**, as explained in OpenStax’s normal-distribution confidence interval guidance.
- A **95% confidence level** uses **1.960**, according to the same OpenStax guidance.
- A **99% confidence level** uses **2.576**, according to OpenStax’s confidence interval guidance.
These are reference values, not sample-size recommendations or a prescription for every CRO platform. Ask which method the agency uses and why it suits your outcome and experiment design.
Require a written plan before launch
An experiment plan makes it harder to redefine success after seeing the results. Request it before implementation, not as an appendix written after a winner appears.
The plan should state:
- The observed customer problem and proposed explanation.
- The primary outcome that will determine the decision.
- Guardrail outcomes, such as errors, refunds or lead quality.
- Eligibility rules, exclusions and allocation logic.
- The sample-size approach, expected duration and stopping rule.
- How multiple variants, outcomes and subgroup analyses will be handled.
- Who checks implementation and authorizes launch or rollback.
Use established methods, such as those documented in the NIST/SEMATECH e-Handbook of Statistical Methods, rather than accepting proprietary labels whose assumptions the agency cannot explain.
Monitoring is different from declaring a winner
Teams should monitor experiments for broken pages, payment failures and missing events. That does not justify repeatedly checking an ordinary fixed-horizon significance test and stopping as soon as the result looks favorable.
An agency proposing early stopping should explain its sequential method and the rules it follows. A conventional fixed-horizon approach should follow its planned analysis point. Neither approach excuses changing the primary outcome after inspecting the data.
Also ask how the team investigates unexpected allocation imbalances, inconsistent exposure and tracking differences between variants. More traffic does not repair an invalid experiment.
Recognize how uplift gets overstated
The American Statistical Association’s statement on p-values explains that statistical significance does not measure an effect’s size or practical importance. Buyers should therefore demand more than a significant result badge.
Watch for these presentation problems:
- **Relative lift without the baseline.** Request the underlying outcome rates and absolute difference so the change has business context.
- **Only successful tests shown.** Ask how the published example fits into the wider experiment history, including negative and inconclusive outcomes.
- **Subgroups selected after analysis.** A promising device, channel or customer segment discovered afterward should usually become a follow-up hypothesis, not an unqualified success claim.
- **A before-and-after comparison presented as a controlled test.** Promotions, traffic mix, inventory and seasonality may offer alternative explanations.
- **Annualized revenue presented as realized revenue.** A projection requires assumptions about persistence, exposure and commercial conditions; it is not booked income.
- **A proxy outcome presented as profit.** More form submissions may mean more low-quality enquiries. More orders may accompany lower margins or higher returns.
Request the estimated effect and its uncertainty interval. Then ask what decisions would be sensible at the less favorable end of that interval. This shifts the conversation from celebrating a point estimate to managing commercial risk.
Ask for evidence without demanding confidential data
A credible agency can demonstrate its process without exposing a client’s customer records. Ask for a redacted experiment plan, implementation checklist and results report from the same project.
The report should connect the hypothesis to the actual decision: shipped, rejected, repeated or left unresolved. It should also explain departures from the original plan and distinguish exploratory findings from planned analyses.
If a public case study omits sample size, runtime or uncertainty, record **“not published”** and identify the agency case-study URL you checked. Do not reverse-engineer missing inputs from a headline or treat missing public detail as proof of misconduct. Instead, ask whether the agency can provide supporting documentation privately.
Read our agency evaluation methodology alongside your shortlist criteria, but evaluate experiment evidence separately from general agency credentials.
Put testing responsibilities into the scope
CRO can fail operationally even when the statistical plan is sound. Establish who owns analytics repair, design, development, quality assurance and deployment.
If research points to navigation or comprehension problems, compare the proposed research scope with that of UX and UI agencies. If implementation is the bottleneck, clarify whether the CRO provider supplies engineering or coordinates with web development agencies.
For retail businesses, our guide to choosing an ecommerce marketing agency can help frame the broader brief. Keep the CRO remit explicit: which commercial outcomes it owns, which dependencies it manages and which decisions remain with your team.
Prefer commitments to sound research, reliable execution and transparent reporting over guaranteed winning tests. An agency that can explain why a result should not be trusted is demonstrating a capability buyers need—not admitting failure.
