This website uses cookies

Read our Privacy policy and Terms of use for more information.

The new button won on Tuesday. By Friday the result disappeared, payment errors were higher, and the team had checked significance 19 times. Improving a store takes more than choosing the prettiest variant. First understand the shopper's difficulty. Then state what you expect to change and run a design that can answer it. Finally, check the experiment's validity and the business cost of the result. These 18 terms separate observation, experimentation and statistical inference. Examples are illustrative; statistical claims depend on the selected method, assumptions and stopping rules, while interaction recordings require appropriate privacy controls.

🧭 Jump to a term

Diagnose the experience: Conversion friction · CTA · Heatmap · Session replay · Usability testing · CRO · Test hypothesis

Design a comparison: A/B testing · Control group · Randomization · MVT

Trust and interpret the result: MDE · Statistical power · SRM · Statistical significance · Confidence interval · Guardrail metric · Novelty effect

⭐ Know these first

Start with CRO, A/B testing, Statistical significance, Statistical power, MDE. Then follow the grouped learning order below.

📎 How to read this page

What it means gives the precise meaning. Operator translation gives the version you might hear in a real ecommerce meeting. In real life shows an illustrative example. Watch out and the confusion boxes show where a familiar term can mislead.

📈 Read the relationships first

These combinations are diagnostic hypotheses, not proof of causality. Compare the same period and scope, then investigate the mechanism.

Primary metric ↑ + guardrail worsens

Usually means: The apparent win may carry an unacceptable trade-off. Check next: prespecified thresholds and actual economic impact.

Early lift ↑ + later lift fades

Usually means: Novelty, timing or chance may explain the pattern. Check next: planned duration, repeat exposure and uncertainty.

Strong result + allocation mismatch

Usually means: Measurement or assignment may invalidate inference. Check next: Diagnose SRM before celebrating the effect.

Diagnose the experience

01 · 🔵 Operations

Conversion friction

🧠 What it means
Effort, uncertainty or obstacles that make completing a chosen action harder. Friction can be practical or psychological, such as errors, unclear delivery terms or unnecessary steps; removing it must still preserve necessary checks.

💬 OPERATOR TRANSLATION

“Every little “hang on” between wanting the product and paying for it.”

🛍️ In real life
A surprise delivery fee and unclear returns policy make a shopper pause before completing the order.

Shopper drags checkout cart over small obstacles

🔗 Related: CTA · Heatmap · Session replay · ↑ all terms

02 · 🔵 Operations

CTA = Call to action

🧠 What it means
A prompt asking the user to take a specific next step, usually implemented as text, a link or a button. Its wording and placement should match the actual action and what happens next.

💬 OPERATOR TRANSLATION

“The button should say what it does, not audition for a motivational poster.”

🛍️ In real life
“Add to cart” opens a cart containing the selected product, rather than taking payment immediately.

Shopper puzzles over poetic button while plain button waits

🔗 Related: Conversion friction · Heatmap · Session replay · ↑ all terms

03 · 🔵 Operations

Heatmap

🧠 What it means
An aggregated visual representation of recorded interaction, attention proxies or scrolling on a page. The exact measurement depends on the tool; clicks and mouse movement are not direct readings of shopper intent.

💬 OPERATOR TRANSLATION

“A colorful picture of activity, not a mind-reading certificate.”

🛍️ In real life
A click heatmap shows repeated clicks on a product image that visitors expect to enlarge.

⚠️ Watch out

Device mix and tracking coverage can skew the picture; protect personal data.

Analyst stares at colorful map while shopper asks simple question

🔗 Related: Conversion friction · CTA · Session replay · ↑ all terms

04 · 🔵 Operations

Session replay

🧠 What it means
A reconstruction of recorded interactions during a visit, using captured page and event data. It can help diagnose usability problems, but it is not necessarily a literal video and needs careful consent, masking and access controls.

💬 OPERATOR TRANSLATION

“Watching someone fight the form without inviting their card details to the screening.”

🛍️ In real life
A masked replay shows a shopper repeatedly correcting a rejected postcode field.

Analyst watches shopper wrestle blank form while sensitive details hidden

🔗 Related: Conversion friction · CTA · Heatmap · ↑ all terms

05 · 🔵 Operations

Usability testing

🧠 What it means
Observing representative users attempting defined tasks to identify difficulties in understanding or using an experience. It provides diagnostic evidence about problems; a handful of participants does not estimate population-wide conversion lift.

💬 OPERATOR TRANSLATION

“Ask someone to buy the thing and resist explaining the button every ten seconds.”

🛍️ In real life
Participants try to find a suitable size and complete checkout while researchers observe where they hesitate.

Researcher bites tongue while shopper searches for button

🔗 Related: Conversion friction · CTA · Heatmap · ↑ all terms

06 · 🟢 Core

CRO = Conversion rate optimization

🧠 What it means
A disciplined process of improving the share of eligible visitors who complete a chosen action, while preserving business outcomes. It combines diagnosis and testing; maximizing purchases by destroying contribution is not a useful victory.

💬 OPERATOR TRANSLATION

“Make it easier to buy, without paying customers to make the chart look good.”

🛍️ In real life
A store tests clearer delivery information and checks completed purchases alongside contribution and cancellations.

⚠️ Watch out

Define the conversion and denominator; traffic mix can move the observed rate without a site improvement.

Manager cheers checkout while accountant checks empty till

🔗 Related: Conversion friction · CTA · Heatmap · ↑ all terms

07 · 🔵 Operations

Test hypothesis

🧠 What it means
A falsifiable prediction connecting a proposed change, an expected mechanism and a measurable outcome for a specified audience. It guides design and analysis; “make the button blue” is a change, not a complete hypothesis.

💬 OPERATOR TRANSLATION

“The sentence we write before the result so hindsight doesn't get promoted to strategy.”

🛍️ In real life
For first-time visitors, clearer delivery dates should reduce uncertainty and increase completed purchases without harming contribution.

Researcher writes plan before opening results envelope

🔗 Related: Conversion friction · CTA · Heatmap · ↑ all terms

Design a comparison

08 · 🟢 Core

A/B testing = Split testing

🧠 What it means
Comparing two variants by assigning eligible experimental units to groups and measuring a predefined outcome. Random assignment and stable exposure help estimate the effect of the change, rather than comparing unrelated weeks.

💬 OPERATOR TRANSLATION

“Two versions get a fair audition instead of the boss declaring a favorite.”

🛍️ In real life
Eligible shoppers are randomly assigned to the existing delivery message or a clearer version during the same test.

Two storefront versions audition while boss puts opinion aside

🔗 Related: Control group · Randomization · MVT · ↑ all terms

09 · 🔵 Operations

Control group

🧠 What it means
The experimental group receiving the baseline or comparison experience. It provides the reference against which treatment outcomes are evaluated; it must be comparable through the assignment and measurement design.

💬 OPERATOR TRANSLATION

“The shoppers who keep yesterday's experience so tomorrow's claims have something to stand on.”

🛍️ In real life
One group sees the original product page while the treatment group sees the revised sizing guide.

Shopper receives original shop setup as researcher measures new one

🔗 Related: A/B testing · Randomization · MVT · ↑ all terms

10 · 🔵 Operations

Randomization

🧠 What it means
Assigning experimental units to conditions using a random mechanism. The unit might be a user, household or other entity; consistency is needed to avoid contamination, while randomization balances differences in expectation rather than guaranteeing identical groups.

💬 OPERATOR TRANSLATION

“A coin gets the job so our favorite customers don't all land in the new design.”

🛍️ In real life
A stable assignment sends each eligible user to one version throughout the test.

Researcher flips coin instead of handpicking shoppers

🔗 Related: A/B testing · Control group · MVT · ↑ all terms

11 · 🔵 Operations

MVT = Multivariate testing

🧠 What it means
An experiment varying several elements and testing combinations, often to estimate main effects and interactions. The design and number of combinations determine traffic needs; testing many pieces at once is not automatically more informative.

💬 OPERATOR TRANSLATION

“We changed three things, created eight combinations and discovered our traffic isn't a research university.”

🛍️ In real life
A full factorial test of two headlines, two images and two buttons has eight combinations.

⚠️ Watch out

Match the design to available sample size and the question; a simple A/B test may be more useful.

Researcher juggles eight tiny storefront combinations

🔗 Related: A/B testing · Control group · Randomization · ↑ all terms

Trust and interpret the result

12 · 🟢 Core

MDE = Minimum detectable effect

🧠 What it means
The effect size a test is planned to detect with specified power and significance settings. It helps calculate required sample size; it is not necessarily the smallest effect the business would care about.

💬 OPERATOR TRANSLATION

“The improvement our traffic budget can realistically hear, not the one our slide desperately wants.”

🛍️ In real life
The team designs a test for a 10% relative conversion increase; at 2% baseline that is 0.2 percentage points.

Small researcher uses big ear to hear faint checkout bell

🔗 Related: Statistical power · SRM · Statistical significance · ↑ all terms

13 · 🟢 Core

Statistical power

🧠 What it means
The probability that a specified test detects a stated true effect under its assumptions. It is a design quantity depending on effect size, variance, sample size and significance threshold, not a score added after a disappointing result.

💬 OPERATOR TRANSLATION

“Whether our experiment can hear the improvement we're asking it to find.”

🛍️ In real life
A team plans sufficient traffic for 80% power to detect its chosen effect at the prespecified threshold.

⚠️ Watch out

Low power makes an inconclusive result unsurprising; it does not prove no effect.

Researcher listens for tiny bell while crowd makes noise

🔗 Related: MDE · SRM · Statistical significance · ↑ all terms

14 · 🔵 Operations

SRM = Sample ratio mismatch

🧠 What it means
A statistically unexpected discrepancy between planned and observed group allocation. It can indicate assignment, eligibility or telemetry problems and requires diagnosis before trusting the experiment's outcome estimate.

💬 OPERATOR TRANSLATION

“The test promised equal groups and quietly seated half the guests somewhere else.”

🛍️ In real life
A planned 50/50 split yields 60/40 after enough observations to trigger the allocation check.

⚠️ Watch out

A small numerical imbalance is not automatically SRM; use the appropriate statistical check.

Researcher counts uneven queues behind equal signs

🔗 Related: MDE · Statistical power · Statistical significance · ↑ all terms

15 · 🟢 Core

Statistical significance

🧠 What it means
A result meeting a predefined statistical criterion under a chosen testing model. In conventional testing, it concerns compatibility with a null hypothesis; it does not establish commercial importance or the probability that the hypothesis is true.

💬 OPERATOR TRANSLATION

“The result passed the statistical gate; it still has to impress the business.”

🛍️ In real life
A prespecified test reports p=0.03 against a 0.05 threshold, after the planned analysis and validity checks.

⚠️ Watch out

Repeated peeking and many comparisons need appropriate methods; p<0.05 is not “95% certain it works.”

Researcher opens statistical gate while accountant holds another gate

🔗 Related: MDE · Statistical power · SRM · ↑ all terms

16 · 🔵 Operations

Confidence interval

🧠 What it means
A range produced by a statistical method with a stated long-run coverage rate under its assumptions. It communicates uncertainty around an estimate; a conventional 95% interval is not a 95% probability statement about a fixed parameter.

💬 OPERATOR TRANSLATION

“The honest part of the result that says our best guess comes with elbow room.”

🛍️ In real life
An estimated purchase-rate change of +0.2 percentage points has an interval from −0.1 to +0.5 points.

💡 Why it matters

An interval crossing zero can still contain commercially meaningful gains and losses.

Analyst holds uncertain ruler between small loss and small gain

🔗 Related: MDE · Statistical power · SRM · ↑ all terms

17 · 🔵 Operations

Guardrail metric

🧠 What it means
A predefined outcome monitored to prevent an improvement in the primary metric from causing unacceptable harm elsewhere. A guardrail needs a threshold and decision rule, rather than simply appearing on a dashboard.

💬 OPERATOR TRANSLATION

“The rule that stops “more orders” from meaning “more refunds and less money.””

🛍️ In real life
A checkout test targets purchases while requiring no unacceptable increase in payment errors or refund rates.

Celebrating merchant stopped by rail protecting coin jar

🔗 Related: MDE · Statistical power · SRM · ↑ all terms

18 · 🔵 Operations

Novelty effect

🧠 What it means
A temporary behavioral response caused by an experience being new, rather than a durable benefit. Effects can fade or change as users learn the design, so time patterns and repeat exposure deserve attention.

💬 OPERATOR TRANSLATION

“The new button got attention; that doesn't mean it found a lasting career.”

🛍️ In real life
A redesigned navigation attracts unusual clicks on launch, then settles after repeat visitors become familiar.

Shopper admires shiny new button then walks past later

🔗 Related: MDE · Statistical power · SRM · ↑ all terms

🔀 A/B testing vs multivariate testing

An A/B comparison evaluates two experiences. A multivariate design varies several elements and combinations, usually with greater traffic demands.

🔀 Significance vs confidence interval vs business impact

A statistical criterion, an uncertainty range and commercial value are separate judgments. A tiny precise effect can still be unhelpful.

🔀 MDE vs power

MDE names the target effect size; power describes the detection probability under the design. Neither is a promise that a test will win.

🔀 Heatmap vs replay vs usability testing

Aggregate activity, a reconstructed visit and observed tasks offer different evidence. None by itself proves a conversion change.

🤔 Still confused?

Follow this thread: Conversion friction → A/B testing → MDE. That sequence moves from the basic object or relationship to the decisions and checks it supports.

Sources and scope

Primary documentation checked on 27 September 2026. Platform features and eligibility can change; examples and cartoon situations are illustrative.

Reply

Avatar

or to participate