Interactive demo

A real verdict, without signing up

This is the results view, running on the sample experiment that ships in every account. Switch metrics, open a comparison, toggle an arm in the histogram — everything is live.

Sample data. A fictional checkout assistant, with figures chosen to be internally consistent — the intervals, p-values, and histogram all agree with each other, because you are going to check.

Verdict

“concise” wins — better answers, cheaper, and it converts

3 of 4 metrics moved significantly across 12,048 judged events over 24 days. The traffic split checks out, so the result stands.

Experiment
Checkout assistant prompt
Model
claude-haiku-4-5
Window
24 days
Split check
SRM p = 0.62 — passing

Pick a metric

Answer quality

The judge scores each sampled answer 1–5 against your rubric: is the question actually answered, and answered correctly.

Higher is betterPrimary metricScored by the judge

Control (control (detailed)) averaged 3.41 over 3,012 rows. Every rewrite below is compared against it.

Answer quality: each rewrite compared with the control
RewriteEffect vs controlp (raw)p (FDR)n per armOutcome
concise+9.3%+0.317 [0.261, 0.373]< 0.001< 0.0013,012Wins
bulleted+6.7%+0.23 [0.174, 0.286]< 0.001< 0.0013,012Wins
warm+0.44%+0.015 [−0.041, 0.071]0.5960.5963,012No difference

Test: Permutation test with bootstrap CI. Significance is judged on the FDR-corrected p-value, not the raw one — testing four metrics against three rewrites is twelve chances to get lucky.

The distribution, not just the average

Every judged answer’s score, by arm. An average can move because a few answers got much better or because most got slightly better — those are different products, and only the shape tells you which one you have.

1
2
3
4
5

Judge score, 1–5

“concise” and “bulleted” push mass from 2–3 up into 4–5. “warm” tracks the control almost exactly — which is why it wins on conversion and not on quality.

What’s actually running

control (detailed)control

You are a helpful checkout assistant. Answer thoroughly, covering shipping, returns, and payment options where relevant…

25% of traffic · 3,012 exposures

concise

You are a checkout assistant. Answer in at most three sentences. Lead with the answer, then one line of context…

25% of traffic · 3,012 exposures

bulleted

You are a checkout assistant. Answer as a short bulleted list, one fact per bullet, most important first…

25% of traffic · 3,012 exposures

warm

You are a friendly checkout assistant. Acknowledge the concern before answering, and close by offering further help…

25% of traffic · 3,012 exposures

This experiment is waiting in your account

Every new account opens on this same read-only sample, so you can look around before sending a single event. Two lines of SDK and your own experiment sits beside it.

Free plan · no credit card required

Want it walked through instead? Book 30 minutes — or read how the statistics work. Questions go to [email protected].