A/B testing splits a campaign's leads across copy variants at each step and shows which performs best, sometimes shifting traffic to the winner automatically. Buyers care because copy is the one lever they fully control once lists and infrastructure are fixed. Agencies also use test results to justify retainers to clients.
How it works
At each step the sequencer assigns a lead to a variant (randomly, evenly, or by weight), sends it, and attributes opens, clicks and replies back to that variant. "Auto-optimise" modes then stop sending the losing variants. The design questions are: how many variants, how leads are allocated, which metric decides, and when the decision is made.
| Tool | Variants per step | Allocation | Winner metric | Source |
|---|---|---|---|---|
| Instantly | Up to 26 (A–Z) | Balanced over campaign lifetime; daily-equal option | Auto-optimise on reply, click or open rate | help |
| Smartlead | Up to 10 on manual split (min 10% each) | Equal, manual %, or "AI distribution" | Open, click, reply or positive reply rate | help |
| EmailBison | Not stated | Step-level, subject and body | Auto-winner | homepage |
| Woodpecker | Up to 5 | Not stated | Not stated | pricing |
| lemlist | 2 only | 50/50 random | Manual choice | help |
| Salesforge | Not stated | Not stated | Not stated | Growth plan only, $80/mo (pricing) |
lemlist also supports A/B testing an entire sequence rather than a single step. That answers a gap in Instantly, where variant assignment is random at each step and Variant A in step 1 cannot be tied to Variant A in step 2.
Who does it well
Smartlead has the most thoughtful optimisation target: positive reply rate, using its AI reply categorisation. Instantly has the most variants and the simplest UX, but its auto-optimise accepts open rate as a target. Open rate is the least reliable metric in the stack: Apple Mail Privacy Protection pre-fetches pixels, and Instantly itself offers a global switch to disable open tracking for deliverability reasons. lemlist is the most honest about statistics. It tells users to wait for 50–100 leads per variation and leaves the decision to them, but caps the test at two variants.
Where implementations differ
- Statistics. No vendor among the leaders documents a significance test, confidence interval or minimum detectable effect. With cold reply rates in the low single digits (Reply benchmarks), 26 variants on a 2,000-lead campaign gives each variant about 77 leads and two or three replies. Any "winner" from that is mostly noise.
- Metric. Opens (broken), clicks (rare in cold email, and tracking links hurt deliverability, see Custom tracking domains), replies, positive replies. Nobody optimises on meetings booked or pipeline. That data sits in the CRM (CRM sync).
- Interaction with spintax. Random variation inside a variant (Spintax and message variants) adds noise that no tool separates out.
- Interaction with rotation. If variant A happens to go out from healthier mailboxes, it wins on deliverability, not copy. No tool controls for sending mailbox or recipient provider.
Risks and abuse
The main risk is false learning: teams "optimise" toward noise and change working copy. A subtler one: auto-optimise on opens rewards subject lines that trigger image pre-fetch or curiosity-bait, which can raise complaint rates. There is no meaningful abuse vector beyond that. Testing is benign.
In this market, A/B testing is marketing theatre. The vendor that says "this result is not significant yet; you need about 1,400 more sends" will look less exciting in a demo and be trusted far more by GTM engineers and serious Lead-gen agencies.
What this means for an entrant
- Ship multi-variant tests per step and whole-sequence tests. Both are expected.
- Default the winning metric to positive replies, with meetings booked as an option once CRM sync is in place. Hide open rate as a decision metric by default.
- Show significance, sample-size estimates and a "not enough data" state. It costs almost nothing to build and no leader does it.
- Stratify allocation by sending mailbox and recipient provider so deliverability differences do not masquerade as copy wins. This is a technical edge that is hard to retrofit.
- Feed test results into the AI copy loop (AI campaign builder): generate variants, test, keep the winner. That is where experimentation becomes a moat rather than a checkbox.