# A/B testing


An A/B test sends two to ten versions of the same message to a slice of the audience, measures
click-through per version, and then sends the winner to everyone who was deliberately held back.
Assignment is deterministic, the holdback never overlaps the test group, and promotion can be
manual or automatic.

All examples use `https://app.openpush.ai`, the OpenPush API base URL.

## When to use this page

Use an A/B test when you have a real question about copy: does urgency beat curiosity, does naming
the reward beat teasing it, does a shorter body get more taps. Use it on audiences large enough
that a difference could be visible — a few hundred devices per arm at minimum.

Do not use it as a way to send different copy to different segments. That is just two messages.

## Prerequisites

- The app's REST API key.
- An audience big enough to split. With the default 25% test share and two arms, a 10,000-device
  audience puts about 1,250 devices in each arm and 7,500 in the holdback.
- The receipt ladder wired up if you want click data — clicks arrive from the SDK via
  `/v1/ingest`, not from the provider.

## Creating a test

Add `variants` and, optionally, `ab` to an ordinary send.

```bash
curl -X POST https://app.openpush.ai/v1/apps/acme-app/messages \
  -H "X-OP-API-Key: $OP_REST_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "name": "Season 4 launch",
        "title": "Season 4 is live",
        "body": "Three new maps and a ranked reset.",
        "include_segments": ["seg_7b1c9e4a2f60"],
        "variants": [
          {"name": "Neutral",  "title": "Season 4 is live",
           "body": "Three new maps and a ranked reset."},
          {"name": "Deadline", "title": "Season 4 — ranked resets tonight",
           "body": "Claim your placement match before midnight."}
        ],
        "ab": {"test_pct": 20, "auto": {"enabled": true, "after_h": 6, "min_per_arm": 500}}
      }'
```

Arms are given ids server-side in order: `A`, `B`, `C`, and so on. You never supply an arm id.

### The `variants` array

| Arm field | Type | Notes |
|---|---|---|
| `name` | string | Truncated to 80 characters. Defaults to `Variant A`, `Variant B`, … |
| `title` | string | Falls back to the message's `title` when omitted |
| `body` | string | Falls back to the message's `body` when omitted |
| `languages` | object | Per-language copy for this arm, same shape and same base-language fallback as a normal send |

Two to ten arms. Anything else is `400 variants must contain between 2 and 10 arms`, and a
non-object entry is `400 every variant must be an object`.

Liquid in every arm — including each arm's per-language title and body — is compiled and validated
when you create the message, not when it sends. A syntax error in arm C's Portuguese body fails the
create call with the field named.

> **Only `title`, `body` and `languages` actually differ per arm.** An arm may carry `image_url`
> and `deep_link` and they will be stored on the message, but the payload delivered to a device
> uses the **message-level** image and deep link for every arm. If you need to test an image or a
> destination, test it as two separate messages.

### The `ab` object

| Field | Type | Default | Notes |
|---|---|---|---|
| `test_pct` | integer 1–100 | `25` | Share of the audience that gets a test arm. The rest is the holdback |
| `auto.enabled` | boolean | `false` | Whether to promote a winner automatically |
| `auto.after_h` | number | `24` | Hours to wait after message creation before auto-promoting |
| `auto.min_per_arm` | integer | `100` | Minimum sends per arm required before auto-promotion will fire |
| `winner` | string | — | Set by promotion; you do not normally supply it |
| `promoted_at` | number | — | Set by promotion |

A `test_pct` outside 1–100 is `400 ab.test_pct must be between 1 and 100`, and a non-integer value
is `400 ab.test_pct must be a whole percentage`.

Setting `test_pct: 100` means the whole audience is in the test and there is **no holdback** —
which is a legitimate way to split traffic evenly with nothing left to promote to. Attempting to
promote such a message returns `409 this experiment has no holdback to promote`.

## How devices are assigned

Assignment hashes the message id together with the subscription id and turns the result into a
uniform number between 0 and 1. Devices above the test share are the **holdback wave**; the rest are
spread evenly across the arms.

Three properties follow from that, and all three matter:

- **It is deterministic.** The same device on the same message always lands in the same place — on
  the first attempt, on a retry, on a quiet-hours release, and on a redrive after a process died
  mid-campaign. A device cannot receive arm A now and arm B on the retry.
- **It is device-disjoint.** The holdback is exactly the devices that were not in the test. Nobody
  gets the test copy *and* the winner copy for the same message.
- **It re-rolls per message.** Because the message id is part of the hash, a device that was in
  arm A last week is not systematically in arm A this week.

During the test send, holdback devices are simply not sent to — no delivery row is created for
them, and they consume nothing from their frequency cap. The message's reported `Audience` still
counts the full matched audience, so the funnel tells you the truth about how many people the
message is ultimately for.

## Reading the results

The message report carries the experiment:

```bash
curl https://app.openpush.ai/v1/apps/acme-app/messages/msg_2f7ba0c41d93 \
  -H "X-OP-API-Key: $OP_REST_KEY"
```

```json
{
  "id": "msg_2f7ba0c41d93",
  "testing": true,
  "winner": null,
  "significance": 96.4,
  "variants": [
    {"id": "A", "name": "Neutral", "title": "Season 4 is live", "body": "…", "languages": {}},
    {"id": "B", "name": "Deadline", "title": "Season 4 — ranked resets tonight", "body": "…", "languages": {}}
  ],
  "ab": {"test_pct": 20, "winner": null, "promoted_at": null,
         "auto": {"enabled": true, "after_h": 6, "min_per_arm": 500}},
  "arm_stats": [
    {"id": "A", "name": "Neutral",  "wave": "test", "sent": 1043, "accepted": 1039,
     "clicked": 58,  "failed": 4, "ctr_value": 5.56},
    {"id": "B", "name": "Deadline", "wave": "test", "sent": 1051, "accepted": 1044,
     "clicked": 91,  "failed": 7, "ctr_value": 8.66}
  ]
}
```

| Report field | Meaning |
|---|---|
| `testing` | `true` while the message has variants and no winner has been promoted |
| `winner` | The promoted arm id, or `null` |
| `significance` | Confidence percentage that the two test arms genuinely differ, or `null` |
| `variants` | The arm definitions as stored |
| `ab` | The experiment configuration, including `winner` and `promoted_at` once promoted |
| `arm_stats` | One row per arm per wave |

Each `arm_stats` row carries `id`, `name`, `wave` (`test` or `winner`), `sent`, `accepted`
(provider-accepted), `clicked`, `failed`, and `ctr_value` — clicks as a percentage of sends,
rounded to two decimals.

### Significance

`significance` is a two-proportion comparison of click-through between the arms, expressed as a
confidence percentage. It is computed **only** when there are exactly two test arms and both have
sent at least one message; otherwise it is `null`.

Treat it as a guardrail, not a verdict. A reading of 96.4 means the difference between 5.56% and
8.66% is unlikely to be noise at these sample sizes. A reading of 61 means you are looking at
noise, whatever the bar chart suggests. With three or more arms, no significance figure is
produced at all and you are on your own.

Two honest caveats: clicks only exist if your app posts click receipts, and the comparison uses
clicks over *sends*, not over delivered-and-displayed notifications.

## Promoting a winner

Promotion sends the winning arm's copy to the holdback wave.

```bash
curl -X POST https://app.openpush.ai/v1/apps/acme-app/messages/msg_2f7ba0c41d93/promote \
  -H "X-OP-API-Key: $OP_REST_KEY" \
  -H "Content-Type: application/json" \
  -d '{"winner": "B"}'
```

The arm id is case-insensitive. The response is the message's delivery report after the winner
wave has been queued, in the same shape as a send response.

| Error | Meaning |
|---|---|
| `404 unknown message` | No such message on this app |
| `409 winner must name an existing A/B arm` | The id is not one of this message's arms |
| `409 a winner has already been promoted` | Promotion is one-shot |
| `409 this experiment has no holdback to promote` | `test_pct` was 100 |

Promotion re-resolves the audience and then excludes every device that already has a delivery row
for this message. That means the test group is never double-sent, and devices that joined the
segment since the test went out will receive the winner. If that matters to your measurement, note
the timing.

Quiet hours and frequency caps apply to the winner wave exactly as they do to any send.

## Automatic promotion

Set `ab.auto.enabled` and the scheduler will promote for you:

```json
{"ab": {"test_pct": 20, "auto": {"enabled": true, "after_h": 6, "min_per_arm": 500}}}
```

The rules are strict, and all must hold:

1. At least `after_h` hours have passed **since the message was created** — not since the last
   click arrived.
2. Every test arm has sent at least `min_per_arm` messages.
3. No winner has been promoted yet.

When they hold, the arm with the highest click-through rate wins, resolved deterministically if two
arms tie. Note what is *not* in that list: **auto-promotion does not consult `significance`.** If
two arms are within noise of each other after six hours, it will still promote one. Set
`min_per_arm` high enough that the CTR comparison means something, or promote by hand.

If the conditions are never met — the audience was too small to reach `min_per_arm` — the message
simply stays in testing and the holdback never receives anything. Check on it.

## Limits

- **2 to 10 arms.**
- **`test_pct` is 1–100**, whole numbers only.
- **Only title, body and per-language copy vary per arm.** Image, deep link, custom data, action
  buttons, TTL, priority and collapse key are message-level and identical across arms.
- **Promotion is one-shot.** There is no un-promote and no second winner.
- **Significance is two-arm only**, and is not consulted by auto-promotion.
- **The measured metric is always click-through.** You cannot optimise for anything else, because
  nothing else is measured — there are no conversions, goals or custom outcomes on a message.
- **There are no holdout groups.** The holdback wave here is part of the experiment and receives
  the winner; it is not a permanently-suppressed control group for measuring the lift of pushing at
  all.
- Retrying a create-message call creates a **second, independent experiment** with a different
  message id and therefore a different assignment. There is no idempotency on message creation.

## FAQ

**Can I test send time instead of copy?**
Not as an A/B arm. Arms differ only in copy. Send two messages at different times to different
halves of an audience if you need that.

**What is the difference between the holdback and a holdout group?**
The holdback is the part of the audience saved for the winner — it is *going* to receive a message.
A holdout group in the classical sense is a control cohort that is never messaged so you can
measure the lift of messaging at all. OpenPush does not have the latter.

**Do holdback devices count against their frequency cap during the test?**
No. They receive nothing during the test wave, so nothing is consumed. The winner wave is capped
normally.

**What happens if I promote while the test wave is still sending?**
Devices that already have a delivery row are excluded from the winner wave, so the test group is
safe. But your arm statistics are still moving, so you are deciding on partial data.

**Can I use best-hour delivery with an A/B test?**
Yes. The audience is resolved and arms are assigned first; per-device timing is applied afterward.
Be aware that spreading a test over 24 hours means your arm statistics fill in slowly, and
`auto.after_h` counts from message creation regardless.

**Why is `arm_stats` empty right after I create the test?**
Rows appear per arm per wave once deliveries exist. If it stays empty, check that the audience was
not entirely capped or held.

## Related

- [Sending messages](sending-messages.md) — the rest of the message body
- [Personalization](personalization.md) — Liquid inside arm copy
- [Segments](segments.md) — sizing the audience before you split it
- [Best-hour delivery](best-hour-delivery.md)
- [Messages API reference](../api-handbook/02-messages.md)
