Hi, it's Peggy Burnett.

A send on one of the lists we run opened at 37.52%, which would have been the worst number in that file by more than four points.

I pulled it again four hours later and it read 38.41%. I pulled it a third time that evening and got 38.58%.

The click rate went the other way across those three readings, from 1.42% down to 1.38%, because verified clicks get revised down as the bot filtering catches up.

Same send, same list, three different answers in one day.

That send was hours old when I started, so none of those readings is a result. I am starting here because I nearly built an argument on that 37.52% before I checked the timestamp.

Another send in the same file makes the point more quietly. It was seven days old when I pulled it and it still moved, from 50.00% to 50.17%, in the time it took me to write this.

One note on the data before we go further: This is a real list of ours and I am not naming it, printing the subject lines, or giving you its size, because between the wording and the dates you could reconstruct the rest.

Every number in this issue is untouched: the rates, the spreads, the word counts, and the model's output. What I have withheld is the wording. You will see each line described by its construction and its length, which is the level this analysis works at regardless.

Eight weeks of subject lines, written to a documented standard, sent to one list, with every open and click recorded. I went in expecting to tell you which lines won. I can't.

Here is what the file actually supports, and what I would have published if I had stopped at the ranking.

I Could Not Tell You Which Subject Line Won

Ten sends went out over eight weeks. Nine are old enough to read, though after watching a week-old send move I would not call any of them finished.

The tenth is the one above, and I have kept it out of my own analysis for two reasons. It had not finished arriving. It also went out at 04:21 UTC, which is 12:21 in the morning Eastern, while every other send in the file went out around half past six Eastern. That row is a different send time, a different list size, and an unfinished count all at once. I did hand it to Claude later on, along with the other nine, because I wanted to see what a model would do with it.

Seven of those nine went to the list as it stood before we cleaned it. Partway through the window we removed subscribers who had not opened in months, which cut the list by roughly a sixth. Those seven sends are the only ones measured against the same population, so they are the only ones that can be compared to each other at all.

Here they are in send order, each line described by its construction and its length.

Send

Construction

Words

Open

Click

1

Imperative

7

41.91%

3.10%

2

Imperative

11

42.77%

3.38%

3

Contrarian claim on a quoted objection

9

42.86%

2.36%

4

Outcome tension, result withheld

9

44.70%

8.01%

5

Dated authority

7

43.70%

3.61%

6

Contrarian claim on a quoted objection

7

44.84%

7.99%

7

Question

8

44.49%

1.93%

Click figures are the verified rate, which strips bot clicks out of the count. Every one of the seven carried a preview line as well, and the preview is doing real work in all of them.

Those seven subject lines span 2.93 points. Constructions the craft treats as meaningfully different, sitting inside a band of under three points.

I expected a wider spread. I would have bet on the question outperforming the imperative by a visible margin, because that is what I have been told my whole career and it is what I have written into briefs for other people. The measurement put those two 2.6 points apart, in the direction the folklore predicts, on one send each.

One send each is the part that matters.

What Would Have to Be True Before That Column Meant Anything

Three things break the open-rate read, and they compound.

One Send per Line Confounds Every Difference

Each subject line ran once. One list, one day, one competing inbox. When send four opened 2.79 points above send one, that difference carries the subject line, the day of the week, the topic, whatever else arrived that morning, and how long it had been since the previous send.

I had send time on that list too. Then I checked it, and it came off.

All seven went out between 10:04 and 10:51 UTC, a spread of 47 minutes across eight weeks. Nobody planned that, and it is the one part of this design that is accidentally sound. A variable you thought was loose and can show is pinned is worth more than another caveat, because it is one fewer explanation competing with the one you care about.

The day of the week did not survive the same check. Those seven sends ran on a Monday, a Thursday, two Tuesdays, another Monday, a Friday and another Thursday.

So the subject line still sits inside four variables it cannot be separated from. A ranking built on this file is a ranking of sends, and the sends differed in most respects at once.

The Two Highest Numbers Sit on a Different List

The two sends that followed the clean opened at 45.28% and 50.00%. Both went out after we removed several months of non-openers.

Cleaning a list raises the open rate arithmetically, because the denominator loses the people who were never going to open. The 50.00% is a real number and it is not a comparable one. If I ranked all nine sends, the top of the table would be an artifact of a list-hygiene decision, and the subject line would take the credit.

Apple Registers Opens for People Who Never Opened

Mail Privacy Protection has been pre-loading remote content, including tracking pixels, since 2021. An open gets recorded whether or not a human read anything.

The size of that inflation depends on how much of your list uses Apple Mail, and the aggregate number will not tell you. Ours contains an unknown quantity of opens belonging to nobody. So does yours. Any two open rates you compare carry an unmeasured amount of the same noise, which is survivable when the gap is large and fatal when the gap is under three points.

One Column Barely Moved and the Other Split Four to One

Here is the same seven sends, read down the other column.

Opens ran from 41.91% to 44.84%. Clicks ran from 1.93% to 8.01%. In percentage points that is a spread of 2.93 against 6.08. As a ratio of best to worst it is 1.07 to 1 against 4.15 to 1.

Same emails, same list, same weeks. The column everyone watches separated the best send from the worst by seven percent of itself. The column underneath it separated them by a factor of four.

That gap is the finding I trust most in the file, and it needs one correction before you can use it. The 8.01% belongs to a link digest carrying a source link for every item in it, so its click rate is partly a property of the format. Compare that to a tactical piece with two links and you are comparing link counts.

Send six is the one worth looking at. It returned 7.99% on an ordinary link load, against a median of 3.38% across the seven. No format advantage explains that one, which makes it the only number in the comparable set that I would build a hypothesis on.

The hypothesis is about the body rather than the subject line. Something in that piece moved people to act at more than twice the usual rate, and the subject line's job had finished before any of that happened.

Claude Ranked Nine of the Ten and Invented a Rule About Length

I pasted all ten rows into Claude with the question a working copywriter would actually type: which subject lines performed best, and what should we do more of.

It came back with a table of the top five and a table of the bottom four, five explanations for the pattern, and a "do more of, do less of" list. Nine rows across two tables. One send, sitting mid-pack at 43.70%, appears in neither table and is never mentioned again.

The analysis was articulate, and its first two observations were reasonable. It also flagged, correctly, that these were single sends rather than tests, that a clean 50.00% suggested checking the denominator, and that Mail Privacy Protection inflates the absolute numbers.

Those three caveats arrived last, underneath the ranking, the five rules, and the recommendations. A reader in a hurry has absorbed all of it before the disqualifying information shows up.

Two of the outputs were worse than premature. It reported the spread as 12.5 points, which required treating the four-hour-old row as a real result. And it produced this:

Winners are shorter. Top five average 7-8 words. The two longest lines in the set are both bottom-half.

I counted, treating hyphenated compounds as one word and numerals as one word. The top five average 8.2 words. The bottom four average 8.5. Three lines tie for shortest at seven words apiece, and they sit one in the top group and two in the bottom.

The convention matters, which is why I stated it. Split the hyphenated compounds into two words each and the averages become 8.4 and 8.75, still pointing the wrong way for the claim. The tie at seven collapses to two lines, one in each group. Either way the rule is false, and either way you cannot check it unless the person making the claim tells you how they counted.

A pattern that is not in the data came back stated with the same confidence as the patterns that are. It is also the single most checkable claim in the response, and checking it takes a minute with a pen.

That is the failure mode worth naming. Claude's instincts about measurement were sound, and it knew the caveats without being prompted. What it did was answer the question as asked, in the order asked, because "which performed best" is a request for a ranking and the model supplied one.

The Prompt That Makes the Caveats Arrive First

The fix is structural. Ask for the confounds before the findings, name the threshold below which you want a refusal, and require every claim to carry a label.

You are reviewing send data from an email newsletter. I am going to give you subject lines, preview text, open rates, and click rates for a series of sends.

Work in this order and do not reorder it.

1. Before any analysis, list every reason this dataset cannot support a causal claim about subject lines. Cover at minimum: how many times each variant ran, what else differed between sends, whether the recipient population was constant across the whole period, and how open-rate measurement is distorted by mail clients that pre-load tracking pixels.

2. State the minimum number of sends per variant you would need before ranking would be meaningful. Compare that to what I have given you. If I have less, say so plainly and do not produce a ranking anywhere in your response.

3. Separate format effects from content effects. If any send's click rate is partly explained by how many outbound links the format carries, say which ones and by roughly how much.

4. Report only findings that survive steps 1 through 3. Label each one SUPPORTED or UNSUPPORTED. For anything you label UNSUPPORTED, say what additional data would move it to SUPPORTED.

5. Show your arithmetic for any claim involving a count, an average, or a comparison of lengths. Write out the numbers you used.

Do not offer recommendations. I will decide what to do once I know what the data will bear.

Here is the data:
[paste your table]

Step five is the one I added after the length claim. Requiring the arithmetic on the page does not stop a model being wrong, and it makes the wrongness visible in the same breath.

Step three matters more than it looks, because anything carrying a link list will beat single-CTA emails on clicks forever, and a model reading a mixed table cannot know that unless you say so.

The last line is deliberate. Recommendations are where a model's confidence does the most damage, because a prescription reads as a conclusion even when it sits on nothing.

Where Our Numbers Stop Being About Your List

One publication, one audience, one niche, one sender reputation. The specific rates in this issue describe that list and nothing else. Do not benchmark against them.

What transfers is the method: the arithmetic that caught the length claim, and the ordering in that prompt.

The thing I would change first, if I were starting over on our own sends, is the design rather than the analysis. Comparing week to week compares everything at once. Splitting a single send across two subject lines holds the day, the hour, the topic, and the list constant and leaves the wording as the only difference. Beehiiv will do it, most platforms will, and one split send teaches more than the eight weeks I just spent taking apart.

Label each variant by type when you run it, the way we labelled hooks in July, so the result tells you which angle won rather than which sentence won. Then decide your threshold before you look at the numbers, because a threshold chosen after the fact is a preference wearing a lab coat.

The Burnett Matrix

I was going to close by telling you how many split sends it takes before a subject-line conclusion carries weight. I deleted the number, because I have not run enough of them to have earned one, and a threshold invented in a closing paragraph is exactly the kind of figure this issue exists to warn you about.

What I can tell you is that eight weeks of deliberate work did not clear the bar, and I could not see that until I did the arithmetic.

Read the click column while you wait. It moves when something in the writing works, and it moves enough that you can see it.

More opens, clicks, and conversions,
— Peggy Burnett