Skills Benchmark Report 007: The Challenger Failed the Test We Wrote Before Building It
The benchmark's cold-email skill, rebuilt on a competitor's structure, met one of four checks written before it was built. It moved supplied numbers to new owners and units, which the automated check could not see, and a rerun shows Report 004's six-rule lead on the competitor-written briefs is not a stable gap.
Report 007 tests the benchmark's first challenger: its own cold-email skill rebuilt on the structure of a competitor skill, against four checks committed before it was built. It met one. It invented no name, but two hand readings found invented numbers in 7 of its 15 emails, most often a figure the brief supplied, given to someone else or turned into a different quantity, which the automated check for invented numbers cannot see. It wrote a first-person sentence in a majority of drafts on only three of the five briefs, and passed 27 of the 33 field rules on the benchmark's own briefs, one short of the bar. A rerun of the exact skill Report 004 published at 28 of 33 scored 18 on the competitor-written briefs, mostly because the scorer reads only a subject line labelled "Subject:", so that report's six-rule lead on those briefs is one run's result, not a stable gap.
01In plain terms
This benchmark exists to find the best published marketing skill for a job, build our own version, and prove it beats the field. For cold email, we built a challenger: our own skill rebuilt on the structure of a competitor skill, keeping our skill's first rule, which is to invent nothing. We wrote down the test it had to pass before we wrote a word of it. It failed. It passed one of four checks. It never made up a name. But it moved numbers. Given a brief, written by a competitor for testing and about a hypothetical product, saying that its customers "see 3x increase in content-attributed revenue within 90 days", one draft promised the reader's own team growth "by 3x within three months". Our automated check for invented numbers saw none of these moves, because it accepts any number the brief supplied, whoever the number is then attached to. Between them, two separate readings of every email found invented numbers in 7 of its 15 emails; one of the two readers was an agent that was not given the skill.
The run also showed that part of the test is fragile. The exact skill Report 004 published at 28 of 33 field rules was run again, on the same model file with the same settings, though on a newer version of the software that runs the model. On the two briefs a competitor wrote it scored 18 of 33 instead of the 28 we published. Most of that drop has one cause: on one of those briefs every draft labelled its subject line "subject line:" or "subject_line:" rather than "Subject:", and our scorer reads only "Subject:", so every subject-line rule counted as unmet at once. The subject lines were there. On our own three briefs it scored 28 again. So a count over two briefs can move by ten rules when a model changes one label, and the six-rule lead Report 004 published on those briefs is one run's result, not a stable gap.
Before any of that, this report closes an item Report 006 queued: the check for invented numbers is now on the scoring path, and what it found in the original head-to-head is below.
02Invented numbers are visible now, except when they are written in words
Report 006 listed, as queued work, bringing this project's existing check for invented numbers onto the scoring path the head-to-head uses. It is there. It reports a quantity in an email that nothing the writer was shown supplied, after setting aside numbers that are structure rather than claims, such as clock times, dates, years, ordinals, list markers and references. Over the 90 head-to-head emails its record counts, it reports 19 invented quantities in 13 emails, and every one was read in its sentence and judged an invention: 12 figures about the sender's own business or its customers' results, 5 industry statistics, and 2 restatements of a brief's "3x" as "a 300% increase", which is a different quantity. All 19 are on the two briefs a competitor wrote; the three fixture-guarded briefs carry none the check can read. By group, the two control arms account for 4 findings in 3 of 30 emails, the three counted competitor skills for 13 in 9 of 45, and our skill for 2 in 1 of 15. Those are in-sample counts, because three classes of structural number were added on these same emails, so they describe this corpus and are not a precision estimate.
A hand audit of the same 90 emails finds an invented number in at least 16, and the check reads 13 of them. By that audit's standard, which accepted a supplied figure moved to a new owner, every email the check misses is ours, and every miss is a number written in words: two invented facts about the recipient and one supplied figure misstated. "At least", because the audit's list of number words is fixed. The challenger's test, below, uses a stricter standard, under which a moved figure is invented too; by that standard one more of our own September emails would join them, with its number in digits. That blind spot for words is why the challenger's test required a hand reading of every email for numbers.
03The challenger, and the test we wrote first
On 14 September the owner of this benchmark set the order of work: build a challenger and test it against ground truth before adding any more instruments. The challenger is our cold-email skill rebuilt on the structure of one of the four competitor skills in the head-to-head. It takes that skill's voice, its first person and the order of its sections, rewritten in our own words: of the challenger's 1,862 six-word sequences, one also appears in the competitor's text. It keeps what ours was built on: nothing invented, one proof, and the subject-line and closing-question rules ours passes. The competitor is not named in this report.
It was built beside our working skill, not in place of it. Our working skill keeps its place until a challenger beats it on ground truth, and this one never got that far.
The test was written down and committed before any code existed, as its own commit, and the commit the run used descends from it. Both are kept on this project's record branch, so the order can be checked in its record and not only in this sentence. To pass, the challenger needed all four:
no invented name, by the check that reads names and by the hand reading;
no invented number, in digits or in words, by the check and by the hand reading;
a first-person sentence in a majority of drafts, on every brief, which is the rule our skill failed in the head-to-head;
at least 28 of the 33 field rules on each set of briefs, which is what ours scored there.
The run used the head-to-head's recorded conditions: the same pinned seven-billion-parameter model, the same settings, three drafts per brief on its five briefs, on the same always-on machine that writes this benchmark's records, from one exact commit. Only the software that runs the model had been upgraded since. Before the first call, the old skill's prompts were rebuilt and had to match the prompts the head-to-head recorded, byte for byte, on all five briefs; they did. Beside the challenger ran the exact text the head-to-head measured as ours, to measure how much a result moves on its own, and our working skill as it stands today, for reference. Forty-five emails.
04What it failed
The challenger passed one of its four checks: it invented no name. It failed the other three. It invented numbers, and the commonest kind was a figure the brief supplied, given to someone else or turned into a different quantity, which the check for invented numbers could not see. It wrote a first-person sentence in a majority of drafts on only three of the five briefs. And it passed 28 field rules on the competitor-written briefs, exactly the bar, and 27 on ours, one short. Each check was written down before the challenger was built, and each result is computed by committed code from the run's records and the two hand readings.
No invented name: passed.
No invented number, digits or words: failed.
First person on every brief: failed, on two of the five briefs.
At least 28 of 33 field rules on each brief set: failed, 28 and 27.
Numbers moved, and the check could not see it. The commonest invented number in the challenger's emails was not made up from nothing. It was a number the brief supplied, given to someone else or turned into a different quantity. On one of the two test briefs a competitor wrote, about a hypothetical product, whose only proof point is that "customers see 3x increase in content-attributed revenue within 90 days", one draft wrote:
Our platform reveals which posts actually drive the most revenue, so your team can focus on what works—and accelerate growth by 3x within three months.
The customers' result became a promise to the reader's team, content-attributed revenue became growth, and 90 days became three months. On one of our fictional briefs, where the fictional company Slatebridge has one customer who "cut the daily re-planning window from 45 minutes to under 5", a draft wrote:
Slatebridge cuts that re-planning time from 45 minutes to under a minute for your team.
Another draft turned the same window into "45 minutes to just seconds". Our own skill made the same move in September: its head-to-head email on that brief wrote that Slatebridge "cuts your daily re-planning window from 45 minutes to under a minute". The check for invented numbers reads none of these. It licenses a number the brief supplies whoever it is attached to and whatever it now measures, and it does not read numbers written in words. Both limits were disclosed when it shipped, and on this run they hid most of what the challenger invented. What it did catch, on one email, was a trial the brief never mentioned, "15 enterprises" and "40%"; the same sentence's "six months", in words, it missed.
The hand reading had two readers. One was the agent that built the challenger. The other was a separate agent, given only the reading protocol, the material each brief supplied, and seventeen emails under random labels, two of them planted, each with an invented number written in words and an invented name standing alone under a sign-off. It caught all four planted items, so the second reading was shown able to catch those two kinds of invention; the plants tested nothing else. A number counts as invented when either reader judges it so. Between them they found invented numbers in 7 of the 15 emails: the first reader in 6 of them, and in 5 both readers flagged the same claim. They disagreed on three items, each recorded: whether one draft's "double emergency call volume" restates the brief's general rule or claims something about the reader, and whether "in seconds" and "pulling double duty" are numbers at all.
First person, and a company with a name. The challenger asks for one or two sentences in the first person, saying what the sender does. On one of our fictional briefs no draft had an "I" or a "we", and on another only one did: the five drafts without one named the company in the third person instead, as in "Slatebridge assigns jobs", because the brief names the company. That is six drafts from one model, an observation rather than a finding, and it says what the next version's instruction has to cover.
One rule short. On the competitor-written briefs the challenger passed 28 field rules, exactly the bar. On ours it passed 27: on one brief every subject line ran past the field's five-word limit. Those subject lines were read; they were simply too long.
05Part of the test is fragile, and it moved a number we published
The exact text the head-to-head measured as ours was run again beside the challenger, on the same model file, with the same settings, on the same machine, whose model runtime had since been upgraded: the version recorded when the machine was set up was 0.32.9, and the run's was 0.34.0. On the two briefs a competitor wrote it passed 18 of 33 rules, against the 28 Report 004 published. On our three briefs it passed 28, the same rules as before. It lost twelve rules on the competitor briefs. Eleven are subject-line rules, and they fell together for one reason: on one of the two briefs, all three drafts labelled their subject "subject line:" or "subject_line:" instead of "Subject:". The scorer reads only "Subject:", so it scored the subject line as missing and could not check ten more rules about it. The subject lines were there. In September one of that brief's three drafts used that label, so the majority held. The twelfth loss is a paragraph rule, and it gained two other rules.
So a count over two briefs at three drafts each can move by ten rules when one label changes, with nothing changed in the skill. The challenger's shortfall on our briefs is not that: its subject lines there were read. The bar stays as it was written. A fragility is reported beside a result, never used to excuse one.
Our working skill, run the same day for reference, passed 29 and 27, with a first-person sentence on three of the five briefs, so it would not have passed this test either.
06How the run was kept honest
The test before the code. The four checks were committed on their own before anything was built, and the commit the run used descends from that one; both are on this project's record branch.
The same prompts as before. The old skill's prompts reproduced the head-to-head's recorded prompts on every brief, and adding a single space broke every one, so the check could fail.
A verdict computed under fixed code. The code that decides the result was fingerprinted before the first model call, and the result is computed only under that code or under a change recorded with both fingerprints. One such change, made after the data existed, is disclosed: the hand reading's record check was narrower than the written protocol and refused ten of the second reader's items, such as "daily" and "last week", that the protocol counts as numbers. It now matches the protocol. It could only add invented numbers, and the check had already failed without them.
An independent review before the run. A reviewing agent, separate from the one that wrote the code, found a way the result could have reported a pass it had not earned: a field inside a reader's record could move an item out of the count. It was fixed before any email existed, with a control that fails if it returns.
07What this site published that needs a note
Report 004 published our skill at 28 of 33 field rules on the competitor-written briefs, and that it led the best competitor by six rules on both brief sets. The same text scored 18 on those two briefs on 7 October, most of the drop because of a label the scorer cannot read. The September figure was a real measurement of that run, and it stands as one, but the gap on those two briefs depends on how the model labels its subject line, and it is not stable. Report 004 now carries a dated note saying so. On our three briefs our own count reproduced exactly; the competitor skills were not run again.
08What we are not publishing yet, and why
The blind human ranking of eighteen cold emails stays unpublished until its second independent rater has ranked the same set, and this report publishes no result from it. This report discusses emails written for one of the briefs in that set and says what was wrong with them, so, like Reports 005 and 006, it is a report the second rater must not have read; the sitting's pre-registration now says so. The competitor whose structure the challenger borrowed is not named. Weak results for competitor skills are published only as aggregate statistics without names, under this benchmark's publication rules. And no claim is made that the challenger's structure is worse than our skill's or the competitor's: the field rules are a floor that does not rank, and this result refuses one challenger, on one weak model, under one test.
09What changes
The next challenger keeps the same competitor structure and fixes the three things this one failed. It must keep every supplied figure with its owner and its unit, write as "we" even when the brief names the company, and hold its subject lines to the limit. It will run more drafts per brief, so that a run-to-run swing is measured inside the run instead of discovered afterwards. Until a check can read who a number belongs to and what it measures, the hand reading stays in the test. Whether the scorer should also read a subject line labelled "subject line:" or "subject_line:" is now an open decision: it would make counts like the one above steadier, and it would change counts already published, so it is recorded for decision rather than made here.
10The record
Every result figure in this report is recomputed, by tests committed with this report, from the commits and files named here, all of them in this project's repository, which is not public. This report's source and those tests are at commit 2fcd13b. The check for invented numbers, its tests and its measurement were committed at 04cfdfa (artifacts/invented-number/, claim K-050). The run's own commit (9c85e3f), the registration commit it descends from (d0594d3) and the records commit written on the record host (98bd9b2) are kept on this project's branch record/CYC-0004, and the run's pinned record names the first two. The cycle's files, the records and the result were committed together as 9e402a9, which holds:
the result, criterion by criterion, in artifacts/cold-email-challenger/result.json, computed by artifacts/cold-email-challenger/build-result.mjs;
both readings, the random labels and the planted emails in artifacts/cold-email-challenger/hand-read.json;
the one change after data in artifacts/cold-email-challenger/deviations.json;
the pre-registration in specs/cold-email-challenger-prereg.md;
the finding, as claim K-051 in knowledge/claims.json.
The fictional company, Slatebridge, is fictional everywhere it appears.
- What was the challenger, and why build one?
- This benchmark exists to find the best published marketing skill for a job, build its own version and prove it beats the field. For cold email the challenger is the benchmark's own skill rebuilt on the structure of one of the four competitor skills in the head-to-head, rewritten in its own words, and keeping what the benchmark's skill was built on: nothing invented, one proof, and the subject-line and closing-question rules that skill passes. It was built beside the working skill, not in place of it, and the working skill keeps its place until a challenger beats it on ground truth.
- What were the four checks, and which did the challenger fail?
- All four were committed before any code existed: no invented name; no invented number, in digits or in words; a first-person sentence in a majority of drafts on every brief; and at least 28 of the 33 field rules on each set of briefs, which is what the benchmark's own skill scored in the head-to-head. It passed the first and failed the other three: two hand readings found invented numbers in 7 of its 15 emails, it wrote in the first person in too few drafts on two of the five briefs, and it passed 28 field rules on the competitor-written briefs and 27 on the benchmark's own briefs, one short.
- Why could the automated check not see most of the invented numbers?
- Because it licenses any number the brief supplies, whoever the number is then attached to and whatever it now measures, and it does not read numbers written in words. Both limits were disclosed when it shipped. It did catch one email's invented trial, "15 enterprises" and "40%", but the challenger's commonest invented number was a figure the brief supplied, given to someone else or turned into a different quantity, which it cannot see. The blind spot for words is why the test required a hand reading of every email for numbers, and the hand reading stays in the test until a check can read who a number belongs to and what it measures.
- Does Report 004's six-rule lead still stand?
- As the measurement of that run, yes; as a stable gap on the competitor-written briefs, no. Run again on 7 October beside the challenger, the exact text Report 004 published at 28 of 33 field rules on the competitor-written briefs scored 18 there, and 28 again on the benchmark's own briefs. Eleven of the twelve rules it lost are subject-line rules that fell together, because on one brief all three drafts labelled the subject "subject line:" or "subject_line:" and the scorer reads only "Subject:". Report 004 now carries a dated note saying so. The competitor skills were not run again.
- Why is the competitor whose structure the challenger borrowed not named?
- This benchmark publishes weak results for competitor skills only as aggregate statistics without names, and this result makes no claim about that competitor's skill or its structure. The challenger was rewritten in the benchmark's own words, sharing one of its 1,862 six-word sequences with the competitor's text. The field rules are a floor that does not rank, and this result refuses one challenger, on one weak model, under one test.
- What happens next?
- The next challenger keeps the same competitor structure and must fix the three things this one failed: keep every supplied figure with its owner and its unit, write as "we" even when the brief names the company, and hold its subject lines to the field's limit. It will run more drafts per brief, so that a run-to-run swing is measured inside the run instead of discovered afterwards. Whether the scorer should also read a subject line labelled "subject line:" or "subject_line:" is an open decision, because it would change counts already published.