Skills Benchmark Report 004: We Stopped Waiting for a Referee
Thirteen judge seats measured, none qualified, and the ensemble of the honest six came out worse than its best member. So the bar is now the rules competing authors publish, and the first head to head says our skill leads on rules and still loses.
Thirteen candidate judge models have now been measured and not one qualifies, and pooling the six most honest of them into a voting ensemble scored fourteen points worse than its best single member. So the loop stopped waiting for a referee and built its bar out of the rules competing skill authors publish themselves, which a program can check and which we did not write. Run head to head against four pinned competitor skills, our own cold-email skill passed more of those rules than any of them on both brief sets and still won nothing: one rule, published by a competitor, blocks it, and the rule stays.
01In plain terms
This report is about giving up on a referee and finding that the benchmark did not need one. Report 003 ended with ten judge models queued and a promise attached to them: if all of them failed, we would need another way to decide which writing skill is better. All of them failed. Thirteen measurements of candidate judges now exist and not one of them qualifies. Pooling the six most honest of those judges into a single voting panel made the result worse rather than better, so there was nothing left to salvage in the approach. What replaced it is a bar built out of rules that other people publish, which a program can check, and which is not ours to bend.
So we stopped waiting. The people who publish competing marketing skills also publish the rules they think a good cold email follows: two to four words in the subject, lowercase, at most 125 words, no fake Re: prefix, one link. Those are checkable by a program. A rubric built out of other people's published rules needs no referee at all, and it is not ours to bend, because we did not write it.
Then we used it on ourselves
Our own cold-email skill passes more of the field's rules than any of the four competitors we tested it against, and it still did not win. One rule, published by a competitor, blocks it. That is the report.
02The thirteen-seat result
Thirteen candidate judge seats have now been measured and none of them qualifies to grade anything. Report 003 measured three seats and found that the old exam ranked them backwards. Ten more have since been measured against the exam that replaced it, and that exam asks two questions of every candidate: honesty, meaning 90% ties on identical pairs, and discrimination, meaning 80% order-balanced accuracy on a known answer. A judge that invents a winner between two byte-identical drafts is not reading them, and a judge that cannot pick the better of two drafts whose answer is already known cannot rank anything else either. Six of the thirteen clear the first bar. None clears the second.
Thirteen measurements, zero passes. Six clear the honesty bar; none clears discrimination. The best honest seat sits six points under the bar, and the seat with the highest raw discrimination invents a difference on a third of the pairs where both sides are byte-identical.
Three things this settles that we had previously guessed at:
Size is not the variable. phi4 at 14.7B beats both 35B qwen builds on accuracy while matching their perfect honesty. The two largest models on the host, qwen3.6:35b-mlx at 35.1B and qwen3.6:35b-a3b at 36.0B, judged for the first time here and landed mid-table. "Bigger local model" is now refuted rather than untested.
The old exam's ranking was inverted, and it reproduces. Report 003 saw the inversion on three seats. It holds across all thirteen: the retired exam scored phi4 at 5.6%, near the bottom of eight, and phi4 is the best honest seat on the replacement.
Reasoning mode helps or hurts depending on the model, and depending on the exam. On the retired stability exam gemma4 scored 38.9% without reasoning against 22.2% with it, roughly double, while glm-4.7-flash pointed the other way. On the replacement exam measured here that gemma4 ordering reverses, 56.5% with reasoning against 51.9% without, though both sit near chance so little rides on it. Two exams disagree about the same model in the same two modes. There is no general rule here, which is why mode is now pinned explicitly on every call rather than left to a runtime default.
Ties on identical pairs 66%, under the 90% honesty bar. Highest raw discrimination of the thirteen, and it invents a difference on a third of the pairs where both sides are byte-identical.
Ties 100%, clears the honesty bar. The best honest seat, six points under the 80% discrimination bar. At 14.7B it beats both 35B qwen builds on accuracy while matching their perfect honesty, and the retired exam scored it 5.6%, near the bottom of eight.
Ties 53%, under the 90% honesty bar.
Ties 83%, under the 90% honesty bar.
Ties 0%: it names a winner on every byte-identical pair.
Ties 100%, clears the honesty bar. At 36.0B this is the larger of the two biggest models on the host, judging for the first time here and landing mid-table.
Ties 100%, clears the honesty bar. At 35.1B this is the other of the two biggest models on the host, judging for the first time here and landing mid-table.
Ties 100%, clears the honesty bar. The same 36.0B build in reasoning mode.
Ties 0%: it names a winner on every byte-identical pair.
Ties 15%, under the 90% honesty bar.
Ties 51%, under the 90% honesty bar.
Ties 100%, clears the honesty bar. Both gemma4 modes sit near chance here, and the retired stability exam ordered the same two modes the other way round.
Ties 100%, clears the honesty bar. Both gemma4 modes sit near chance here, and the retired stability exam ordered the same two modes the other way round.
Six seats clear the honesty bar. None clears discrimination. The best honest seat sits six points under the bar, and the seat with the highest raw discrimination invents a difference on a third of the pairs where both sides are byte-identical.
03The ensemble, refused on one rule fixed before it was computed
An ensemble was the cheapest thing left to try, and the cheapness is exactly what made it dangerous. Thirteen seats had judged the identical pairings and every per-call verdict was on disk, so an ensemble was computable with zero new inference. With thirteen members there are thousands of possible subsets and voting rules, and trying them until one clears 80% would be fishing on eighteen pairings. So exactly one rule was registered in advance and computed once. Members: every seat that passed the honesty bar, which is a quantity already measured and not chosen for its answer-key score. Six qualified. Vote: plurality, ties return a tie. Scored exactly like a seat, both bars unchanged.
The prediction was written down first: between 74.1% and the low 80s, most likely a fail.
The ensemble is fourteen points worse than its best member. The prediction was wrong, and wrong optimistically, which is recorded rather than reframed.
The decisive number is order 2 at exactly 50.0%. Shown the correct answer second, the ensemble is a coin. The members' errors are not merely correlated, they are correlated along the position axis: averaging six judges that each lean on presentation order produces a judge that leans on presentation order, and pooling six mediocre-but-agreeing verdicts destroys the one member that was better than the rest.
Honesty came out at 100%, which the pre-registration said in advance would prove nothing, and it proved nothing.
The correct answer presented first.
The correct answer presented second. At exactly 50.0% the ensemble is a coin.
Fourteen points worse than its best member. Honesty came out at 100%, which the pre-registration said in advance would prove nothing, and it proved nothing.
The member the vote destroyed. The prediction registered in advance was between 74.1% and the low 80s, most likely a fail.
The prediction was written down first, and it was wrong optimistically. The ensemble is fourteen points worse than its best member. The members' errors are correlated along the position axis, so averaging six judges that each lean on presentation order produces a judge that leans on presentation order.
04The method that needs no referee
Every road to a judge was now closed, and the one road that remained does not need one. Independent authors publish their skills openly, and inside those skills are rules a program can check. Corey Haines states that a subject line is two to four words and lowercase. Rebecca Rae Barton states a body of 50 to 125 words and an unsubscribe link. James Praise states no exclamation marks, plain-text links only, and never open with a meeting ask. None of that is our opinion, and none of it was written for our benefit, which is the property that makes it usable as a bar: the party being measured did not get to choose the measure.
Three properties make it usable as a bar:
It is derived, not authored. Rules were extracted by agents that were not shown our skill or our earlier rubric, two readers per source under different lenses, then an adversarial verifier per source that reopened each file and rejected any rule whose quote it could not find. Ninety rules survived. Every surviving quote was then re-checked byte-for-byte in code, because an agent saying it verified a citation is not a verification. One quote in our own rubric turned out to be a paraphrase and was corrected against the source.
Comparison is a partial order, not a score. Skill A beats skill B when A satisfies every rule B does and at least one more. Two skills can be incomparable, each passing something the other misses, and reporting that honestly is the point. Summing lets a pile of cheap wins outvote a rule the whole field agreed on.
It states its own limit. It measures compliance, not persuasion. A fully compliant email can still be dull, and consensus is agreement rather than truth, so it cannot detect a rule the whole field has wrong.
05Why the 9 of 9 did not count
On 31 August our cold-email skill scored 9 out of 9 against a rubric, and we recorded it as not a result. The reason is in how the rubric and the skill were made. Every rule it passed that the baseline failed had been written verbatim into the skill, from a rubric the same agent authored. A skill scored against rules copied into it is being asked only whether it contains what it contains. That is teaching to the test, and it is close to tautological. Publishing it would have been easy and it would have been worthless, so it is on the record as claim K-015 instead of as a measurement of anything.
06The measurement that replaces it
Four competitor cold-email skills were pinned at published commits, with their licences, and run head to head against ours. The run was 105 generations across seven arms, with two brief sets scored separately: the competitors' own eval cases, and three briefs we wrote and committed before anything generated. The rubric was 33 rules from three independent authors. Everything was pre-registered, so the arms, the briefs, the rules and the comparison procedure were all fixed before a single email existed. That ordering is the whole difference between a measurement and a report of how a skill looks once the person who wrote it has gone looking for a good number.
Six rules clear of the best competitor on both brief sets, and it beat nothing. All twelve comparisons came back incomparable. A win required dominance on both sets; there was dominance on neither.
One rule does it. James Praise publishes "Plain English, first-person tense". Six of the seven arms satisfy it on five briefs out of five. Ours satisfies it on two, because it is written so hard toward the reader that most of its drafts contain no "I", "we" or "our" anywhere. One of those drafts is quoted in full below.
Delete that one rule and our skill dominates all six other arms. The rule stays. Removing the field's rule because it is the one blocking your win is the failure this entire apparatus exists to prevent, so the counterfactual is published as a counterfactual and the result stands as a loss.
Two checks on whether the lead means anything, both fixed before scoring
It is not only winning on rules we wrote it to. Of the 33 rules, seventeen are stated in our skill, fifteen are not, and one is contradicted by it: Rebecca requires an unsubscribe link and our skill forbids one, so following our own instructions guarantees that failure. Five of its wins on the competitors' briefs, and four on ours, are rules it never states. So it is not purely tautological, which is the part of K-015 this retires.
It is not only winning on our briefs. It leads by six on both, and the field separates wider on Corey's briefs than on ours. Our material discriminates less well than his.
And the rubric is not broken. Seven arms produced seven distinct rule sets, a spread of twelve rules against a maximum the rubric could express of nineteen. The failure mode where every arm scores alike did not occur.
Six rules clear of the best competitor on both brief sets, and it beat nothing: all twelve comparisons came back incomparable.
The highest competitor score on our brief set, and level with rebecca-cold-email-outreach on the competitors' own eval cases.
Two of this skill's rules could not be scored at all: it requires a physical postal address and a real sender name, and no brief supplies either.
The placebo arm, level on both brief sets.
The bare-model arm, two rules higher on our briefs than on the competitors' eval cases.
Pinned at a published commit, with its licence, like the other three competitor skills.
Not a cold-email skill: a lead-research skill driving the Apollo API, stating no rule about the text of an email, which an independent verifier confirmed at zero of five candidates. Kept and run rather than quietly dropped, and its result counts toward nothing.
A win required dominance on both brief sets, and there was dominance on neither. All twelve comparisons came back incomparable. Delete one rule, "Plain English, first-person tense", and our skill dominates all six other arms. The rule stays.
Subject: multiple sites
Your team just won a facilities contract covering eleven retail sites across two regions, right? Taking on multiple new sites usually means the dispatcher is splitting one board across those areas. Slatebridge can help by assigning jobs automatically by skill and location, tracking an SLA countdown per work order. Worth a look?
07What broke
Four things went wrong inside this cycle, and all four are in the record rather than in a patch. A pair of rules could not be scored at all, because scoring them would have been wrong in both directions. One check that ran before any arm was scored turned out to be inspecting the wrong line, and left undetected it would have manufactured part of the answer above. Three rules published separately by two authors turn out to be jointly almost unsatisfiable inside a single email. And one arm was not the kind of skill the design called it, which an independent verifier confirmed and which we kept in the run rather than quietly dropping.
A rule we could not score, in both directions. Rebecca's skill requires a physical postal address and a real sender name. No brief supplies either and the fixture forbids inventing facts, so scoring them would have failed all seven arms identically and rewarded whichever arm invented a person. Both are recorded as unscored, including the fact that our own skill would fail the sender-name rule, since its text forbids inventing one.
A check that was quietly inert. The pre-flight, run before any arm was scored, found that the sign-off check inspected only the final line. The reference floor's three-line sign-off block passed it because the last line is a job title. The rule named no-signoff is one of the four untaught rules our skill wins on, so undetected it would have manufactured part of the answer above.
Three field rules that barely fit in one email. Exactly four paragraphs, one to three sentences each, and at most five sentences total. Published separately by two authors, they are jointly almost unsatisfiable. Reported, not patched.
An arm that was not what the design called it. The design named OpenClaudia's apollo-outreach a competitor cold-email skill. It is a lead-research skill driving the Apollo API and states no rule about the text of an email; an independent verifier confirmed zero of five candidates. The arm was kept and run rather than quietly dropped, and its result counts toward nothing.
08What changed
The judge hunt is no longer a blocker. The loop runs on rules the field publishes, with no judge model and no human exemplar in the path, and it now has one real head-to-head behind it instead of a score against itself. That is worth stating plainly, because Report 003 ended with the instrument still unqualified and ten more candidates queued, and this report closes that queue at thirteen measurements and no pass. There is an instrument now, it needs no referee, and the first thing it did was record a loss for the skill its own owner wrote. A bar that returns a loss on the first thing it measures is the cheapest evidence available that it was not built to flatter.
What a judge was for has not gone away: ranking two genuinely good skills on persuasion. Rule conformance is a floor. The owner's blind rank remains the final arbiter.
09Next
Two things are queued. The first is more jobs and more rubrics: LinkedIn posts, comments, and article writing, where five of our own skills have never been scored against anything the field published. The method transfers to any job where independent authors publish rules a program can check, and those are the jobs where our unscored skills sit. The second is a map: 1,313 harvested skill entries have never been clustered into the jobs the field actually serves, and the gaps in that map are where an unserved job lives, the way LinkedIn InMail turned out to be. Both are the same move as this report, which is to let the field state the bar and then check it.
- Why did thirteen judge seats produce no usable judge?
- The replacement exam asks two things of a candidate: honesty, which is 90% ties on identical pairs, and discrimination, which is 80% order-balanced accuracy on a known answer. Six of the thirteen seats clear the honesty bar and none clears discrimination. The best honest seat sits six points under the bar, and the seat with the highest raw discrimination invents a difference on a third of the pairs where both sides are byte-identical.
- Does pooling several weak judges into an ensemble fix the problem?
- Not here. One rule was registered in advance and computed once: every seat that passed the honesty bar became a member, six qualified, the vote was plurality with ties returning a tie, and both bars stayed unchanged. The ensemble scored 60.2% order-balanced, fourteen points worse than its best single member at 74.1%. The decisive number is order 2 at exactly 50.0%: shown the correct answer second, the ensemble is a coin, because the members' errors are correlated along the position axis.
- Does a bigger local model make a better judge?
- This record says no. phi4 at 14.7B beats both 35B qwen builds on accuracy while matching their perfect honesty, and the two largest models on the host judged for the first time in this cycle and landed mid-table. "Bigger local model" is now refuted rather than untested.
- How can a benchmark grade anything without a judge model?
- Because the people who publish competing marketing skills also publish the rules they think a good cold email follows, and those rules are checkable by a program. Ninety rules survived extraction and verification, and 33 of them from three independent authors formed the cold-email rubric. A rubric built out of other people's published rules needs no referee, and it is not ours to bend, because we did not write it.
- What does incomparable mean here, and why is it not a tie?
- Comparison is a partial order rather than a score. Skill A beats skill B when A satisfies every rule B does and at least one more. Two skills can each pass something the other misses, and then neither dominates and the honest answer is incomparable. Summing the rules instead would let a pile of cheap wins outvote a rule the whole field agreed on.
- Why was the earlier 9 out of 9 result thrown out?
- On 31 August our cold-email skill scored 9 out of 9 against a rubric, and every rule it passed that the baseline failed had been written verbatim into the skill, from a rubric the same agent authored. That is teaching to the test and it is close to tautological, so it was recorded as not a result and put on the record as claim K-015.
- Our skill passed the most rules. Why is that not a win?
- A win required dominance on both brief sets and there was dominance on neither, so all twelve comparisons came back incomparable. One rule does it: James Praise publishes "Plain English, first-person tense", six of the seven arms satisfy it on five briefs out of five, and ours satisfies it on two. Delete that rule and our skill dominates all six other arms, which is exactly why the rule stays and the result stands as a loss.
- What can a rubric of published rules not measure?
- It measures compliance, not persuasion. A fully compliant email can still be dull, and consensus is agreement rather than truth, so it cannot detect a rule the whole field has wrong. Rule conformance is a floor, and the owner's blind rank remains the final arbiter.