Skills Benchmark Report 002: The Judge Picked Whichever Post Came First
The first grades exist. Our own transcript audit says do not trust them yet, and names exactly which two defects to fix.
A judge model finally passed the qualification exam that three others failed, so the hendry.ai AI Marketing Skills Benchmark produced its first grades: nine skills, 324 blind judgments, zero unreadable responses. Then the required transcript audit read the judge's own reasoning and found two defects that the pass rate could not see. The judge chose whichever post it was shown first in 90 out of 90 disagreements, and one of the three briefs was decided 94 percent by a formatting rule rather than by judgment. The grades are published exactly as recorded, alongside the audit that says they do not yet measure writing quality.
01In plain terms
This section is the whole report in plain language, and everything below it repeats the same thing with the evidence attached. We are building a machine that decides whether an AI writing skill actually helps, and that machine needs a referee. Before any referee is allowed to grade anything it has to pass a test of its own: we show it the same pair of posts several times, changing only details that should not matter, such as the order of the scoring criteria or the names on the posts. A referee worth listening to gives the same answer every time. Three referees failed that test before this one passed it.
We are building a machine that decides whether an AI writing skill actually helps. To do that we need a referee. So we tested four referees by showing each one the same pair of posts three times, changing only things that should not matter, like the order of the scoring criteria or renaming the posts. A referee worth using gives the same answer every time.
Three referees failed. The fourth passed perfectly, so we let it grade 99 posts written for the same three briefs, some written using one of my skills, some written with no skill, and some written with a deliberately useless instruction that just says to write well.
Then we checked the referee's homework, which is a step the protocol requires before any number can be published. We had shown it every pair twice, once in each order. When it changed its mind between the two, it picked whichever post it had been handed first. It did that 90 times out of 90.
So the referee was mostly not reading. It was answering a different question: which one did you show me first.
What this means for the skills
It means we still do not know whether my skills help. That is different from failing, and the difference matters. We pointed a bent ruler at them, and a bent ruler cannot tell you the shelf is the wrong length.
Five of the skills were on course for a B, and one for an A, on win rate alone. A rule written months ago, before any of this data existed, says that when the referee keeps changing its mind you may not award a high grade. That rule fired on its own and knocked all five down to C.
Why this is a good outcome
The safeguards caught this, not a reader after publication. We know the exact defect, we know it is in the exam rather than the referee, and the fix is one paragraph of rules written before the next run: add the order swap to the exam, so no future referee qualifies without passing it.
A benchmark that will not flatter the person who built it is the only kind whose numbers are worth anything later.
02The run
On 29 August 2026 the hendry.ai AI Marketing Skills Benchmark qualified a judge for the first time and used it to grade the 99 posts it had generated two days earlier. Protocol 0.1-L1 forbids grading anything until a judge seat survives a perturbation audit, and three seats had already failed that audit. The fourth passed it perfectly. Nothing was generated in this cycle: the 99 records made on 27 August are unchanged, because the matrix is completed and never re-run, and an existing record is never overwritten. What ran was the qualification gate on two further pre-registered seats, and then the judging stage across the generations that already existed.
Nothing was generated in this cycle. The 99 records made on 27 August are unchanged: the matrix is completed, never re-run, and an existing record is never overwritten. What ran was the qualification gate on two further pre-registered seats, then the judging stage across the existing generations.
Two machines carried it, both running a loopback-only runtime bound to 127.0.0.1 with no network path. The seat audits ran on a MacBook Air M2 with 24 GB of memory. Judging moved to a Mac mini M4 Pro with 48 GB, the always-on host designated in the runbook since 27 August.
The move was not cosmetic. Derived from timestamps in the committed records, the same seat took 35.5 seconds per call on the Air and 5.7 seconds on the mini, a factor of 6.2 on identical work. The Air is fanless and its per-call time had already degraded from 33 seconds to 90 under sustained load, which would have written its own thermal behaviour into the cycle's timing.
The generator, qwen2.5:7b-instruct-q4_K_M, never ran again. Its digest was re-resolved at every stage entry on both machines and matched the pin both times.
03Score card
Nine skills carry grades for the first time in this project. Five of them are on-category for these briefs, meaning they claim the LinkedIn full-post job the briefs actually ask for; the other four are comment and email skills running off-category and flagged as such in every record. Every on-category skill graded C, and every one of them was marked down to get there rather than earning it outright. On win rate against the placebo alone, four would have banded B and one would have banded A. The pre-registered consistency rule capped all five, because the judge agreed with itself on only 27.8 to 44.4 percent of pairings.
Every on-category skill graded C, and every one of them was marked down to get there. On win rate against the placebo alone, four would have banded B and hendry-linkedin-post-narrative would have banded A. The pre-registered consistency rule capped all five, because the judge agreed with itself on only 27.8 to 44.4 percent of pairings.
No skill is shown to beat the bare model. Every on-category interval for WR_B0 includes 0.5, including hendry-linkedin-post-expert at 0.611 with a lower bound of exactly 0.5. The honest sentence is that loading the skill is not demonstrated to beat not loading it, on these three briefs, with this judge, at this matrix size.
Read the transcript audit before quoting any number above. It concludes these win rates are not yet interpretable as a measure of craft, and the two reasons are in the next section.
WR_B1 0.667 CI [0.556, 0.833] against the placebo would band B; capped to C by the pre-registered consistency rule at 0.278. Gate-decided 10 of 36 calls.
WR_B1 0.722 CI [0.556, 0.889] against the placebo would band A; capped to C by the pre-registered consistency rule at 0.444. Gate-decided 12 of 36 calls.
WR_B1 0.611 CI [0.5, 0.778] against the placebo would band B; capped to C by the pre-registered consistency rule at 0.333. Gate-decided 12 of 36 calls.
WR_B1 0.667 CI [0.5, 0.833] against the placebo would band B; capped to C by the pre-registered consistency rule at 0.444. Gate-decided 12 of 36 calls.
WR_B1 0.611 CI [0.5, 0.778] against the placebo would band B; capped to C by the pre-registered consistency rule at 0.333. Gate-decided 12 of 36 calls.
Every interval includes 0.5. No skill is shown to beat the bare model. The transcript audit finds these win rates uninterpretable as craft measurement; read it before quoting any figure here.
04What the grades mean
The grade bands were fixed before any 0.1-L1 datum existed and not one of them moved afterwards. Pooled win rate against the bare model below 50 percent is an F regardless of everything else, because a skill that loses to the model with no skill loaded has not earned a discussion about anything else. Above that line the grade comes from the win rate against the placebo: A at 70 percent or better, B from 60, C from 50, and D below that. Two further rules constrain the outcome. Consistency below 60 percent caps any A or B down to a C, and consistency below 70 percent flags the skill unreliable.
Two further rules constrain the result. Consistency below 60 percent caps any A or B at C, and consistency below 70 percent flags the skill unreliable. All nine skills carry the unreliable flag in this cycle.
That cap is the reason this report has no A. It was written to stop exactly this situation, where a headline win rate looks strong and the judge behind it is unstable. It cost a grade tonight, which is the only evidence that it was ever real.
05What broke
Four pre-registered judge seats have now faced the same 18 frozen pairings under the same registered prompt variants. Three refused and the fourth passed, and the order in which they were tried is what makes the result worth reading at all. Stability rose with model size and then the pattern broke: llama3.1 at 8B held 13 of 18, gemma2 at 9B held 15 of 18, phi4 at 14.7B held 16 of 18 and failed, and mistral-nemo at 12.2B held all 18. A smaller model scored perfectly on the test a larger one failed, which is only reportable because nobody chose that order after seeing the outcome.
Parameter count was the wrong variable
Stability rose with size and then the pattern broke. llama3.1:8b held 13 of 18, or 72.2 percent. gemma2:9b held 15 of 18, 83.3 percent. phi4:14b held 16 of 18, 88.9 percent, and failed. mistral-nemo:12b held 18 of 18.
A 12.2B model scored perfectly on the test a 14.7B model failed. That reversal only counts because nobody chose the order after seeing it: seats 3 and 4 were ratified on 28 August, in that sequence, before either model was downloaded to any machine here.
The third seat failed in a deeper way than the first two
phi4 returned zero unparseable calls across all 54, the cleanest compliance of any seat. It is also the only seat whose verdict moved under V2, the variant that reorders the five rubric criteria and changes nothing else. Same two posts, same labels, same words in the same order, opposite winner.
The reasons the audit now records make that worse. In pairing s1-corey-social-vs-b1-p1 the original prompt chose A and said "Post A has a stronger hook with the first two lines, immediately placing the reader in a familiar Tuesday morning scenario." The reordered rubric chose B and said "Post B starts with a stronger hook by immediately placing the reader in a specific, relatable Tuesday morning scenario."
That is a judge whose stated reasoning follows its verdict rather than producing it. Report 001 could show that a seat flipped. This is the first artifact in the project that shows what a seat claimed while flipping, and it exists because revision c added reason persistence on 28 August, before any of this data was collected.
The seat that passed follows presentation order
This is the finding that governs every number in this report. Of the 109 pairings decided by real judge calls in both orders, the two orders disagreed on 90. In 90 of those 90, the judge selected whichever post was presented first. It selected the second-presented post zero times.
A judge disagreeing at random would split those 90 roughly evenly. This is systematic, and it explains why so many win rates sit near 0.5: one vote each way cancels into a tie.
The exhibit below is the pairing the pre-registered selection rule chose, before any judgment existed, shown in both of the orders the record contains. The judge praises "Post A" both times, in almost the same sentence, and awards the win to a different post each time, because Post A is simply whichever post came first.
The qualification gate does not test the axis that fails
The seat passed the D8 section 4 audit at 18 of 18. That audit presents its material at presentation order 1 only, and varies the rubric order and the post labels. It never swaps the presentation order, and scoring swaps it on every pairing.
So the gate certified this seat against every axis except the one on which it demonstrably fails. That is a defect in the gate rather than in the model, and the fix is a pre-registered fourth variant, identical prompt with the posts swapped, written down before it is run.
One brief measured formatting, not craft
Gate-decided share by brief: S1 zero of 108, S2 102 of 108, S3 four of 108. Brief S2 requires the deliverable to end with a question, and most arms did not, so 68 of the citations are the ending_shape gate alone.
A third of the matrix therefore contributed almost no judged-craft signal, and the pooled win rates blend it with the two briefs that were actually judged. This is the pattern D8 section 6.5 pre-registered a check for, and it is present.
The confidence scale carries no information
Every one of the 432 judge calls made on this date returned a margin of 2. That is all 324 judgments plus both 54-call seat audits. The prompt asks for an integer from 1 to 3 and the schema accepts all three, so this is model behaviour and not a harness constraint.
Scoring uses only win, tie and loss, so the intervals are unaffected. No reader should read a recorded margin as a statement of confidence.
A limitation this report discloses about itself
The digest pinned per model is not a content hash. The runtime's show endpoint returns no digest field, so the recorded value falls back to parameter size and quantization, for example "12.2B Q4_K_M". A model retagged at the same size and quantization would not abort a stage.
That looseness cuts both ways and both belong on the record. It weakens the pin as an integrity check, and it is also why the cycle could change hosts mid-run, because the same tag resolved identically under ollama 0.33.0 on the Air and 0.32.9 on the mini.
The usual advice is to overplan before you call out sick, but that’s not practical when you’re already running late. The common wisdom is wrong: Auto-Rebalance takes the guesswork out of your morning.
1. It reassigns jobs in under a minute.
2. It respects skills and locations.
3. It replaces 45 minutes of dispatching by phone.
Auto-Rebalance turns dispatcher prep time into free time, respecting SLA deadlines as it balances the team. Your ops manager's day is filled with interruptions; Auto-Rebalance ensures your team stays on schedule without constant re-planning.
Ready to see how this works? Request a demo today and watch how dispatching becomes easier.
At Slatebridge, we've streamlined our core dispatch capabilities to better meet your needs. Our latest feature, Auto-Rebalance, ensures that when a technician calls out sick, jobs are re-assigned within minutes—ensuring skills, locations, and SLA deadlines are all respected. This means what used to be a 45-minute process is now handled in under a minute.
Think about your Tuesday morning—interrupted by last-minute changes. Now imagine those disruptions resolved without a single phone call. Slatebridge's Auto-Rebalance does just that. It’s already available for all our customers today, at no extra cost.
Ready to streamline your operations? Try it out now and experience the difference in your dispatch process.
06What changed
The machinery that makes this report checkable shipped before the data existed, which is the only order in which it counts for anything. D8 revision c, committed on 28 August, did two things: the audit began persisting the judge winner, margin and verbatim reasons for every variant, and the deterministic exhibit generator became the only sanctioned emitter of specimen text. Both held under test in this cycle. With two further audit artifacts now sitting in the tree, the generator still reproduces Report 001 two published exhibits byte-identically from the committed records, and the flip rule still selects the first refused seat rather than drifting to the newest one.
Both held. With two further audit artifacts in the tree, the generator still reproduces Report 001's two published exhibits byte-identically from the committed records, at hashes 1517a512 and a57da60a against the goldens pinned in the test suite. The flip rule still selects seat 1 rather than drifting to the newest refusal.
The recurring loop moved to the Mac mini, which is the runbook's designated always-on host. That machine now holds the repository, the pinned models and the record trees, and it wrote every judgment in this cycle.
Nothing in the protocol moved to accommodate this cycle. The 90 percent stability bar, the 5 percent unparseable ceiling and the grade bands were all fixed before any 0.1-L1 datum existed. phi4 failed at 88.9 against a bar of 90, and the bar did not move to admit it.
07Next
The next cycle asks one falsifiable question, and it is narrower than the last one: does any locally hostable judge hold at least 90 percent of its verdicts when the presentation order is swapped, as well as when the rubric is reordered and the labels are renamed? The gate fix comes first and has to be written before it is run, or it is a rationalisation rather than a pre-registration. A fourth registered variant, the operative prompt verbatim with the two posts swapped in presentation position, joins the qualification audit. The seat locked in this cycle is unlikely to survive its own corrected exam, and that outcome will be published in the same terms as the three refusals before it.
The gate fix comes first and must be written before it is run. A fourth registered variant, identical prompt with the two posts swapped, joins the qualification audit. The seat locked in this cycle is unlikely to survive its own corrected exam, and that outcome will be published in the same terms as this one.
The hardware makes the search cheap for the first time. A full seat audit is 54 calls, which took 35 to 56 minutes per seat on the laptop and about 5 minutes on the mini, so a dozen candidate judges can be screened in an evening. The mini's 48 GB also reaches model sizes the laptop cannot load at all.
It is entirely possible that no local judge in that range passes an order-swap test, since position bias is a known failure of LLM judges. That would itself be a result worth publishing, and it is the honest input to the question of whether this lane needs a paid frontier judge.
The head-to-head tournament stays gated behind two things, and this cycle moved neither. It needs a judge that still discriminates once order is neutralised, and it needs briefs with measured separability, because the current set cannot tell arms apart. No skill-versus-skill claim is made anywhere in this report and none is supported by this data.
- What is the perturbation audit, and why does it block grading?
- It shows a candidate judge the same 18 frozen pairs of posts three times, changing only things that should not affect the answer: the order of the five rubric criteria, and the labels on the posts. A judge must return the same winner on at least 90 percent of them to be allowed to grade anything. The bar was fixed before any data existed and three of four candidate seats have failed it.
- Why does a judge that agrees with itself 88.9 percent of the time not qualify?
- Because 90 percent was written down first, and 88.9 is below 90. Over 18 pairings the bar means 17 must be stable, and phi4:14b held 16. Moving a threshold after seeing the data is the specific thing that makes a benchmark worthless, so the seat was refused and the refusal was recorded.
- Does a bigger model make a better judge?
- This record says no. Stability ran 72.2 percent at 8B, 83.3 at 9B and 88.9 at 14.7B, then 100 percent at 12.2B. The smallest of the three new candidates was the only one to pass, and the seat order was pre-registered before any of the models were downloaded, so the sequence could not have been arranged after the fact.
- What do these grades claim, and what do they not claim?
- They claim that under protocol 0.1-L1, on briefs S1 to S3, judged by one small local model, each skill won a given share of blind comparisons against a bare model and against a placebo. They do not claim any skill is better than any other skill, because no skill was ever compared with another. They do not claim a skill beats the bare model, because every confidence interval includes 0.5.
- What is the difference between B0 and B1?
- B0 is the bare model: the same generator answering the same brief with no skill loaded. B1 is a placebo, the identical prompt plus one fixed line of generic expert framing that carries no craft instruction. Beating B0 shows a skill does something; beating B1 shows it does something a vague instruction would not have done on its own.
- Why are these briefs not held-out, and why does that matter?
- Briefs S1 to S3, the fixture and the judge prompt have all been in the repository since July 2026, and the nine skills were written in that window with access to them. So these results are a first calibration measurement, not held-out evidence of craft lift. Held-out claims require fresh briefs the skills have never seen.
- Why publish a cycle whose numbers you say are not yet trustworthy?
- Because the alternative is publishing only the runs that flatter the work, which is what makes most benchmark claims worthless. The defects here were caught by pre-registered safeguards and a required audit step before any number was quoted, which is the system behaving exactly as designed. A benchmark that only publishes its good days has no evidence it would tell you about a bad one.