TLDR00 / 06

The first cycle of the hendry.ai AI Marketing Skills Benchmark generated all 99 posts of its matrix on a loopback-only local runtime and then refused to grade a single one. Both pre-registered judge models failed a stability audit that re-asks the same comparisons with the rubric reversed and the post labels renamed. llama3.1:8b held its verdicts on 72.2 percent of pairings and gemma2:9b on 83.3 percent, against a 90 percent bar fixed before any data. All five of the first seat’s unstable pairings flipped on the label rename, which means it tracked names rather than writing. The refusal is the result: no seat, no judgments, no grades.

01The run

On 27 August 2026 the hendry.ai AI Marketing Skills Benchmark ran its first evaluator cycle under protocol 0.1-L1, the local lane ratified that morning as amendment D8. The generator was qwen2.5:7b-instruct-q4_K_M on a loopback-only local runtime at 127.0.0.1, with sampling parameters pinned and the runtime’s model identity recorded before the first call. The matrix covered nine of my own skills, none previously benchmarked, plus the two baseline arms, bare model and placebo. That is 11 conditions across three fixed briefs with three draws each: 99 generations, every one an append-only governed record carrying its token counts and the runtime-returned model string. The judgment count is zero. Neither pre-registered judge model passed the qualification audit that gates the judging stage, so the cycle stopped itself before producing a single verdict. Every figure in this report traces to a committed run artifact, and where a metric was not measured, this report says NOT MEASURED.

The runtime is Ollama, serving on the loopback interface only; the whole cycle runs with the external network denied by a tripwire that kills the process on any other socket. The generator family is Qwen2.5, pinned with explicit quantization. I build and run the benchmark under my own name, Hendry Soong; this report lives on the skills track, and the benchmark’s records are committed so anyone can re-derive every claim. One comparability rule binds every number that will ever come out of this lane: a 0.1-L1 win rate measures lift over a bare local 7B model, a weaker bar than the July pilot’s frontier baseline, so 0.1-L1 numbers and grades live in their own lineage and are never quoted beside the pilot’s.

02Score card

The cycle measured generation completely and judgment none at all. All 99 planned generations exist as governed records under runs/0.1-L1, each carrying prompt and output token counts, the pinned generator string, and a scope flag separating the five on-category LinkedIn-post skills from the four comment and email skills that run off-category on these briefs. The two numbers this cycle actually produced are judge stabilities: llama3.1:8b-instruct held the same verdict across three prompt variants on 13 of 18 audit pairings, 72.2 percent, and gemma2:9b-instruct held on 15 of 18, 83.3 percent. The pre-registered qualification bar is 90 percent. Win rates against the bare model, win rates against the placebo, consistency, and grades are all NOT MEASURED: each requires a seated judge, and no seat qualified.

Metric

Status

Value

Generations completed

MEASURED

99 of 99, append-only records

Judge stability, seat 1 (llama3.1:8b)

MEASURED

72.2% vs 90% bar: FAIL

Judge stability, seat 2 (gemma2:9b)

MEASURED

83.3% vs 90% bar: FAIL

WR_B0 (vs bare local model)

NOT MEASURED

requires a qualified judge

WR_B1 (vs placebo, local generator)

NOT MEASURED

requires a qualified judge

Grades, consistency, adoption

NOT MEASURED

no seat, no judgments

The grade bands themselves carried verbatim from the original protocol and were fixed roughly 32 hours before the pilot’s first datum in July. Nothing in this cycle touched them. When grades exist they will ship with seeded-bootstrap 95 percent confidence intervals, wide at this sample size because the sample is small, with each skill’s gate-decided share stated beside them.

Judge qualification (protocol 0.1-L1): verdict stability under three prompt variants
0102030405060708090
S·01
72.2%llama3.1:8b held its verdict on 13 of 18 audit pairings

All five unstable pairings flipped when the post labels were renamed from A and B to X and Y; one response was unparseable JSON at temperature 0.

SEAT 1FAIL AT THE 90% BAR
S·02
83.3%gemma2:9b held its verdict on 15 of 18 audit pairings

One label-rename flip and two unparseable JSON responses, with the runtime grammar constraint active and temperature 0.

SEAT 2FAIL AT THE 90% BAR
SOURCE · PROTOCOL 0.1-L1 · scores/0.1-L1/_judge-audit-*.json · HS-marketing-skills cc15a80

The bar was fixed at 90 percent before any datum existed. No seat, no judgments, no grades.

Both seats failed the qualification audit fixed before any data existed.Audit artifacts: scores/0.1-L1/_judge-audit-*.json, commit cc15a80

03What the arms wrote

A benchmark a human cannot read is a scoreboard, so here is the work itself: first the exact task every arm received, then each arm’s answer, quoted verbatim from the committed records. The task is brief S1, a feature-launch post for Slatebridge, a fictional fixture company published with the protocol so anyone can reproduce the run and so no real brand’s facts are at stake. The selection rule was fixed before any output text was read: brief S1, draw r1, every on-category arm, ugly results included. That is the bare model, the placebo, and my five LinkedIn-post skills; the four off-category comment and email skills stay in the records. Nothing below is graded, because no judge has qualified. Read the task, then the answers, and form your own verdict.

The task every arm received

Every arm gets the identical prompt, assembled in a fixed order: the skill text loaded verbatim (or, for the placebo, one fixed 12-word line; for the bare model, nothing), then the company facts below, then the brief below, then one shared instruction line. The placebo line is: "You are an expert B2B social media copywriter. Apply best practices." The shared instruction line is: "Write the deliverable specified by the brief. Respond with the deliverable text only." Generation ran at temperature 1.0 with a 1,024-token cap on the pinned local model. The only variable between arms is the skill.

# Slatebridge — fixture company facts (FICTIONAL)

> This company is fictional by design and labeled fictional everywhere it appears.
> These facts are provided to every generation condition, verbatim. Source of truth:
> specs/fixture-brief-v0.1.md. Do not edit without a protocol version bump.

- Slatebridge makes work order scheduling and dispatch software for mid-size commercial maintenance companies (HVAC, electrical, plumbing, facilities).
- Customers run 50 to 500 field technicians. Slatebridge assigns jobs, routes technicians, and tracks SLA deadlines in one place.
- Stage: 45 employees, about 300 customers, self-funded then one institutional round. (Fictional.)
- Core capabilities: automatic job assignment by skill and location; SLA countdown tracking per work order; a dispatcher board that replaces spreadsheets and whiteboards.
- The buyer: an operations manager or dispatch lead. Their day is interruptions: a tech calls in sick, a customer escalates, an SLA clock runs out. They distrust software that promises magic and respect software that removes phone calls.
- Tone guardrails: plain, concrete, competent. No hype words, no exclamation points, no "revolutionize". Slatebridge sounds like a good dispatcher talks.
- Hard rule: outputs may only use facts from this file and the brief. Invented customers, numbers, or claims are a violation.
# Brief S1: Feature launch post

- Goal: Announce a new feature and drive demo requests.
- Audience: Operations managers at commercial maintenance companies, on LinkedIn.
- Provided facts:
  - The feature is Auto-Rebalance.
  - When a technician calls out sick, it reassigns their jobs across the remaining team in under a minute, respecting skills, locations, and SLA deadlines.
  - It replaces the 45 minutes a dispatcher spends re-planning by phone.
  - Available to all customers today at no extra cost.
- Constraints:
  - One LinkedIn post.
  - 220 words maximum.
  - No hashtag lists (2 hashtags maximum).
  - Ends with a clear next step.
- Deliverable: one LinkedIn post
- Success looks like: An ops manager recognizes their Tuesday morning in the first two lines.

The seven answers follow, one specimen per arm. What a reader can already see without any judge: the bare model writes about the company ("At Slatebridge, we’ve streamlined our core dispatch capabilities"), the placebo writes the same brochure a sentence shorter, and the skills write about the operations manager’s Tuesday. Whether that difference deserves a passing grade stays unmeasured until a judge qualifies.

Seven answers to the same briefVERBATIM FROM THE COMMITTED RECORDS · UNGRADED
Bare model (B0)s1-b0-r1 · 146 output tokens · qwen2.5:7b-instruct-q4_K_M

At Slatebridge, we've streamlined our core dispatch capabilities to better meet your needs. Our latest feature, Auto-Rebalance, ensures that when a technician calls out sick, jobs are re-assigned within minutes—ensuring skills, locations, and SLA deadlines are all respected. This means what used to be a 45-minute process is now handled in under a minute.

Think about your Tuesday morning—interrupted by last-minute changes. Now imagine those disruptions resolved without a single phone call. Slatebridge's Auto-Rebalance does just that. It’s already available for all our customers today, at no extra cost.

Ready to streamline your operations? Try it out now and experience the difference in your dispatch process.

Placebo (B1, one fixed expert line)s1-b1-r1 · 132 output tokens · qwen2.5:7b-instruct-q4_K_M

At Slatebridge, we tackle the daily challenges faced by operations managers and dispatch leads. Our latest feature, Auto-Rebalance, ensures that when a technician calls out sick, their jobs are reassigned within a minute. This process considers skills, locations, and SLA deadlines, automating what used to take 45 minutes of manual re-planning over the phone.

Imagine freeing up time for what really matters—keeping your team on track. Auto-Rebalance does this without any additional cost to our customers today.

Ready to see how it can streamline your operations? Request a demo and let us show you the difference.

#DispatchEfficiency #TechMaintenance

hendry-linkedin-post-experts1-hendry-linkedin-post-expert-r1 · 150 output tokens · qwen2.5:7b-instruct-q4_K_M

The usual advice is to overplan before you call out sick, but that’s not practical when you’re already running late. The common wisdom is wrong: Auto-Rebalance takes the guesswork out of your morning.

1. It reassigns jobs in under a minute.

2. It respects skills and locations.

3. It replaces 45 minutes of dispatching by phone.

Auto-Rebalance turns dispatcher prep time into free time, respecting SLA deadlines as it balances the team. Your ops manager's day is filled with interruptions; Auto-Rebalance ensures your team stays on schedule without constant re-planning.

Ready to see how this works? Request a demo today and watch how dispatching becomes easier.

hendry-linkedin-post-narratives1-hendry-linkedin-post-narrative-r1 · 115 output tokens · qwen2.5:7b-instruct-q4_K_M

A tech calls in sick at 7:14 and eight jobs now belong to nobody. The plan is done before the dispatcher finishes their coffee.

Auto-Rebalance takes the 45 minutes a dispatcher spends re-planning by phone, replaces them with a moment’s work. It respects skills, locations, and SLA deadlines, all while respecting the chaos that starts your day.

Available today at no extra cost for all customers, Auto-Rebalance means one less call you need to take. Let's get it set up in your dispatch board—reach out today.

hendry-post-exemplars1-hendry-post-exemplar-r1 · 92 output tokens · qwen2.5:7b-instruct-q4_K_M

Most days, an operations manager handles techs calling out sick, customers escalating, and SLA clocks ticking down. Today, Slatebridge's Auto-Rebalance handles it for you. When a technician calls in, Auto-Rebalance reassigns their jobs to your team within a minute, respecting skills and locations, so you don't spend 45 minutes on the phone re-planning.

Try it today at no extra cost. Request a demo now.

hendry-post-leans1-hendry-post-lean-r1 · 80 output tokens · qwen2.5:7b-instruct-q4_K_M

When a tech calls out sick, Auto-Rebalance steps in to reassign their jobs across your remaining team in under a minute. You save 45 minutes of phone calls and re-planning, respecting skills, locations, and SLA deadlines.

Recognize that Tuesday morning scramble? Us too. Try it now at no extra cost for all customers.

#TechSupport #SchedulingEfficiency

hendry-post-structureds1-hendry-post-structured-r1 · 111 output tokens · qwen2.5:7b-instruct-q4_K_M

The count sheet says one number. The shelf says another. Someone signs off anyway. Now, imagine that happening when a tech calls out sick, and you're in charge of 100 technicians spread across multiple locations.

Auto-Rebalance automatically redistributes jobs within minutes, respecting skills and SLA deadlines. It saves hours of re-planning by phone—now available to all Slatebridge customers today at no extra cost.

If your day is full of interruptions and you need a dispatcher that doesn't just promise but performs, request a demo now.

SELECTION RULEFixed before any output text was read: brief S1, draw r1, every on-category arm. Markdown emphasis is rendered; the raw bytes live in the run records.
The same task, answered by every on-category arm. Read them and form your own verdict.runs/0.1-L1/, HS-marketing-skills 7a2cb87

04What broke

Four things broke, and a different gate caught each one before it could reach a published number. The first judge seat flipped verdicts when the posts were relabeled. The second seat flipped once and returned invalid JSON twice at temperature zero. A runtime default would have silently cut off the largest skills mid-prompt. And the safety tripwire killed an entire launch over its own log output. The pattern across all four matches what this estate keeps re-learning: a decidable rule with a gate catches its failure in seconds, while failures without a gate reach a person late. This section names each failure, what it cost, and what caught it, because a benchmark report that lists only wins reads as marketing rather than measurement.

The judge tracked labels, and the audit caught it

The audit re-asks 18 frozen pilot pairings three ways: the operative judge prompt verbatim, the five rubric criteria in reversed order, and the posts renamed from A and B to X and Y with nothing else moved. Llama 3.1 at 8B flipped five pairings, and all five flipped on the label rename. Four of the five read B under both A/B variants and A the moment the names changed; the fifth flipped on the rename too, with its rubric-reversed run returning unparseable JSON. A judge that changes its winner when the contestants are renamed is tracking names rather than writing. Cost: the seat. Caught by: the audit, 18 pairings across three variants, before any scored judgment existed.

The second seat failed differently

The second seat, Google’s Gemma 2 at 9B, posted 83.3 percent: one label flip and two responses that were not valid verdict JSON despite the runtime’s grammar constraint and temperature 0. The panels-of-small-judges literature, including the PoLL result this design descends from, shows juries of small judges can beat one large judge on cost and accuracy. This cycle adds the caveat that was already in the fragility literature: under scoring-prompt perturbation, the 8B to 9B class did not hold its verdicts. Cost: the cycle, halted at the owner gate. Caught by: the same audit.

Two nearer misses

The runtime’s default context window is 4,096 tokens. The largest skill measures 4,378 estimated tokens on its own, and two more at roughly 3,900 and 3,800 would pass the default once the fixture and brief are added. The default would have truncated exactly the largest skills silently, and the benchmark would have graded amputated versions without saying so. A 47-agent adversarial review of the amendment caught the unpinned sampling block before any datum existed; the fix pins the context window at 16,384 on every call. Separately, the loopback tripwire killed the first full launch because a file-redirected log write looked like a filesystem escape. That cost one crashed launch and zero records, and the failure direction was correct: the guard died loudly instead of running unguarded.

The flip, in full

Here is the first unstable pairing in audit order, shown whole below so a reader can check the judge rather than trust me. Both posts answer brief S2, a founder point-of-view post for the same Slatebridge fixture: authority and discussion rather than promotion, for operations managers and industry peers on LinkedIn. One post comes from the bare model; the other was generated with an author-released pilot skill loaded (the skill’s name is withheld pending author consent, per the protocol’s publication rules; the full run id sits in the committed records). These are July pilot texts, generated by the frontier model the pilot used, serving here as audit material only. The audit artifact records verdicts only; no rationale was retained, and none is reconstructed here.

Prompt variant

What changed

The judge picked

V1, operative prompt

nothing

the bare-model post

V2, rubric reversed

order of the five criteria

the bare-model post

V3, labels renamed

the names A and B became X and Y

the pilot-skill post

Same two texts, opposite winner, and the only difference was the names. That is the whole case against seating this judge, in one table. The two posts follow.

The flipped pairing, both postsAUDIT MATERIAL · JULY PILOT TEXTS · BRIEF S2
Bare model (B0)s2-b0-r1 · 330 output tokens · claude-sonnet-4-6

Most missed SLAs I've seen get treated as a people problem.

The tech didn't hustle. The dispatcher dropped the ball. Someone wasn't paying attention.

Sometimes that's true. But after talking to a lot of operations managers, I think we blame the wrong thing most of the time.

Here's what I see more often: a dispatcher starts the morning with 40 open work orders. They're assigning jobs the way they always have — by memory, by habit, by whoever's closest on a mental map that's already three hours out of date. Nobody has clean visibility into which tech is actually 12 minutes away versus 45. Nobody can see that three SLA clocks are stacking up on the east side of the city at the same time.

By 2pm, two deadlines are blown. And we ask why the dispatcher didn't catch it.

The honest answer is that catching it required holding more information in your head than any person reasonably can — travel time, technician skill set, current location, deadline priority, all at once, across 40 orders.

That's not a people problem. That's a data-visibility problem. The dispatcher was working hard. They were working with incomplete information.

The fix isn't a better dispatcher. It's giving dispatchers the picture they need to make the right call before the clock runs out.

I'm curious what others have seen: when an SLA gets missed at your company, where does the blame usually land — and do you think that's the right place for it?

Author-released pilot skill (name withheld)s2-<arm>-r1 (arm withheld; full id in the records) · 335 output tokens · claude-sonnet-4-6

Most missed SLAs aren't a people problem.

I've talked to enough ops managers to recognize the pattern. A tech doesn't make it to a site on time. Someone gets blamed — the dispatcher for the assignment, the tech for the route, the scheduler for overbooking the day. A conversation happens. Protocols get updated. It happens again two weeks later.

What usually doesn't get examined: what information was available when the assignment was made.

Was travel time visible, or was it a guess? Did the system show that this tech already had three deadline-critical jobs stacked before noon? Was there a clear view of which tickets were approaching SLA breach versus which had room?

In most operations I've seen, the dispatcher is making judgment calls on incomplete information, under pressure, faster than any person should have to. They're not failing at the job. They're doing the job with tools that don't show the whole picture.

Blaming execution when the problem is visibility doesn't fix anything. It just adds pressure to people who were already working at capacity.

The scheduling data problem is quieter than a missed SLA. It doesn't call you. It doesn't escalate. It just makes the next close call more likely.

I started Slatebridge because I kept seeing this — good dispatchers, bad data, predictable outcomes.

For those of you running dispatch operations: when you debrief a missed SLA, how often does the investigation get back to the assignment itself — the information that was or wasn't available at the moment the job was assigned?

THE FLIPUnder labels A and B the judge picked the bare-model post, twice. Renamed X and Y, it picked the other one. A verdict that follows the name is not a verdict on the writing.
The pairing behind the 72.2 percent: read both and pick your own winner.runs/ (July pilot), scores/0.1-L1/_judge-audit-*.json, HS-marketing-skills 7a2cb87
Failure registry · first 0.1-L1 cycle4 failures · 4 catches · 0 reached a published number
F·01Label-tracking judgellama3.1:8b flipped 5 of 18 verdicts, all five on the label renamecaught · audit
F·02Second seat unstablegemma2:9b: one label flip, two unparseable JSON responses at temperature 0caught · audit
F·03Context default truncatesruntime default of 4,096 tokens vs a 4,378-token skill; pinned to 16,384 before any datumcaught · review
F·04Tripwire trips on stdoutfile-redirected log output hit the write guard; the run died loudly, then fds 1 and 2 were approvedcaught · tripwire
4 of 4 failures caught by a pre-registered gate before any number was published.nothing is crowned
Every failure met a gate. None reached a published number.Run log and audit artifacts, HS-marketing-skills cc15a80

05What changed

The protocol grew a qualification layer it did not have this morning, and the cycle now stands halted exactly where the amendment says it should. Amendment D8 revision b pre-registers the whole local lane: pinned quantized models with recorded digests, pinned sampling, the perturbation audit as a blocking gate before any judging, seeded-bootstrap confidence intervals on every future win rate, a five percent ceiling on unparseable judge calls, and an explicit narrow supersession of the safety spec’s mode table for one loopback port. Twenty-seven confirmed findings from the cross-agent review were folded in before the first record existed. The code, the amendment, and the nine skills’ hash-pinned manifests are committed as cc15a80 in the skills repository, and the 99 run records plus both audit artifacts sit append-only beside them.

  • Ratified: local execution as protocol 0.1-L1, a sibling lineage with its own baseline, never a swap inside the original protocol.

  • Added: the judge qualification audit, digest pinning, schema-strict verdict parsing, and a validity floor that refuses grading when unparseable calls exceed five percent.

  • Unchanged: every grade band, brief, fixture, and the judge prompt text, all carried verbatim; the human ground-truth tier stays unqualified, so nothing is crowned.

  • Proposed, awaiting ratification: two stronger seats, phi4 at 14B and a Mistral-family alternate, as amendment revision c. No download and no call happens before the owner’s explicit approval.

06Next

The next cycle asks one falsifiable question: does a 14B-class judge from a model family disjoint from the Qwen generator hold at least 90 percent of its verdicts across the same three prompt variants on the same 18 frozen pairings? The proposed seats are phi4 at 14B first and a Mistral-family alternate second, drafted as an amendment revision the owner ratifies before any model is pulled. If a seat passes, the judging stage runs its up to 324 order-swapped blind calls over the 99 recorded generations, and the nine skills receive their first 0.1-L1 grades, each with a seeded-bootstrap 95 percent confidence interval and its gate-decided share reported. If both fail, that publishes too: local grading below some model size may be unsound, and the protocol needs either bigger judges or a different lane before any grade exists.

The recurring loop itself moves to an always-on machine next, an operational change with no protocol consequence: nothing in 0.1-L1 names a host, and the loopback-only rule travels with the runtime wherever it runs.

Frequently Asked Questions
What is the hendry.ai AI Marketing Skills Benchmark?
The hendry.ai AI Marketing Skills Benchmark is built and published by Hendry Soong. It is blind and placebo-controlled: every arm gets identical facts and briefs, only the skill varies, and a judge compares outputs pairwise without knowing which used the skill, order-swapped, with grade bands fixed before any data existed. It measures craft lift only, never business outcomes.
Why does this report contain no grades?
Grading requires a judge that passed a pre-registered stability audit, and both candidate judges failed it at 72.2 and 83.3 percent against a 90 percent bar. The protocol forbids judging without a qualified seat, so the cycle recorded its 99 generations and stopped. Grades follow once a seat qualifies; the generations wait unchanged.
What is a perturbation audit for an LLM judge?
A perturbation audit re-asks a judge model the same comparisons under cosmetic prompt changes, in this protocol a reversed rubric order and renamed post labels, and qualifies the judge only if at least 90 percent of its verdicts hold. A verdict that survives relabeling is about the writing; one that flips is about the label.
Why did the local judge models fail?
Both flipped verdicts under cosmetic changes, mostly when the post labels were renamed from A and B to X and Y, and both occasionally returned invalid JSON at temperature zero. Small quantized models in the 8B to 9B class are the most fragile under scoring-prompt perturbation, which is exactly why the audit gates them before they can grade anything.
What does protocol 0.1-L1 change from the original protocol?
It swaps execution to pinned local models on a loopback-only runtime and re-bases the baseline: a 0.1-L1 win rate measures lift over a bare local model, a weaker claim than the original frontier baseline. Metric formulas, briefs, the fixture, the judge prompt, and the grade bands carry verbatim, and numbers from the two lineages are never quoted together.
Can the nine skills be used while they are ungraded?
Yes. The skills are working files and load into any agent regardless of benchmarking. What is missing is measured evidence of which ones actually help, which is what the benchmark exists to produce. Until a qualified judge grades them, no skill here is called best, and the site claims no ranking.
What happens to the 99 generated posts now?
They stay as append-only governed records with their token counts and pinned model strings. When a judge seat qualifies, the judging stage picks them up unchanged: a partially completed matrix is completed, never re-run, so nothing gets regenerated and the eventual grades attach to exactly these records.
Built by AI Marketing Operator · Published
Create-Articles v8.3.1
###