In plain terms

Marketing teams adopt AI marketing skills on faith and on popularity. A skill ships, a thread praises it, the install count climbs, and almost nobody checks whether the output is actually better than what the same AI produces without it. This section is where that gets checked, in a way you can audit rather than take on trust. Every skill is tested blind against two things: the same AI with no skill loaded at all, and the same AI handed a deliberately useless instruction that only tells it to work well. A skill has to beat both to have earned anything, because beating silence is easy and beating vague encouragement is the harder test.

The remit is every job an AI marketing skill claims to do, not writing alone. The cycles published so far cover written deliverables, because that is where the skills market started and where my own set started. Ads, briefs, research, analysis and the rest enter the benchmark the same way: each category arrives under its own protocol version, with its own briefs and its own bands, and never folded into a number earned on something else.

The grades come from a referee that must first pass its own test. We show it identical comparisons with irrelevant details changed, and if it does not give the same answer every time, it is not allowed to grade. Referees have failed that test more often than they have passed it, and when they fail, no grades get published.

Every number here traces back to a stored record you can look up, and the pass marks were fixed in writing before any data existed. When something does not work, the report says so.

What this track is

This is the home of the AI Marketing Skills Benchmark, which I built and run under my own name, Hendry Soong. An AI marketing skill is a packaged instruction file that a model loads before it does the work, the format popularised by Anthropic’s Agent Skills, and the market for them runs on faith and popularity: skills ship, get praised, get adopted on a star count, and almost never get measured. This benchmark measures them, across every job a marketing skill claims to do rather than writing alone. Each benchmark report published under this track records one evaluation cycle: what ran, what it scored, what broke, and what changed, with the generated work quoted verbatim and every number traceable to a committed run record. Reports are numbered and permanent; a refusal to grade publishes with the same prominence as a grade.

How a skill gets measured

Every skill is tested against the same fictional company on the same briefs, in three arms that differ only in one variable. The bare model gets the facts and the brief and nothing else. The placebo gets one fixed 12-word expert line on top. The skill arm gets the full skill text, loaded verbatim. A judge model then compares outputs pairwise, blind to which arm wrote what, with every pairing judged twice in swapped order; disagreement scores as a tie. Win rates against the bare model and against the placebo become grades, on bands that were fixed before any data existed and have never moved. The benchmark measures craft lift only, whether the skill makes the deliverable better at the job it claims, and never claims anything about conversions, reach, or revenue.

The judge itself has to qualify first. Before any grading, a candidate judge faces a perturbation audit: the same comparisons re-asked with the grading rubric reversed and the post labels renamed. A judge that changes its winner when only the names change is judging labels, and it is refused the seat. The architecture follows the published evaluation literature, from panel-of-small-judges results (PoLL) to the blind order-swapped pairwise design with bootstrapped confidence intervals (Arena-Hard). Where a small judge fails the audit, that failure is reported rather than smoothed over.

The state of the benchmark

As of 6 September 2026 the benchmark runs protocol 0.1-L1: generation on a pinned local model over a loopback-only runtime, with sampling parameters and model digests recorded before the first call. Five reports are published. No judge is seated, no skill here is called best, and the rubric that replaced the judge has now been measured on two jobs and works as a floor rather than as a ranking. The map of the field is finished. A second instrument, one that measures how a skill is built rather than how well it writes, is pre-registered and has not been run yet.

Reports 001 to 004: the judge search and what replaced it

The first cycle generated all 99 posts of its matrix, nine of my own skills plus the two baseline arms across three briefs with three draws each, and then refused to grade any of them, because both pre-registered judge models failed a stability audit. Report 002 published the first grades together with the transcript audit that limits them: the seated judge chose whichever post it was shown first in 90 of 90 disagreements. Report 003 ran a control with a known answer and found the qualification exam ranking judges backwards, so the exam was retired. Report 004 closed the search at thirteen seat measurements with none qualifying, refused a pooled ensemble of the six most honest seats for scoring worse than its own best member, and replaced the referee with a rubric assembled out of the rules that competing skill authors publish about their own job. That rubric needs no judge model, because a program can check it and I did not write it.

Report 005: the rules are a floor, not a ranking

Report 005 ran that method on a second job to find out whether it generalises, and it does not. Eleven arms — five of my LinkedIn skills, four competitor skills pinned at their commits, a bare model and a placebo — produced 99 posts, scored by program against 14 checkable rules drawn from five published skills by four distinct authors. The whole field lands inside a spread of two rules. Seven arms tie at 11 of 14, including the largest competitor in the bracket and three of mine, and six of those seven produce a byte-identical set of passes. The bare model and the placebo score 10. The lowest arm, at 9, is one of mine. Every skill in that bracket is worth about one rule more than using no skill at all.

The instrument was checked before the arms ran rather than after. Two reference texts written first, one deliberately rule-ignoring and one deliberately compliant, were separated by nine rules, so the rubric can express a spread of nine and the field produced two. Of the 14 rules, seven are inert because every arm passes them, two are dead because no arm passes them, and five discriminate at all. The reason is what the field publishes: of 99 rules harvested from six source skills, about 64 percent are shape — character counts, hashtags, line breaks, the see-more fold — and about a tenth touch content at all. That split is a classification made at derivation time rather than a count, and the report says so; what is counted is the scoring. Nothing here tests whether a post is worth reading.

One source skill turned out to be written for X rather than LinkedIn. It contributed 34 of those 99 rules, its frontmatter mentions LinkedIn zero times, and the corrected base is 65 rules from four LinkedIn authors. The same class of error reached the published results table in Report 004. This time it was caught before the rubric existed.

The published zero that reverses to thirty

Report 005 reverses a null this project published, and the reversal is the most useful thing in it. A check called metric ownership had reported zero misattribution violations across 240 generations in two pre-registered runs, and the earlier reading concluded the failure was not inducible on this generator. The check credits the reader with a result only when a second-person subject governs a verb drawn from a closed list of 43. Two of the three briefs carried proof points built on "close" and "recover", neither of which was on the list, so 160 of the 240 generations could not have failed whatever the model wrote. Rescored with those verbs restored, the same 240 drafts yield 30 violations: 10 of 120 on the unchanged skill and 20 of 120 on the arm that carries the ownership rule, a direction inverted from the rule's intent. Neither pre-registered run clears its own threshold — run 1 p = 0.0535, run 2 p = 0.6179, pooled post-hoc p = 0.0776 — so it is recorded as a direction and not a finding. What it is not is a null. The control now lives in code: a positive control builds the canonical violation out of each brief's own proof point and asks the real detector to catch it, and a brief the control cannot fire on is reported unscorable instead of passing for free.

Five more figures in that report were wrong and are corrected in place, each quoting the original wording, because an independent audit recomputed every figure against the primary artifacts and six did not survive. The closing sentence claiming everything had been recomputed was cut, because it was not true. In the same cycle, nine of my fourteen skills were measured against my own published skill spec and violated it, and nothing had been enforcing that spec; the check is in the build now and reads 14 of 14.

Six instruments, none of which ranks craft

Every instrument tried so far has failed to rank two good skills against each other on craft. Six distinct instruments across eighteen measurements, recorded in this project's own decision log, They are thirteen local judge seats, the qualification exam that measured them and ranked them backwards, the ensemble pooled from the six most honest of them, the cold-email union rubric against a human reader, the LinkedIn union rubric against itself, and the ownership check above. What they share is one shape: reference-free scoring of outputs, by small local models, against rubrics written here. That is a narrower claim than "ranking is impossible", and the narrower one is the correct one. A frontier judge has never been tried, and nothing measured so far closes that question.

The field map: 855 skills, 508 in scope, 21 clusters

The map needs no referee and it is done. Every SKILL.md file in the eight admissible repositories was read at its pinned commit: 855 real files, of which 508 serve a go-to-market job. Those 508 were assigned to a 21-cluster taxonomy built from three independent organising lenses and merged into one, with zero unmatched and zero non-canonical assignments. The largest cluster, building site pages and blocks, holds 60 skills. The smallest real cluster, writing a cold outbound message, holds 7 — fewer than the 13 skills in the bucket for things that fit nothing. An independent count agrees: of the 508, 12 write email copy and 5 write LinkedIn copy. This is a search and owned-web field, and it barely writes outreach at all, which is where both of this project's confirmed finds came from.

The map also corrects a number this site published. Report 004 said 1,313 harvested skill entries had never been clustered. 1,313 is 855 real files plus 458 symlinks, all 458 of them inside a single repository, and 73 of the 458 resolve to agent personas and slash commands rather than skills at all. Unique by content across the eight repositories: 843. The correction changes no conclusion in Report 004, and it is published because the number was printed here.

Twenty-eight candidate unserved jobs were then proposed and searched for across all eight checkouts, first by keyword and then, on 5 September, by a second skeptic working from the taxonomy inward. None was refuted: 25 survived outright, and three proposed refutations were overturned by an independent reader on the bytes. The stage was validated before it was believed — two positive controls were injected, a cold email citing a customer result and a technical SEO audit, and both were refuted by both methods, four times out of four. None of the 28 is called unserved in public, because a job nobody has found is not the same as a job nobody serves.

The construction instrument: 15 properties out of 79 proposed

Registered and frozen on 6 September 2026: a second instrument that measures the skill as an engineered artifact rather than the copy it produces — resident context cost, trigger design, output contract, portability, self-verification. Every property is specified to be computed by a deterministic program from the bytes of the skill package, with no model judging anything anywhere in it. Nothing has been measured against it yet, because the detectors are not pinned in code. Construction is not craft. A skill can be immaculately engineered and write dull copy, and this instrument will never say otherwise; reading a construction score as a craft ranking is a misreading. It goes first for one reason: a skill that is never selected never runs, so its craft is unreachable.

Seven lenses read the field only, under an explicit prohibition on opening my own skills or the house spec, and proposed 79 properties. Each went to an adversary who implemented the detector independently, ran it against its own positive and negative controls and against at least 15 real corpus skills, and tried to break it. Fifteen survived and 64 were rejected. Seven of the fifteen are labelled causal and may carry a recommendation; eight are labelled unknown and carry no advice at all until an outcome variable exists; none cleared the bar for convention, and no property may be relabelled after anyone sees where my skills sit on it.

One whole lens produced zero survivors. All 11 self-verification properties failed, because readers working from the field alone could not turn "does the skill check its own work" into anything a program scores without a judgement call — while my own spec makes a final validation pass mandatory and a gate enforces it. That is a property enforced here and not measurable here, and it is written into the pre-registration rather than left out of it.

The property everyone reaches for first, whether the description says when to use the skill, is useless on this corpus at a 98.4 percent base rate. The failure that bites instead is collision: 33 in-scope skills share a normalised trigger condition with at least one other, and three separate SEO-audit skills carry byte-identical conditions. All three pass every trigger property, and installing them together makes selection strictly worse.

Three properties are disclosed as not blind, because they were computed on the field and on my own skills before the pre-registration existed and are the reason the study exists. All fourteen of my descriptions run between 910 and 1,023 characters against a field median of 360 across the 501 in-scope skills that carry one, every one of mine above the 90th percentile of the field, against a 1,024-character cap that no competitor among the 508 exceeds. Two of those three figures were themselves wrong when first stated — a length ratio measured with two different rulers, and a rate published as one number when it was one of two defensible readings — and are corrected where they were recorded: 2.66x on a single reader, with all fourteen of mine above the field's 90th percentile.

What is open right now

  • No construction property may be published until its detector is pinned in code with its controls as tests. The measurement across the 508 field skills and my 14 has not run.

  • The outcome variable for that instrument does not exist. Until it does, construction is hygiene rather than a ranking, and the pre-registration commits to publishing that result either way.

  • Two findings from this cycle are held back. Both rest on a blind human ranking by one rater at n = 18, and the pre-registration gates any reading of them on a second independent sitting that has not happened. Seven of the cycle's nine findings do not depend on it.

  • Every closed list in the loop is now suspect after the verb-list failure. First in the queue are the two dead LinkedIn rules, because from outside a rule no arm satisfies looks exactly like a rule that cannot fire.

  • The instrument nobody has tried reads outcomes rather than text: replies and signups instead of wording. It is pre-registered and unrun, and the arithmetic is unforgiving. At a 3 percent baseline reply rate and 120 sends per arm, the smallest difference detectable at all is 13.4 percent, a 4.5x lift, and no copy edit does that; six replies against three at that size is p = 0.4994, which is chance. Ranking six arms at a 50 percent lift needs roughly 15,000 sends. Outcomes do not replace the rubric. They replace the crown, one pair at a time.

The track record · as of 7 September 2026protocol 0.1-L1 · every row traces to a committed artifact
01Generations99 of 99 recorded, append-only (runs/0.1-L1)recorded
02Judgments324 blind calls over 162 pairings, 0 unreadable against a 5% ceilingrecorded
03Gradesnine skills graded, then the transcript audit found the judge read position, not craftdisclosed
04Judge seatsthirteen measured, none qualifies; the pooled ensemble was refused before it was computedretired
05Benchmark Report 002the first grades, published with the audit that limits themreport 00206Benchmark Report 001the first cycle, told in full, generated examples includedfirst report07Benchmark Report 003the referee exam was ranking judges backwards; 432 blind callsreport 003
08Head-to-head, cold email105 generations, seven arms, four competitor skills pinned at their commits; ours leads by six rules and dominates none of themmeasured
09Benchmark Report 004thirteen seats fail, the ensemble is refused, and the first real head-to-head publishes a lossreport 004
10Head-to-head, LinkedIn posts99 generations, eleven arms, fourteen checkable rules from five published skills; the whole field lands inside a spread of twomeasured
11The ownership null, reversedzero violations published across 240 generations, then rescored to thirty when the detector was found unable to fire; twenty of them in the arm carrying the rulecorrected
12The field, mapped855 SKILL.md files read at pinned commits across eight repositories, 508 in scope, assigned to 21 clusters with zero unmatched; needs no judgemapped
13Benchmark Report 005the rules are a floor, not a ranking, and a zero this project published recomputes to thirtylatest report
Refusals publish here with the same prominence as results.nothing is crowned
The track record to date. Every row traces to a committed artifact, and the report rows link to the reports.HS-marketing-skills d0e19d1, runs/0.1-L1, judgments/0.1-L1, scores/0.1-L1, scores/0.1-N1, artifacts/head-to-head-cold-email, artifacts/linkedin-post-cycle, artifacts/ownership-rule and artifacts/gtm-jtbd-map/jtbd-directory.json
Frequently Asked Questions
What is the AI Marketing Skills Benchmark?
The AI Marketing Skills Benchmark is a fully agentic and autonomous AI Marketing project built and published by Hendry Soong. It measures whether loading an AI skill produces better marketing output than the same model without it, using blind pairwise judging, a placebo control, and grade bands fixed before any data existed. Every report ships with its run records.
What counts as an AI marketing skill?
A packaged instruction file, usually markdown, that an AI model loads before producing marketing work: post writers, comment writers, email writers, and in principle any other marketing job a skill is written for. The benchmark’s remit is every one of those jobs. The cycles run so far test instruction-style skills on two jobs, LinkedIn posts and cold email; other formats and categories get their own protocol versions, with their own briefs and their own bands, and are never scored on a number earned elsewhere.
Are there grades yet, and does anything rank skills?
There are grades, published with the audit that limits them, and nothing currently ranks two good skills against each other. Report 002 graded nine skills over 324 blind judgments and the required transcript audit then found the seated judge chose whichever post it was shown first in 90 of 90 disagreements. Thirteen seat measurements have been run and none qualified. The rubric built from rules that competing skill authors publish needs no judge model and works, but Report 005 measured it on a second job and it separates almost nothing: eleven arms inside a spread of two rules, with a bare model one rule off the top. It is a floor, not a ranking.
Has the benchmark ever published a result that was wrong?
Yes, and correcting it in public is the point of the track. Report 005 reverses a null this project itself published: a check reported zero misattribution violations across 240 generations, and that zero came from a detector that could not fire, because it matched achievement verbs against a closed list of 43 and two of the three briefs used verbs absent from it. Rescored, the same 240 drafts yield 30 violations, 20 of them in the arm carrying the rule the check exists to enforce. Six further figures in that report were wrong and are corrected in its own text with the original wording quoted.
Are the benchmark results comparable across protocol versions?
No, and the reports never mix them. A win rate under the local protocol 0.1-L1 measures lift over a bare local model, a weaker bar than the original frontier baseline, so each lineage keeps its own numbers and grades are always version-scoped.
Whose skills get tested?
The cycles cover the hendry.ai skills, published with names regardless of how they score — the set now runs to fourteen, and in Report 005 the lowest-scoring arm in the bracket is one of mine. Author-released competitor skills from the July pilot appear in aggregate or anonymized until their authors are contacted, per the protocol’s publication rules. Named competitor grades require consent first.
Can I reproduce a report’s numbers?
Yes. Each cycle’s run records, judgments, audit artifacts, and scores are committed append-only with pinned model identities and a published fixture, so every figure in a report can be re-derived from the records it cites. The protocol document itself publishes with the first graded results.
Published
###