AI Marketing Skills, Measured
The hendry.ai AI Marketing Skills Benchmark: blind, placebo-controlled, pre-registered, and published with its receipts. This track carries its reports.
What this track is
This is the home of the hendry.ai AI Marketing Skills Benchmark, which I built and run under my own name, Hendry Soong. An AI marketing skill is a packaged instruction file that a model loads before writing, the format popularised by Anthropic’s Agent Skills, and the market for them runs almost entirely on faith: skills ship, get praised, get adopted, and almost never get measured. This benchmark measures them. Each benchmark report published under this track records one evaluation cycle: what ran, what it scored, what broke, and what changed, with the generated posts quoted verbatim and every number traceable to a committed run record. Reports are numbered and permanent; a refusal to grade publishes with the same prominence as a grade.
How a skill gets measured
Every skill is tested against the same fictional company on the same briefs, in three arms that differ only in one variable. The bare model gets the facts and the brief and nothing else. The placebo gets one fixed 12-word expert line on top. The skill arm gets the full skill text, loaded verbatim. A judge model then compares outputs pairwise, blind to which arm wrote what, with every pairing judged twice in swapped order; disagreement scores as a tie. Win rates against the bare model and against the placebo become grades, on bands that were fixed before any data existed and have never moved. The benchmark measures craft lift only, whether the skill makes the writing better, and never claims anything about conversions, reach, or revenue.
The judge itself has to qualify first. Before any grading, a candidate judge faces a perturbation audit: the same comparisons re-asked with the grading rubric reversed and the post labels renamed. A judge that changes its winner when only the names change is judging labels, and it is refused the seat. The architecture follows the published evaluation literature, from panel-of-small-judges results (PoLL) to the blind order-swapped pairwise design with bootstrapped confidence intervals (Arena-Hard). Where a small judge fails the audit, that failure is reported rather than smoothed over.
The state of the benchmark
As of 28 August 2026 the benchmark runs protocol 0.1-L1: generation on a pinned local model over a loopback-only runtime, with sampling parameters and model digests recorded before the first call. The first cycle generated all 99 posts of its matrix, nine of my own skills plus the two baseline arms across three briefs with three draws each, and then refused to grade them: both candidate judge models failed the perturbation audit, at 72.2 and 83.3 percent verdict stability against a 90 percent bar. Zero judgments exist, so zero grades exist, and no skill here is called best. The full story, with the generated posts and the flipped pairing shown whole, is in Benchmark Report 001 below.
Arm | What it gets |
|---|---|
B0, bare model | company facts + brief, nothing else |
B1, placebo | B0 plus one fixed 12-word expert line |
S, skill | B0 plus the full skill text, loaded verbatim |
- What is the hendry.ai AI Marketing Skills Benchmark?
- The hendry.ai AI Marketing Skills Benchmark is built and published by Hendry Soong. It measures whether loading an AI skill produces better marketing output than the same model without it, using blind pairwise judging, a placebo control, and grade bands fixed before any data existed. Every report ships with its run records.
- What counts as an AI marketing skill?
- A packaged instruction file, usually markdown, that an AI model loads before producing marketing work: post writers, comment writers, email writers, and similar. The benchmark currently tests instruction-style skills on LinkedIn-post briefs; other formats and categories get their own protocol versions.
- Why are there no grades yet?
- Grading requires a judge model that passed a pre-registered stability audit, and both candidates failed it: they changed verdicts when the post labels were renamed. The protocol forbids seating an unreliable judge, so the first cycle published its generations and its refusal instead. Grades follow once a judge qualifies.
- Are the benchmark results comparable across protocol versions?
- No, and the reports never mix them. A win rate under the local protocol 0.1-L1 measures lift over a bare local model, a weaker bar than the original frontier baseline, so each lineage keeps its own numbers and grades are always version-scoped.
- Whose skills get tested?
- The first cycles cover the nine hendry.ai skills, published with names regardless of how they score. Author-released competitor skills from the July pilot appear in aggregate or anonymized until their authors are contacted, per the protocol’s publication rules. Named competitor grades require consent first.
- Can I reproduce a report’s numbers?
- Yes. Each cycle’s run records, judgments, audit artifacts, and scores are committed append-only with pinned model identities and a published fixture, so every figure in a report can be re-derived from the records it cites. The protocol document itself publishes with the first graded results.