TLDR00 / 09

Run on a second job, the referee-free rubric does not separate skills: eleven arms and ninety-nine generated LinkedIn posts land inside a spread of two rules on a fourteen-rule bar, six arms tie on an identical rule set, and a bare model with no skill scores ten against a best of eleven. The other headline is a reversal of this project's own published result: a null of zero ownership violations across 240 generations was an instrument that could not fire, because the detector matched achievement verbs against a closed list of 43 and two of three briefs used verbs absent from it, so 160 of the 240 generations could not have failed whatever the model wrote. Rescored, the same drafts yield 30 violations, 20 of them in the arm that carries the rule.

01In plain terms

A bar built out of rules that competing skill authors publish about their own job is a floor, not a ranking. Report 004 assembled that bar to replace a judge model: a program can check those rules, and this project wrote none of them. On cold email it separated seven arms into seven different results. Run on a second job, LinkedIn posts, it does not separate anything. Eleven arms land inside a spread of two rules on a fourteen-rule bar, and a bare model with no skill scores ten against a best of eleven. A second result reverses a number this project already published. Two pre-registered runs, 240 generations in total, returned zero ownership violations, which looked like a model that never makes the mistake and was a detector that could not fire. Those same 240 drafts yield thirty violations.

Run again on a second job, LinkedIn posts, it does not separate anything. Eleven arms, ninety-nine generated posts, fourteen checkable rules, and the whole field lands inside a spread of two. Six arms tie on a byte-identical rule set. A bare model with no skill at all scores ten, and the best score in the bracket is eleven. Every skill here is worth about one rule more than using no skill.

The second finding is a reversal of something this project already published. Two pre-registered runs, 240 generations in total, returned zero ownership violations, and that zero was an instrument that could not fire. Rescored, the same 240 drafts yield 30 violations, 20 of them in the arm carrying the rule the check exists to enforce.

Six figures in this report were wrong. All six are corrected in the text below, with the original wording quoted.

02Eleven arms, ninety-nine posts, and a spread of two rules

The rubric does not separate LinkedIn skills. Eleven arms — nine skills and two controls — wrote ninety-nine posts across three briefs, three draws each, and a program scored every post against fourteen rules drawn from what competing skill authors publish about writing LinkedIn posts. No model graded anything. The top of the table passes eleven of the fourteen and the bottom passes nine, so the whole field fits inside a spread of two rules. Six arms come out byte-identical on which rules they pass. A bare model handed no skill at all passes ten, one rule below the leaders and one above the arm placed last. The bar can express more: on two reference posts written before any arm ran, one made to break the rules and one to follow them, it returned a gap of nine.

The run was pre-registered at specs/linkedin-post-head-to-head-prereg.md: eleven arms, three briefs, three draws each, ninety-nine posts, every rule scored by program with no model anywhere in the loop. Eleven arms is nine skills plus two controls, and this report says arms rather than skills throughout for that reason — counting the bare model and the placebo as skills would inflate the field being described by two.

The pre-registration claim is stated exactly as the document supports it and no further. That file says it was written "before the run's outputs were scored". It does not say it was written before any arm generated, and an earlier draft of this report said so anyway. On a report whose weight rests on pre-registration, the difference matters: "before scoring" rules out fitting the reading to the results, which is the thing that matters here, and "before generating" would additionally rule out fitting the briefs to the arms, which this document cannot support.

Rules passed, out of fourteen:

  • corey-social — 11

  • hendry-linkedin-post-narrative — 11

  • hendry-post-lean — 11

  • hendry-post-structured — 11

  • kostja-linkedin-posts — 11

  • openclaudia-linkedin-content — 11

  • openclaudia-social-content — 11

  • b0, a bare model with no skill — 10

  • b1, a generic instruction placebo — 10

  • hendry-linkedin-post-expert — 10

  • hendry-post-exemplar — 9

The top is a seven-way tie that includes a competitor whose repository carries more than 46,000 stars and three of our own skills. The bottom is one of ours, below the bare model. Across eleven arms there are five distinct rule sets, and six arms produce a byte-identical set of passes.

Why the spread is two and not nine

The obvious objection is that the rubric is broken, and it was tested against that objection before a single arm ran, on the rule this project fixed after the judge hunt: compute what the metric returns for the null case before you read it.

Two reference texts were written first, a deliberately rule-ignoring post and a deliberately compliant one. The rubric separated them by nine rules. It can express a spread of nine. The real field produced two.

The pre-flight also caught two defects in the instrument before they could contaminate a result. One check used a multiline regular expression and would have failed every post ending in a hashtag block, which the same rubric requires elsewhere. And the ideal reference text itself broke a rule, where the check was right and the reference was wrong.

So the instrument works and the field is flat. Of the fourteen rules, seven are inert because every arm passes them, two are dead because no arm passes them, and five discriminate at all. The two dead rules are worth naming: a 75-character line length, published by one author, and paragraphs of at most two sentences, published by another in two separate skills. Nothing any arm in the bracket produces satisfies either one, including the posts written by the skills of the authors who published them.

03What the field publishes about LinkedIn posts

The rules competing skill authors publish about LinkedIn posts are almost entirely rules about shape. Six source skills yield 99 rules that a program can check on the post text alone, and a hand classification made when those rules were derived puts 64 percent of them in the shape category: character counts, word counts, hashtags, emoji, line breaks, the see-more fold, formatting, links. About 10 percent touch content at all, and most of even that tenth is a forbidden-phrase list wearing content clothing — no vague attributions, no sycophantic tone, no generic positive conclusions, no questions in hooks. So the published standard for a LinkedIn post is a specification for how it should look rather than for what it should say, and the bar this report scores against is assembled out of that.

That shape-versus-content split is a judgement made when the rules were derived, and it is the one figure in this report that is a classification rather than a count. It is reported as such. What is counted, and what the conclusion rests on, is the scoring: seven inert rules, two dead, five that discriminate, on a rubric that separated its own reference texts by nine.

This is less extreme than cold email, where the same hand classification found no content rules in the union rubric at all. It is the same shape. Essentially nothing in what five independent authors publish, across six skills, tests whether a post is worth reading. A skill can top this rubric and still be dull. The rules are a floor any competent writer clears, so they separate a bad post from a good one and not good skills from each other.

A rule source that was not a LinkedIn source

Thirty-four of the 99 checkable rules, 34 percent, came from a skill for X rather than LinkedIn. Its frontmatter reads "Write long-form X (Twitter) posts and threads", it mentions LinkedIn zero times, and the rules derived from it included splitting a post into numbered tweets and putting ASCII art in code blocks so it renders in monospace on X.

The cause was ours. A scan workflow listed it as a LinkedIn source because it targets long-form posts, which is true and irrelevant, and the platform was never checked.

The corrected base is 65 checkable rules from four distinct LinkedIn authors across five skills. That is also a correction to how the sources are counted: the derivation record and the knowledge claim both said "five authors", and OpenClaudia contributes two of the five skills. Two skills from one repository are not independent sources, and independence is the whole basis on which a union rubric claims to be the field's bar rather than one author's opinion. The consequence is small and stated anyway: of the fourteen scored rules, the one that most depends on it is the two-sentence paragraph rule, whose two stated sources are those same two OpenClaudia skills, so it rests on one author rather than two. It is also one of the two dead rules, so nothing in the result turns on it.

The first draft of the corrected-base sentence said "five genuine LinkedIn authors" and so repeated the error eighty-three lines after correcting it.

Every number in the LinkedIn result is computed on the corrected base of 65. The figures describing what the field publishes — 99 rules, 64 percent shape, 34 percent from the wrong platform — are computed on the uncorrected base of 99 by construction, because their whole purpose is to describe what was harvested before the scoping correction removed a third of it.

This is the same class of error as a lead-research skill entering the cold-email bracket in Report 004, with one difference that matters: that one reached the published results table, and this one was caught before the rubric existed. The check that caught it is now the standing rule — read what a skill actually writes before it becomes an arm.

04A supplied fact with the wrong owner

A supplied fact can be used exactly as given and still make a false claim about the reader, because accuracy says nothing about who owns the result. On one brief our cold-email skill was handed "customers see 3x increase in content-attributed revenue within 90 days" and wrote a line crediting that increase to the reader's team. The figure was real and unaltered. Only its owner moved, from our customers to the reader. Invent-nothing rules miss it, because they ask whether a fact is real and never whose it is. The rule is now stated in the skill and enforced by an artifact-level check called metric_ownership, which flags one draft across the 105 generations of the cold-email head-to-head. A looser first version raised three flags across those same generations and two were false, a 67 percent false-positive rate; the rule was narrowed and both false cases are now regression tests.

"your team saw a 3x increase in content-attributed revenue in just 90 days. worth a look?"

Nothing was invented. The number was supplied and it was used accurately. What changed was the owner of the result: our customers became the reader's team. Every invent-nothing rule in the skill and in the field rubric passes this sentence, because both ask whether a fact is real and neither asks whose it is.

The skill now states the rule. That alone would not be enough. This project has already measured that a phrase added verbatim to a skill's banned list still appeared in 2 of 9 later drafts, so a ban written inside a skill does not bind. The rule is therefore also an artifact-level check, metric_ownership, which reads the brief, extracts the metrics it supplies, and fails any sentence that puts the reader in front of one with no first-person marker in the clause. Run across all 105 generations of the cold-email head-to-head, it flags one, and it is that one.

The instrument's own history belongs in the record. A looser first version of the rule flagged three, and two were false: a sentence whose subject was the vendor, and one whose subject was a peer. A 67 percent false-positive rate on flags is an instrument defect, not a finding, and this project has shipped five of those already. The rule was narrowed to require a second-person subject governing a verb of achievement, and both false cases are now regression tests.

The residual limits are stated rather than hidden. Three of the five briefs in that set supply a metric at all, so 40 percent of the corpus cannot test this rule, and a misattribution phrased without an achievement verb is not caught. The first draft of that sentence said two of five and 60 percent. The measuring tool prints its own accounting on every run — checked 9, unscorable 6 per arm, which is 40 percent — so this was a figure typed beside a number rather than read off it.

A fourth limit was not stated at all, and it is the one that mattered. One of those three briefs supplies figures that describe the reader's own current state rather than a result belonging to anyone else. There is nothing to misattribute, so its cells pass for free. Counting a free pass as a pass is what the next section is about.

One occurrence is a defect worth fixing. It is not evidence that our skill does this more often than anyone else's, and it is not reported as one.

05The zero that recomputes to thirty

Two pre-registered runs tested an ownership rule across 240 generations on the pinned local model, over two arms: the unchanged skill and the skill carrying the rule. This project published zero violations in all 240 and concluded that the failure is not inducible on this generator. The zero was not a result. It came from a detector that could not fire. The detector credits the reader with a result only when a second-person subject governs a verb of achievement, and those verbs came from a closed list of forty-three. Two of three held-out briefs build their proof points on close and recover, and neither verb was on the list, so 160 of the 240 generations could not have failed whatever the model wrote. Rescored with those verbs restored, the same 240 drafts yield thirty violations, twenty of them from the arm that carries the rule.

Two pre-registered runs, 240 generations on the pinned local model, both arms. The first version of this section reported what the scoring returned: zero ownership violations in all 240, and concluded that the failure is not inducible on this generator — a claim about the whole class of artifact-level rules, not just this one. Every supporting sentence was true as far as it went. In run 1, 111 of 120 drafts reproduced the supplied metric and 116 of 120 contained "you" or "your". Every ingredient was present. The model simply never made the mistake.

It made the mistake thirty times.

The check credits the reader with a result only when a second-person subject governs a verb of achievement, and the verbs come from a closed list of forty-three. The proof points of two of the three held-out briefs read "Sites we work with close their queries 18 days faster" and "Agencies using it recover 3.2 hours". Neither close nor recover was on the list. One hundred and sixty of the two hundred and forty generations could not have failed whatever the model wrote.

Rescored with those verbs restored, the same 240 drafts yield 30 violations. Verbatim, from the arm that carries the ownership rule, against a brief whose owner field reads "the sites the sender works with, not the recipient":

"hi there, your team closes queries 18 days faster on average."

That is the defect the rule was written to prevent, reproduced by the arm carrying the rule twenty times across the two runs, nine of them in almost exactly those words.

Violations by arm, out of 120 generations each:

  • v2, the unchanged skill — 2 in run 1, 8 in run 2, 10 in total

  • v3, the skill carrying the ownership rule — 9 in run 1, 11 in run 2, 20 in total

The direction is inverted from the rule's intent, and neither pre-registered run clears its own threshold: run 1 p = 0.0535, run 2 p = 0.6179, pooled post-hoc p = 0.0776. So this is recorded as a direction and not a finding, on the same convention this project applied to its own fifth-place finish. What it is not is a null.

The sentence that should have caught it was already in the report

"Every ingredient was present" verified that the metric appeared and that the second person appeared. It never verified the only thing that mattered, which is that the detector could fire at all. That is this project's own null-case rule — compute what the metric returns for the null case before you read it — applied one level too low. It was applied to the measurement and not to the instrument.

So the control now lives in the code rather than in a paragraph. A positive control builds the canonical violation out of the brief's own proof point and asks the real detector to catch it. A brief the control cannot fire on is reported unscorable and can no longer return a free pass. It costs one published denominator and no published conclusion: head-to-head brief T2 supplies figures describing the reader's current state rather than anyone's result, so its 21 cells were passing for free and are now correctly unscorable. Widening the verb list moved nothing on the published head-to-head — still exactly one flag, same arm, same draw — so recall was bought without spending precision.

The thirty falsifying drafts are committed at artifacts/ownership-rule/missed-violations.json, because they were generated into a gitignored directory. The evidence against the published null was not in the repository, and only the null was.

This is not the first instrument in this project to return a clean number nobody recomputed. Five separate defects each produced a fake zero on a single day in August, and this is the second to invert a published finding. Thirteen judge seats, two rubrics we assembled ourselves, and now a regex list. The pattern is not the oracle. The pattern is that an instrument gets scored once, by the person who wants the number.

What survives, and why the original conclusion still holds

Run 1 caught a defect in our own experimental design, and it is published rather than fixed quietly. Its briefs named the owner inside the proof sentence — "Our customers cut detention charges by 41 percent" — while the brief where the defect originally occurred uses a bare subject, so run 1 pre-solved part of the problem it was built to test. Run 2 was registered as a different question before it ran, with run 1 left standing, because re-running with easier briefs until the failure appears is fitting the test to the answer.

The original conclusion holds, now for a better reason. An artifact-level truth rule has to be tested where the failure occurs, on a strong generator on the sending rail. Not because the local model cannot produce the failure — it produces it constantly — but because a rule this project cannot yet measure reliably at home should not be certified at home. That is the sending partner's own design position, stated in their brief before any of this: their gates check the artifact, the skill only writes it.

06We measured our own house standard and it failed

Nine of our fourteen skills broke our own published skill spec, and the violations went unnoticed because for weeks the spec existed only as a written document. It is the house standard for how a skill is written: a description no longer than 1,024 characters, no process verbs inside that description, and a validation pass inside every skill. A program written this session turned those rules into a check, and at commit e037ea5 it failed nine of the fourteen skills. Every one of the nine had process verbs in its description, and three of those also ran past the character cap, at 1,256, 1,410 and 1,746. A governed version bump in the same session fixed all nine, and the gate now reports fourteen of fourteen to spec. On its first run the gate itself was wrong twice, both times flagging something that was already correct.

The spec caps a skill's description at 1,024 characters, forbids process verbs there on the recorded ground that a description summarising the method gets followed instead of the body, and requires a final validation pass inside every skill. It sat in the repository for weeks as a document rather than a gate, and the defect surfaced only because the owner asked whether the house standard had been applied.

Measured with a program written this session: nine of our fourteen skills violated our own spec. Past tense, and the tense is the point. The gate is in the build now, so the sentence a reader can check today is that the skill spec gate reports green at fourteen of fourteen skills to spec. The nine violations are what the gate found at commit e037ea5, and they were fixed under a governed version bump in the same session. All nine carried process verbs in the description. Three of the nine also ran over the character cap, at 1,256, 1,410 and 1,746.

And then the gate corrected the record it was built to enforce

Our own handoff notes said those nine skills have no validation pass. They all have one — seven to ten numbered items each, as the final section, under headings that simply do not use the spec's own words: "Check before you deliver", "Final pass", "Final pass, run it before you deliver".

The first version of the gate matched the literal phrase in a heading, reported nine skills as missing the section, and was one step away from a fix that would have appended a second checklist to skills that already had one. That is a real behaviour change wearing a governance fix, which is the exact thing this exercise exists to prevent. The spec specifies a section's role, not its title, so the check is now structural: a numbered list of at least four items in the last or second-to-last section, and no numeric self-score, which the spec calls validation theater because the model cannot compute the number it asserts.

The gate disclosed a second limitation on its first run. A whole-word verb lexicon cannot tell whose method a verb describes, and it fired falsely on "which LinkedIn enforces", where the subject is the platform. That case is adjudicated once, in a file, with the quote that justifies it, and the gate fails if the quote ever leaves the description.

A rule you have to remember is not a control. And a control's first job is to be checked against the thing it is measuring, because it can be wrong in the direction that makes you act.

07The field, mapped end to end for the first time

Eight admissible skill repositories have now been read end to end at their pinned commits and mapped onto a single 21-cluster taxonomy. Report 004 closed by saying the harvested entries had never been clustered, and this is that map. The field holds 855 real SKILL.md files, 508 of which do a go-to-market job, and all 508 were placed, with zero unmatched and zero non-canonical assignments. The weight of the field sits in search and owned web. Building site pages and blocks is the largest cluster at 60 skills, planning the growth motion has 44, and winning organic search rankings has 41. Writing a cold outbound message is the smallest real cluster at seven, fewer than the 13 skills that fit no cluster at all. Outreach is the thinnest thing this field does.

Every SKILL.md in all eight admissible repositories was read at its pinned commit and labelled with the job it does. Of the 855 real files, 508 serve a go-to-market job. Those 508 were assigned to a 21-cluster taxonomy built from three independent organising lenses and merged into one, with zero unmatched and zero non-canonical assignments.

The distribution, in skills:

  • Build site pages and blocks — 60

  • Plan the growth motion and campaigns — 44

  • Win organic search rankings — 41

  • Write and edit marketing copy — 36

  • Post and grow a social audience — 31

  • Lift conversion and capture leads — 29

  • Run paid ad campaigns — 27

  • Run lifecycle and retention messaging — 25

  • Get others to promote you — 23

  • Measure, attribute and test — 23

  • Fill and work the sales pipeline — 20

  • Research buyers, rivals and demand — 20

  • Earn coverage, links and listings — 18

  • Define positioning, ICP and brand voice — 18

  • Produce video and image creative — 17

  • Wire up marketing tooling and automation — 16

  • Decide pricing, packaging and offers — 15

  • Make the site crawlable and indexable — 14

  • Other, does not fit — 13

  • Get cited by AI answers — 11

  • Write a cold outbound message — 7

An independent count agrees: of those 508 skills, 12 write email copy and 5 write LinkedIn copy.

Both of this project's confirmed finds — a missing InMail writer, and a missing writer for the case where you have no case study — came out of that seven-skill cell. Each was found by scanning for one job. The map is the same finding at the level of the whole field, and it explains them rather than repeating them: this is a search and owned-web field, and it barely writes outreach at all.

Twenty-eight candidate unserved jobs were then proposed and searched for by an adversarial skeptic across all eight checkouts. None was found. That is reported as a candidate list and not as a finding, for two reasons. The first is that it is admissible at all only because the stage was validated: two positive controls were injected — a cold email citing a customer result, and a technical SEO audit — and both were refuted by both search methods, four times out of four. A refutation stage that never refutes anything is not evidence, and this one was made to prove it could. The second is that one search method is not enough, and the first skeptic demonstrated why: it found two skills an earlier pass had missed, because that pass grepped line by line and the recipient noun sat five lines above the subject line. A second, method-diverse skeptic is owed before any of these is called unserved in public.

A number we published was wrong: 1,313 is 855 files plus 458 symlinks

Report 004 closed by saying 1,313 harvested skill entries had never been clustered, and that figure has been carried in the project's own handoff notes ever since. Re-derived this cycle from fresh checkouts of all eight repositories at their pinned commits, 1,313 is 855 real files plus 458 symlinks. Every one of the 458 is in a single repository. Counting them as separate skills inflates that repository from 388 to 846 and the field from 855 to 1,313.

A first draft of this paragraph got the mechanism wrong, inside the paragraph announcing that a number was wrong, and the corrected version is the sharper finding. It said the repository re-exports its skills under four hidden directories. All 458 SKILL.md symlinks are under one of them alone; the other three hold zero. What those three hold instead is 1,115 symlinked directories, which a file-level walk never sees at all — so the traversal hazard is larger than the count being corrected, not smaller. And 73 of the 458 are not skills in any sense: they resolve to agent personas and slash commands, none of which is named SKILL.md.

Unique by content across the eight repositories is 843.

The correction changes no conclusion in Report 004. It is published because a number we put in print was wrong, and because the method that produced the wrong number is the one worth retiring: the original count came from a directory listing rather than from a file the counter had opened. Two of the eight repositories had already returned phantom directory trees from the GitHub API during an earlier sweep, including a skill that does not exist at the commit that supposedly contains it. Discovery now clones at the pinned commit and records the resolved SHA.

08Six corrected figures, and two findings held back

Six figures in Skills Benchmark Report 005 were wrong. Each is corrected in the article's own text, with the original wording quoted, so the correction is checkable rather than invisible. An independent audit recomputed every figure against the primary artifacts and found six that did not survive: a zero that recomputes to thirty, a 60 percent that recomputes to 40, and four more. An earlier draft ended with a sentence claiming every quoted figure had been recomputed from those artifacts rather than copied from an earlier summary. It is cut, because it was not true. Two findings are held back. Both concern a blind human ranking of six cold-email arms, and both rest on one rater at n=18. One is about whether the cold-email rubric predicts a human reader, the other about how our own skill ranked. Neither will be published until a second independent sitting lands.

An earlier draft of this report ended with the sentence "every quoted figure in this report was recomputed from those artifacts rather than copied from an earlier summary." It is cut, because it was not true. An independent audit recomputed every figure against the primary artifacts and found six that did not survive:

  • A 60 percent that recomputes to 40.

  • A four-directory symlink attribution that recomputes to one directory.

  • A present-tense spec violation that recomputes to zero on the committed tree.

  • A pre-registration claim stronger than the pre-registration supports.

  • "Five authors" reprinted eighty-three lines after being corrected to four.

  • A zero that recomputes to thirty.

The lesson is the one this report is already about. A closing sentence asserting that everything was recomputed is not a control. It is the same shape as the instrument that reported a zero it could not have escaped. What makes a figure trustworthy is a second reader who recomputes it and is allowed to disagree, which is what happened here, and is why these corrections exist rather than shipping.

Two findings held back

Two findings from this cycle are not published. Both concern a blind human ranking of six cold-email arms, and both rest on one rater at n=18. The pre-registration says rater-versus-rater agreement gates any reading of that correlation, and the second sitting has not happened.

They are named here so the omission is visible rather than silent. One is about whether the cold-email rubric predicts a human reader. The other is about how our own skill ranked. Neither will be published until a second independent sitting lands, whichever way it comes out.

Seven of the nine findings in this cycle do not depend on that sitting. Those are the seven reported above. Every one of them is recorded in the project's claims file with the artifact it rests on.

09What changes, and what is measured next

The referee-free rubric keeps the job it can do and loses the one it never had. As a floor it works: it catches a bad post, it caught a real defect in one of our own skills, and it costs almost nothing to run on every draft forever. As a ranking of good skills it is finished, and that is measured rather than argued, on two jobs now. Choosing a winner moves to outcomes instead, replies and signups rather than wording, one champion against one challenger at a time, pre-registered before any data exists. The next measurement leaves craft aside and asks about construction: fifteen properties of the skill package itself, derived without looking at our own skills and frozen before measurement, with nothing scored against that list so far.

The loop still cannot rank two good skills on craft. That is the honest state of it.

Outcomes replace the crown, not the rubric

The instrument nobody has tried reads outcomes rather than text: replies and signups, behaviour rather than wording. The arithmetic is unforgiving and is computed before any data exists. At a 3 percent baseline reply rate and 120 sends per arm, the smallest difference detectable at all is 13.4 percent, a 4.5x lift. No copy edit does that. Six arms at a 50 percent lift needs roughly 15,000 sends.

So outcomes do not replace the rubric. They replace the crown, one champion-versus-challenger pair at a time, and the pre-registration for that is written and committed at specs/outcome-loop-v0.1-prereg.md — including the number most likely to be misread when it arrives: six replies against three, on 120 sends per arm, is p = 0.4994, which is chance.

The next instrument measures how a skill is built, not how well it writes

The measurement that runs next asks an objective question instead of a contested one: how is a skill built. Its size, its context cost, how it declares when to trigger, whether it specifies a parseable output, whether it is coupled to one model or harness, whether it checks its own work. Every property is specified to be computed by a deterministic program from the bytes of the skill package, with no model judging anything anywhere in it. The list of fifteen was frozen before measurement and the detectors are not yet pinned in code, so nothing has been measured against it yet.

Construction is not craft, and the pre-registration says so in advance so the result cannot be read the other way. A skill can be immaculately engineered and produce dull copy, and nothing in this study will detect that. It is measured first for two reasons. Craft has now defeated every instrument tried so far, all of them one shape of the same idea, while construction has objective answers. And for an agent skill the ordering is not merely convenient: a skill that is never selected never runs, so its craft is unreachable.

The field is the 508 in-scope skills at their pinned commits, ours is the fourteen recorded by hash, and the reading is our position as a percentile of that field. Two of the three figures that motivated the study were themselves defective when first stated — one computed with two different rulers, one published as a single number when it was a choice between two defensible readings — which is the sharpest available argument for pre-registering the rest of it before anything is scored.

Frequently Asked Questions
Does the LinkedIn result mean marketing skills do not matter?
Not quite. It means the field's published rules are a floor rather than a ranking. Measured against them, every skill in this bracket is worth about one rule more than using no skill at all. Eleven arms produced ninety-nine posts, scored on fourteen checkable rules built from five published skills. A bare model with no skill scores ten of fourteen; the best score in the bracket is eleven. Six of the eleven arms tie on an identical rule set, including a competitor whose repository carries more than 46,000 stars and three of our own, and one of ours finishes at nine, below the bare model.
Why is a rubric that separates almost nothing not simply a broken rubric?
Because it was tested against that objection before any arm ran. Two reference texts were written first, one deliberately rule-ignoring and one deliberately compliant, and the rubric separated them by nine rules. It can express a spread of nine. The real field produced two. Of the fourteen rules, seven are inert because every arm passes them, two are dead because no arm passes them, and five discriminate at all. The two dead rules are a 75-character line length and paragraphs of at most two sentences, which nothing any arm produced satisfies, including the posts written by the skills of the authors who published them.
Why was a result this benchmark already published reversed?
Because the instrument that produced it could not have failed. The published result was zero ownership violations in 240 generations across two pre-registered runs, and the conclusion drawn was that the failure is not inducible on this local model. Rescored after the defect was found, the same 240 drafts yield 30 violations: 10 of 120 for the unchanged skill, and 20 of 120 for the arm that carries the ownership rule. The direction is inverted from the rule's intent, and neither run clears its own threshold, at run 1 p = 0.0535 and run 2 p = 0.6179, so it is recorded as a direction and not a finding.
What does it mean that the detector could not fire?
It means the check was scoring drafts it had no way to fail. The rule credits the reader with someone else's result only when a second-person subject governs a verb of achievement, and those verbs came from a closed list of 43. Two of the three held-out briefs had proof points reading "close their queries 18 days faster" and "recover 3.2 hours". Neither close nor recover was on the list. So 160 of the 240 generations could not have failed whatever the model wrote, and the zero they returned described the word list rather than the writing.
What did the map of 508 go-to-market skills find?
It found a search and owned-web field that barely writes outreach at all. Every SKILL.md in eight repositories was read at its pinned commit and labelled with the job it does, and 508 of the 855 serve a go-to-market job. Those 508 were assigned to a 21-cluster taxonomy with zero unmatched and zero non-canonical assignments. The largest clusters are building site pages and blocks at 60 skills, planning the growth motion at 44, and winning organic search rankings at 41. Unique by content across the eight repositories, the 855 files are 843 distinct skills.
Why does it matter that cold outbound is the smallest cluster in the field?
Because it says where an unserved job is likely to be. Writing a cold outbound message is the smallest real cluster in the field, at 7 skills of 508, smaller than the 13-skill bucket for skills that fit nothing. An independent count agrees: 12 of the 508 write email copy and 5 write LinkedIn copy. Both of this project's confirmed finds, a missing InMail writer and a missing writer for the case where you have no case study, came out of that seven-skill cell. The map explains those finds rather than repeating them.
How can someone outside the project check these results?
Every figure traces to a committed artifact, and the rules are scored by program with no model in the loop. The LinkedIn run was pre-registered at specs/linkedin-post-head-to-head-prereg.md: eleven arms, three briefs, three draws each, ninety-nine posts. Competitor skills are pinned at published commits, and every finding is recorded in knowledge/claims.json with the artifact it rests on. Nine of our fourteen skills violated our own skill spec when it was first measured; npm run verify now prints SKILL SPEC GATE: GREEN (14/14 skills to spec). The 30 drafts that falsified a published null are committed at artifacts/ownership-rule/missed-violations.json, because they were generated into a directory git ignores.
Why trust a benchmark that has just admitted it was wrong?
Because the corrections are the output rather than an apology attached to it. An independent audit recomputed every figure in this report against the primary artifacts, and six did not survive: a 60 percent that recomputes to 40, a four-directory attribution that recomputes to one, a present-tense spec violation that recomputes to zero, a pre-registration claim stronger than the pre-registration, "five authors" where the answer is four, and a zero that recomputes to thirty. All six are corrected in place, with the original wording quoted. The closing sentence claiming everything had been recomputed was cut, because it was not true.
Built by AI Marketing Operator · Published
Create-Articles v8.3.1
###