Skills Benchmark Report 006: The Rubric Could Not See a Made-Up Name
The field's own rules could not tell a cold email that invented a podcast name from the same email in the brief's words. A rule that reads the brief finds invented names in 14 of 90 counted emails, including under the benchmark's own invent-nothing rule, and what this site published is corrected in five places.
A cold email in this benchmark invented a podcast name its brief never gave, and none of the 37 rules in the cold-email rubric could tell it from the same email in the brief's own words, because every rule the rubric scores reads the email alone. Report 006 adds a rule that reads the brief. Over the 90 counted emails of the cold-email head-to-head it makes 31 findings, 30 of them inventions in 14 emails, counted on the same emails the rule was corrected on. At least 12 of the 54 counted emails written under the benchmark's own invent-nothing rule carry an invented name, and the new rule reads 5 of those 12. What this site published is corrected in five places, each with its original wording quoted.
01In plain terms
Report 005 showed that the rubric this benchmark assembled from other authors' published rules works as a floor and not as a ranking. This report is about something that floor could not see at all: a writer making up a name. Given a fictional brief saying a prospect had spoken "on a trade podcast", one cold email in this benchmark wrote "I listened to your recent appearance on the TradeTalk podcast". The brief never named the podcast. None of the rubric's 37 rules could tell that email from a copy with the brief's own words put back, and not by accident: every rule it scored had to be checkable on the email text alone, and an invented name reads exactly like a real one until it is held against what the writer was given.
So the benchmark now has a rule that reads the brief. Run over the 90 cold emails of the head-to-head that its record counts, it makes 31 findings, and every one was read in its sentence: 30 are inventions, in 14 of the 90 emails. Those are the rule's counts, taken on the same emails the rule was corrected on, so they describe this corpus and are not a precision estimate for new material; a hand audit of three of the briefs finds seven more emails with an invented name that the rule cannot read.
The number that matters most is about this benchmark's own guard. The fictional company file carried by every prompt for the three briefs this benchmark wrote ends with a hard rule: outputs "may only use facts from this file and the brief". This project's record found no fabricated statistics in the emails written under that rule, and cannot say whether the rule is why. Those emails were not free of invented names: at least 12 of the 54 counted emails written under it carry one.
The rule's own checks failed more than once before it shipped: its acceptance test passed for the wrong reason, and an independent reviewer broke its first version on purpose. Those two failures are below, with what now prevents each. Passages this site already published are also wrong, in five places, and they are corrected at the end with their original wording quoted.
02What the rubric could not see, and why it was blind by construction
The cold-email rubric from Report 004 holds 37 rules taken from competing skill authors' published text: the 33 it scored, each checkable by a program on the email alone, and four it recorded without scoring. That scope is what made the rubric independent of us, and it is also exactly what made invention invisible. Whether "the TradeTalk podcast" is a fact or a fabrication is not a property of the email. It is a property of the email against the brief, and no scored rule was allowed to read the brief. The derivation record said so at the time: a rule against inventing tools, links and offers was rejected because it "[r]equires the allow-list of supplied URLs and offers from the brief", and a personalisation rule because "the program has no recipient facts to compare against".
The blindness is demonstrated rather than argued. Take the email that named the podcast and change one phrase back to the brief's own words: "the TradeTalk podcast" becomes "a trade podcast". The rubric scores the two emails identically, rule for rule. Neither is a clean email: both fail the same 8 of the 33 scored rules. The new rule fails the first and passes the second. That comparison is a permanent test in the benchmark's build, so the claim that no rubric rule could see this is checked on every verification run rather than stated once.
The email was found by reading the emails by hand while checking an outside review of this project, not by any instrument in the experiment. That is the strongest argument this report has for building one.
03A rule that reads the brief
The rule, invented_name, flags a name in an output that the writer's supplied material never gave it. Three decisions define it. The supplied material is everything the writer was shown as fact: the brief's fields and, on the three briefs this benchmark wrote, the fictional company file every prompt carried, because a name that the brief's audience line or the company file supplied was given to the writer, whatever the brief's list of facts says. A capital letter counts as a name only where capitalisation carries a signal: a camel-case word such as TradeTalk wherever it stands, except inside the label that opens a list item, or a capitalised word in running prose that does not start a sentence. And a name is invented when one of its words was never supplied, or when all of its words were supplied but never side by side.
That last basis exists because of a real case. A second email for the same brief wrote "the Trade Talks podcast". Both words are supplied: "trade" is in the brief, and "talks" is in the fictional company file's tone line, which says the company "sounds like a good dispatcher talks". A check that compared words one at a time passed it. The invention is the combination, and the rule now reports a capitalised pairing the supplied material never contains, unless one of its words is a name the material itself gave. "Dispatcher Board" passes, because "dispatcher board" is the brief's own phrase.
Acronyms, honorifics, the pronoun I, placeholders like [First Name], links and email addresses are never read as names. Subject lines, headings, table rows, lines written entirely in title case, and bold or italic text are not read as running prose, because capitals there are formatting; a camel-case name is still read in them. A brief with too little text to test against is reported as unscorable rather than passed.
04What it found in the cold emails
The cold-email head-to-head produced 105 emails: five briefs, seven arms and three draws each. Its record says one of the seven arms must not be counted toward any claim, so every count of invented names in this report is over the other six arms and their 90 emails. Over those 90, the rule makes 31 findings, one for each time a name occurs, and every occurrence was read in its sentence. Thirty are inventions, in 14 of the 90 emails. The one that is not is "Chief Technology Officer", written to a brief addressed to CTOs, which is the brief's own word spelled out. Those counts are taken on the same emails the rule was corrected on, as the next sections explain, so they describe this corpus and are not an estimate for new material.
The three briefs this benchmark wrote, whose company file forbids inventing: 54 emails; the rule reads an invented name in 5 of them, in 6 findings.
The two competitor-authored briefs, which carry no such rule: 36 emails; the rule reads an invented name in 9 of them, in 24 findings.
These are the rule's counts, and a finding is one occurrence, so a company the brief never supplied, named four times in one email, is four findings. A hand audit of the guarded briefs, described below, finds seven more emails with an invented name that the rule cannot read.
What gets invented is mostly who is writing. In the rule's findings, eight emails give the sender a name the brief never supplied, and the hand audit adds seven more; four give the sender a company the brief never supplied. Three emails name the podcast the brief left unnamed, three different ways, from three different arms. Two name a product the brief never mentioned, two name a conference the recipient supposedly spoke at, and one attributes a statistic the brief never gave to a real research firm it never mentioned.
By arm, reported only in aggregate, because this benchmark's publication rules publish weak results only as aggregate statistics without names: the rule reads an invented name in 4 of the 30 emails from the two controls, a bare model and a placebo; in 10 of the 45 from the three counted competitor skills, at least 17 with the hand audit's seven; and in none of the 15 from our own cold-email skill. That last number is not a win. Fifteen emails is not a ranking, and this rule sees names and not numbers: this project's record shows the same skill fabricated two statistics on the competitor-authored briefs, which no part of this rule can see.
05Clean of fabricated statistics, not of invented names
Every prompt for the three briefs this benchmark wrote carried a fictional company file that ends with a hard rule: outputs "may only use facts from this file and the brief", and "[i]nvented customers, numbers, or claims are a violation". This project's own record counted fabricated statistics across all 105 emails once before and found a clean split: twenty on the two competitor-authored briefs, which carry no such rule, and none on the three that carry it. That count stands. The reading taken from it, that an instruction binds when it sits in the brief and not when it sits in the skill, was never tested, because the guarded and unguarded briefs also differ in who wrote them, in their format and in how many facts they supply. For names, the guarded briefs were not clean.
On two of the three guarded briefs the rule finds nothing. On the third, the rule reads an invented name in 5 of its 18 counted emails: three name the unnamed podcast and three give the sender a name, one email doing both. A hand audit of every capitalised word in the guarded emails that the supplied material does not contain then found what the rule cannot read: 7 more guarded emails carry a first name the brief never supplied, alone on the line under the sign-off. So at least 12 of the 54 counted emails written under the guard carry an invented name, and the rule reads 5 of those 12. "At least", because that audit listed single words, and a name built only from supplied words outside running prose would not appear in it.
Why the podcast brief drew invented names in running prose, while the other two drew none the rule can read, is something this corpus cannot answer. The other two also leave things unnamed: a contract, and a service-status notice. Three briefs are not enough to separate the brief from chance.
06How much of this the rule can see, and how we know
Two numbers describe an instrument like this, and neither is an estimate for new material. Precision is how many of its findings are real: in the counted emails, 30 of its 31. That figure is counted on the same emails the rule was corrected on, twice: once after its first version's findings were read one by one and some were not inventions, and again after the independent review described below. So it says nothing about emails the rule has not seen. Recall is how many emails carrying a real invention it flags: on the guarded briefs, the hand audit puts that at 5 of at least 12 emails, or 6 of at least 13 invented names. The first honest precision estimate will come from the next corpus the rule runs on, with what counts as a false finding written down before any finding is read.
What the rule does not see, stated so it cannot be mistaken for a clean result:
a plain name that starts its line or its sentence, including a first name alone on the line under a sign-off, which is the recall miss above, and a name directly after a colon;
a plain name on a line written entirely in title case: a greeting such as "Dear Jane", a signature, a heading, a table row, or a list of names after a label;
an acronym, or a word inside the title-case label that opens a list item;
an invented number, which is a separate check this benchmark has and does not yet run on this scoring path;
an invention that names nothing, such as "your excellent episode";
a supplied name used in the wrong role, which is what Report 005's ownership check exists for, and only for numbers.
Each of those blind spots is pinned as a test in the benchmark's build, so a change that starts reading one of them fails until the rule's specification and its measurement are updated with it. Three of the pins, for invented numbers, inventions that name nothing and supplied names in the wrong role, were added only after a fact-check of this report found its first draft claiming tests that did not exist.
07The rule's own checks failed more than once before it shipped
The rule was built against an acceptance criterion fixed in the cycle's registration before the rule existed: its control had to fire on the real email, or the work was not to be approved. On the unchanged code the control failed, with the verdict "unknown check type"; with the rule it passes, with the verdict that "TradeTalk" is not in the brief. Across the 18 counted emails written for that brief, the rule flags 5, and every name it flags there is an invention. That is the part that worked. Other parts did not: one test passed on its first run for the wrong reason, and the mutation check described below found two behaviours no test covered. Two of the failures are described here, the first caught inside the build and the second by an independent reviewer.
The first failure was the acceptance test. It checked that the rule's verdict on the "Trade Talks" email mentioned "Trade Talks", and it passed. But the verdict quotes the sentence of every name it flags, and that sentence also held an invented sender name the rule did flag. The test was matching the quotation. The rule could not see "Trade Talks" at all, because both of its words are supplied, and the phrase basis described above exists because that test was rewritten to check what was flagged rather than what was quoted.
The second failure was the control meant to prevent exactly this kind of result. Report 005 published a zero that came from a detector that could not fire, and the fix it described was a positive control: before trusting a clean result, prove the detector can fire on this input. The first version of this rule carried one, and an independent reviewer with read-only access broke it on purpose. The control planted made-up names that no brief could ever contain, so it declared itself able to fire on a brief reading only "N/A", and tested nothing about the brief in front of it. An exemption for names whose initials spell an acronym the brief uses, added to spare "Chief Technology Officer", also passed a made-up test name, "Vanessa Park", on a brief that says "VP". The script that builds the measurement stamped every finding as an invention whatever its audit said. And the test file passed seven of the reviewer's single changes, each deleting or weakening one behaviour.
Each defect became a failing test before it was fixed. The control now builds one of its three probes out of the brief's own words, so a property of that brief can make it fail, and the control has a control: four deliberately broken detectors, blind to plain names, to camel-case names, to names built from supplied words, and to everything, must each be refused. The acronym exemption is gone, and "Chief Technology Officer" is reported and audited as not an invention, which is a visible cost instead of an invisible miss. The audit is now recorded per occurrence, so a changed or added sentence cannot republish unread. Twenty-seven single changes to the rule, its control and its builder, each deleting or weakening one stated behaviour, now each fail the tests, and the script that shows it is committed beside the measurement.
08What the loop ran, and what it did not
This project's goal includes a loop that keeps running with its owner only at approval gates. The benchmark's unattended driver carried its first cycle to the approval gate on 8 September; that cycle was committed and released on 13 September, and this rule is the driver's second. The cycle was registered by an earlier session, before the rule existed, with the exact paths it was allowed to touch and the criterion that would decide it. Once the rule was written, the driver took it from there with nobody between the steps: it checked that the work stayed inside those paths, ran the full verification gate, applied the records, staged exactly the declared paths, ran the gate again on the exact staged content, and stopped, holding a diff and a drafted commit. The commit was made at that gate, and the driver then released its locks. It does not push.
Two steps still sit outside the driver: registering a cycle, and the build. Cycles are registered by a working session, as this one was. The driver's build step may only run a node or python3 command, so an agent cannot yet be that step, and this rule was written in an interactive working session before the driver took over. Letting an agent be the build is a question of containment before it is a question of code, and it is on this project's list. The corrections below were not driven by the loop either, and no gate in the loop compares the live pages with the record yet.
09What this site published that was wrong
Every correction below changes a passage this site published, and the original wording is quoted so the correction can be checked rather than taken on trust. None of them changes a conclusion of the page it appears on. Two of them were recorded as wrong on 7 September, committed on 8 September, and stayed uncorrected on the site for the five days from that commit to this report, which is a failure in its own right. Two are contradicted by what this report measures. And one makes the same claim as a line in this project's own handoff notes, which stood from 5 September until it was corrected on 13 September while the site kept it, and as a line in its decision log, which is corrected with this report.
Report 004, the spread. It read: "Seven arms produced seven distinct rule sets, a spread of twelve rules against a maximum the rubric could express of nineteen." All three counts included the arm that the head-to-head's record says must not be counted toward any claim, though its pre-registration read the spread across all seven arms. Over the six counted arms, on the competitor-authored briefs those figures describe, there are six distinct rule sets and a spread of nine. The finding that the rubric separates real skills stands; its size was overstated. Report 005 repeated the count as "separated seven arms into seven different results", and is corrected with it.
Report 004, the sender-name rule. It explained why a competitor's rules requiring a postal address and a real sender name went unscored, and read: "No brief supplies either and the fixture forbids inventing facts, so scoring them would have failed all seven arms identically and rewarded whichever arm invented a person." For the address rule that still holds. Its claim that scoring the sender-name rule would have failed every arm identically is contradicted by this report: over the six counted arms, the rule above reads an invented sender name in 8 of the 90 emails, from 4 of the 6 arms, and the hand audit adds 7 more on the guarded briefs. Scored, that rule would have split the arms by which ones invented a person. The decision not to score it stands, and more strongly than stated.
Report 005, the ownership check's reach. It read: "Three of the five briefs in that set supply a metric at all, so 40 percent of the corpus cannot test this rule". The same report then added a positive control under which one of those three briefs can no longer be scored: its figures are the sender's own campaign numbers, not a result anyone could be credited with, and its proof point carries no verb the check recognises. With each brief passed as text, that leaves two of the five, so 60 percent of the corpus cannot test the rule, and the check reports 6 emails checked and 9 unscorable for every arm. More passages move with it. The passage after the quoted sentence gave the tool's count as "checked 9, unscorable 6 per arm, which is 40 percent", which was the count before the control. The report's closing account of its own audit names "a 60 percent that recomputes to 40" among six figures that did not survive, in its text, in a list and in its questions; under the control that figure recomputes back to 60, so five of the six were wrong, and the report's two sentences saying six figures were wrong are corrected with it. The skills hub restates that count twice, as "Five more figures in that report were wrong" and "Six further figures in that report were wrong". Both count beyond the zero that was one of the six, so the second was one too many even before the control; both now read four, and both are corrected with it. And the report twice described the brief's figures as "the reader's own current state", once without "own", when they are the sender's own campaign figures.
Report 005, the field rubric and invention. It read: "Every invent-nothing rule in the skill and in the field rubric passes this sentence, because both ask whether a fact is real and neither asks whose it is." The field rubric has no invent-nothing rule at all: every rule it scored reads the email alone, so none of them can ask whether a fact was given, which is the blindness this report measures. Our skill's rule does ask it. The point that no rule asked whose a fact is stands.
The skills hub, the frontier judge. It read: "A frontier judge has never been tried, and nothing measured so far closes that question." As written that is false. Before this benchmark moved to local models, a Claude Sonnet 4.6 judge gave 40 verdicts comparing three competing social-post arms with each other on 18 July 2026. On 28 July a panel of three frontier models, one each from Anthropic, OpenAI and Google, judged 90 pairings of one of our comment skills against a placebo, each seat judging both orders; the Anthropic seat was later dropped for changing its verdict too often when only the order changed, and that run's individual verdicts were not committed to the record. And in early August, blind LLM panels ranked skills against each other in three rounds of a champion loop, and the record does not say which model sat on them. What is true is narrower: no frontier judge has been tried as the instrument that ranks two good skills against each other under protocol 0.1-L1, against the answer key built from the field's published rules, and nothing measured so far closes that question. The same paragraph describes every instrument tried as "reference-free scoring of outputs, by small local models, against rubrics written here". The two rubrics were assembled here from rules other authors publish and are scored by programs rather than models, and the ownership check reads the brief, so what the instruments share is scoring outputs against rubrics assembled here, by small local models or by programs.
10What we are not publishing yet, and why
Three things are held back on purpose. The two findings from the blind human ranking of eighteen cold emails stay unpublished until a second independent rater has ranked the same set, as that sitting's pre-registration requires, and this report publishes no result from that ranking. It does quote emails from that set and say what they invented, and Report 005 quoted one and said what it got wrong, so the second rater must be someone who has read neither report nor any page, post or record that summarises what they say about those emails, a condition now recorded in that sitting's pre-registration. Invention results for competitor skills are published only in aggregate, because this benchmark's publication rules publish weak results only as aggregate statistics without names, and the arm its record does not count is left out of every count of invented names. And no precision figure for the new rule is offered for material it has not seen, because none exists yet; the in-sample figure above is not one.
11What changes
The rubric keeps its job as a floor, and it now has a second instrument beside it that can see what a text-only rule cannot: whether a name was given or made up. That rule is not added to the rubric, because the rubric is by construction a list of other authors' published rules and this rule is ours. It belongs beside the rubric as its own column, and giving it that column heads the invention items in this project's backlog. Its first run on emails it was not corrected on will be its first honest precision estimate, with what counts as a false finding written down before any finding is read. Three of the items queued behind it are: reading a name standing alone under a sign-off, which is the measured recall miss; bringing the existing check for invented numbers onto this scoring path; and deciding whether to narrow or supersede this project's own claim that an instruction in the brief binds, which the names above show does not hold for names.
The corrections expose one gap as well. This project's own rule already makes a report due whenever the published record and the committed one disagree, and the five days above are that rule not being applied. No gate performs that comparison yet.
12The record
The rule, its tests, its specification and the measurement were committed together at f752005: the rule in scripts/scout/claim_check.mjs; its tests, including the blindness comparison, in checks/test_invention_rule.mjs; its specification in specs/invention-rule-v0.1.md; and in artifacts/invention-rule/, the measurement over all 105 emails with a verdict on every occurrence, the acceptance test's before-and-after record, and the mutation check. The finding about the guard is recorded as K-049 in knowledge/claims.json. Commit c593235, the one that follows it, carries the three blind-spot tests added after the fact-check, the mutation check re-recorded against them (27 of 27 changes caught, 44 tests passing), the revised wording of K-041, K-047, K-049 and the specification, the second-rater condition in specs/blind-rank-sitting-prereg.md, and the corrections above. The emails themselves are carried verbatim, with a digest per file, in artifacts/claim-evidence-bundle.json. The fictional company, Slatebridge, is fictional everywhere it appears.
- Why could the rubric not see an invented name?
- Because every rule it scored had to be checkable by a program on the email text alone, and whether a name is a fact or an invention is not a property of the email. It is a property of the email against the brief. The cold-email rubric holds 37 rules taken from competing skill authors' published text, 33 of them scored, and it scores an email that invented a podcast name exactly as it scores the same email with the brief's own words put back, rule for rule.
- How often did the cold emails invent names?
- Over the 90 emails the head-to-head's record counts, the new rule makes 31 findings, one per occurrence of a name, and 30 are inventions, in 14 emails. A hand audit of the benchmark's own three briefs finds seven more emails with an invented name that the rule cannot read. The most common invention is who is writing: in the rule's findings, eight emails give the sender a name the brief never supplied and four give the sender a company it never supplied. Those counts are taken on the same emails the rule was corrected on, so they describe this corpus and are not an estimate for new material.
- Were the emails written under the benchmark's own invent-nothing rule free of invented names?
- No. The fictional company file carried by every prompt for the benchmark's own three briefs says outputs may only use facts from that file and the brief. This project's record found no fabricated statistics in the emails written under that rule, and cannot say whether the rule is why, because those briefs also differ from the others in author, format and how many facts they supply. But at least 12 of the 54 counted emails written under it carry an invented name, and the new rule reads 5 of those 12.
- What does the invention rule not catch?
- A plain name that starts its line or its sentence, including a first name alone under a sign-off; a plain name on a line written entirely in title case; acronyms; invented numbers, which are a separate check; inventions that name nothing, such as "your excellent episode"; and a supplied name used in the wrong role. Each of those blind spots is pinned as a test, so a change that starts reading one of them fails until the rule's specification and its measurement are updated with it.
- Which competitor skills invented names?
- Results for competitor skills are published only in aggregate, because this benchmark's publication rules publish weak results only as aggregate statistics without names. In aggregate, the rule reads an invented name in 4 of the 30 emails from the two controls, in 10 of the 45 from the three counted competitor skills, at least 17 with a hand audit's seven more, and in none of the 15 from our own cold-email skill. That last number is not a win: this project's record shows the same skill fabricated two statistics, which this rule cannot see.
- What did Report 006 correct?
- What this site published, in five places. In Report 004, the spread of rule sets, which counted an arm the head-to-head's record says must not be counted toward any claim and which Report 005 repeated, and a sentence saying a sender-name rule would have failed every arm identically. In Report 005, the share of its corpus that cannot test the ownership check, which is 60 percent rather than 40, with the passages that depend on it, and a sentence saying the field rubric asks whether a fact is real. And on the skills hub, the claim that a frontier judge has never been tried. Each correction quotes the original wording.