AI Marketing Framework: Eight Engines in a Governed Marketing Loop
A governed AI marketing architecture with eight engines, shared Context, explicit contracts, validators, and human gates.
The 2026 AI Marketing Framework connects eight marketing engines around governed Context. Orchestrate coordinates their contracts, Replicate tests portability, and Governance + Verify and Protocol + Interop apply across the loop. My work spans L2 to L4 across personal, multi-brand, and enterprise environments; public claims, canonical Context changes, CMS writes, and publication remain human-gated.
01Why architecture matters now
Architecture matters in 2026 because agents are arriving faster than anyone can govern their work. AI spending has outrun the operating capacity to use it. Only 30% of CMOs report being ready to scale. That is a readiness gap, not a funding gap, and no additional tool closes it. What closes it is one governed operating system for every agent to run inside. Owned inputs, declared handoffs, and a named person on every consequential output. The alternative is already visible. Capable agents optimise their own corner, hand off to nobody, and leave the result unattributable. The rest of this article is the shape of that operating system, and an honest account of how much of it currently runs.
The readiness gap
The model has eight numbered engines on a shared Context layer. Orchestrate and Replicate work across them, while Governance + Verify and Protocol + Interop run through every layer.
The 2025 thesis survives, but production exposed handoffs hidden by its three linear layers. The model therefore became a closed loop, while open protocols began covering the connective standard that the old edition could only describe. The frozen 2025 edition records that earlier model.
The martech catalogue barely grew in 2026 after a decade of expansion, but agents are multiplying on the same shelf. Each can still optimise in isolation.
Gartner reported marketing budgets at 7.8% of revenue. CMOs allocated 15.3% of that budget to AI, while 30% reported readiness to scale (Gartner 2026 CMO Spend Survey). Money has moved into AI faster than operating readiness.
How the architecture joins the tools
MCP and A2A can connect tools and agents. The Framework adds shared Context, explicit contracts, source records, and approval gates. Without those controls, each added service creates another handoff that someone must map, review, test, and maintain.
In May 2026 Harvard Business Review argued that siloed operating models constrain AI-enabled marketing (Taite, Winsor, and Fernandez, 2026). A governed marketing loop with explicit handoffs offers one operating answer.
I wrote this for the operator who also leads and remains answerable for the output. An enterprise may use the model to consolidate work, while a small team may use it to extend one operator. Portability provides the practical test: Context, contracts, validators, permissions, and gates should survive a move to another brand.
02The coordination challenge
The coordination challenge is that marketing work is split across services that never see each other's state. Content drafting sits in one tool, analytics in another, paid media in a third. Nobody owns what passes between them. The Framework's answer is to name those handoffs and attach them to engines. Create, Measure and Distribute each declare what they send and what they receive. An operator can then inspect who passed what to whom, and where it broke. Adding a ninth tool does not help, because tool count was never the constraint. The constraint is that the handoff between tools is nobody's artefact. Until somebody owns it, every improvement stays trapped inside the tool that produced it.
The 2026 census counted 15,505 tools, only 0.79% above the 15,384 recorded a year earlier by the State of Martech 2026 census, after a hundredfold rise since 2011. The census measures the catalogue, not coordination, shared state, or handoff quality.
Eighty-eight per cent of organisations use AI in at least one function. About 6% meet McKinsey’s two high-performer criteria: at least 5% EBIT impact attributed to AI and significant reported value (McKinsey, November 2025). Those high performers redesigned individual workflows more often than their peers.
That comparison links reported value to operating redesign rather than another isolated model or service. A connected workflow is therefore the practical unit, with Context, contracts, evidence, and gates making each connection inspectable.
This reading measures use in one or more functions. It does not show enterprise impact.
Attribute 5% or more of EBIT to AI and report significant value; this group redesigns workflows more often.
Gartner CMO survey, May 2026. Spend has moved faster than operating readiness.
AI use is widespread. High performance and scale readiness remain uncommon. Architecture separates access from operating capability.
03From three layers to a closed loop
The 2025 edition's three layers became one closed loop. Foundation, Execution and Optimisation read as a sequence, and a year of production showed what a sequence hides: shared inputs, manual coordination, and the route back from evidence to the next decision. A loop makes that route part of the model rather than an assumption about good behaviour. Two things fell out of the rework. Some boxes turned out to be inputs rather than actions, and belonged underneath the loop instead of inside it. The rest had to prove they deserved the name engine, which produced a test still applied to every candidate. The shape changed because production broke the old one. A diagram that hides a handoff is not simpler, only less honest.
Three engine boxes became Context
Define produced voice.md, Understand produced icp.md, and Position folded into messaging.md. Several production engines read those files before they act, so I moved the files beneath the loop as governed Context instead of pretending they were one-time steps.
An approved change to voice.md reaches the next content run without a new prompt. The source keeps an owner, review history, and drift checks, while each task receives only the runtime view it needs.
The verb test
That correction produced a reusable diagnostic. For any proposed engine, name its verb, the concrete artefact or state change it leaves behind, and the contract that separates it from the boxes beside it.
Box | What does it do? | Verdict |
|---|---|---|
Create | Produces validated content artefacts under a declared brief | Engine |
Listen | Detects market signals | Engine |
Optimise | "It optimises" | A label. The real verb is Iterate: read Measure, propose Context updates |
Nurture | Sends, and also converts | Two verbs. Sending went to Distribute, conversion to Convert |
Define, Understand, Position | Produce the governed voice.md, icp.md, and messaging.md files | Governed Context files with named owners, review history, and several downstream engine readers |
Grow | No verb, no artefact, no evidence | Removed |
The test reduced eleven boxes to eight marketing engines. It exposed files posing as steps, two verbs sharing one name, and one box with no owned output at all.
Data belonged on the map
I had treated clean data as pre-Framework work, which hid fragmented inputs from the architecture. Data is now engine 00, with a designed contract for normalisation and identity resolution and a named owner for its handoffs.
Start with an inventory of the records you already hold, including their owners, update dates, and cross-system identifiers. Add a pipeline when a real signal requires it, rather than building a warehouse before you can name the first useful handoff.
Orchestrate names the missing coordination layer
I also routed work between engines by hand, yet coordination had no box. Orchestrate now names that partly built layer and chains explicit handoff contracts between bounded engines.
Each engine receives one job, a small toolset, and relevant Context; I use the simplest structure that can finish the work. Production also exposed the missing return path from Measure to Context, which changed the target from a hand-built line into a governed loop. That keeps coordination visible without turning every handoff into another engine.
The wider operating scope
Prompt engineering shapes the instruction for one model call. Context engineering works at the next scope: the work of curating the right information for the model at each turn (Anthropic, September 2025).
Loop engineering covers a bounded task, and graph engineering connects work, state, evidence, and gates across a workflow. Harness engineering names the governed environment around all four scopes.
Addy Osmani summarised the practical effect: a strong surrounding system can outweigh model choice (Osmani, April 2026). LangChain reports moving a coding agent from the top 30 to the top five on a public benchmark by changing that surrounding system rather than the model (LangChain, March 2026).
The operating system in practice
Hendry Context holds approved truth, constraints, and provenance. Each engine receives a governed runtime view through a contract, while Orchestrate manages the work-and-control graph between engines.
A bounded loop may run inside one engine. Verify already checks article artefacts; Measure will read business outcomes and counter-metrics, and Iterate can turn that evidence into a proposed Context change for an authorised owner.
The workflow runtime executes those relationships without inheriting authority. For this article, a human enters every external number and the validator reads its prose rules from shared/context/voice.md, preserving the same controls when the model changes.
Four design scopes inside one governed harness
call · turn · task · workflow04The engines
The model has eight numbered marketing engines, a shared Context layer, and two system engines, Orchestrate and Replicate. An engine is a contract, not a product. Each declares four things: the input it takes, the verb it performs, the output it leaves behind, and the route that output travels next. A vendor tool may implement a contract, and several tools may implement the same one, but no tool defines the engine. That is what keeps the vocabulary stable while the software underneath it keeps changing. It also decides what counts as an engine and what is only a label on a box. Numbering the eight is not a running order. It is a way to say which part of the work you mean, and who owns it.
One architecture across several environments
Since the 2025 edition, I have built these contracts across personal, multi-brand, and enterprise projects. Hendry.ai is the public record I control; work across several environments shaped the Framework.
The enterprise work includes GTM workflows, Data and Signal pipelines, governed Context, an enterprise website, and Create + Verify. Its evaluation layer exercises Measure, while a marketing team now uses the mktr MCP path through Teams, Copilot, and Claude.
The next build stream: reusable skills
A separate workstream packages marketing practice as reusable skills for Claude, Codex, and Gemini. Its internal evaluation system checks whether a skill follows its own claims, adds value beyond a plain instruction, and behaves reliably across jobs.
That work is still in development. The planned Hendry.ai section will organise finished skills by the job a reader needs to complete.
The longer-term target is autonomous execution inside that section: you choose a job, and the system runs and verifies a finished skill. Explicit permissions and stop conditions still define what it may change or publish. Its final name is still open.
Nothing signs a skill file
Distribution is the unsolved part, and it is unsolved for everyone. A security audit of one public skills registry found 341 malicious entries among the 2,857 skills it examined, most of them from a single coordinated campaign disguised as popular utilities. No signing or attestation mechanism has shipped for skill files, so a registry listing carries no evidence about what the file actually does when it runs.
A packaged capability is executable text arriving from somewhere else, which puts it in the trust class of a dependency rather than a document. That sets the ceiling on how any skills surface can be distributed.
My control for it is unglamorous and manual. Skills move one way, carried by me, audited before adoption, with a digest binding the version that was reviewed. It does not scale past one person, and it is the strongest control available at this layer today.
Workstream | Framework expression | Evidence boundary |
|---|---|---|
Hendry.ai content system | Context, Listen, Create, partial Convert and Orchestrate | Public receipts link to maintained artefacts and versioned operating records |
Enterprise GTM and data | Data, Listen, Signal, and Orchestrate | Private delivery records; capability is described without employer names, payloads, or internal results |
Enterprise website programme | Create, Distribute, Convert, Governance + Verify | Private delivery records; the project name and internal implementation details are omitted |
Grounded content and evaluation | Context, Create + Verify, and Measure | Documented privately, with transferable methods separated from employer systems and payloads |
Marketing skills workstream | Reusable skills inside Create, with Verify and Measure checks; a later Hendry.ai section will test Replicate | Private work in progress; the section name, finished-skill set, named grades, proprietary prompts, and evaluation payloads are omitted |
MCP agent used by a marketing team | Context, Protocol + Interop, Distribute, and Replicate | mktr is public; the organisation, private Context, and usage data are omitted |
How the taxonomy changed
Define, Understand, and Position became Context sources. Nurture split because sending belongs to Distribute and response handling belongs to Convert; Grow disappeared because it had no discrete work, owned output, reviewable evidence, or downstream handoff.
Data now owns normalisation, Signal maintains scoring state, and Orchestrate coordinates contracts. The deepest public Hendry.ai evidence remains Create and Listen, with partial Convert and Orchestrate work recorded in its publication ledger.
That ledger controls direct public proof. It does not grade my broader capability across working environments, and private work appears here only at a de-identified capability level without employer names, payloads, or internal results.
Create has the deepest production evidence: 66 of the 116 Operator Log entries, across four sub-engines for articles, images, compilation and social, alongside 81 extracted principles as of June 2026. The remaining entries belong to other engines and to cross-cutting architecture work. Those receipts show how the engine changed, including failed approaches rather than only its current state.
Each engine is a maintained folder of instructions, scripts, and rules. Only job-specific material loads.
Anthropic calls this Agent Skills with progressive disclosure (Anthropic, October 2025). A skill packages repeatable capability inside an engine; it does not become a ninth marketing engine.
A run also starts from a brief. That mirrors spec-driven development and its rule that intent is the source of truth (GitHub, September 2025). The engine holds the capability, while the brief sets the current job and its limits.
05How the system connects
The eight engines connect because the loop closes on Context, and a person closes it. Measure gathers business evidence. Iterate turns that evidence into a reviewable proposal. Context changes only when an authorised owner approves that proposal. That single rule is what makes this a system rather than several tools pointed at the same brand. Every automated step gets one place to read from and one guarded place to write to. It also keeps disagreement traceable, because a changed output leads back to an approved change and to the person who approved it. Speed comes from the loop running unattended between gates. Safety comes from the one gate the loop cannot open by itself. The layers described below are how that rule gets enforced.
How the operating layer runs
Hendry Context is the governed source for approved brand truth, operating constraints, provenance, ownership, and change history. Runtime context is the relevant view assembled for one job. Drift checks compare that view with its owned sources.
A bounded agent loop acts, observes, verifies, and stops under a declared budget. A workflow graph coordinates those loops with deterministic steps, checkpointed state, and human gates.
IBM published its Loop Engineering explanation on 17 July 2026. Its goal, action, observation, and adjustment cycle describes the bounded inner execution pattern used inside a node. The date records chronology, not ownership of the idea.
Graph engineering is broader than LangGraph
I use graph engineering for explicit topology among work, state, loops, evidence, budgets, and gates. The wider label remains unsettled. LangChain's own account says the label surfaced in July 2026 while the underlying graph approach is established.
Graph view | What it connects | Question the relationship must answer before routing |
|---|---|---|
Work and control | Tasks, agents, validators, gates, state, and allowed transitions | What may run next? |
Context and knowledge | Named entities, external claims, source records, provenance, and the semantic relationships among them | What is known and relevant? |
Improvement | Operating loops, evaluators, outcomes, counter-metrics, and audits | Which operating loops exchange outcomes, counter-metrics, constraints, approvals, or vetoes with one another? |
Evidence and trace | Checkpoints, versions, artefacts, approvals, and terminal reasons | What was recorded, and which declared reason routed it? |
These four graph views answer distinct operating questions and should not be collapsed into one data model. Microsoft GraphRAG extracts entities, relationships, and claims for retrieval. I classify that as a context-and-knowledge use case.
Carlos Perez develops the current graph-of-loops interpretation: operational, quality, business, and governance loops can feed or constrain one another. Outside evidence anchors those relationships. The design does not require a graph database; Hendry Context remains governed files with explicit source and artefact lineage.
LangGraph is one runtime adapter
The eight-engine ring is the marketing-domain loop. A bounded execution loop may run inside one engine, while Orchestrate owns the work-and-control graph and Verify accepts or rejects artefacts and transitions.
Versioned evaluations compare capability across runs; Measure reads business outcomes and counter-metrics; Iterate proposes learning. A graph can invoke a human gate. Owners and the permission system keep authority.
Official LangGraph documentation describes shared state, nodes, edges, loops, checkpoints, and resume behaviour. The LangGraph overview covers durable execution and human review, while LangGraph supplies execution mechanics around the owned controls.
Hendry Context remains authoritative across every runtime adapter. Its research, validators, permissions, and human approval therefore stayed outside the LangGraph pilot's control. The Build Log 002 pilot routed deterministic nodes, bounded loops, checkpoints, and gates without acquiring new authority or redefining graph engineering.
The loop on a real job: a competitor moves
Suppose a competitor changes its pricing. That is a composite job, in the jobs-to-be-done sense from Part 1. The target loop divides it into smaller contracts.
Listen catches the move, and Signal scores its materiality and exposed accounts. Create drafts a comparison page, sales notes, and a positioning change. A person approves the response first.
Convert can assemble the landing page around the visitor. Measure then reads outcomes and counter-metrics, while Iterate drafts an evidence-backed change for messaging.md. Another approval must precede any Context update.
In the public Hendry.ai implementation, Listen and Create run fully, while Convert and Orchestrate are partial. Other implementations exercise Data, Signal, Measure, MCP distribution, and cross-environment replication, so I close a missing handoff manually rather than pretending one environment contains the whole loop.
The ratchet turns failure into a check
The ratchet turns a verified failure into a durable rule or review step. Osmani argues that each durable AGENTS.md rule should trace to a specific failure (Osmani, April 2026). The Operator Logs preserve the incident that justified each control.
This article supplied a recent example. An earlier draft passed every word-level check and still sounded mechanical. I added six cadence and specificity rules to voice.md, including the executable PAT-022 check below.
PAT-022 Sentence-length variance
No more than two consecutive sentences within +/- 3 words.
AI clusters at 14 to 22 words. Break the metronome.PAT-021 and PAT-026 still require human review because a pattern count cannot decide whether a paragraph sounds natural. The Context file now records both the automated check and the manual craft standard.
The front half stacks signal so the back half can personalise
Data supplies normalised records, and Listen detects what changed. Signal adds state: a decaying score for each account, enriched by each new detection. That accumulated record becomes the input for personalisation.
Signal remains designed in the Hendry.ai implementation, so Listen carries some scoring there. Enterprise Data and Signal work supplies the practical boundary: once the public implementation adopts it, Create and Distribute can receive an account, buyer, and moment instead of a generic brief.
Detection and generation stay separate because a threshold monitor behaves differently from a validated content template. Signal joins detection to the next action.
The external dimension: GEO
Your brand now needs accurate representation in ChatGPT, Perplexity, Gemini, Claude, and similar systems. The 2025 edition called this AIO. I now use GEO, or generative engine optimisation, while AEO also remains in use.
Google says AI Overviews reaches more than two billion monthly users. Bain reports that about 60% of searches on traditional search engines now end without the user moving to another website. Those figures make AI answers a measurable discovery surface.
The instruments arrived this summer. Google launched generative-AI performance reports in Search Console in June 2026, and the report gives impressions inside AI features and no clicks, no click-through rate, and no position. Impressions count appearances and stop there.
The IAB published Measuring Visibility in the AI Era on 3 August 2026, which sorts the question into presence, prominence, portrayal, and persuasion, and then does the more useful thing: it separates decision-grade measurement from directional measurement and sets disclosure requirements for the vendors selling it. More than twenty tools now measure AI visibility with methods that disagree with each other. A shared vocabulary is what makes their answers comparable.
Listen checks how the brand appears in those answers today. Until Measure is built, I track AI presence, share of model, and citation share manually, and the gap between those two sentences is mine rather than the industry’s now that the primitives exist. The target Measure engine will own those metrics alongside human clicks, as described in The AI Marketing Measurement Problem.
The closed loop
8 ENGINES · 1 CONTEXT LOOP06The two spines: Governance + Verify and Protocol + Interop
Two concerns refused to sit inside any single engine. The model draws them as spines, running the length of the loop. Governance + Verify asks whether an output is allowed to stand, and who signed for it. Protocol + Interop asks how one engine reaches another without a bespoke integration each time. Orchestrate is designed to enforce both at the boundaries between engines, which is exactly where they usually go missing. Neither spine is finished. The first is partly built and already operating on real work. The second is still mostly design. That asymmetry is stated rather than smoothed over, because a spine described as complete when it is not is the failure it was drawn to prevent.
Governance + Verify
Built controls include deterministic validators, edit boundaries, and approval gates. An unsourced external claim fails, and a human enters every number. Those controls already run on article artefacts.
The target design adds a write-protected evaluator distinct from the existing article validator. It will remain fixed during each declared comparison but may be versioned between runs. Circuit breakers will cap spend and send velocity.
External MCP or A2A access needs identity, a trust registry, and a guardian for unknown agents. Provenance must travel with each input and output. Those are design requirements, not claims about current production coverage.
Two of them stopped being aspirational while I was writing this edition, and both close a specific hole rather than adding a policy. MCP’s authorization specification now requires a server to reject tokens that are not audience-bound to it, and forbids passing a client token downstream, which shuts the confused-deputy path.
A2A v1.0 added Signed Agent Cards, cryptographic verification of an agent’s identity before interaction across organisational boundaries. Nothing I run consumes either one today. The primitives shipped, and the adoption is mine to do.
Simon Willison named the lethal trifecta in June 2025: private data, untrusted content, and outbound communication create a dangerous combination. I treat separation of those capabilities as a build rule.
OWASP publishes a Top 10 for agentic applications. Its Top 10 for LLM applications lists prompt injection as LLM01 (OWASP, 2025). Governance therefore belongs in engine contracts and tool permissions, not only in policy copy.
The legal edge
Article 50 of the EU AI Act sets transparency duties for specified AI interactions and generated or manipulated content. Those duties became applicable on 2 August 2026, with a backstop of 2 December 2026 for Article 50(2) marking and detection on systems already placed on the market. Both dates now sit inside release controls rather than on a roadmap.
The Commission published the final Code of Practice on Transparency of AI-generated Content on 10 June 2026. It is the readiest specification to design a release against, and about 190 organisations had signed by the end of July. It covers Article 50(2), (4), and (5), asks providers for machine-readable marking, and gives deployers a published icon set for visible labels.
In the United States, the FTC's fake-reviews rule covers AI-generated fakes and permits civil penalties (FTC, 2024). California's AI Transparency Act became operative on the same day as the EU duties, and it reaches image, video, and audio rather than text, and only providers above one million monthly users. What it names plainly is the failure mode: from 1 January 2027, a large platform may not knowingly strip provenance data from the content it distributes.
The C2PA specification defines an open standard for recording content provenance and authenticity. These controls belong before publication.
Protocol + Interop
The protocol spine carries Context and actions between tools and agents. On 9 December 2025 the Linux Foundation formed the Agentic AI Foundation, anchored by Model Context Protocol, goose, and AGENTS.md. That moved important connective projects into neutral governance.
Google's Agent2Agent protocol remains a separate Linux Foundation project. The Linux Foundation launched the Agent2Agent Protocol project under its own open governance in June 2025. No single vendor owns both protocol paths.
Easier protocol connection makes governance, portable Context, and explicit permission boundaries more important to each operator. I keep the approved sources and portable history outside any runtime adapter. Orchestrate moves documented contract traffic without owning authority.
Why not buy one all-in-one marketing platform?
An all-in-one marketing platform bundles creation, analytics, automation, and delivery from one vendor. It can simplify setup, contracting, and support when its workflow and data rules fit the team.
I make a different trade-off. I keep the Context layer and the Operator Logs in governed, exportable artefacts outside any vendor platform.
Explicit contracts let the same sources serve different tools, so I can replace one tool without rebuilding the system’s operating memory. This is a control preference, not a claim that every all-in-one product traps data.
When agents become readers and buyers
Andrej Karpathy names the deeper shift in his original 2025 talk on software in the age of AI. Agents now read content alongside humans and conventional software. Formats such as llms.txt give that second audience a direct, machine-readable route to approved product information.
Your content therefore has a second audience. Product facts, evidence, usage terms, ownership, and update dates need a machine-readable structure that agents can inspect. The protocol spine serves that material without changing its owner.
Transaction infrastructure is also emerging. In April 2025 Visa announced Intelligent Commerce, Mastercard announced Agent Pay, and PayPal released agentic-commerce developer tools.
OpenAI and Stripe followed in September 2025 with Instant Checkout in ChatGPT and the open-sourced Agentic Commerce Protocol (OpenAI, 2025). In April 2026, Google contributed its Agent Payments Protocol to the FIDO Alliance. FIDO now develops it through an open, community-led standards process (FIDO Alliance, 2026).
The Visa, Mastercard, PayPal, OpenAI, and FIDO releases document APIs and protocols for agent discovery, comparison, recommendation, and payment. Adoption remains an open measurement question.
A useful near-term test asks whether an agent can find, compare, and recommend your product. A later acceptance test covers the agent that completes the purchase.
Make the product feed, claims, and evidence machine-readable now. Add autonomous purchase flows only when the protocol, liability, and review model settle.
Open protocols reduce connection work. Keep Context, permissions, and evidence portable and under accountable ownership.
07The autonomy progression
Autonomy is set per engine, never per company. Listen can scan the market continuously, because a wrong reading costs a second look. Create must stop for a person before publication, because a wrong claim is public. In this system the setting ranges from L2 to L4, engine by engine. The ceiling is drawn by consequence, not by confidence in the model. One company-wide autonomy number hides that spread, and usually flatters it. Autonomy is also only one axis of three. Each engine is read alongside how well it interoperates and how much governance covers it. Raising autonomy without verification buys speed and no assurance. The progression is uneven on purpose, and the gates stay where the consequence is.
Autonomy is one operating axis
The 2025 edition framed autonomy as a climb from L1 to L5. Production made that framing too simple because engines advance at different rates. The rungs still help describe who decides.
Andrej Karpathy calls this the decade of agents (Dwarkesh Patel podcast, October 2025). He favours partial autonomy with verifiable work. Code editors already expose that range from autocomplete to full agent mode.
My current work spans L2 to L4. Bounded local work reaches L4 because agents propose and act within repository, permission, retry, and validation guardrails while I monitor and spot-check the result. Public claims, canonical Context changes, CMS writes, publication, deployment, and permission changes still stop for explicit approval.
Where bounded L4 is justified
Fast, measurable feedback can support more autonomy in bounded optimisation. Braze says its Decisioning Studio is generally available and uses reinforcement learning for continuous, individual decisioning.
Google says AI Max is out of beta and the automatic Dynamic Search Ads upgrade begins in February 2027. These are product examples, not evidence that every marketing engine belongs at L4.
Governance must rise with autonomy. Thoughtworks calls before-action controls guides and after-action controls sensors (Bockeler, April 2026). Context rules guide the run; validators and human review sense failures afterwards.
METR now measures the 50%-reliable task horizon on a rolling basis against its 228-task Time Horizon 1.1 suite. The leading model sat at about five hours when that suite launched in January 2026, and the doubling time across models released since 2023 is roughly four months.
Read the limitation before the headline. METR states it directly: a 50% horizon of X hours does not mean tasks under X hours can be delegated, because much production work needs far higher reliability, and horizons at 99% cannot be fitted with the benchmarks that exist. Longer horizons make gate placement matter more; they do not remove gates.
Where the system is now, gaps included
The demonstrated L4 boundary applies only to bounded local work; it does not describe one complete eight-engine deployment. Hendry.ai publishes its deepest direct evidence for Create and Listen, while private implementations appear only as de-identified capability records.
Measure must produce evidence, and Iterate must convert it into a governed proposal. An authorised owner approves any change to Context. Drafting and validation stop before each consequential transition.
Current quality gates do not yet screen external inputs for prompt injection. That defence in depth is required before engines read untrusted content, hold private Context, and publish outward with more autonomy.
The Operator Logs already record provenance. Step-level traces must next show what each agent read, called, and produced. Versioned evaluations can then run on every governing-document change instead of only on demand.
A bounded loop in practice
Karpathy's autoresearch demonstrates a useful boundary. An agent changes a training script, runs for five minutes, and checks a fixed val_bpb metric. It keeps or discards the change under that budget.
The agent operates only inside a sandbox with a fixed budget, permitted files, and a person-defined stop condition. It may alter training code while the evaluator remains fixed. That separation makes experiments comparable.
Verify follows the same rule during a declared run: the acting loop cannot rewrite its evaluator. Versioned evaluations compare capability between runs. The evaluation owner may recalibrate them between declared comparisons, with the change recorded before another run begins.
Measure has a different job because it reads external outcomes and counter-metrics. It sends that evidence to Iterate. Separating internal evaluation from business evidence reduces opportunities for self-grading.
OpenAI describes an internal beta with daily users and external alpha testers built entirely by coding agents. The reported repository reached roughly one million lines, with three engineers directing agents over five months (OpenAI, February 2026). Humans still set the boundaries and reviewed the system.
Evidence sets the autonomy limit
State of Martech 2026 found the same adoption order across measured categories: analytical, then generative, then autonomous. The authors describe a trust gradient. Fully autonomous action currently remains the least adopted.
McKinsey's State of AI Trust in 2026 found nearly two-thirds of surveyed organisations name security and risk as the leading barrier to agentic scale. That ranked ahead of regulation or technical limits. I therefore require evidence and a safe stop before raising an engine’s autonomy.
The same survey found that responsible-AI maturity improved only slightly, with fewer than one-third reaching mature controls. That prerequisite sits in engine 00 (Data) and Governance + Verify.
The human gate is a control that wears out
Every gate in this article assumes a person reads what they approve. That assumption now has a legal counterpart. Article 14 of the AI Act requires oversight of a high-risk system to be designed so the person doing it stays aware of automation bias, the tendency to over-rely on the output.
A rule written that way is an admission. Regulators expect attention to degrade under repetition, which is what anyone who has approved the same prompt fifty times already knows. Designing around it means the gate cannot be the only control.
So the load sits on the deterministic side. Permissions, protected files, exact digests, and a validator that returns an exit code do not get tired at the ninety-third prompt. Human authority decides what may happen; the machinery decides what can.
Three operating axes
autonomy × interoperability × governanceThe autonomy ladder
L1 to L5 · the human role climbs in step08The agentic divide
The agentic divide separates two ways of adopting agents. One group runs them end to end. The other attaches them to the stack it already owns. This is an architectural split before it is a spending split. The first group wires a new agent into a shared contract, so its input has a source and its output has an owner. The second adds a capable component and one more handoff nobody supports, then finds the gains stay local. Executive surveys capture the money side of this divide. They also capture intent more than realised results. The architectural side is the part a team can act on this quarter, and it is settled before procurement. What settles it is one question: what must a new agent connect to before it runs.
BCG's AI Radar 2026 surveyed 2,360 executives, including 640 CEOs. Trailblazers put over half into agents. They were about twice as likely as Followers to deploy agents end to end.
Respondents committed over 30% of AI investment to agents. Surveyed organisations expected total AI spending to approach 1.7% of their annual revenue. These are forward-looking survey responses, not audited outcomes.
Ninety-four per cent of surveyed organisations expected to keep investing without 2026 returns. Nearly all surveyed CEOs expected returns that year. Those expectations need separate evidence on organisational readiness, realised outcomes, failure rates, and complete operating cost.
Disconnected adoption also raises the Integration Tax. A new tool strengthens the architecture only when it serves a shared contract; otherwise, it adds another unsupported handoff.
Controls required before expansion
Name the data, evaluator, controls, and owner first. Record how a failed run stops and which person may approve a restart. That is the readiness test used here.
Enterprise and SME read it differently
S&P Global's September 2025 survey found a positive net staffing balance at smaller firms and a negative one at large firms. That supports two possible operating uses, not fixed outcomes for every company.
A large enterprise may use the Framework to consolidate specialised work across several existing marketing teams and systems. A smaller company may use it to extend one accountable operator. S&P's June 2026 update shifted the overall outlook towards modest net reduction, so the size split remains a dated 2025 result.
Replicate tests whether a small team can reuse a production architecture without starting at version one. It does not assume the same staffing effect in every organisation.
09The business case
The business case is not a promised multiple. It is the ability to test one workflow against a baseline. An evaluator judges the result, the full operating cost is counted, and the run is repeated. Gartner's budget-and-readiness gap is what makes that discipline necessary, because the money moved before the operating evidence did. Published upside figures work as hypotheses and fail as proof. They describe what might happen, not what was measured. Engine contracts are what make the test possible at all. They fix what the work was, what it produced and where it stopped, so two runs can be compared instead of two impressions. Isolated AI usage cannot support a return claim for a whole operating system.
BCG cites estimates of triple ROI, speed, and volume. BCG also estimates 5 to 10% incremental growth and 15 to 20% efficiency gains. Those figures describe potential upside, not audited outcomes.
Test returns by engine
Treat the BCG estimates as hypotheses for individual workflows. Keep the human gate, record the baseline, and measure the complete operating path. Scale only after the gain survives repeated runs.
A connected architecture makes those tests comparable. Engine contracts identify the work, while Orchestrate records handoffs and stops. Without that structure, isolated AI usage cannot support a credible return claim for the entire operating system.
Count the full cost
Model, API, review, integration, and failure-recovery costs belong in the denominator. Gartner expects over 40% of agentic-AI projects to be cancelled by 2027 because of cost and unclear value. Gartner also calls out "agent washing," where chatbots are relabelled as agents (June 2025).
Start with one engine and one visible artefact. Preserve the validator receipt, human-review time, token cost, failed attempts, and the reason for each retry. Call it a return only when the same accounting holds across another run.
Decision | Evidence and receipts retained for later human review | Scale only when |
|---|---|---|
Start one engine | Baseline, artefact receipt, validator result, full cost | The workflow produces a useful artefact and its gates stay green |
Connect the loop | Repeatable handoffs, outcome measure, failure record | The gain survives across more than one run and the Integration Tax stays visible |
Raise autonomy | Fixed evaluator, retry budget, human-gate receipts | Failures route to a safe stop and the operator can still read the evidence |
10The roles it creates
This architecture creates three accountable roles: the Director, the AI-Marketing Engineer, and the Agent Manager. They are drawn by answerability, not by channel, tool or seniority. Each one owns a different kind of failure. A wrong decision, a broken contract and an unreviewed exception are three separate problems, and an agent system produces all three. Putting a name against each is most of the design. None of this describes a bigger team. In a small operation one person can hold two of the roles, provided the boundary between them stays written down somewhere. What the model rules out is a marketing function where the artefacts get automated and the accountability quietly does not.
Three accountable roles
The Director sets intent and remains accountable for business decisions. An AI-Marketing Engineer owns Context, engine contracts, and technical quality.
The Agent Manager reads validator and run receipts, supervises exceptions, and holds named gates. These roles describe responsibilities; one person may cover several in a small team.
Human work therefore shifts from editing every artefact to setting intent and reviewing the system that produced it. The Harvard Business Review article cited earlier describes the same operating change (Taite, Winsor, and Fernandez, 2026).
The related marketer-roles analysis develops those responsibilities. BCG and MIT SMR found 58% of agentic-AI leaders expect governance changes within three years (BCG and MIT SMR, November 2025). Those changes need named owners, not only a revised org chart.
The Director, AI-Marketing Engineer, and Agent Manager make ownership explicit across intent, technical quality, and gates. Durable rubrics and exception reviews turn those names into operating responsibilities.
Voice rules became durable when I encoded them as checks and review steps. I call the accountable leader the AI-Native CMO. That role owns the boundary between marketing intent, model behaviour, engine permissions, and human accountability.
The talent profile: the Pi-Shaped Marketer
The role combines two depths: marketing judgement and technical fluency. Marketing depth distinguishes a material, evidence-backed claim from a merely plausible model-generated sentence. That judgement identifies what the buyer needs next.
Technical depth turns that judgement into Context files, engine contracts, validators, and observable handoffs. Prompt craft covers only one part of that work.
A Pi-Shaped Marketer need not become a full-stack engineer. The operator must set boundaries, read receipts, and recognise when a specialist is required.
11Operational evidence and its limits
The Framework rests on working artefacts, and it does not yet run as one complete eight-engine loop in any single environment. Both halves of that sentence are load-bearing. Individual contracts are built and operated. Several can be opened and inspected by a reader rather than taken on trust. The full loop cannot be inspected, because it has not been assembled in one place. Saying so now costs less than being corrected later. Everything that follows is written to hold those two claims apart. Evidence anyone can verify never gets folded into a maturity claim only I can see. The parts are real, the whole is unproven, and the gap is dated rather than hidden.
Readiness: what you can check, and what I operate
Two columns, because one field carrying both questions is what let this page claim more than it could show. The first is an evidence state: what a reader can open right now. The second is a maturity grade: what is operated. They are graded on separate vocabularies on purpose, and the rules below are the ones this table is held to.
Engine | What you can check | What I operate | Next artifact |
|---|---|---|---|
Listen · verified 2026-08-10 | NOT SHOWN — Runs inside enterprise work, and production firing has not happened yet. | partial — Twelve sources catalogued, eight tested, three in flight and one awaiting tool validation. Collection is a documented protocol an agent executes against connector-attached tools rather than a single program, and the inventory says "tested" rather than "verified end to end". Eight of twelve tested is coverage, not proof. Nothing here has run unattended against a market in production. | A de-identified write-up of the twelve-source collection architecture. |
Convert · verified 2026-08-10 | NOT SHOWN — There is nothing to show, because nothing was built. | designed — Specification and negative findings only. A verified inventory established that just two pages embed a form while every landing page merely links to one, and that a sitewide campaign-parameter extractor already persists the entry URL in a cookie surviving a multi-click funnel — which the form never reads, so the fields arrive empty. This row is a zero and stays a zero until something ships. | The lead-routing specification, published with its negative finding intact. |
Data · verified 2026-08-10 | NOT SHOWN — Almost all of it runs inside enterprise work and is structurally uninspectable. | built — Confidence is computed by a fail-closed ladder rather than typed by a human, across two axes, where a named confirmer certifies a claim only if their declared authority covers that claim type. A four-state claim model separates approved, held, ruled-false and unclaimable, and a ruling must name what would make it claimable, so it has an exit. An append-only ledger proves its own append-only property by walking every committed revision. | A provenance layer built against a wholly fictional subject, where an uncited claim is a build error — the repository and a live page, not a write-up. |
Signal · verified 2026-08-10 | NOT SHOWN — Runs inside enterprise work. | built — A figure gate whose design decision matters more than its code: verify that a small known set of governed figures is not contradicted, rather than classify every number on a page, because exhaustive classification floods the allowlist and trains rubber-stamping. Plus an undeclared-divergence adjudicator, on the principle that a deliberate evidence-backed divergence is legitimate and the defect is a divergence nobody adjudicated, and a drift work order whose population is discovered from the filesystem rather than typed by hand. One honest zero stays visible inside this row: emit-eligibility predicates are written, tested and deliberately unwired, with a test asserting that no module imports them. | The readiness scorer itself, open-sourced with a fictional company’s scores — the framework that grades this table, made runnable by the reader. |
Create · verified 2026-08-10 | built — The entity graph behind this site is fetchable right now, and so is the JSON-LD on every page. Identifiers name entities rather than pages, so the graph stays stable across pages and is joinable. It has never been run through an external validator, and nothing measures whether any AI system actually cites it. | — | |
Distribute · verified 2026-08-10 | NOT SHOWN — Runs inside enterprise work. | built — delivery machinery only — A catalogue generated from the canonical artifacts themselves so it can never become a second source, gated on every design token resolving to a canonical snapshot and every declared path existing on disk, and published by a script that refuses to run when the target is the source repository. An audience-by-language fallback matrix is resolved by one pure module imported by both the build and the gate, so the two cannot disagree. No channel outcome data is attached anywhere. The variant matrix is an A/B by construction with no measurement, and the one notification path has never sent a message because the deployment sets no provider key. | A delivery-contract spec, naming no channel and no account. |
Measure · verified 2026-08-10 | NOT SHOWN — Runs inside enterprise work. | built (instrumentation + eval) · partial (outcomes) — Instrumentation and evaluation are built: an eval suite that certifies the context rather than any one consumer, with a withheld answer key the retriever never sees, an ungrounded control arm, and a rule that a pass extracting zero claims reports "nothing verified" rather than green. It names its own blind spot — that the generator and the grader share a rule set. A deploy-boundary watcher timestamps the instant a build first serves, which is not the instant publish was pressed; on one release those were about nine minutes apart. Outcome measurement is partial and the row says so. The click-attribution work diagnosed why roughly half of click events arrive with an empty destination, designed a destination-independent taxonomy, then shipped under two percent of it — and the write-up records the mechanism as untested rather than claiming it failed. | A pre/post baseline kit runnable against any public site, carrying zero numbers and a mandatory "missing and unrecoverable later" section. |
Iterate · verified 2026-08-10 | NOT SHOWN — Runs inside enterprise work. | partial — One genuine closed loop runs in production, and it went outcome to root cause to invariant to enforcement: heatmaps came back blank below the fold, root-caused to content revealed by the page’s own scripting behind a fallback that could never fire in a replay, first "fixed" with a reset that silently killed every hover state and verified by a screenshot structurally incapable of seeing hover, then narrowed correctly and frozen as a pre-push guard. One loop. The ceiling is real and this row does not claim more. | The mutation battery as a standalone tool, published with the post-mortem in which every finding is against itself. |
Orchestrate · verified 2026-08-10 | NOT SHOWN — Runs inside enterprise work. | built — Publish-time gate chains where refusal deliberately lives at publish rather than at serve, because a server that refuses to serve leaves the consumer with nothing — and the governance gate runs the same extractor the server runs, so a re-implementation cannot drift and then certify a lie. The local gate prints what it did NOT verify on green as well as on red, because a caveat shown only on failure is a caveat nobody reads on the day it matters. Deploy verification fingerprints the served bytes rather than a declared version, because a declared version is a claim you can forget to bump. In one cluster there is no CI anywhere, so the honest answer to "what stops someone shipping without running the gates" is nothing mechanical. And the engine conformance registry is declaration-only, imported by no tool. | A single-file publish-gate reference implementation, generic and naming nobody, shipped with its tests. |
Replicate · verified 2026-08-10 | NOT SHOWN — The one deployment sits behind a login and is noindexed by design. | built (site-architecture porting) · none (multi-tenant replication) — Exactly one completed port exists, and it was executed as a pipeline rather than a redesign: a fetch-once asset manifest under a hard rule that nothing may be re-fetched or colour-guessed once local, a token file as the single colour and spacing authority, programmatic mark substitution across all markup, six pages generated from one template driven by per-page data dictionaries, and repackaging as a CMS theme. It then survived a second-generation rebrand propagated by script and re-audited at three viewports, which is the test that separates a port from a one-off. Be strict about the second half: no second brand, no multi-tenant routing and no second deployment exists. A hand fork of a route tree is not replication. | A fictional Brand A architecture ported onto a fictional Brand B in public, then rebranded again by script to show the port survives a change wave. |
The rules this table grades itself by
- No link, no grade. A row with no openable artifact is capped at NOT SHOWN, so every cell that claims more is a URL you can open.
- Monotonicity. What you can check may never claim more than what I operate.
- The zeros stay. Rows that are honestly empty remain visible, including the four named on the zeros page.
- Date the row, not the page. Every row carries the date it was last verified.
The four zeros, in full
Four things here were built and deliberately not switched on. They are not a backlog and not failures — each is a decision, and in three of the four the decision is enforced in code rather than remembered. Nothing below moves a grade. It is here because a table where every row is built is a brochure, and because the rows that read built are only worth believing if the empty ones are visible too.
Orchestrate — A conformance registry no tool imports. A registry declaring what counts as trusted output for the engines that consume it, with each consumer registered against the contract it is meant to meet. Nothing imports it. It is declaration-only, and one registered consumer is recorded in it as non-conformant — so the registry currently documents a gap rather than closing one.
Signal — Gates written, tested, and deliberately unwired. Emit-eligibility predicates, implemented and covered by tests, deciding when a figure is allowed to leave the system. They are not wired in, and that decision is enforced rather than remembered: a test asserts that no module imports them. Wiring them in would fail the suite, which is the point.
Convert — The lead-routing specification, and the finding under it. A grouping taxonomy shared deliberately between the form field and the analytics event, so a CRM record and an analytics event could be joined later, plus a fallback ladder for the case where the vendor declines URL-parameter pre-fill. None of it was built. The specification also carries a negative finding worth more than the design: a sitewide campaign-parameter extractor already persists the entry URL in a cookie that survives a multi-click funnel, and the form never reads it — so those fields arrive at the CRM empty.
Listen — A competitive-listening pack with zero outputs. A prompt-and-schema pack for competitive listening, with a signal schema requiring a source link, an access date and a counter-position score. It has produced no outputs. The machine-readable state file says so, which is why this row reads partial rather than built.
Working artefacts
The live mktr connector serves scrubbed Context for agent clients while the owning repository retains write authority. The Create article path has also run from captured inputs to a published artefact, with a human approval gate before publication.
Validation began as a prompt asking the model to inspect its own draft. A passing draft still contained violations a person found immediately. The current validator reads the same governing documents as the writer and returns independent evidence.
The Measure-to-Context learning path remains open. Hendry’s Operator Logs contain 116 entries and seven deep dives tied to shipped artefacts as of June 2026. They preserve failures and fixes, but none closes the missing Measure-to-Context edge.
A derived learning index now compiles that record into machine-readable form. Research records, session entries, wrap-ups, and authority entries become one index that a clean checkout can rebuild byte for byte, and a session cannot close while the index is stale. Every claim in a compiled operator-log candidate has to resolve to an entry in it.
Nothing at run time reads that index, by design. That is also why it does not close the Measure-to-Context edge: it records what the system did, not what the market did.
What the LangGraph pilot proved
The Build Log 002 pilot placed checkpoints, conditional routes, and a digest-bound human gate around the maintained article workflow. The unchanged builder produced byte-identical JSON twice. Existing validators then inspected Context, article structure, lineage, and images.
The first run stopped on an 8/9 image result before any live-link or CMS step could execute. After Create-Images v4.5.1 repaired the stale filename rule on 25 July, both article heroes passed 9/9.
LangGraph routed the work and recorded the stop; it did not rewrite the builder, change a validator result, or acquire publication authority. Its proof covers one work-and-control slice around an existing article workflow, rather than execution across all eight marketing engines or a closed improvement loop.
A fresh bounded run on 5 August 2026 repeated that shape after the workspace moved: seven repository surfaces pinned clean, an owner-approved packet, a resume-freshness check that passed, and a terminal state of accepted before artefact generation, with no writes outside the lab.
How sources bind to claims
The maintained builder pairs every inline external link with one source record, one section, and a specific support claim. Automated checks can prove that the map is complete and reproducible; a person still judges whether the source is current, fair, and strong enough for the wording.
Approval applies to the exact article bytes reviewed by the owner. A later change needs another review, while CMS draft creation and publication remain separate authority decisions with their own readback.
The publishing path itself has now moved. Content is a tracked file tree rather than rows in a database, the build is hermetic and needs no secrets, and publication happens by merging an approved pull request that required status checks have passed. What a reader can check from outside is the output: every page here is generated from those files. What cannot be checked from outside is the enforcement, because the repository is private, and this article does not ask anyone to take that on trust.
What remains unproven
Silver and Sutton argue that AI systems will learn from their own experience (Silver and Sutton, 2025). The Operator Logs provide local experience as recorded failures, fixes, and outcomes; they guide later changes rather than training the system automatically.
Replicate has now had its first real test: the governed architecture was ported to another brand as a pipeline rather than a redesign, and it survived a second-generation rebrand propagated by script. What it has not been is multi-tenant. There is no second brand running concurrently, no tenant routing, and one completed port is not a general claim. Many marketing outcomes also lack a single ground-truth test, so Measure will remain human-gated longer than content creation and autonomy will rise only when the evidence supports it.
12Getting started
Start in four steps, in this order: govern the sources, prove one workflow, define the engine contracts, then measure capability. The order is most of the advice. Governance comes first because every later step reads from the same files, and an ungoverned source spreads its errors faster once agents are involved. One workflow comes next, because a single proved path teaches more than four half-built ones. Contracts and measurement follow, since both need something real to describe. Here, voice.md anchors the first step and Build Log 002 anchors the second. Autonomy is the last thing to raise, not the first. It moves only when the evidence shows the run stops safely without a person watching it.
1. Govern the sources
Start with owned files for voice, ICP, messaging, factual claims, and change authority. In this system, voice.md gives every writer, reviewer, and validator one owned, versioned source for voice decisions. Run one small artefact and inspect where generic or unsupported language still enters.
Change voice.md only through its approved path. Rerun the identical workflow and confirm that every affected downstream artefact changes in the expected place. A maintained file gives the rule an owner and a diff.
2. Prove one workflow
Before selecting a runtime, choose one visible artefact such as the Framework article and name the accountable human who judges it. Define the baseline, evaluator, retry budget, complete operating cost, and human gate before the first run. Those controls make the result interpretable.
Preserve ordinary failures instead of discarding them. The Build Log 002 pilot stopped on an 8/9 image receipt, even after its JSON matched twice. Add another engine only after several readable receipts repeat across runs.
3. Define the agent contracts
Give Create, or any engine you add, an input, one verb, an output contract, tools, stop conditions, and an accountable owner. Use a deterministic node when identical inputs should yield the same transformation. Bounded agent loops handle judgement-heavy work.
An external evaluator, independent of the acting loop, should decide whether that loop may continue under the declared rubric. Connect engines through explicit contracts and open protocols. This keeps the Integration Tax visible and favours shared Context over isolated tools.
4. Measure capability
Track artefact quality, failure class, complete cost, elapsed time, autonomy level, and recorded human interventions per completed run. A record of supervised autonomy for Create and prompt assistance for Listen identifies a system state in practice. A generic AI-use statement lacks that detail.
Hold the evaluator stable while testing a change. Retain failed receipts alongside green ones, then add AI presence to the marketing scorecard. Part of your audience now meets the brand through an AI answer rather than a search result.
Four implementation steps
13The opportunity
The opportunity is the distance between access and readiness. AI access is widespread. Only 30% of surveyed CMOs report readiness to scale, and McKinsey's high-performer band sits near 6%. That distance is an operating gap rather than a technology gap. Operating gaps close with architecture, not with another model. Shared Context, engine contracts, validators, permissions and human gates are what turn scattered AI usage into a system a business can stand behind. The evidence for that now spans personal, multi-brand and enterprise work. That breadth matters, because an architecture holding in only one environment is a preference. Adoption already happened. Readiness is the part still unbuilt, and it is the part worth building first.
Hendry.ai publicly proves Listen and Create, with partial Convert and Orchestrate. Private work adds GTM, Data and Signal, enterprise web, Measure, and MCP deployment evidence without exposing employer systems.
Replicate must carry the governed Context, contracts, validators, permissions, and human gates to another brand. A transfer that needs a fresh architecture fails the test.
You can inspect whether each new agent connects to evidence and an accountable owner. If it does not, it adds another pile of parts.
The implementation receipts remain in the Operator Logs, including failed runs and rule changes. Build Log 001 shows how the content path was built. The next evidence should instrument the Replicate transfer from its first change.
- What is the AI Marketing Framework?
- The AI Marketing Framework is Hendry Soong’s target operating model for connecting eight marketing engines around a shared Context layer. Governance + Verify and Protocol + Interop apply across the system. Orchestrate coordinates engine contracts, while Replicate tests whether the model transfers to another brand.
- Why did the model change from eleven engines to eight?
- Each engine must name a verb, a concrete artefact or state change, and a distinct contract. Define, Understand, and Position became Context files; Nurture split between Distribute and Convert; Optimise became Iterate. Grow was removed, while Data, Signal, and Orchestrate filled production gaps.
- What changed for the 2026 edition?
- Production turned the three-layer 2025 model into a governed loop. Define, Understand, and Position became Context; Nurture split between Distribute and Convert; Data, Signal, Orchestrate, and Replicate gained explicit contracts. Governance + Verify and Protocol + Interop now apply across every engine.
- How autonomous is the system today?
- The demonstrated upper bound is L4 for bounded local work: agents propose and act inside repository, permission, retry, and validation guardrails while I monitor and spot-check. Public claims, canonical Context changes, CMS writes, publication, deployment, and permission changes remain explicit human gates. The Hendry.ai proof ledger is not a complete inventory of private or enterprise implementations.
- Is this just prompt engineering?
- Prompt engineering shapes one model instruction, while context engineering assembles the model-facing state for a turn. Marketing skills package repeatable task instructions, references, and checks inside the governed system; they do not replace its sources, permissions, validators, or human gates.
- Is graph engineering the same as LangGraph?
- Graph engineering is broader than LangGraph and maps how work, state, evidence, budgets, and authority connect across runtimes. LangGraph is one tested work-and-control adapter; Hendry Context, validators, permissions, and people retain authority. Knowledge graphs remain a separate, optional information model.
- What is the Integration Tax?
- It is the cost of stitching disconnected tools and agents together by hand. Shared Context and explicit contracts reduce repeated handoff work, although every added service still carries integration cost.
- How is this different from an all-in-one marketing platform?
- An all-in-one platform bundles several marketing functions under one vendor, contract, and support desk. That can reduce setup. This architecture instead keeps Context and operating history in governed, exportable artefacts, so one tool can be replaced without rebuilding the system’s memory.