TLDR00 / 13

AI tends to arrive tool by tool, each piece optimising its own corner while few teams build the layer that connects them. Every tool you need exists, and the open standards to connect them arrived in 2026; what separates teams now is whether they can architect those parts into one governed system. This updated framework organises marketing into eight engines on a shared Context layer, connected as a closed loop, governed at a human-gated level, and sequenced toward autonomy. With connection now a commodity, the value moves up the stack to governing and contextualising the agents. The load-bearing bet is portability: that the architecture, not the headcount, is what carries the system to a new brand, and the next edition exists to test it.

01Why architecture matters now

What changed since the 2025 edition: the thesis held. The model moved from three linear layers to a closed loop. The connective standard I said was missing now exists. Read the frozen 2025 edition.

The martech shelf went flat in 2026, the first time in a decade, even as the agents stacked on it keep multiplying. Too many agents is just the 2026 version of too many tools, and it fails the same way: every part optimises alone and nothing holds them together.

The returns show it. A review of more than 300 publicly disclosed AI initiatives found 95% of organisations getting zero return on generative AI, with only about 5% of integrated pilots reaching real value (MIT NANDA, "The GenAI Divide: State of AI in Business 2025," July 2025). The headline is contested, and the report itself calls its figures directionally accurate from interviews rather than official reporting, but the gap is real. Most companies have the tools. Few reach the profit line.

Budgets stay flat at 7.8% of revenue, and CMOs now fund AI from that same envelope (Gartner 2026 CMO Spend Survey). The money is committed; the delivery lags. 45% of martech leaders say AI agents fail to meet expectations (Gartner, October 2025). Half blame stack readiness. Half blame talent.

Every tool you need exists, and so do the open standards to wire them together. What separates teams now is whether they can architect those parts into a working system. Add more agents without that architecture, and you repeat the mistake a decade of buying more tools already made.

A year ago that was a contrarian claim. In May 2026 Harvard Business Review made the same case, plainly: "The issue isn't the tools. It's the operating model" (Taite, Winsor, and Fernandez, 2026). The diagnosis held. What changed is the answer, which the rest of this piece works through.

This is written for the operator who also leads, the person building an AI-native marketing function and answerable for how it performs. Enterprises will read it as a consolidation model and smaller teams as an enablement one, a split the later sections take up. The wager underneath all of it is portability: that the same loop redeploys to a new brand on its architecture, without rebuilding the team behind it.

The whole shape is small enough to hold in one line: eight numbered engines on a shared Context layer, two system engines that orchestrate and replicate them, and two spines, governance and protocol, that cut through every layer.

02The coordination challenge

AI tends to arrive tool by tool. The content team uses one for drafts, analytics another for reporting, paid media a third for ad copy, each optimising its own corner. Connecting those corners into one system is the rare part, as the numbers below show.

The martech landscape lists 15,505 tools in 2026, up just 0.79% from 15,384 the year before, according to the State of Martech 2026 census. The count went flat after a hundredfold run since 2011. More tools never solved the problem.

88% of organisations use AI in at least one function, but only about 6% are high performers, the ones attributing 5% or more of EBIT impact to AI (McKinsey, State of AI, November 2025). Those high performers are about three times as likely to have redesigned their marketing workflows rather than bolting AI onto the old ones. The reason the rest stall is the architecture around the model, not the model itself.

The value is in how the engines connect, the operational layer the high performers built and the rest skipped.

0102030405060708090
R·01
88%use AI in at least one function

Adoption is near-universal. It is no longer what separates the winners.

ORGANISATIONS USING AINEAR-UNIVERSAL
R·02
6%are high performers

Attribute 5% or more of EBIT to AI: the thin band that redesigned the workflow instead of bolting AI on.

HIGH PERFORMERSTHE THIN BAND
R·03
95%see zero return on generative AI

MIT NANDA review of 300+ disclosed initiatives. The headline is contested and directional, but the gap is real.

ZERO RETURNTHE GAP
SOURCE · McKINSEY 2025 · MIT NANDA 2025

Adoption is near-universal. Returns are rare. Architecture separates the two.

Adoption is near-universal; returns are rare. Architecture separates the two. Sources: McKinsey 2025, MIT NANDA 2025.

03From three layers to a closed loop

The 2025 edition organised marketing into three layers: Foundation, Execution, and Optimisation. That model came from the classic frameworks: STP, AIDA, RACE, AARRR, the flywheel. They are linear, and they assume a human pace.

A year of production broke that shape. Here is what building it taught me.

The engine is the file

In the old model, Define, Understand, and Position were engines. In production, Define produced voice.md. Understand produced icp.md. Position folded into messaging.md. Every other engine reads those files. Change voice.md, and the next content run adapts with no new prompt. Those three were never steps in a pipeline. They were a shared brain.

The verb test

Out of that I got a test you can run on any architecture, including your own. Point at a box and ask what the agent does. Name the verb and the concrete artifact it produces. If you cannot, the box is a label, a file, or two engines sharing one name.

Run it on the 2025 model and eleven boxes became eight engines.

Box

What does it do?

Verdict

Create

Produces content artifacts

Engine

Listen

Detects market signals

Engine

Optimise

"It optimises"

A label. The real verb is Iterate: read Measure, propose Context updates

Nurture

Sends, and also converts

Two verbs. Sending went to Distribute, conversion to Convert

Define, Understand, Position

Produce voice.md, icp.md, messaging.md

Files, not engines. They became the Context layer

Grow

No verb, no artifact, no evidence

Removed

The test is strict on purpose. A box that cannot name its verb is where a bloated architecture hides.

Data was always king. I just never put it on the map

I treated data as the step before the framework. Clean your data, then run the engines. That assumption stayed invisible in the published architecture. Without a real Data engine, agents work on fragmented inputs. The practical move is to inventory and use the data you already have rather than wait to build the perfect pipeline first. And because this is a loop, there is no mandatory front door: you do not have to start at Data. Start wherever you already have signal and let the cycle pull the rest into place. Orchestration had the same problem. In the old system I routed work between engines by hand. That coordination was never named as an engine. Now it is Orchestrate, the coordination layer that already runs the engines as a chain through shared protocols and handoff contracts. Each engine runs as its own bounded workflow, one capable agent with a small toolset reading the shared Context layer, and Orchestrate chains those workflows through contracts instead of running one big agent swarm. The simplest structure that can do the job is the default.

Linear frameworks match how teams are organised: strategists plan, operators execute, analysts measure. The org chart is linear, so the framework was linear. Agentic systems do not work that way. They respond, adapt, and loop. The rebuilt system is a closed loop because production forced it.

The discipline has a name now: harness engineering

In 2025 the craft was called prompt engineering. By late 2025 it was context engineering, the work of curating the right information for the model (Anthropic, September 2025). By early 2026 the field named the larger thing: harness engineering. The agent is the model plus everything you build around it. As Addy Osmani put it, "a decent model with a great harness beats a great model with a bad harness" (Osmani, April 2026).

You rent the model. You build the harness, and the harness is where the compounding lives: the context files, the engines, the evaluator, the gate. LangChain moved its coding agent from outside the top thirty to the top five on a public coding benchmark by changing only the harness, on the same model (LangChain, March 2026). Anthropic's engineers report the same effect. A frontier model falls short of building a real app from a high-level prompt until you add the scaffolding around it (Anthropic, November 2025).

This framework is a marketing harness. The Context layer holds the memory and the rules. The engines are the tools. Measure is the evaluator, and the human gate is the operator at the controls. An engine is a contract, not a prompt, for the same reason the field is moving past prompts. What compounds is the system around the model rather than the words you feed it.

Call it microservices for marketing if you like, but a marketing harness is more than a coding harness with the words swapped. It carries three rules a software architecture does not need. Numbers are typed by a human and never generated. A claim that cannot cite a source cannot ship. And the brand voice is compiled into checks instead of a written brief, so the same voice survives a thousand runs and a change of model. Those three rules are what make machine-generated marketing safe to publish under your own name.

From prompt to harness

the discipline, renamed twice
2025Prompt engineeringWording the model. Get the phrasing right and hope it holds.
late 2025Context engineeringCurating the right information for the model.
2026Harness engineeringThe model plus everything you build around it: context files, engines, evaluator, gate.
A decent model with a great harness beats a great model with a bad harnessyou rent the model · you build the harness
A decent model with a great harness beats a great model with a bad harness. You rent the model; you build the harness.

04The engines

The system is eight numbered engines, a Context layer beneath them, and two system engines that coordinate and replicate the whole thing. Each engine is a contract: an input, a process, an output, and where it routes. Nothing here is a tool.

Three of the old Foundation engines became the Context layer. Nurture split between Distribute and Convert. Optimise became Iterate. Grow was removed. Data, Signal, and Orchestrate were added. Eleven became eight, plus a shared brain and a command centre.

Two engines run in production today, Create and Listen. Convert ships one of its five parts, the landing-page infrastructure; Orchestrate is the coordination layer that already runs the pipeline. Both are partly built, and everything else is designed, which the tags mark. Create sits deepest, on the most production evidence: over a hundred versions, 116 operator-log entries and 81 extracted principles as of June 2026.

The build pattern matches where the labs landed. Each engine is a folder of instructions, scripts, and rules the agent loads when it needs them, which Anthropic calls Agent Skills with progressive disclosure (Anthropic, October 2025). Each run starts from a brief, which mirrors spec-driven development and its rule that intent is the source of truth (GitHub, September 2025). The engine holds the skill. The brief sets the intent. The model just fills in the words.

Canonical engine model · 20268 engines · Context layer · 2 system engines
−1ContextThe shared brain. voice, icp, messaging, products, plus the rules agents run under. Formerly Define, Understand, Position.law layer
00DataExtracts, normalises, and resolves identity across systems. The prerequisite I always assumed.designed
01ListenDetection. A fresh scan of the market each pass: what is moving now. Scan, filter, route. No memory between runs.built
02SignalState. Keeps a running, decaying score per account, stacks Listen's detections onto it, and enriches. The synthesis that becomes personalisation.designed
03CreateProduces content artifacts. The deepest engine in the system.built
04DistributePushes content to channels. Formerly Amplify, absorbed Nurture's sending.designed
05ConvertCaptures the response: forms, qualification, routing. Its agentic role is real-time page assembly from Signal. Absorbed Nurture's conversion logic.partial
06MeasureEvaluates performance. The immutable judge.designed
07IterateReads anomalies from Measure, proposes Context updates. Formerly Optimise.designed
ΩOrchestrateCoordinates work between engines, and owns the protocol and trust boundary.partial
ΠReplicateDeploys the whole loop to a new brand. Proof the system is brand-agnostic.designed
Built: Listen, Create · Partial: Convert (1/5), Orchestrate · the rest designedeach engine earns its place
The canonical engine model: a Context law layer, eight numbered engines, and two system engines. Built: Listen and Create. Partial: Convert and Orchestrate. The rest designed.

05How the system connects

The eight engines run as one loop. Each engine feeds the next, and the loop closes when Iterate writes back to the Context layer that everything reads from.

The loop on a real job: a competitor moves

Take a job every team knows: a competitor changes its pricing, and you have to respond. That is a composite job, in the jobs-to-be-done sense from Part 1, and the system runs it as a chain of atomic ones. Listen, the one engine here that runs today, catches the move. Signal scores how material it is and which accounts are exposed. Create, which also runs today, drafts the counter: a revised comparison page, new sales talking points, a positioning note. A human approves it. Distribute pushes it out. Convert assembles the landing page around the visitor. Measure, the immutable judge, reads whether the response recovered ground. Iterate turns that result into an edit to messaging.md, which writes back to Context, so the next time Listen catches a similar move the system starts smarter.

That last step is the whole point. A standalone agent answers the competitor once and forgets. The loop answers, measures, and updates the brain. Signal, Distribute, Convert, Measure, and Iterate are designed or only partly built, so today a human still closes that loop by hand. The architecture is what will let it close on its own.

Practitioners gave this loop a name in 2026: the ratchet. A mistake the system makes becomes a permanent fix in the harness, so the rules only tighten. "Every line in a good AGENTS.md should be traceable back to a specific thing that went wrong" (Osmani, April 2026). The operator logs are exactly that. A failure becomes a rule, the rule becomes a check, and that failure does not ship again.

The ratchet, this week, on this article. The draft you are reading passed every word-level check and still read like a machine. The fix was not a longer wordlist. I added six cadence rules to voice.md, the Context file that every Create run reads, so the next article cannot repeat the mistake. One of the six:

PAT-022  Sentence-length variance
No more than two consecutive sentences within +/- 3 words.
AI clusters at 14 to 22 words. Break the metronome.

A failure became a rule. The rule became a check. The Context layer compounds exactly there, which is why an engine is a file the whole system reads.

The front half stacks signal so the back half can personalise

Data builds the raw material. Listen detects what moves. Signal does the thing Listen cannot: it holds state. Listen is a fresh scan each time. Signal keeps a running, decaying score for each account, stacks new signal onto it, and enriches it. The synthesis that comes out is personalisation. Signal is still designed, so today the built Listen carries some of that scoring itself until Signal takes it over. When that stacked signal feeds Create and Distribute, generic output becomes a message shaped to one account, one buyer, one moment.

This detection-generation split is the honest read on how the system is organised. The left side detects. The right side generates. Synthesis bridges them. A monitor with thresholds is a different machine from a template with validation, which is why the engines do not collapse into each other.

The external dimension: GEO

Your positioning, now a property of the Context layer, carries a dimension that did not exist five years ago. Your brand has to be cited correctly inside the AI systems people now ask, like ChatGPT, Perplexity, Gemini, and Claude. The 2025 edition called this AIO. The industry settled on GEO, generative engine optimisation, with AEO as a synonym. Only the label changed. The shift is not small: Google's AI Overviews now reach around two billion monthly users, and a majority of searches end with no click to a website (Semrush, 2026). Listen monitors how your brand appears in AI answers, and Measure now tracks AI presence, share of model, and citation share alongside human clicks, the shift I work through in The AI Marketing Measurement Problem.

The closed loop

8 ENGINES · 1 CONTEXT LOOP
writes back reads reads 00Data 01Listen 02Signal 03Create 04Distribute 05Convert 06Measure 07Iterate Context THE SHARED BRAIN
THE LOOPEight engines run as one loop around the Context layer; Iterate writes each pass back to Context so the next begins smarter.
Eight engines run as one loop around the shared Context layer. Each feeds the next, the cycle returns to the start, and Iterate writes each pass back to Context so the next begins smarter.

06The two spines: Governance + Verify and Protocol + Interop

Two things in 2026 run through every layer rather than sitting at one step, so I treat them as spines instead of engines: structures that cut across the system. They are concerns, and Orchestrate is the engine that enforces them at the system edge.

Governance + Verify

The first spine keeps the system trustworthy. Inward, it constrains my own agents: an immutable judge that cannot be rewritten, circuit breakers that cap spend and send velocity, and editable surfaces that fence what an agent may touch. Outward, new in 2026, it gates other people's agents: identity, a trust registry, and a guardian function for when an unknown agent reaches in. Under both sits provenance by construction. An unsourced claim does not pass the gate, and a human types every number.

The security case got concrete in 2026. Simon Willison named the lethal trifecta: any agent that combines access to private data, exposure to untrusted content, and the ability to send data out can be turned into a leak by one injected instruction (Willison, June 2025). The frontier move is to treat that as a build rule rather than a warning, and keep the three apart by design. OWASP now publishes a Top 10 for agentic applications, and its companion Top 10 for LLM applications ranks prompt injection as the number-one risk (OWASP, 2026). Governance is no longer a policy page. It is now part of the architecture.

Governance now has a legal edge as well. The EU AI Act's Article 50 requires AI-generated content to be disclosed and machine-readable, a duty that lands in 2026 (the deadline moved from August to 2 December under the Digital Omnibus) (EU AI Act, Article 50). In the US the FTC's rule against fake and AI-generated reviews already carries civil penalties (FTC, 2024), and content-provenance standards like C2PA Content Credentials are moving from pilot to practice. A harness that types every number and sources every claim is most of the way to compliant by construction. The rest is disclosure, and disclosure is a setting, not a rewrite.

Protocol + Interop

The second spine connects the system. The agent protocols that wire tools to tools and agents to agents shipped and moved to neutral ground. On 9 December 2025 the Linux Foundation formed the Agentic AI Foundation, anchored by Anthropic's Model Context Protocol, Block's goose, and OpenAI's AGENTS.md, with platinum members including AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft, and OpenAI. Google's Agent2Agent protocol is a separate Linux Foundation project, governed independently. This is the connective architecture the 2025 edition said was missing. It now exists, and no single vendor owns it.

When connection is a commodity, connectivity is no longer a strategy; the moat moves up the stack, from connecting the tools to governing and contextualising the agents that run them. Context stays the law and the memory, the part that is mine and portable. The interop layer carries the traffic, and Orchestrate owns that boundary.

Why not just buy a suite?

The honest answer concedes the suite case first. Agentforce, Adobe, and the others now speak the same open protocols, ship governance, and come with an SLA, a vendor to call, and a roadmap you do not have to build. For a team that will not build its own harness, that is the faster and safer path, and a real one. What a suite cannot give you is the right to leave with your compounding asset intact. The Context layer and the operator-log history are the moat, and inside a suite they belong to the vendor. Keep the context portable and the same brain runs across surfaces the suite does not own, and you can walk when the terms change. Interop buys you exactly that: the leverage stays yours.

Andrej Karpathy names the deeper shift in his 2025 talk on software in the age of AI. There is now a third kind of reader, the agent, sitting beside humans who need a screen and computers that need an API. Agents read through formats built for them, like llms.txt. Your content has a second audience now, and the protocol spine is how you serve it.

That second audience can also buy. In late April 2025 the card networks opened to agents, with Visa Intelligent Commerce, Mastercard Agent Pay, and PayPal all launching within days of each other (Visa, 2025). In September, OpenAI and Stripe shipped Instant Checkout in ChatGPT and open-sourced the Agentic Commerce Protocol (OpenAI, 2025). In April 2026, Google contributed its Agent Payments Protocol to the FIDO Alliance, where it is now being developed through FIDO's open, community-led standards process (FIDO Alliance, 2026).

Be honest about the maturity. OpenAI scaled back its native in-chat checkout in March 2026 and refocused on discovery. The proven, durable behaviour today is the agent that finds and recommends your product. The agent that completes the purchase on its own is real infrastructure, still early. So build for agent discovery now, and make your product, your feed, and your claims machine-readable. The autonomous purchase is the next boundary to build for, once it settles.

The moat moves up the stack
Speak the open protocols, keep the context portable, and you sit beside the platforms instead of inside one of them.
When connection is a commodity, the moat moves up the stack to governing and contextualising the agents.

07The autonomy progression

Architecture answers what connects to what. It does not answer how much AI should control at each point. Should Listen run on its own around the clock? Should Create publish directly or draft for approval? Each of these sits on a spectrum rather than a yes-or-no switch, and different engines belong at different points.

The 2025 edition framed this as a single climb from L1 to L5. That framing was too clean. Andrej Karpathy makes the sharper case. He calls this the decade of agents, not the year of agents (Karpathy on the Dwarkesh Patel podcast, October 2025), and argues for partial autonomy with a human verifying the work. He points to the autonomy slider already in the tools people use: autocomplete to command to full agent mode in a code editor, or search to research to deep research in an answer engine. Keep the agent on a tight leash, and make each step easy to verify. Autonomy is one of three axes, not a single climb.

The autonomy axis still has rungs; I just stopped treating them as a single ladder you climb to the top. The five below run from prompt-assisted to goal-directed orchestration, and the human role climbs in step, from creator to director. I run at supervised autonomy today, L2 to L3, where the agent executes and a human approves the decisions that matter. Most teams hold at that band today. The end state I build for is a durable, human-gated L3, not L5. I keep the gate on purpose. Most vendor guidance treats that gate as scaffolding to remove once confidence grows. I treat it as load-bearing, and production practice sits closer to my posture: across the industry, human review remains the most common way teams evaluate agent output, years into the agent era.

One caution on the numbers: this ladder measures how much the agent decides, not how mature a team's adoption is. Other maturity models count the team; this one counts the autonomy.

The climb is a barbell. Ad-buying already reaches L4: Braze's OfferFit-based decisioning agents replaced static A/B tests with per-user reinforcement learning, and Google is migrating search campaigns off keyword-only targeting to AI Max, with full Dynamic Search Ads migration now pushed to February 2027. Across the market, most other marketing work holds at a governed L2 to L3. Where a domain has a clean, fast feedback signal, autonomy runs ahead; where it does not, the human gate stays.

A move up the autonomy axis is only safe with a matching move up the governance axis. The two travel together, with a gate between every step. Thoughtworks gives the two controls names: guides that steer the agent before it acts, and sensors that catch it after (Bockeler, April 2026). My context rules are the guides. My validation gates and the human review are the sensors.

The runway is real and it keeps getting longer. METR found the task length an agent can finish on its own, at 50% reliability, had been doubling roughly every seven months, reaching about an hour for the best model as of March 2025 (METR, March 2025). Longer horizons do not remove the gate. They make where you keep the human the most important design decision in the system.

Where the system is now, gaps included

I have not closed the learning loop yet, the part where Measure writes back into Context on its own. Create and Listen run deep. The measurement and learning tail, Measure and Iterate, is designed but not yet built. The system runs with a human in the loop. It drafts, validates, stops for a person, then publishes. That human gate is a deliberate feature at this stage. I am not waiting to remove it.

Two more gaps are on the board because the frontier says they should be. The quality gates are built, and they do not yet screen inputs for prompt injection; a system whose engines read outside content, hold private context, and publish outward needs that defense in depth before autonomy rises. The operator logs record provenance, and step-level tracing of what each agent read, called, and produced is the next instrumentation layer, so the evals can run on every change to the governing documents instead of on demand.

The closest working model of the loop I want is Karpathy's autoresearch, which Fortune called the Karpathy Loop in March 2026. An agent changes a real training script, runs for a fixed five minutes, checks one fixed metric, keeps or discards the change, and repeats about a hundred times overnight. It produced a real, measured speedup. The agent is autonomous, but only inside a sandbox a human designed. It can change the code. It cannot change the judge. That is the exact rule this framework puts on Measure, and watching the most credible practitioner land on the same rule is the strongest signal I have that the design is right.

OpenAI shipped the same shape at production scale. One internal product reached about a million lines of code with no human-written lines, built by agents that three engineers steered, under the rule "humans steer, agents execute" (OpenAI, February 2026). Autonomy at the keys, a human at the wheel. The gate, working.

The market moves the same way, slowly. The State of Martech 2026 census found AI adoption follows one order in every category: analytical first, generative second, autonomous a distant third. The authors call it a trust gradient, not a capability gradient. Fully autonomous action is the least-adopted rung. McKinsey's State of AI Trust in 2026 found nearly two-thirds of organisations name security and risk as the top barrier to scaling agentic AI, ahead of regulation or technical limits. The market reaches for autonomy last, for the reason I gate it. Trust is earned with evidence.

Three axes, not one ladder

autonomy × interoperability × governance
AutonomyHow much the agent decides without a human.
InteroperabilityHow many tools and agents it reaches, through which protocols.
GovernanceWhat gates, limits, and judges constrain it.
Autonomy is one of three axes, not a single ladder. Interoperability and governance travel alongside it.

The autonomy ladder

L1 to L5 · the human role climbs in step
L1
Prompt-assistedA human drives. Single asks, full review. The human is the creator.
L2
Workflow automationChained steps with logic. The human reviews the output.
L3
Supervised autonomy · where I run todayThe agent executes; a human approves the decisions that matter.
L4
Guided autonomyThe agent proposes and acts inside guardrails; the human spot-checks and monitors.
L5
Goal-directed orchestrationThe agent sets strategy from objectives across the whole function. The human directs.
THE GATEThe end state I build for is a durable, human-gated L3, not L5. Most teams hold at L2 to L3; ad-buying is the one domain already at L4.
The end state is a durable, human-gated L3, not L5. Most teams hold at L2 to L3; ad-buying is the one domain already at L4.

08The agentic divide

The framework describes a path. Reality is uneven. BCG's AI Radar 2026, a survey of 2,360 executives including 640 CEOs, splits the field by how CEOs spend. Trailblazer CEOs put more than half of their 2026 AI investment into agents and are about twice as likely as Followers to deploy them end to end. Across the board, CEOs have committed more than 30% of 2026 AI investment to agentic AI, and total AI spend is set to roughly double, toward 1.7% of revenue.

The intent is striking, and it is intent, not results. 94% of CEOs say they will keep investing even if 2026 brings no payoff, and about 90% expect agents to deliver measurable returns this year. Read those as forward-looking executive surveys, since none of it is an audited outcome. The gap between that confidence and the readiness underneath it is where the risk sits.

The divide compounds through the Integration Tax. A team without the architecture pays that tax on every new tool it buys, so each purchase widens the gap instead of closing it. Architecture is what lets more tools become more leverage. Teams that compound and teams that stall split on this one fact of plumbing, not maturity.

Readiness is the gate

Confidence runs ahead of governance. McKinsey's 2026 trust survey found the average responsible-AI maturity score rose only slightly year over year, and just under a third of organisations reach a mature level on strategy, governance, and agentic-AI controls. Gartner puts the upstream data gap at 57% of organisations whose data is not AI-ready (Gartner, 2025 Hype Cycle for Artificial Intelligence). The 2025 edition called this prerequisite a Level 0.5 phase of data and governance work. In the rebuilt model it is no longer a footnote. It is engine 00, Data, on the map, and the governance spine running through everything above it.

Enterprise and SME read it differently

For large enterprises, this framework is a replacement model. AI consolidates specialised roles and drives efficiency. For smaller companies, it is an enablement model. S&P Global found AI's staffing effect splits by size: a net positive balance at small and medium firms, and a net reduction at large ones (S&P Global, September 2025). A June 2026 update from S&P later recalibrated that toward a modestly negative overall outlook, so read the split as the 2025 picture. One operator now wields what used to need a department. This is why I built Replicate: portable architectures that let a small team deploy a production system without starting at version one.

09The business case

Can this deliver returns? The evidence says yes, with a condition. CMOs now put 15.3% of their budget into AI, but only 30% say they are ready to scale it (Gartner, 2026). The spend is committed, and the architecture to use it well is the gap. BCG's March 2026 scenario report cites an earlier BCG study estimating that an AI-first marketing organisation can triple ROI, speed, and volume, with 5 to 10% incremental growth and a 15 to 20% efficiency gain, framed as upside potential rather than a guarantee (BCG, March 2026).

The condition is the 95% from MIT. Agents bolted onto legacy processes do not pay off. The value shows up when an agent owns an end-to-end journey, and it disappears when the agent improves one isolated step. Architecture and orchestration are what turn usage into results.

And be honest about the bill. An agentic task can cost an order of magnitude more than a single model call, and the model API is often only a tenth of the real spend once tools, retries, and orchestration are counted. Gartner expects more than 40% of agentic-AI projects to be cancelled by 2027 on cost and unclear value, and calls out "agent washing," chatbots rebranded as agents (Gartner, June 2025). The discipline that survives that cull is the one this framework is built on: prove value on one engine, keep the human gate, and put the real token cost in the denominator before you call it ROI.

The talent profile: the Pi-Shaped Marketer

The talent gap now demands two depths. The T-shaped generalist with a single specialty is no longer enough, because operators need deep marketing judgment and deep technical fluency in AI systems at once. The technical leg has moved on from prompt craft. In 2026 it means context engineering, retrieval and grounding, and orchestrating agents over open protocols.

The Integration Tax

One operator managing fifty disconnected agents is slower than a human team. The cost of stitching tools together by hand grows with every tool added. Prioritise a shared data and context layer over a shelf of standalone point tools. That tax is the hidden reason "more AI" so often produces less.

10The roles it creates

If engines are contracts and agents run them, the human work changes shape. Harvard Business Review describes the same shift: marketers become "directors of work" who set intent and review outputs, and "the manager's role shifts from reviewing deliverables to reviewing systems" (Taite, Winsor, and Fernandez, 2026). BCG and MIT Sloan Management Review found 58% of agentic-AI leaders expect their governance structures to change within three years (BCG & MIT SMR, November 2025). Someone has to hold those new decisions. Three roles fall out.

The honest caution is that drawing the boxes is the easy part. Capability and governance are the hard part, and an org chart supplies neither. The human role moves from approving every output to writing the rubric the system enforces forever. Voice rules became durable when I encoded them as checks. Every part of the operating model becomes durable the same way. The leader who owns this system end to end is what I have called the AI-Native CMO.

the roles it creates
DirectorSets intent; writes the laws the system enforces against itself.
AI-Marketing EngineerOwns the Context layer, engines, and constraints. HBR calls it the brand code.
Agent ManagerSupervises agents, reads the judge, holds the gates. Owns the accountability question.
BRAND CODEWhoever owns the brand code owns the quality of everything downstream.
Whoever owns the brand code owns the quality of everything downstream.

11The proof

None of this is a slide. The protocol and context spine ships as a live connector: a read-only context layer served over the open protocol, where many consumers can read and none can change the source. The source content is scrubbed before it is served, and the connector is authless with expiry rather than gated. Install it once, and the latest published context is what they get.

One artifact carries the signature of the whole system: a safe surface over a controlled foundation. The forward path has run end to end on real work, capture in and a published artifact out, with the human gate intact. Part of that path is the validation gate, which started as checks a model ran on its own output and became executable code after a draft sailed through every prompt-level check while carrying violations a human caught in minutes. The validator now runs the voice and structure rules as deterministic code, compiled at runtime from the same governing documents the writers read, and an article ships only when it exits clean. The learning loop that feeds Measure back into Context is the part still open. The receipts live in the operator logs: 116 entries and 7 deep dives as of June 2026, each tied to something that shipped.

Those logs are training data. David Silver and Richard Sutton argue the next era of AI is one where systems learn from their own experience, not only from human text (Silver and Sutton, 2025). A marketing system that records every failure, fix, and outcome is building exactly that. The operator logs are the experience the loop compounds on.

Be clear about the scope of that proof. This is one operator's system, not a hundred-person department's, and that is the honest size of the evidence. The bet of Replicate is that the architecture, not the headcount, is what carries: the same loop redeployed to a new brand without starting at version one. The claim is falsifiable, and porting it is how it gets tested.

There is broader evidence underneath this. Several of these engines I have already begun standing up in other production settings, so the parts have run beyond this one system. What the 2026 edition adds is the connective architecture: the account of how those parts link into a single loop. The engines were the easier proof; how it all connects is the advance.

In marketing, Measure is the last gate to fall, because marketing has no ground-truth reward signal the way code has tests. The teams that earn autonomy last are the ones who solved measurement first.

12Getting started

Four principles separate the high performers from the rest.

1. Foundation before automation

Do not automate execution without a solid Context layer. AI content without AI-informed audience understanding produces generic output at scale. Get voice, ICP, and positioning right first. The engine is the file.

2. Start focused, prove value, then scale

Begin with one or two engines at workflow-automation level. Show returns. Expand once trust is established. The trap is trying to stand up all eight engines at once.

3. Design for agents, not around them

Build agent-native workflows now, and build them on open protocols so external agents can reach your surfaces. Avoid the Integration Tax: favour a shared data and context layer over isolated tools.

4. Measure capability, not just activity

Track autonomy by engine. "Supervised autonomy for Create, prompt-assisted for Listen" tells you something. "We use AI" tells you nothing. Add AI presence to the scorecard, since some of your audience now reads through an AI instead of a search result.

Four principles

01Foundation before automationGet voice, ICP, and positioning right first. The engine is the file.
02Start focused, prove value, scaleOne or two engines, show returns, expand once trust is earned.
03Design for agents, not around themBuild agent-native workflows on open protocols.
04Measure capability, not activityTrack autonomy by engine. Add AI presence to the scorecard.
Four principles separate the high performers from the rest.

13The opportunity

88% adoption, 6% high performers. 15,505 tools, the shelf flat for the first time in a decade. 45% of agents missing expectations, only 5% of integrated pilots reaching real value. The numbers have not moved toward the promise.

The gap between the promise and the delivery is an architecture problem. The teams that are winning solved it. The rest are still accumulating parts. A year ago that argument cut against the grain. It no longer does, and the work now is to build the loop and tell the story straight.

The question is whether you architect the transformation on purpose, or stumble into another pile of parts.

The argument is here. The receipts are in the operator logs, every version and failure behind the system, and Build Log 001 shows how the content engines were actually built.

Frequently Asked Questions
What is the AI Marketing Framework?
An architecture for AI marketing systems. It organises marketing into eight engines on a shared Context layer, connected as a closed loop, with two cross-cutting spines for governance and protocol, and two system engines that orchestrate and replicate it.
Why did the model change from eleven engines to eight?
Each engine has to pass a verb test: what does this agent do? Define, Understand, and Position were not engines, they were the Context files every engine reads. Optimise became Iterate. Nurture split into Distribute and Convert. Grow had no function and was removed. Data, Signal, and Orchestrate were added.
What changed in 2026 specifically?
The connective standard arrived. Agent protocols became an open standard, so connecting tools stopped being the hard part. The value moved up the stack to governing and contextualising agents. A second audience also appeared: agents that read and transact alongside humans.
How autonomous is the system today?
It runs with a human in the loop. It drafts, validates, stops for approval, then publishes. The build, measure, learn, iterate loop is designed, not yet closed. The human gate is a deliberate feature at this stage.
Is this just prompt engineering?
No. Prompt engineering is wording. Harness engineering is the system around the model: the Context layer, the engines, the evaluator, and the human gate. That system is what compounds, and it is the part a competitor running the same model cannot copy.
What is the Integration Tax?
The cost of stitching disconnected tools and agents together by hand. It grows with every tool added, so one operator managing fifty disconnected agents is slower than a human team. A shared data and context layer is what removes it.
How is this different from an all-in-one marketing suite?
A suite speaks the same open protocols and ships governance, so for a team that will not build its own harness it is the faster path. What it cannot give you is the right to leave with your compounding asset intact. The Context layer and the operator-log history are the moat, and inside a suite they belong to the vendor.
Published
###