TLDR00 / 05

The L1 to L5 Autonomy Model measures how much an engine decides, not how mature a team is. Autonomy is one of three axes alongside interoperability and governance. Most marketing work operates at L2 to L3; the target here is a durable, human-gated L3.

01The Framework

Most arguments about AI in marketing are really arguments about where approval sits, and they stay unresolvable while everyone uses different words for it. This model supplies the vocabulary. It puts five numbered settings on one axis, so a team can state what a given engine is allowed to do and be understood without a paragraph of explanation. That is its whole function. It is not a scorecard, and a higher number is not a better result, since every step up moves a decision from a person to a system. Read a level as a statement about one engine at one moment. Settings are meant to move, in both directions.

The L1 to L5 Autonomy Model measures how much an AI marketing engine decides and executes before human approval. It does not measure how mature a team is. The axis is adapted from two sources: SAE International’s L0–L5 framework for autonomous vehicles (the industry standard for discussing machine autonomy), and University of Washington research that adapted autonomy levels for AI agents.

I combined both frameworks to create a model specifically for AI marketing systems. The question it answers: how does an Operator choose the right autonomy setting for each engine and workflow?

The core insight: Autonomy is a control setting, not a destination. At L1, a person directs each task. At L3, the engine executes while a person approves consequential decisions. L4 and L5 describe higher settings, but the target here is a durable, human-gated L3.

02The Five Levels

The five levels describe one engine at a time. A team is never at a level; a workflow is. Each step up moves a different decision from a person to the system: the wording of an output, the sequence of tasks, the choice to publish, the allocation of spend, and finally the strategy itself. Two things deliberately sit outside the scale. How well systems connect to each other is measured separately, and so is how tightly their actions are governed, because a highly autonomous engine with weak governance is not a high score. It is an incident waiting to be written up.

Level

Name

Description

Example

Feasibility (2026)

L1

Prompt Assistant

Single prompts, human reviews all output

“Write 5 email subject lines”

Widely available

L2

Workflow Automation

Chained prompts with conditional logic

Brief → draft → SEO check → schedule

Available with setup

L3

Supervised Autonomy

AI executes workflows, human approves key decisions

AI drafts campaign; marketer approves before publishing

Emerging

L4

Guided Autonomy

AI proposes and executes within guardrails

AI adjusts ad spend within budget limits

Early adoption

L5

Goal-Based Orchestration

AI determines strategy from objectives

“Increase MQLs 20%” → AI selects channels, content, timing

Frontier

Understanding Each Level

L1: Prompt Assistant. This is where most teams start. You open ChatGPT, write a prompt, get output, review it, edit it, use it. The AI is a tool; you’re doing the work. According to McKinsey’s 2025 State of AI report, 88% of organizations have adopted AI but only 6% are high performers seeing attributable business impact.

L2: Workflow Automation. Multiple prompts chain together with logic. A brief triggers a draft, which triggers an SEO check, which triggers scheduling. Tools like Zapier and Make enable this, but humans still review every output.

L3: Supervised Autonomy. The engine executes an entire bounded workflow while a person approves consequential decisions. This is the durable target for most marketing work: the human gate stays where brand, budget, or risk demands it.

L4: Guided Autonomy. AI proposes and executes within guardrails you set. “Spend up to $500/day on ads, target these audiences, optimize for conversions.” The AI makes decisions within boundaries.

L5: Goal-Based Orchestration. You give the AI an objective (“Increase MQLs by 20%”), and it determines strategy: which channels, what content, when to publish, how to optimize. This is frontier technology.

03Evidence and Current Reality

The evidence puts the market lower on the ladder than the vendor language suggests. Real deployments cluster at L3, where systems execute and a human still approves. L4 exists, but it is arriving through acquisition rather than through in-house builds. L5 remains theoretical in marketing and lives only in narrow domains elsewhere. The more useful finding is not the level but the spread. Scaling, where it has happened, stays confined to a narrow slice of the business. It has not spread across the whole operation. A large share of agent implementations fall short of expectations. The causes split evenly between unready technical stacks and missing skills. Both are assembly problems rather than model problems.

Where is the market today? According to McKinsey’s 2025 State of AI report, 23% of organizations are scaling agentic AI systems in at least one business function, with an additional 39% experimenting. But most scaling efforts occur in only one or two functions.

Level

2025 Evidence

Example

L3

Google’s AI Max manages bidding, targeting, and ad creation within a unified campaign structure

System executes; human approves campaign launch

L4

Braze’s acquisition of OfferFit for $325M signals arrival of true L4 systems

AI optimizes messaging within brand guardrails

L5

Still largely theoretical for marketing; exists in narrow domains like algorithmic trading

Full marketing strategy from objectives

Gartner’s October 2025 research found 45% of AI agent implementations don’t meet expectations. Half cite technical stack readiness. Half cite talent gaps. This is the Pile of Parts Problem in action. BCG’s 2025 research confirms only 5% of companies are “future-built” and generating substantial AI value at scale.

04Where Are You?

You are probably at more than one level at once. That is why the question has to be asked per workflow, not per company. A single team can run Create at L3 while Measure sits at L1. Averaging those into one org-wide score hides both the workflow ready to move up and the one already running past its controls. So take the diagnostic below one engine at a time. Pick a workflow. Answer honestly about what the AI actually does in it today. The level is the last question you can say yes to. What comes back is a starting position, not a grade. The point is knowing which workflow to move next, and why.

Question

If Yes → This Workflow Is At

Do you use AI only for single prompts (writing, ideation)?

L1

Do you have automated workflows that chain AI tasks?

L2

Does AI execute campaigns while you approve before publishing?

L3

Does AI make and execute decisions within guardrails you set?

L4

Does AI determine strategy from goals you provide?

L5

The honest assessment: If a workflow begins with “I ask ChatGPT to write things and then I edit them,” that workflow is at L1. Other workflows may sit elsewhere. The useful question is whether a higher setting creates enough value to justify stronger controls.

05How to Set Autonomy Safely

You cannot raise autonomy on its own. Every step up has to be matched in two other places: how well your systems exchange data, and how tightly the resulting actions are governed. Move the setting without moving those and the failure is not underperformance. It is a system acting confidently on decisions nobody checked. Three operating requirements make an increase safe, and none of them are model choices or tooling choices. They are questions of who owns what. Missing any one caps the level you can safely run, however capable the underlying engine is, which is why the ceiling most teams hit is organisational rather than technical.

Requirement

Why It’s Needed

Without It

Operator Function

Someone must design the system architecture

Tools remain disconnected; no one builds workflows

Connected Data Layer

AI needs data to flow between systems

Integration Tax consumes all productivity gains

Clear Guardrails

AI needs boundaries to operate within

Risk of uncontrolled actions; stakeholder trust erodes

BCG’s 2025 AI research found that the quarter of executives who created significant value did so by focusing on a small set of AI initiatives and scaling them swiftly. The practical move is to change one workflow at a time and keep the evidence gate intact.

Pro tip: Start with one workflow. Identify a high-volume, low-risk process, move it from L1 to L2, and prove value. Raise it to L3 only when the controls and evidence can support supervised execution. One reliable L3 workflow is worth more than ten disconnected experiments.

As Harvard Business Review noted, “AI won’t replace humans, but humans with AI will replace humans without AI.” The Autonomy Model makes that relationship explicit for each engine: who directs, who executes, and where approval sits.

For the complete framework, see the AI Marketing Framework. For the role that governs the setting, see The Operator Function.

Frequently Asked Questions
What is the L1 to L5 Autonomy Model?
The L1 to L5 Autonomy Model measures how much one AI marketing engine or workflow decides. L1 is prompt-assisted, L3 is supervised autonomy with human approval, and L5 is goal-directed orchestration. It is an autonomy axis, not a measure of team maturity, and it sits alongside interoperability and governance.
What level are most marketing teams at?
Most marketing work operates at L1 to L3, often at different settings inside the same team. One workflow may be prompt-assisted while another runs with supervised autonomy. Diagnose the engine or workflow rather than assigning one maturity score to the organisation.
What is the realistic target for most teams?
A durable, human-gated L3 is the target here. At L3, an engine executes a bounded workflow while a person approves decisions that carry brand, budget, or risk. Higher levels remain useful descriptions, but they are not a universal destination.
How do I move a workflow from L1 to L3?
Choose one bounded workflow. Connect the data and context it needs, define what the engine may change, add evidence and approval gates, then move from prompt assistance to workflow automation. Raise it to supervised autonomy only after the controls hold under real use.
What's the difference between L4 and L5?
L4 Guided Autonomy lets an engine propose and execute within explicit guardrails, such as a fixed ad-spend limit. L5 Goal-Directed Orchestration lets the system select a path from an objective. Both require governance; neither removes accountability or makes L5 the default target.
How does the Autonomy Model relate to ROI measurement?
Each autonomy setting needs an evidence standard. L1 can measure time saved per task. L3 should measure workflow efficiency, output quality, and approval load. A higher setting is justified only when the added value exceeds the cost of governance and risk.
Built by AI Marketing Operator · Published
###