How products depend on AI, how agents act, and where human judgment stays in control.
Everyone seems to be moving toward AI-native.
Except some mean adding a copilot. Some mean rebuilding the product around a model. Some mean letting agents execute workflows. And some mean reducing the need for human intervention.
We keep putting these ideas on the same ladder:
AI-assisted → AI-first → AI-native → agentic → autonomous.
It looks like evolution. But the labels change the question halfway through.
One describes the relationship between a person and a tool. Another describes a product strategy. Another describes how a system chooses its next action.
There isn't one degree of AI. There are several.
The vocabulary extends well beyond AI-native and agentic. Teams use these labels to describe different changes, sometimes in the same sentence.
Here is a working guide. These are useful interpretations, not standardized definitions every vendor follows.
| Family | Terms | What they tell us |
|---|---|---|
| Human-AI relationship | Assisted, augmented, delegated | How the person shares the work with AI. |
| Product integration | Added, enabled, powered, embedded, integrated | How AI participates in an existing experience. |
| Product orientation | AI-first, AI-native | How product priorities and design depend on AI. |
| Execution | Copilot, agent-assisted, agentic | Whether AI supports a human-directed path or dynamically directs work. |
| Agent-centered design | Agent-native | Whether agents are first-class participants in the product. |
| Independence | Autonomous, bounded autonomy | Which work can proceed without routine human intervention. |
| Execution boundary | Approved tools, resources, environments and budgets | Where the system may operate, regardless of how it chooses its steps. |
| System construction | Neural, generative, neuro-symbolic, compound, multi-agent | What kinds of components make the system work. |
| Oversight | Human-in-the-loop, on-the-loop, out-of-the-loop | Where people approve, supervise, or do not routinely intervene. |
| Deployment | Cloud, local, edge, hybrid | Where computation and data processing happen. |
AI-assisted usually means the human performs the task while AI helps with a step. AI-augmented emphasizes expanded capability, perhaps finding patterns or comparing alternatives the person would struggle to handle alone. Delegated means handing over a bounded outcome rather than requesting every individual step.
Those can overlap. A researcher can delegate a literature scan and then use its results to augment their judgment. Delegation says who owns the task for now; it doesn't establish how the implementation works.
AI-powered says AI contributes materially to a capability, but tells us little about the rest of the product. Embedded or integrated says AI is built into the workflow rather than accessed as a separate destination. An embedded assistant may still wait for a human prompt every time.
Agent-assisted puts an agent inside human-led work. Agentic describes dynamic model-directed execution. The difference is useful, but the vocabulary is loose: a product can be both.
The mistake isn't using these terms. It is treating them as interchangeable measures of advancement.
Start with two imaginary products.
The first is an expense platform. It stores receipts, routes approvals, and reimburses employees. Add AI to explain a rejection or extract details from a receipt. Remove that capability and the platform still has a recognizable job.
The second is a tool whose entire purpose is to turn a conversation into a working prototype. Remove the AI and you haven't removed a feature. You have removed the central proposition.
That difference is more useful than counting how many screens have a sparkle icon.
I would call it AI dependence: how fundamental is AI to the experience and value the product promises?
The familiar vocabulary can help here, provided we treat it as working language rather than an industry standard.
| Label | A useful reading | The limit of the label |
|---|---|---|
| AI-added | AI is an optional addition to an established experience. | An optional feature can still be technically sophisticated. |
| AI-enabled | AI makes a meaningful capability possible or materially better. | The rest of the product may remain conventional. |
| AI-first | AI is a starting assumption in product decisions. | A strategic priority doesn't prove architectural dependence. |
| AI-native | The central experience is designed around AI capabilities and limitations. | Dependence doesn't tell us what AI may execute. |
These are overlapping descriptions, not four certified levels. A product may be AI-first in strategy while still AI-added in what it actually ships. An established product can redesign a central workflow around AI. Its age doesn't settle the question.
Even the removal test needs care. Removing AI might make a system economically impractical without making it technically impossible. A recommendation product may survive with crude rules but lose much of its value. Dependence has degrees too.
The point is to expose the dependency, not win an argument about the badge.
Now imagine the expense platform has a second capability.
An employee asks it to resolve a rejected claim. It checks the policy, inspects the receipt, asks for a missing detail, and chooses whether to prepare an appeal or correct the submission.
The underlying product can remain an established expense system. The new workflow can still be agentic.
Here the question is different: who determines the path through the task?
Anthropic's engineering account distinguishes predefined workflows from agents whose models dynamically direct their processes and tool use.[1] That is a useful architectural distinction, even though the word "agent" is used more broadly elsewhere.
A fixed pipeline can contain several model calls, retrieval, tools, and a critique loop. If code determines the route, the number of calls doesn't turn it into an agent.
Conversely, an agent doesn't need a grand multi-agent architecture. A model choosing actions, observing results, and adapting its next step can be enough.
AI-native describes how fundamentally AI defines the product. Agentic describes how the system chooses its path through a task.
Neither answers whether the system is allowed to finish that task without you.
Suppose the expense agent can investigate a rejection and prepare a corrected claim. It must still ask before submitting.
It remains agentic. Its autonomy is bounded.
Now suppose a conventional automated process reimburses every claim under a threshold after fixed checks. It operates without individual approval, but it doesn't dynamically choose a plan. It may not need AI at all.
That gives us two counterexamples to the ladder: an agent with limited execution authority, and an independent process with no agentic reasoning.
Autonomy describes the degree of independence within a specified boundary.
The boundary matters more than the adjective. Independent for which actions? For how long? With what spending limit? When must the system stop or escalate?
A coding agent might freely read files, run tests, and edit a local branch while requiring approval to merge. Calling it simply "autonomous" erases the most consequential part of its design.
Assisted, augmented, and delegated are useful descriptions of the human's relationship to the work. They are not clean architectural categories either. You can delegate a task to a fixed workflow or to an agent. The same system can assist with one task and own another.
We need to describe the actual arrangement.
There is no contradiction between a product being AI-native and people governing its decisions.
AI can define the central experience while humans set policy, approve consequential actions, and retain control over releases. An AI-native coding environment might let agents investigate and implement changes, while engineers approve merges and deployments.
That is AI-native with human approval gates. If the model chooses and adapts the implementation path, it is also agentic. Approval doesn't erase agency.
Likewise, an established product can be AI-enabled with bounded execution. Its AI capability may operate independently inside a restricted environment while the broader product remains conventional.
Consider an agent that can install dependencies only from an organization's internal npm or Python package repository. It may choose a library, run tests, and revise its approach. The allowed source of dependencies is constrained. That tells us about its execution boundary, not whether the product is AI-native.
We need to separate three questions:
| Dimension | Question | Coding example |
|---|---|---|
| Agency | Who chooses the next step? | The agent chooses a library and implementation strategy. |
| Authority | Which actions can execute without approval? | It can edit a branch; merging requires a person. |
| Boundary | Which resources and environments may it use? | Dependencies must come from the internal package repository. |
Autonomy describes the independence that results from that arrangement. It is always independence for particular work under particular conditions.
An approval gate and a resource boundary are different controls. A human-approved installation can still be restricted to an approved registry. An independent installation can be restricted to that same registry. Neither arrangement implies access to production.
The restriction also needs enforcement. Configuring a preferred package source doesn't establish a closed boundary if arbitrary downloads or unapproved registries remain reachable. The profile should describe the controls the system actually encounters.
flowchart TD
G["Human-defined goal and policy"] --> A["Agent chooses and adapts a plan"]
A --> B["Check resource and environment boundary"]
B -->|"Outside allowed boundary"| S["Stop or request policy exception"]
B -->|"Within allowed boundary"| P["Check action authority"]
P -->|"Approval required"| H["Human decision"]
P -->|"Already authorized"| E["Execute and record result"]
H -->|"Approved"| E
H -->|"Declined"| S
E --> A
This is a conceptual control flow. In an implementation, the policy and permission checks should be enforced by the surrounding system rather than depend only on the agent following instructions.
Human-governed is therefore a useful qualifier. It doesn't require a new rung on the ladder. Governance can include approval before action, ongoing supervision, or preauthorized operation inside a defined boundary.
For an initial map, put AI dependence on one axis and model-directed agency on the other. Keep execution authority visible beside each example.
The following are hypothetical configurations, not ratings of particular vendors.
| Model-directed agency ↑ / AI dependence → | Optional AI | Meaningful AI | AI-dependent core |
|---|---|---|---|
| Agentic loop | An optional helper that plans and adapts a minor task. | An expense agent investigating rejected claims. | A coding environment centered on agents that inspect, edit, and test. |
| Model-directed steps | An optional helper selecting its next suggestion. | A model selecting a route through a useful workflow without an adaptive execution loop. | An AI-dependent product with model-selected steps but no feedback-based replanning. |
| Human/code directed | A conventional product with optional text rewriting. | Expense software with a predefined receipt extraction step. | A prompt-to-image product with a fixed generation flow. |
Every row can have its own approval requirements and execution boundaries. The top row does not automatically imply permission to act independently.
Products don't have to travel diagonally across this map.
A team can add model-directed planning to one workflow without rebuilding the product around AI. Another can make AI central to the product while preserving a deliberately human-directed experience.
The map is a conversation aid, not a scorecard. Its cells are not equally spaced measurements, and the top right corner is not a destination every product should pursue.
Agent-native sits near the intersection of AI-native product design and agentic execution. That is a useful starting point, but it needs more substance than adding the two labels together.
I would use agent-native for a product that treats agents as first-class participants. They have defined identities, scoped permissions, access to task state, and ways to hand work back to people. Its workflows anticipate partial completion, retries, interruptions, and review.
An AI-native writing tool can remain a sequence of human prompts. An agent-native coding environment would instead support handing over an issue, tracking the investigation, inspecting changes, and intervening while work is underway.
The design question shifts from "Where do we put the chat box?" to "How does delegated work live inside this product?"
That doesn't mean humans disappear. Review, disagreement, clarification, and cancellation become important parts of the interface.
flowchart TD
P["Product designed around AI"] --> N["Candidate for agent-native design"]
E["Model-directed task execution"] --> N
N --> T["Persistent task state and handoffs"]
N --> B["Agent identity and bounded permissions"]
N --> H["Human review and intervention"]
This is a proposed design interpretation of agent-native, not a certification or a settled industry definition. An AI-native product with a single agent feature would not automatically satisfy it.
There is another question hiding inside "we use AI."
Where?
A feature, a workflow, a whole product, or several connected organizational processes?
An agent might independently handle one narrow task while everything around it remains manual. Another system might assist people across an entire business without receiving permission to act independently anywhere.
Scope is the reach of AI participation. It doesn't automatically increase agency or autonomy.
Feature, workflow, product, and organization are useful scopes to name. They aren't a compulsory sequence. An organization contains many products and workflows, each with a different profile.
That's why "our company is AI-native" is usually too coarse to guide an engineering decision. The useful unit may be the workflow, or even the individual action.
Consider the expense example again:
| Part of the work | Who chooses the path? | What can execute independently? |
|---|---|---|
| Receipt extraction | A predefined processing flow | Extract fields and flag uncertainty. |
| Rejection investigation | The agent chooses checks and follow-ups | Read policy and claim history within access limits. |
| Corrected submission | The agent can prepare the action | Submission requires the employee's approval. |
| Reimbursement | A separate financial control process | Payment follows its own authorization rules. |
One workflow contains several arrangements. A single label loses them.
The product map only describes part of the arrangement. The system underneath can have its own combination of choices.
Neural describes a model family. Generative describes a capability to produce outputs. Neither implies that the system acts independently.
Neuro-symbolic combines neural methods with symbolic representations or reasoning.[2] For example, perception might identify objects and relationships, then a symbolic solver reasons over that representation. An LLM surrounded by a few validation checks is not automatically a meaningful neuro-symbolic architecture.
Compound AI combines interacting components such as models, retrieval, tools, and other processing steps.[3] A retrieval-and-generation pipeline can be compound without being agentic. A neuro-symbolic system can also be compound. These categories overlap.
Multi-agent describes an arrangement involving multiple agents. It says nothing by itself about autonomy, reliability, or whether the decomposition is useful. Several agents can run inside a tightly controlled workflow. A single agent can receive broad authority.
Anthropic's account of its research system illustrates a real multi-agent implementation, but also describes coordination, evaluation, and resource tradeoffs.[4] It is evidence that the pattern can work for particular tasks, not that every application should move toward it.
The architecture diagram therefore uses branches and overlapping relationships. Neural → neuro-symbolic → compound → multi-agent would incorrectly suggest that each replaces the last.
flowchart TD
S["AI system construction"] --> M["Model capabilities"]
S --> C["Component composition"]
S --> E["Execution arrangement"]
M --> N["Neural and generative models"]
C --> NS["Neuro-symbolic integration"]
C --> CA["Compound models, retrieval and tools"]
E --> F["Predefined workflow"]
E --> A["One agent or multiple agents"]
NS -. "can be part of" .-> CA
A -. "can be part of" .-> CA
Oversight belongs beside this architecture. In a human-in-the-loop arrangement, a person participates at a defined decision or approval point. Human-on-the-loop commonly means supervision with the ability to intervene. Human-out-of-the-loop means no routine human participation in the specified execution loop, not the absence of human responsibility for its design and operation.
A dashboard does not establish effective supervision. Can someone understand the state, stop execution in time, and recover from the result? Those are the practical questions behind on-the-loop.
Deployment is separate again. A local model can power a human-directed assistant. A cloud agent can require approval for every write. A hybrid system can keep some processing local and invoke remote services when needed. Edge deployment is not a higher degree of AI than cloud deployment, and local inference alone does not prove that data never leaves the device.
The complete description is closer to this:
AI-native product × agentic workflow × bounded autonomy × human approval for writes × compound architecture × hybrid deployment.
Longer than a badge. Much more informative.
Rejecting one universal ladder doesn't mean nothing evolves.
A product may move from an optional AI feature to workflows that assume AI is present. A person may move from asking for suggestions to handing over an outcome. A system may expand from a single model call to a composition of retrieval, planning, verification, and execution.
The mistake is assuming these transitions happen together, in order, or permanently.
The following diagram shows possible product and work-design transitions. Arrows mean a possible redesign, not chronology, superiority, or an obligation to proceed.
flowchart TD
T["Traditional product"] --> X["AI-enabled feature"]
X --> I["AI integrated into a workflow"]
I --> P["AI-first redesign"]
P --> N["AI-native core experience"]
I --> D["Bounded outcome delegated"]
N --> D
D --> W["Predefined execution path"]
D --> A["Agent chooses the task path"]
A --> L["Human approves consequential actions"]
A --> B["Independent execution within limits"]
A product can enter this picture at a different point. An AI-native product can be built from scratch. A traditional product can gain an agent without first passing through an AI-first redesign. After incidents or changing requirements, a team can reduce autonomy while keeping the same agent architecture.
That last move matters. Reducing authority can be an improvement in the system's design.
If these are degrees, should we score them?
Only when the measure has a clear denominator and a decision it will inform. I wouldn't assign a universal AI-native score or treat agents per workflow as a performance metric.
For a specific workflow, observable behavior is more useful.
| Dimension | Evidence to collect | What the measure cannot prove |
|---|---|---|
| AI dependence | Which promised capabilities disappear or become uneconomic without AI; observed fallback performance. | A larger dependency makes the product better. |
| Agency | Which planning and routing decisions are model-directed, versus fixed in code or supplied by a person. | More decisions should move to the model. |
| Autonomy | Share of eligible tasks completed without intervention, broken down by action class and risk. | Fewer interventions mean correct outcomes. |
| Authority and boundary | Actions requiring approval; enforced tool, resource, environment and budget restrictions. | An agent's stated restrictions match its effective access. |
| Scope | Which workflow steps and systems AI can read or change. | Wider reach creates more useful value. |
| Oversight | Approval coverage, escalation quality, time to detect and halt, recovery success. | The presence of a reviewer ensures effective control. |
| Outcome | Verified task success, rework, end-to-end time, cost per accepted result. | A faster response creates a better business outcome. |
For example, a higher no-intervention completion rate can reflect improved reliability, easier tasks, or missed escalations. Measure accepted results and failures alongside it. The rate alone can't tell you which happened.
Architecture and deployment need fit-for-purpose comparisons: cost, latency, failure isolation, data movement, and maintainability. They don't need maturity points.
The ladder doesn't just confuse language. It smuggles in a product strategy: more AI, more agency, and less intervention must mean progress.
I don't think that follows.
A fixed workflow can be the right choice when the path is known. Human approval can be part of the product's value. AI can remain optional because the user needs a predictable fallback.
Greater independence can save coordination time. It can also increase the amount of work a system performs before someone catches an error. Broader scope can connect a process, while also exposing more systems to a mistaken decision.
These are design tradeoffs. A vocabulary that calls every increase "maturity" makes them harder to discuss honestly.
Governance therefore belongs around the profile: permissions, evidence requirements, monitoring, escalation, and recovery. It is not another degree of AI. It makes the chosen degrees accountable.
And capability is not authority. A model's ability to produce a plausible action doesn't establish the system's right to execute it.
This is the future-facing part of the argument. The following are plausible directions, not a forecast that every product will follow. Existing compound systems, multi-agent implementations, and tool protocols provide evidence of the building blocks.[3][4][5] The implications below are my interpretation.
Products may serve two kinds of users: people and agents. A person needs navigation, explanations, and visible choices. An agent needs discoverable capabilities, structured inputs, task state, and clear permission boundaries. MCP already standardizes connections between AI applications and external tools and data.[5] Connectivity makes an agent-facing product possible; it does not make that product AI-native or establish authority to perform every exposed action.
An expense product could retain its ordinary interface while allowing an external agent to prepare a claim. The application remains responsible for validating the submission and enforcing payment controls. The agent interface changes who can operate the product, not who owns its rules.
More work may move from requests for answers to delegated outcomes. Instead of "summarize these receipts," the user asks "prepare my claim and tell me what is missing." That shift requires persistent work, progress visibility, cancellation, and exception handling. The interface may become a place to supervise tasks as well as start them.
Autonomy may become configurable by action rather than advertised for the whole product. Read records independently. Prepare a draft independently. Submit below a defined threshold. Ask before sending money or changing an external commitment. One system can support several degrees at once.
Composition may matter more than the headline model. Better models can simplify some pipelines. Other tasks will still benefit from retrieval, deterministic calculations, domain models, verification, or specialized agents. Compound AI is already an established system-building idea.[3] My expectation is that successful designs will combine these selectively, rather than expand every workflow into a multi-agent network.
Oversight may shift from watching every step to designing and checking the operating boundary. That is only useful when people can inspect evidence, detect failures, intervene, and recover. Exception-only review can become superficial if the system fails to recognize the exceptions. Less frequent human involvement requires better system evidence, not merely more confidence in the model.
Deployment may remain mixed. Some tasks can benefit from local execution; others require shared systems or remote resources. Workload, data boundaries, economics, and response time will influence the split. There is no reason to assume every agent ends up entirely in the cloud or entirely on a device.
Those directions will look different across domains.
| Domain | A plausible next configuration | The design question that matters |
|---|---|---|
| Coding | Agents own investigations and proposed changes; people review intent and release. | What evidence makes a change ready to merge? |
| Enterprise operations | Agents coordinate work across existing products. | Which system owns transaction state and authorization? |
| Creative tools | AI-native generation with both direct human control and optional delegation. | How does the creator retain editorial control? |
| Personal assistance | Agents prepare and execute selected routine tasks. | Which commitments need explicit approval? |
These are examples of possible configurations, not claims of universal adoption. Each can coexist with simpler tools and conventional workflows.
The future I expect is a mixture of profiles. A system may become more agentic while its execution boundaries become more precise. A product may become easier for agents to operate while remaining conventional underneath. A workflow may become more automated without becoming more dependent on generative AI.
That gives us a better question than "What comes after AI-native?"
Which dimension changes next, what does that enable, and what new obligation does it create?
Assess one product or workflow as it behaves today. Answers stay in this page. This is a working interpretation, not a validated maturity score.
| Placement | Rule |
|---|---|
| Optional AI | Removing AI preserves the central outcome and no important workflow materially depends on it. |
| Meaningful AI | Central outcome survives removal, but an important workflow materially improves with AI. |
| AI-dependent core | Removing AI eliminates the central promised outcome at viable quality and cost. |
| Human / code directed | The model does not choose the next task step. |
| Model-directed steps | The model chooses steps but does not revise its plan from execution feedback. |
| Agentic loop | The model chooses steps and adapts its plan from execution feedback. |
The four questions below are a useful mnemonic. To actually place a product on the map, we need a more concrete questionnaire.
Assess one named product or workflow as it behaves today. Don't answer for its roadmap. If a whole organization has several different arrangements, assess those workflows separately before describing the organization.
Use Yes / No / Not sure for the following checklist. Unknown answers stay unknown; they aren't silently scored as no.
| Dimension | Checklist questions |
|---|---|
| AI dependence | Would removing AI eliminate the central promised outcome at viable quality and cost? Does AI materially improve at least one important workflow? Are core workflows designed around AI uncertainty, evaluation, and fallback? |
| Agency | Does the model choose the next task step rather than follow only a human or coded sequence? Does it revise its plan from execution results or changing conditions? Can it invoke tools to inspect or change its environment? |
| Agent-native design | Do agents have distinct identities and scoped permissions? Can delegated tasks retain state and resume across sessions? Are clarification, review, and hand-back to people built into task handling? |
| Oversight | Can people inspect task state and evidence behind proposed actions? Can a responsible person stop execution before further consequential actions? Is there a defined recovery or compensation process for failed actions? |
| Execution boundary | Are permitted tools and resources explicitly listed? Are package sources, data access and environments technically restricted? Are spending, runtime or action limits enforced? Are attempts to cross those boundaries denied or escalated? |
Then classify execution authority separately for reading/investigation, draft preparation, internal record changes, and external commitments/payments/releases. Each action class can be unavailable, require approval, execute independently, or remain unassessed. Separately record which tools, resources, environments and budgets constrain it. These answers describe the actual arrangement; they do not establish that it is appropriate.
Finally, select the scope, architecture, agent arrangement, and deployment. Those are descriptive tags, not points to add up.
The horizontal axis has three descriptive positions:
Core dependence plus design around AI uncertainty, evaluation, and fallback supports an AI-native interpretation. Dependence alone does not establish that design. AI-first remains an expression of priority rather than a coordinate inferred from runtime behavior.
The vertical axis also has three positions:
Tool use is supporting context, not sufficient proof of an agentic loop. A model can also run a reasoning loop without an external write tool. These working placement rules deliberately avoid rewarding tool count.
An AI-dependent core with AI-native design evidence, an agentic loop, agent identities, persistent task state, and human handoffs becomes an agent-native candidate in this proposed framing. It is a design interpretation, not a certification.
Unanswered placement questions leave that axis unplaced. Conflicting answers pause placement: for example, saying AI is necessary for the central outcome while denying that it materially improves any important workflow. Such conflicts should prompt clarification, not disappear into an average.
An interactive version can move the marker as these answers change. Execution authority stays visible by action class; enforced boundaries, scope and architecture sit alongside the grid. The result is an explainable profile, not a combined maturity score. These rules are an author's heuristic for discussion, not a validated assessment instrument.
The questions to remember are:
Those questions turn a positioning statement into something we can examine.
"We are AI-native" tells me very little about how a product behaves. "AI plans the investigation, can read these systems, stays inside these boundaries, and needs approval before changing a claim" tells me enough to start a useful discussion.
Degrees of AI should describe a system's choices. They shouldn't rank its ambition.
[1] Anthropic, Building effective agents, originally published December 19, 2024. Referenced for its distinction between predefined workflows and model-directed agents. The dependence map and working definitions in this essay are the author's proposed framing, not an established taxonomy.
[2] IBM Research, Neuro-symbolic AI. Referenced for the combination of neural/statistical methods and symbolic knowledge or reasoning.
[3] Berkeley AI Research, The Shift from Models to Compound AI Systems, February 18, 2024. Referenced for compound system construction and its distinction from a single model.
[4] Anthropic, How we built our multi-agent research system, June 13, 2025. Referenced as an implementation account, not a general claim that multi-agent systems outperform single-agent systems.
[5] Model Context Protocol, Specification. Referenced for connecting AI applications to external tools and data. Protocol connectivity should not be confused with application-level execution authority.
The terminology tables are working interpretations. The agent-native definition, evolution paths, measurement suggestions, and future scenarios are the author's synthesis. Arrows in the diagrams are defined in their surrounding text and do not establish a universal maturity ladder.