•
20 min read
•
Agustinus Nathaniel

Not All AI Agents Are Built the Same

https://images.unsplash.com/photo-1669295384050-a1d4357bd1d7?auto=format&q=80

If you’ve been following AI lately, you’ve probably noticed that almost everything is being called an “AI agent.” The label spans products such as ChatGPT, Grok, Claude Code, Codex, Cursor, Devin, OpenClaw, and Hermes Agent, alongside newer products like Dots, Gemini Spark, Muse, and Grok Bot.

Yet they don’t necessarily work the same way. Some primarily interact through conversations, while others can access your computer, write code, browse websites, or interact with applications on your behalf.

Some can continue working after you’ve closed your laptop, while others can even initiate tasks without you explicitly asking.

I’ve been experimenting with several of these systems, including coding agents like Codex, Claude Code, and OpenCode, persistent agents like OpenClaw and Hermes Agent, and more recently, Devin.

What I find interesting is how differently these agents can behave, even when using similar underlying AI models. These differences matter for everyday planning, office routines, and business operations as much as they do for software development.

In my previous publications, AI Basics for Engineers and The AI Development Toolkit, I explored the foundations of language models, agentic workflows, and the tools we use in software development.

But there’s another question I’ve been thinking about: what actually makes these AI agents different, and how should we make sense of the growing ecosystem?

Initially, I tried categorizing them into tiers based on their capabilities, starting from conversational assistants, then agents equipped with tools, persistent agents, and eventually fully managed autonomous agents. It seemed reasonable at first.

But the more I explored how these systems work, the more I realized that the distinction isn’t quite that straightforward. An agent that runs continuously isn’t necessarily smarter than one that doesn’t.

A cloud-hosted agent isn’t automatically more capable than a locally running one. Their design, available resources, and the way their components work together all shape what they can accomplish.

So here’s a practical mental model I’ve settled on to understand today’s AI agent ecosystem.


1. It’s More Than Just the Model 🔗

A common way to compare AI products is by looking at the models powering them. We compare model families such as GPT, Claude, and Gemini.

But the model alone doesn’t determine what an agent can actually accomplish.

A simple equation I find useful is:

AI Agent ≈ Model + Harness

The model interprets instructions, processes information, and proposes responses or actions. Its reasoning can still be mistaken.

The harness is the system surrounding the model that makes those capabilities operational. It manages how information is provided to the model, what tools are available, how actions are executed, and how the results are processed.

A prompt gives the agent instructions. A chat window provides an interface, and a calendar integration provides a tool. These can all be parts of an agent product.

The harness coordinates them with the model, so instructions can turn into actions and their results can inform the next step.

For example, imagine asking an AI to prepare you for tomorrow’s meeting. The model might suggest an agenda.

To prepare one from your actual work, the agent needs access to your calendar, relevant documents, and previous meeting notes. To save the agenda or create follow-up tasks, it also needs tools that can make those changes. The harness coordinates that process.

The same idea applies to fixing a software bug: the model proposes a solution, while the surrounding system provides access to the code and tools to edit and test it.

Instead of generating a single response, the agent can repeatedly decide what to do, execute an action, and observe the result. It continues until it finishes, needs input, or reaches a limit.

This is commonly known as an agent loop, something I’ve covered in AI Basics for Engineers.

The equation is shorthand: the harness depends on several supporting components, which may be bundled into it or provided separately.

The components surrounding an agent 🔗

Here’s a simplified view of the components that influence how an agent operates:

ComponentResponsibility
ModelHow it reasons and understands tasks
HarnessHow it operates and coordinates actions
ToolsWhat it can interact with or execute
RuntimeThe environment where its tools and processes run
Memory & StateWhat information persists between interactions
TriggersWhat initiates its work
PermissionsWhat it’s allowed to access or modify
flowchart TB
    accTitle: The components of an AI agent
    accDescr: Triggers initiate work through the harness. The harness exchanges context and proposed actions with the model, reads and writes memory and state, and invokes tools that return results. Permissions constrain tool actions, and tools execute in a local or cloud runtime.

    Triggers[Triggers] --> Harness[Harness]
    Model[Model] <-->|Context and proposed actions| Harness
    Memory[Memory and state] <-->|Read/write state| Harness
    Harness <-->|Actions and results| Tools[Tools]
    Permissions[Permissions] -.->|Constrains| Tools
    Tools --> Runtime[Runtime: local or cloud]

The harness connects the model to tools and retained context. Triggers start work, permissions constrain actions, and the runtime provides the environment where tools execute.

Not every system implements these components separately. Some are built directly into the harness, while others are provided by external services. But understanding their responsibilities helps explain why different agent products behave differently.

Think of it this way: two people might have similar knowledge and abilities, but their effectiveness can differ depending on the tools, resources, working environment, and responsibilities available to them. The same applies to AI agents.

The model’s capabilities and the surrounding system jointly determine what the agent can do.


2. What Actually Makes AI Agents Different? 🔗

Rather than immediately categorizing different products, I think it helps to understand a few independent characteristics first.

What can the agent access? 🔗

An agent’s capabilities depend partly on the tools and resources available to it. A conversational assistant might have access to web search, connected applications, and document analysis.

A coding agent might also have access to your files, a terminal for running commands, development tools, and repositories where source code is stored. Some agents can even operate a browser or desktop environment much like a human would.

This doesn’t necessarily make one agent more intelligent than another. It means they have different ways of interacting with their environment.

Access and permission also deserve separate attention. An agent may support an action while its configured permissions restrict when it can perform it or require approval first.

Where does the work happen? 🔗

This is an important distinction that I initially overlooked. When using an agent locally, such as Claude Code or OpenCode, the model itself may still run on a remote inference provider. Meanwhile, the tools and commands can execute on your own computer.

So there are effectively two different concerns:

  • Model inference: Where the AI model processes information.
  • Execution environment: Where the agent actually performs actions.

These don’t have to be in the same place. For example, an agent could use a model hosted by Anthropic while executing commands on your laptop.

Alternatively, it could run those commands inside a cloud virtual machine provided by the agent platform. Some products even support both local and cloud execution.

Devin is an interesting example. Its CLI can hand off a local task to a cloud session, carrying over conversation context, the repository and branch, and uncommitted changes. The cloud session starts in a fresh virtual machine; the local running processes don’t move with it.

Claude Code also supports cloud sessions, but the transfer options depend on the interface. Its CLI can start a new cloud task or pull a cloud session into the terminal. Sending an existing local session to the cloud is supported through the Desktop app.

The important point is that where you interact with an agent doesn’t necessarily determine where its work happens. A message sent through Slack could trigger work on a local computer, a remote server, or a provider-managed cloud environment.

Can it continue working without you? 🔗

Some agents are designed around interactive sessions. You give them a task, they work on it, and the interaction eventually ends.

Others support ongoing operation through persistent sessions, scheduled tasks, or external events. For example, you could configure an agent to check for new information every morning, respond to incoming messages, or start investigating when an application reports an error.

This is often associated with the term always-on agent.

But there’s a subtle distinction: always-on doesn’t necessarily mean the AI model is continuously running or thinking. Usually, it’s the surrounding infrastructure that remains available.

When a scheduled task, message, or event arrives, the system invokes the agent to perform the necessary work. Once the task finishes, the agent can return to an idle state.

It helps to separate four capabilities:

  • Persistence: Relevant state survives between interactions.
  • Reachability: The service remains available to receive requests.
  • Proactivity: Schedules or other configured triggers can initiate work.
  • Autonomy: The agent can make decisions and act with limited human intervention.

An agent can remember previous interactions without being proactive. It can also execute a complex task autonomously within a single session.

For a self-hosted system, continued reachability depends on the host staying online. Persistent memory alone won’t keep an agent working after you shut down the computer running it.

Who manages the infrastructure? 🔗

Finally, there’s the question of who provides and maintains the environment. With self-hosted agent systems, you might be responsible for installing dependencies, configuring integrations, provisioning servers, managing credentials, and keeping everything operational.

With managed agent platforms, much of that responsibility shifts to the provider. You configure an agent, grant the necessary permissions, and let the provider handle its execution environment.

Neither approach is inherently better. Self-hosting can provide greater control and flexibility, while managed platforms reduce the operational work required from the user.

This distinction becomes particularly important as agents move beyond occasional conversations into ongoing responsibilities.


3. Making Sense of the Agent Ecosystem 🔗

With those characteristics in mind, we can start identifying a few common patterns among modern AI agent products.

I’d group them into five broad categories:

CategoryPrimary focusExamples
Conversational AssistantsConversation-first interaction with AIChatGPT, Claude, Grok, Microsoft Copilot
Execution AgentsPerforming actions through tools and execution environmentsClaude Code, Codex, Cursor, OpenCode
Persistent Agent SystemsOngoing operation, memory, messaging, and scheduled tasksHermes Agent, OpenClaw
Agent Workspaces & CoordinatorsManaging agent sessions, workspaces, and workflowsT3 Code, Jean, Conductor, Paperclip, Multica
Managed Agent PlatformsProviding integrated agent capabilities with managed infrastructureDevin, Conductor Cloud, Dots, Gemini Spark, Muse, Grok Bot

These categories aren’t official industry standards, nor are they mutually exclusive. They’re simply useful descriptions of the different approaches products take.

Conversational Assistants 🔗

These are the most familiar AI experiences. You open an application, start a conversation, and interact with the assistant through a managed interface.

Modern conversational assistants aren’t necessarily limited to answering questions. They can also access tools, analyze documents, interact with connected services, and perform more substantial tasks.

Their defining characteristic is the conversation-first experience, typically within an environment managed by the provider.

Execution Agents 🔗

These focus on accomplishing tasks by interacting with tools and execution environments. Claude Code, Codex, and OpenCode are good examples in software development.

They can inspect repositories, modify files, execute commands, and iterate on changes.

Tool-based execution is useful beyond coding, too. Depending on the available tools, an agent could organize files, update a spreadsheet, or create tasks from meeting notes. The products listed here emphasize software development, while general-purpose products may offer these actions through other interfaces.

These products also span different execution environments. Claude Code and Codex support cloud execution, allowing tasks running in cloud environments to continue independently of your local computer.

Persistent Agent Systems 🔗

Hermes and OpenClaw place greater emphasis on continuity. They integrate agent execution with persistent sessions, memory, messaging channels, and mechanisms for scheduled or event-driven work.

You might run one on a server and interact with it through Telegram, Discord, or another messaging application. These systems are particularly interesting when you want an agent to remain available and handle recurring responsibilities.

Agent Workspaces & Coordinators 🔗

Some products focus on giving developers a workspace for existing coding agents.

T3 Code describes itself as a control plane for coding agents. It brings multiple agent harnesses into one interface, using your existing subscriptions or credentials, and provides branch and pull-request workflows around their sessions. Its documented setup runs agents on your computer, including when you control them from a web or mobile interface.

In this mental model, I’d place it in the workspace and coordination layer.

Jean also fits here. It manages projects, agent sessions, terminals, and isolated Git worktrees: separate working copies of a code project. You can run it on your laptop or a server you operate, then connect through its desktop app or a browser.

Its remote mode is a useful example of separating the interface from the execution environment.

Conductor spans this category and managed agent platforms. It runs existing coding agents under the hood and provides a shared workspace for managing their work. It supports both local and cloud execution. Conductor Cloud adds isolated virtual computers, called microVMs, managed by the provider and prepared with your code and the software it needs.

Workspace management and automated coordination are related responsibilities. A workspace can help you manage several sessions while you still decide what each agent should do.

Paperclip and Multica approach the problem from another direction. Instead of focusing solely on what an individual agent can do, they help coordinate work across multiple agents. This introduces concepts such as agent identities, tasks, execution runs, assignments, scheduling, and monitoring.

An important distinction is that these systems don’t necessarily replace the underlying agent harness. They can coordinate existing execution agents such as Claude Code, Codex, or OpenCode. Think of them as systems that organize who does what, while the underlying agents perform the actual work.

Managed Agent Platforms 🔗

Products like Devin, Dots, Gemini Spark, Muse, and Grok Bot represent another approach. They combine agent capabilities with provider-managed infrastructure, reducing the amount of setup and maintenance users need to handle.

Some provide cloud computers where agents can interact with applications, browse websites, and execute tasks in the background. Conductor Cloud fits this pattern too: it supplies managed execution environments while retaining the workspace layer described above.

Others focus on creating persistent assistants that can communicate through familiar interfaces such as chat or workplace messaging applications.

Devin, for instance, focuses on software engineering workflows and integrates local execution, cloud sessions, and collaboration through tools like Slack.

Dots and Gemini Spark take a broader approach toward general-purpose and personal work. Muse and Grok Bot also provide personal agents with managed cloud computers.

These products differ in what they support and how their environments are implemented, but they share an important characteristic: much of the infrastructure is operated on the user’s behalf.


4. These Aren’t Levels of Intelligence 🔗

One mistake I initially made was assuming these categories formed a hierarchy, with conversational assistants at the bottom and managed platforms at the top. That would imply each category is an increasingly advanced version of the previous one.

In reality, their capabilities overlap.

Claude Code is primarily an execution agent, but it also supports cloud sessions and scheduled or event-triggered routines. Hermes is a persistent agent system, but it also implements its own execution loop and supports tools.

Paperclip can coordinate other agents without necessarily replacing the systems that execute their tasks. And Devin combines several capabilities into a managed engineering platform.

Each product packages a different combination of capabilities. The categories describe product emphasis, and a single product can fit several of them.

Everyday, office, and business examples 🔗

You don’t need to write software to use this mental model. Consider a few tasks:

  • Daily life: “Help me plan next week.” A conversational assistant could suggest a schedule from details you provide. An agent connected to your calendar could work with actual availability and create events with permission. A recurring trigger could make this a weekly routine.
  • Office work: “Prepare me for tomorrow’s meetings.” An agent could gather relevant documents, summarize previous decisions, draft agendas, and create follow-up tasks. Useful memory might include your preferred agenda format; the necessary tools depend on where your team keeps its documents and tasks.
  • Business: “Give me a weekly summary of sales and customer questions.” An agent could gather figures from a spreadsheet or sales system, summarize support conversations, and draft a report with links to its sources. Running the task every Monday requires a trigger and an execution environment that remains available.

Which of these tasks an agent can carry out depends on its connected tools, permissions, and product support. For a one-off draft, a conversational assistant may be enough.

Recurring work that draws on several systems needs more of the surrounding infrastructure.

Combining capabilities in a workflow 🔗

Suppose you want an AI to help maintain an application. You could use a coding agent to investigate and fix a bug directly. You could configure a persistent agent to respond to monitoring alerts and initiate investigations.

You could use an agent coordinator to distribute investigation, implementation, and verification across several agents. Or you could delegate the work to a managed platform operating in its own cloud environment.

These approaches aren’t necessarily competing with each other. You could combine several of them into a single workflow. The same pattern also applies to a weekly business update:

---
config:
    flowchart:
        rankSpacing: 24
        padding: 10
        subGraphTitleMargin:
            top: 4
            bottom: 16
---
flowchart TB
    accTitle: One possible weekly business-update workflow
    accDescr: A Monday morning schedule triggers an agent system kept online, which starts a task with an agent coordinator. Execution agents gather sales and support updates, draft a summary, and check its figures and source links in sequence. The proposed report goes to a human for review. A single agent could also handle all three execution stages.

    Schedule[Monday morning schedule] --> Service[Agent system kept online]
    Service --> Coordinator[Agent coordinator]
    Coordinator -->|Assigns work| Gather

    subgraph Execution[Execution agents]
        Gather[Gather sales and support updates] --> Draft[Draft the weekly summary]
        Draft --> Check[Check figures and source links]
    end

    Check --> Result[Proposed report]
    Result --> Review[Human review]

This is one possible configuration. The coordinator assigns and tracks the work, while execution agents handle the individual stages.

The agent system could run on a computer or server you keep running, or within a managed platform. A single agent could also handle all three execution stages. The question isn’t which one belongs to the highest tier, but which combination makes sense for the work you’re trying to accomplish.


5. More Capabilities Don’t Necessarily Mean Better Agents 🔗

Another thing I’ve learned from experimenting with different agent systems is that adding more capabilities doesn’t automatically improve the result.

More tools can introduce unnecessary complexity. Persistent memory can be useful, but only if the system manages that information properly. More autonomy can reduce manual intervention, but it also increases the potential consequences of mistakes.

And a more sophisticated architecture doesn’t necessarily translate into better performance for a particular task.

Sometimes, a conversational assistant can draft your agenda from notes you paste in. A coding agent with access to the right repository and tools may likewise be sufficient for a software task.

Other times, having an agent that can operate in the background, retain context, and coordinate work across different systems can be considerably more valuable.

There’s also an operational trade-off. Running your own agent infrastructure provides flexibility but comes with maintenance responsibilities.

Using a managed platform simplifies that experience, but usually means accepting the provider’s limitations, pricing, and security boundaries. Neither approach is universally better.

The right agent is the one whose capabilities, constraints, and operating model fit the problem you’re trying to solve.


Conclusion 🔗

The AI agent ecosystem is evolving quickly, and I don’t expect today’s terminology or product boundaries to remain unchanged. What interests me most isn’t necessarily the growing number of agent products, but how they’re approaching the same underlying problem from different directions.

Some focus on improving the harness and execution capabilities. Others focus on persistence, coordination, or making those capabilities accessible without requiring users to manage the infrastructure themselves. Some products expose these components separately; others bundle them into a single experience.

That’s why I find it more useful to understand the underlying responsibilities rather than memorize a list of agent categories.

The next time you encounter a new AI agent, consider asking:

  • What model and harness does it use?
  • What tools and resources can it access?
  • Where does it execute its work?
  • Can it operate independently or respond to events?
  • Who manages the infrastructure, and what control do I have?

You don’t necessarily need to understand every technical detail behind these questions. But having a basic understanding of them makes it easier to evaluate what an agent actually offers, beyond its marketing or the model powering it.

Start with the work you want to accomplish, then choose the capabilities that matter for it. That’s the mental model I find most useful for navigating the current AI agent ecosystem.


References & Further Reading 🔗

For those interested in exploring the technical details or product implementations mentioned in this article, here are some useful resources.

Previous Publications 🔗

Agent Harnesses & Execution 🔗

Persistent Agents & Coordination 🔗

Managed Agent Platforms 🔗

The product examples reflect publicly documented capabilities as of October 2026. The categories presented here are a practical mental model rather than an official industry classification.