What is an AI agent? Definition, examples and limits
An AI agent is a language model that decides for itself, step by step, which tools to call to achieve a given goal, instead of answering once and stopping like a chatbot: it observes the result of each action and starts again until the task is complete or it gives up.
An AI agent is a program in which a language model decides for itself what actions to take to achieve a goal: it chooses a tool, observes the result, and repeats until it is finished. That is what distinguishes it from a chatbot, which responds and then stops. The term “agentic AI” refers to this way of building systems. As of September 20, 2026, the most useful agents are those that write code and those that conduct research. This page provides a clear definition, real examples, limitations, and what it takes to run one on your own machine.
#What is an AI agent?
The most useful definition fits in one sentence: an agent is a language model that uses tools in a loop. You give it a goal, such as “find out why this test fails and fix it.” It reads the code, runs the test, reads the error, modifies a file, reruns the test, and continues until the test passes or it gives up. No one wrote this sequence in advance: the model decides it step by step, based on what it observes.
The term is a victim of its own success. Many products marketed as “agents” are simply chatbots with a name, or conventional automations in which a model writes text at a fixed step in a script. The decisive question is: who decides what happens next? If it is a predefined program with a fixed sequence of calls, it is not an agent, even if a language model is involved somewhere in the chain.
#Chatbot, workflow, agent: the differences
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
| Chatbot | AI workflow | Agent | |
|---|---|---|---|
| Who decides the steps | No one: one question, one answer | The developer, in advance | The model, along the way |
| Tools | None, or a web search | Yes, in a fixed order | Yes, choose freely |
| Predictability | High | High | Lower |
| Cost per task | A call to the model | A few calls | Dozens of calls, sometimes more |
| Good use | Questions, writing | Well-defined repetitive tasks | Open-ended tasks whose path is unknown |
This is the distinction Anthropic draws in its design guide: “workflows are systems where LLMs and tools are orchestrated through predefined code paths,” whereas “agents are systems where LLMs dynamically direct their own processes and tool use, retaining control over how tasks are accomplished.” The resulting advice applies to everyone: find the simplest solution possible, and increase complexity only when necessary. A well-designed workflow beats an agent for most business needs because it is predictable and inexpensive.
#The four parts of an agent
- A model
- The brain. It must reason through multiple steps and produce well-formed tool calls. It is the component that limits everything else.
- Tools
- Functions the model can call: read and write files, execute a command, search the web, query a database. The MCP protocol has become the standard way to connect them.
- A loop
- The program that chains everything together: send the context to the model, run the requested tool, add the result to the context, and repeat. It also sets the guardrails: maximum number of steps, confirmations, and budget.
- One memory
- In the short term, the context window, which fills up at every step and eventually saturates on a long task. In the long term, note files, a compressed summary of past steps, or a database that the agent rereads from one session to the next, so it does not have to start from scratch with every new conversation.
#The mechanism behind the loop: ReAct
Most production agents follow, without naming it, the pattern described in 2022 by researchers from Google and Princeton under the name ReAct (“Reasoning and Acting”). At each turn, the model produces explicit reasoning (Thought), chooses an action, usually a tool call (Action), and then receives the result (Observation) before starting over. Making the reasoning explicit before the action makes the agent easier to debug: you can see why it chose a given tool, not just what it did.
#Examples that actually work
- Code agents
- The most advanced use case. Claude Code, Cline, OpenCode, or Aider read a repository, modify multiple files, run tests, and fix their errors. The domain is well suited to this: the result can be verified automatically by whether a test passes.
- Research agents
- They launch dozens of searches, read the pages, cross-check sources, and write a summary. Useful for monitoring or a state-of-the-art review, provided you verify the citations.
- Operations agents
- They monitor logs, diagnose incidents, and propose or apply fixes. In production, almost always with human validation before any action.
- Desktop agents
- They control a browser or computer with the mouse and keyboard. Impressive demonstrations, but reliability is still insufficient for unsupervised use.
What these successful cases have in common is an objective success criterion that is quick to verify. A test that passes or fails, a well-formed file, an incident resolved or not. Where success depends on vague human judgment, such as “write a convincing marketing strategy,” an agent loop adds little compared with a simple back-and-forth with the model, because there is nothing objective to observe between steps to guide the next one.
#Why agents fail
The first reason is arithmetic. In a loop, errors multiply instead of canceling out. A model that is 95% reliable at each individual step completes a ten-step task correctly in only 60% of cases, and a twenty-step task in barely 36% of cases, even though no individual step is outright bad.
| Reliability at each step | 5 steps | 10 steps | 20 steps |
|---|---|---|---|
| 90 % | 59 % | 35 % | 12 % |
| 95 % | 77 % | 60 % | 36 % |
| 99 % | 95 % | 90 % | 82 % |
That's why agents succeed where they can verify themselves, for example by rerunning a test, and disappoint elsewhere. The other limitations follow from this: cost, since a task consumes dozens of calls and an ever-growing context; security, since an agent that reads external content can be manipulated by that content; and drift, when the agent diligently pursues a poorly understood objective.
The METR evaluation lab measures this progress with a metric called the “time horizon”: the length of a human task that the best current model can complete with a 50% success rate. According to METR, this duration has followed an exponential trend with a doubling time of about seven months since 2019, and METR measured an acceleration to just over four months since 2023. In practical terms, an agent that failed a one-hour task a year ago has a good chance of succeeding today; the boundary moves, but it does not disappear.
#A local AI agent: what you need
Nothing requires an agent to use the cloud. The loop and tools already run on your machine; you just need to run the model that controls them there as well. Three requirements are stricter than for a simple chat, where you can make do with a smaller, slower model.
| Requirement | Why | Practical reference point |
|---|---|---|
| A model trained for tool calling | Without this, calls are malformed and the loop breaks | Target 14 billion parameters and up: Qwen 3 14B (9 GB), gpt-oss 20B (13 GB), Qwen3-Coder 30B-A3B (19 GB) |
| Lots of context | Each step adds the tool's result to the conversation | At least 32,000 tokens, meaning an additional 4 to 8 GB of KV cache beyond the weights |
| Speed | One task means dozens of round trips | 30 tokens per second and higher; MoE models have an advantage |
In practice, a 16 GB card supports a useful agent for short tasks, while 24 GB provides plenty of headroom. Below that, the agent works as a demo and disappoints in daily use.
One often-overlooked point: the cost of a local agent is not limited to the model weights’ VRAM. Each connected tool adds its description to the prompt on every turn, and that description also consumes KV cache—and therefore VRAM—for the entire loop. With a 14-billion-parameter model and half a dozen well-described tools, expect a few hundred megabytes of KV cache just for the tool definitions, before accounting for the conversation context.
- Ranking: the best local LLMs for agents
- Create a local AI agent: recommended architecture and tools
- Tutorial: a local Python agent with LangChain and Ollama
#Where to start
- 01Use an existing agent before building oneA coding agent in your editor or terminal will teach you in one week what these systems can do and where they stumble.
- 02Choose a verifiable taskA task whose success can be measured without you: a test that passes, a file produced in the correct format, a total that comes out right.
- 03Limit the toolsThree well-described tools are better than thirty. Each tool you add increases the risk of poor choices and makes the context heavier.
- 04Keep control over what is irreversibleRead freely, write with confirmation. Email sending, file deletion, payments, publishing: always require explicit human approval before execution, regardless of the confidence displayed by the agent.
- OpenCode + Ollama: a coding agent in the terminal
- Cline + Ollama: a coding agent in VS Code
- The local Agents kit
#FAQ
What's the difference between an AI agent and ChatGPT?+
What is agentic AI?+
Can an AI agent work without the Internet?+
Will AI agents replace jobs?+
Do you need to know how to code to create an agent?+
What is the ReAct pattern in an AI agent?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.