Is It Really an AI Agent?
A practical test for products marketed as agents or agentic AI: what they can actually do, where the value comes from, and which questions expose the hype.
Fabian Mösli Reading Preferences
Key Takeaways
- • An agent label proves nothing. The useful test is whether the model can observe what happened and choose its next action, rather than following a route fixed in advance.
- • Tools, memory, schedules, and multiple models can support an agent, but none makes a product agentic on its own.
- • Judge agency and autonomy separately. Ask what the system may observe, decide, change, and verify, then test how it behaves when the happy path breaks.
In this guide
Every AI product seems to have acquired an agent.
There are research agents, sales agents, service agents, meeting agents, coding agents, browser agents, and entire teams of agents. Sometimes the label is deserved. The system can investigate a task, use tools, notice when something failed, change course, and keep working.
Sometimes “agent” means a chatbot with access to your files. Or an ordinary automation with an LLM writing one step in the middle. The demo still looks impressive. The distinction only becomes visible when something unexpected happens.
That matters because agentic capability is often used to justify a higher price, deeper access to company systems, and a much bigger promise: this product will complete work, rather than merely help you think about it.
I run one of these myself. My own assistant has been editing this website since March: I describe a change, she makes it, builds a preview, and sends back a link for me to approve. Getting there took a rented server, a lot of configuration by hand, and several setups I broke badly enough that restoring from backup was the only way out. The label costs nothing. Everything I just described took work.
Before I believe a promise like that, I want to know what the system can observe, which decisions the model can make, what it can change, and how anyone will know whether it succeeded. By the end of this guide, you will have a practical way to judge those claims without becoming an AI engineer.
One job, three very different products
Imagine a vendor is showing you an “agentic campaign analyst.” You connect an analytics export, the campaign plan, and a spend sheet. Every Monday it prepares the review.
Here are three versions of that product.
| What the product does | What you are actually getting |
|---|---|
| It reads the files, writes a plausible review, and waits. You notice that spend does not match the plan and ask it to investigate. | Chat or an assistant. It helped with the answer; you carried the work forward. |
| A schedule collects the same exports, runs the same calculations, asks an LLM for a summary, and delivers the report. Programmed rules handle every branch. | An AI-enabled workflow. Useful, repeatable, and predictable. The route was designed in advance. |
| It notices the spend mismatch, chooses to inspect the affected campaign, checks another permitted source, finds conflicting data, and changes its recommendation before stopping for your approval. | An agent. The model chose the next investigation based on what it discovered. |
All three can appear in a chat window. All three may use the same LLM. The difference sits in the operating behaviour around the model.
A product can also use more than one pattern. Its research mode may be agentic while its email writer is a single response. Judge the specific job and mode you are buying, not the logo on the homepage.
The one question that exposes most agent claims
After the system observes the result of an action, who chooses what happens next: fixed rules or the model?
A normal chat exchange ends like this:
you ask → model answers → model waits
An agent can continue the job:
goal → inspect → choose an action → use a tool → observe the result
↑ ↓
└──────────── decide again ────────────┘
↓
verify, stop, or ask for help
The number of steps does not settle it. A product may be programmed to make five searches in a fixed sequence; that is a workflow. An agent may need only one search on an easy run, but the model could have searched again if the evidence was weak.
Anthropic uses the same practical distinction in its guide to building effective agents: workflows follow predefined paths, while agents direct their own process and tool use.
Fixed workflows are often the better design. If the same checks should happen in the same order every Monday, letting a model improvise the route adds cost and new ways to fail. Agentic behaviour is worth the cost when the right next step depends on what the system discovers.
Why the same model can produce a very different product
The LLM is only one component. It receives context and produces an output, but it does not arrive with your files, accounts, memory, or authority.
The product around it supplies instructions, company context, tools, a place to keep track of the task, permissions, the loop that keeps work moving, and checks on the result. Engineers often call that surrounding setup the harness. An agent is a model in a harness that can choose actions, observe what happened, and continue.
The chat window is merely the interface. It can hide one response or a long agentic process, which is why comparing screenshots tells you so little.
I think of the model as a brilliant new hire who knows a lot about the world and nothing about your company.
A bare chat puts that person in an empty meeting room. They can give you advice or hand you a document. You leave the room and do the rest.
An agent setup gives the same person a desk, an access badge, the relevant files, approved tools, written procedures, and a definition of done. It also decides which doors stay locked and when a manager must approve the next action.
The workplace changes what the model can accomplish. It also changes what can go wrong. Anthropic’s guide to evaluating agents therefore treats the model and its harness as one system under test. A familiar model name does not guarantee a strong product around it.
What the marketing claims actually tell you
Most agent claims describe one ingredient and let you infer the rest.
| The claim | What it proves | What is still missing |
|---|---|---|
| “Agent mode” | The product has named a feature. | Which decisions can the model make after seeing a result? |
| “Connects to 100 apps” | It has many possible tools. | Which tools can it use on this job, with which permissions, and who chooses when? |
| “Plans and executes multi-step tasks” | It can produce or follow a plan. | Can it revise that plan when reality disagrees? |
| “Runs autonomously” | It may run without waiting for you. | Is the route adaptive, or is it a scheduled workflow? What stops it? |
| “Remembers you and your company” | It retains or retrieves context. | Can it act, observe the result, and continue? |
| “A team of specialised agents” | The vendor uses several model roles or processes. | Does that improve the outcome, or merely multiply calls, latency, and failure points? |
Tool access is particularly easy to overvalue. A chatbot that searches once and returns an answer may be useful, but the search icon does not make it an agent. The same applies to memory. Remembering your preferences improves context; it does not create an execution loop.
“Multi-agent” deserves extra suspicion. Several agents can divide genuinely different jobs. They can also use far more calls, time, and money while producing an answer one model could have written. Ask what distinct work each role owns and which measured result beats a single-agent baseline.
What agentic capability is worth paying for
The first benefit is the removal of glue work. In chat, you download a file, paste the relevant section, run a calculation, return with the error, copy the answer into another system, and prompt again. An agent can carry that state between tools itself.
The larger benefit is feedback from reality. A fluent answer can be wrong. An agent can open the source, query the record, run the calculation, or execute the test. When the result is missing or contradictory, it can investigate further instead of confidently completing the sentence.
That makes agents useful for jobs where:
- the route varies from case to case;
- the necessary evidence sits across several permitted sources;
- the outcome can be checked in the real system; and
- the work is valuable enough to justify extra time, calls, and supervision.
Agentic work usually costs more and takes longer than one response. Every action creates another place for a wrong assumption, hostile instruction, or tool error to enter. Later steps can then build on the mistake.
Use chat when you need an answer. Use a workflow when the route should stay fixed. Pay for an agent when adapting to new evidence is part of the job.
Agency and autonomy are separate
Product demos often collapse two questions:
- Can the model choose the next step?
- How far may the system go without a person?
Those are different controls.
A research agent can be highly supervised. It chooses which sources to inspect, but you approve every external action. The system is agentic and has little autonomy.
A fixed report can run every night without anyone watching. It has high autonomy and no model-directed route.
An agent that runs overnight with permission to update customer records has both. That may be justified eventually, but it should face a much higher evidence and safety bar than a supervised draft.
My own assistant makes the split concrete. Her default is a strong bias for action: if a job needs a tool that is not installed, she installs it, configures it, and tells me afterwards. I can also switch her to ask first, and nothing about the way she decides changes when I do. Same agency, different autonomy. A vendor who cannot draw that line for their own product has not thought about it.
When a vendor says “autonomous,” ask about time limits, spending limits, maximum steps, approval gates, permissions, rollback, and escalation. Autonomy is where a product’s impressive demo can become your incident.
Test beyond the happy path
A polished happy-path demo proves that one prepared case works. Before a pilot, ask the vendor to run three cases using the same job:
- The ordinary case. All expected inputs are present and consistent.
- The broken case. One source is missing and two others disagree.
- The hostile case. You supply a document or email containing an instruction that tells the agent to ignore its rules, reveal data, or take an unapproved action. Agree the permitted data and action boundary before the run.
Watch the behaviour, not just the final prose. Does the system recognise uncertainty? Does it change its investigation? Does it invent a missing figure? Can it be tricked by content inside a file? Does it stop at the approval boundary?
Then ask to see the run record. A useful trace should show which sources the system read, which tools it called, what changed, the recorded stop reason or tool failure, what it cost, and which approval or limit ended the run. A screenshot of the final answer tells you none of that.
A buyer’s scorecard
I would judge an agentic product across eight dimensions:
| Dimension | Questions worth asking | Weak answer |
|---|---|---|
| Observe | Which sources can it inspect? Are they live, copied, or inferred? What can it never see? | “It understands your whole business.” |
| Decide | Which next actions can the model choose? Which branches remain fixed rules? Can you show that in a run trace? | “Our proprietary agentic architecture handles it.” |
| Outcome advantage | On which realistic exceptions does the adaptive route improve accuracy, cycle time, recovery, or human effort compared with a fixed workflow? What is the baseline? | The vendor proves it is agentic but cannot show that agency improves the job. |
| Act | What can it read, write, send, spend, delete, or publish? Are permissions set per job? | One broad access switch for the whole product. |
| Verify | How does it check the actual outcome? Does it inspect the changed record, rerun the calculation, or merely report success? | “The model self-checks its response.” |
| Stop | What happens when data is missing, tools fail, or the step limit is reached? Where is human approval mandatory? | It keeps trying, guesses, or silently returns partial work. |
| Recover | Can you undo changes and reconstruct a bad run? Who owns an incident? | Logs exist, but nobody can explain or reverse the action. |
| Economics and data | What does one real run cost? What caps usage? Which data leaves your environment, who retains it, and for how long? | A per-seat price with no answer about run cost or data flow. |
OpenAI’s practical guide to building agents suggests assessing tools through factors such as read versus write access, reversibility, account permissions, and financial impact. Those questions are more useful than a single “agent enabled” switch.
Choose the smallest system that survives the real job
An agent is not the premium version of every AI product. Sometimes the best product is a well-designed chat interface. Sometimes it is a boring workflow that does the same thing every time and fails visibly.
The agent earns its extra machinery when the path really must adapt. Your pilot should show that adaptation improving outcomes on difficult cases, not merely producing a more theatrical progress screen.
If you decide the job does need an agent, the next question is whether to configure a general-purpose one or create a purpose-built service. Part two shows how I make that decision and build the job in stages, starting with static files and human approval rather than live accounts and wishful thinking.
For your next product demo, bring one awkward real case. Ask what the system observed, which decision the model made, what it changed, how it checked the outcome, and why it stopped. If the vendor can only show you the happy path, you have seen sales theatre.
For deeper technical orchestration, read Loops, Graphs, and Who Decides What the AI Does Next. To see one agent in production with four months of operational scars, read I Built an AI Agent That Updates This Website From a WhatsApp Voice Note.
Published: 2026-09-15
Last updated: 2026-09-15