"Agent" is used for a wide range of systems, from a chatbot with a calculator to software that works for hours with minimal supervision. A useful definition separates two designs. In Anthropic's guide to building effective agents, published in December 2024, workflows are "systems where LLMs and tools are orchestrated through predefined code paths," while agents are "systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks."[1]
The difference is who decides what happens next. In a workflow, your code does. In an agent, the model does. Everything else in this article follows from that.
Five common workflow patterns
Before reaching for a full agent, the same guide describes five workflow patterns that cover many real needs:[1]
- Prompt chaining: break a task into sequential steps, with programmatic checks between them.
- Routing: classify an input and direct it to a specialised downstream task.
- Parallelization: run subtasks at the same time, either splitting work into sections or running the same task several times and voting.
- Orchestrator-workers: a central model breaks a task down dynamically and delegates the pieces.
- Evaluator-optimizer: one model generates, another evaluates, and they iterate.
Start simpler than you think
The guide's advice on complexity is direct: "find the simplest solution possible, and only increasing complexity when needed." For many applications, it says, "optimizing single LLM calls with retrieval and in-context examples is usually enough." Agents fit "open-ended problems where it's difficult or impossible to predict the required number of steps," and where "flexibility and model-driven decision-making are needed at scale."[1]
Our reading: an agent trades predictability for flexibility. If you can write down the steps, write them down. If a fixed pipeline with one model call per stage solves the problem, it will be cheaper, faster and easier to audit than an open-ended loop. For how a model is connected to the tools an agent uses, see what Model Context Protocol is.
Why agents raise the stakes
A model that only answers can be wrong. A model that acts can be wrong with consequences: it can send the email, change the record or run the command. Two properties of agents make mistakes expensive:
- Errors can compound. Each step feeds on the output of earlier steps, so a wrong assumption early in a run can be carried forward as if it were fact. This is a structural point, and the size of the effect depends on the system.
- Autonomy removes the pause. In a workflow a person can review between steps. In a fully autonomous loop, nobody sees the intermediate states unless you build that in.
The central security risk: prompt injection
The OWASP project's list of risks for LLM applications places prompt injection first. It defines the problem like this: "A Prompt Injection Vulnerability occurs when user prompts alter the LLM's behavior or output in unintended ways." It distinguishes two forms:[2]
- Direct prompt injection: a user's own input changes the model's behaviour, either through deliberate crafting or by accident.
- Indirect prompt injection: the model processes external content, such as websites or files, that contains embedded instructions which alter its behaviour without the user's awareness.
Indirect injection is the one that matters most for agents, because agents read content the user never wrote. A web page, an email or a tool result can carry text that tries to redirect the model. The more tools an agent has, the more a successful injection can do.
Mitigations OWASP recommends
The guidance lists these strategies, and every one of them applies to an agent:[2]
- Constrain model behaviour: give clear role definitions, limit capabilities, and instruct the model to reject attempts to change those instructions.
- Define output formats: specify the expected format, and validate adherence with deterministic code.
- Implement filtering: apply semantic and string-based filters, and evaluate responses.
- Enforce privilege control: restrict the model to the minimum permissions it needs, and handle extensible functions in code rather than leaving the decision to the model.
- Require human approval: use human-in-the-loop controls for high-risk operations.
- Segregate external content: clearly separate and identify untrusted sources to limit their influence.
- Conduct adversarial testing: run regular penetration tests and simulations, treating the model as an untrusted user.
A practical checklist
If you are building or buying an agent, ask:
- Which tools can it call, and what is the worst thing each one can do?
- Which actions are irreversible, such as payments, deletions and outbound messages, and does a person approve them?
- What untrusted content does it read, and is that content kept separate from instructions?
- Is every tool call logged with its inputs and outputs, so you can review what happened?
- Has anyone tried to break it on purpose?
- Could a fixed workflow do this job instead?

