Before you build an agent team, define one good workflow.
A practical way to choose between fixed workflows, tool-using agents, and specialist teams—starting with a decision your business can actually verify.
What the research actually supports
PUBLISHED RESEARCH / DOCUMENTATIONAnthropic distinguishes workflows with predefined code paths from agents that dynamically choose their processes and tool use. Its engineering guidance recommends starting with simpler systems and adding complexity only when it improves measured outcomes. This is implementation guidance, not proof that a particular architecture will work in your company.[1]
An answer needs a way to check the world
PUBLISHED RESEARCH / DOCUMENTATIONReAct studies language models that interleave reasoning and actions to obtain information from external environments. Its evaluations include question answering and interactive benchmarks. Those experiments establish a useful design pattern; they do not demonstrate reliable autonomous operation in an arbitrary enterprise.[2]
A proposed first project
PROPOSED BLUEPRINTTake a recurring internal question: “Can we offer this customer the requested delivery date?” Start with a read-only workflow that retrieves the order, current inventory and dispatch rules, then prepares an answer with links to the records used. Separate the language explanation from the deterministic checks for stock and lead time.
Use one coordinator initially. Add specialists only if there are distinct tools, permissions or evaluation criteria—for example a stock check and a contract check. Give each tool a narrow input schema and make missing records an explicit outcome. A confident sentence should never substitute for a successful lookup.
The proposed output is a decision packet: requested action, supporting records, unresolved constraints, and the person authorised to approve it. An approval should bind to the exact action and record versions. If relevant data changes, repeat the checks instead of reusing an old approval.
How to decide whether it earns more autonomy
PILOT EVALUATIONBuild a small evaluation set from completed cases that includes ambiguous requests, unavailable tools, conflicting records and customer-specific exceptions. Compare the workflow with the existing process on answer correctness, unsupported claims, time to resolution, tool cost and reviewer corrections.
Start in shadow mode: produce proposals beside normal work without sending messages or changing systems. Review failures by stage. A retrieval failure needs different remediation from a correct lookup followed by an incorrect interpretation. Move one reversible action into the pilot only when the measured results justify it. These are pilot criteria, not promised business results.
Sources & publication dates
Primary papers and official documentation consulted for this perspective. Source publication dates differ from the date of this article; undated documentation is identified explicitly.
- Building effective agents — AnthropicSource published 19 December 2024 · Consulted 3 October 2026
- ReAct: Synergizing Reasoning and Acting in Language ModelsSource published 6 October 2022 · Consulted 3 October 2026
Test a first direction.
Bring your own constraints into the demo. Explore an initial workflow, then discuss the evidence, integrations and evaluation a pilot would need.
Explore this workflow