Single-Agent Design Patterns: Five Ways One Model Reasons and Acts
ReAct, Plan-Then-Execute, Search over Actions, Reflection, and Code-as-Action — the control-flow shapes underneath almost every coding agent, research agent, and tool-using assistant you've used

I’m Siddhesh, a Microsoft Certified Trainer, cloud architect, and AI practitioner focused on helping developers and organizations adopt AI effectively. As a Pluralsight instructor and speaker, I design and deliver hands-on AI enablement programs covering Generative AI, Agentic AI, Azure AI, and modern cloud architectures.
With a strong foundation in Microsoft .NET and Azure, my work today centers on building real-world AI solutions, agentic workflows, and developer productivity using AI-assisted tools. I share practical insights through workshops, conference talks, online courses, blogs, newsletters, and YouTube—bridging the gap between AI concepts and production-ready implementations.
Author: Siddhesh Prabhugaonkar — Enterprise AI & Cloud Consultant
Part 2 of the Agentic Design Patterns series. Part 1 covered the four-family map and the cost-vs-latency framework used throughout.
Before any agent talks to another agent, it has to decide how it talks to itself — how it interleaves thinking with doing, how far ahead it plans before it acts, and what it does when it's wrong. That's the single-agent layer, and almost every multi-agent system in Part 3 is really just several instances of one of these five patterns wired together.
Building a tool-using agent and unsure which control loop fits your latency budget? Book a discovery call →
1. Reason & Act (ReAct) Loop
Intent: Interleave an explicit reasoning trace with a concrete tool call, one step at a time, so every action is conditioned on the freshest observation.
This is the pattern most people mean when they say "agent." Originally described in the ReAct paper (Yao et al., 2023), it's the substrate underneath AutoGPT-style loops and most production coding agents evaluated on benchmarks like SWE-bench, where the agent reasons its way through locating, editing, and verifying a fix across a large codebase.
| Strength | Maximally adaptive — the agent reacts to real tool output at every step, recovering from errors mid-flight |
| Cost | Grows with trajectory length; the entire history accumulates in context, so token cost climbs step over step |
| Failure mode | Error compounding — an early hallucinated observation-parse can drift the whole trajectory; without step budgets and stop conditions, loops can wander indefinitely |
| Use it when | The task is exploratory, the environment is only partially known up front, and step count is naturally small (single digits to low tens) |
| Avoid it when | You already know the full task decomposition — paying to re-derive it at every step is wasted spend; see Plan-Then-Execute below |
2. Plan-Then-Execute
Intent: Decouple planning from execution. Generate the full plan once, then work through it — re-planning only when something breaks — instead of re-reasoning about the whole trajectory at every single step.
Where ReAct re-derives "what now?" at every single step, Plan-Then-Execute pays that reasoning cost once and then runs a cheaper execution loop. This matters for long-horizon coherence: a plan made once, up front, doesn't drift the way a step-by-step trajectory can.
| Strength | Long-horizon coherence — the agent doesn't lose the thread over dozens of steps, because the thread was written down once |
| Cost | Lower than ReAct on long trajectories, since execution steps don't re-run full reasoning |
| Failure mode | Brittle to a changing environment — a plan made against stale assumptions can execute confidently into a wall; needs a re-planning trigger |
| Use it when | The task decomposition is knowable up front and the environment is stable enough that step 7 won't invalidate the assumptions behind step 2 |
| Avoid it when | The environment is highly dynamic or adversarial — you'll spend more on re-planning than you saved on execution |
3. Search over Actions
Intent: Instead of committing to one action per step, generate multiple candidate next actions, evaluate the resulting states, and backtrack when a branch dead-ends — tree or graph search over the action space.
This pattern trades tokens for solution quality on tasks that look like search problems — puzzle-solving, code generation with multiple candidate solutions, planning under uncertainty. Because branches are largely independent, this is one of the rare "expensive" patterns that parallelizes well: you can evaluate several candidates concurrently rather than paying for them sequentially.
| Strength | Explores alternatives a single greedy trajectory would never generate; recovers from local dead ends by construction |
| Cost | 3–10×+ baseline, scaling with branching factor and depth — the most expensive pattern in the single-agent family |
| Failure mode | Combinatorial blow-up without a good evaluator/pruning function; a weak state-evaluator makes the search worse than a single greedy pass |
| Use it when | The task has a genuine search structure — multiple plausible paths, a way to score partial progress, and the value of a better answer justifies the spend |
| Avoid it when | The problem is largely sequential with one obviously correct next step — search adds cost without adding quality |
4. Reflection & Self-Correction
Intent: Generate a candidate output, critique it against the goal (or against a verifier), and retry conditioned on that critique — a verbal reinforcement loop instead of a gradient update.
This is the pattern behind Reflexion-style agents and most "self-healing" coding agents: run the tests, read the failure, retry with the failure as grounding context. The critical detail that separates this from a naive retry loop is that the critique gets stored, not just used once — so the second attempt is grounded in a specific, named mistake rather than a vague "try again."
| Strength | Catches its own mistakes without external supervision beyond a verifier; noticeably improves output quality on tasks with checkable success criteria |
| Cost | Roughly 2–4× baseline — you're paying for at least one extra generate-critique-retry cycle |
| Failure mode | Self-critique without an external grounding signal (a test suite, a validator) can hallucinate confidence in a still-wrong answer — reflection needs something objective to bounce off |
| Use it when | You have a cheap, reliable verifier — unit tests, a schema validator, a grader model — to ground the critique |
| Avoid it when | The only critic available is the same model marking its own homework with no external signal; you'll pay for a retry loop that doesn't actually improve accuracy |
5. Code-as-Action
Intent: Instead of emitting one structured tool call per step, let the model write and execute a script — collapsing what would be many round-trips into a single composable action.
Instead of "call tool A, wait, call tool B, wait, call tool C," the model writes for item in fetch_all(): process(item) once and the interpreter runs the whole loop. This is the pattern behind Open Interpreter, Code Interpreter-style tools, and increasingly, production coding agents that treat the shell itself as the action space. Because a for-loop over 50 files is one action instead of 50 round-trips, this collapses path depth dramatically for anything with repetitive structure.
| Strength | Collapses many sequential tool round-trips into one action — the single biggest latency win in this list for repetitive, structured tasks |
| Cost | Comparable to or cheaper than ReAct, since fewer round-trips means less repeated context re-transmission |
| Failure mode | The composability that makes this powerful also makes it the highest-risk pattern to run un-sandboxed — arbitrary code execution needs real isolation, not a system-prompt warning |
| Use it when | The task has repetitive or compositional structure (batch processing, data transformation, multi-file edits) and you have a real sandbox to run in |
| Avoid it when | You don't yet have execution isolation in place — see Part 5 on Sandboxing & Permissioning before shipping this pattern to production |
Choosing between the five
| Pattern | Best fit | Relative cost | Path depth |
|---|---|---|---|
| ReAct Loop | Exploratory tasks, partially known environment | 1–3× | Deep |
| Plan-Then-Execute | Known decomposition, stable environment | 1–2× | Medium |
| Search over Actions | Genuine search structure, scorable partial progress | 3–10×+ | Deep, parallelizable |
| Reflection & Self-Correction | Checkable success criteria available | 2–4× | Medium |
| Code-as-Action | Repetitive/compositional structure, sandboxed execution | 1–2× | Shallow |
A useful rule of thumb: start with ReAct because it's the simplest to reason about, then graduate to a more specialized pattern only once you can point to the specific failure mode ReAct is producing. Plan-Then-Execute fixes drift on long trajectories. Reflection fixes silent wrongness on checkable tasks. Search fixes greedy dead-ends. Code-as-Action fixes round-trip latency on repetitive work. Don't reach for the fancier pattern until the plain loop has actually shown you the problem it doesn't solve.
Next in the series
Part 3 — Multi-Agent Patterns: what happens when you wire several of these single-agent loops together — orchestration, debate, role-based teams, and handoff.
Part 4 — Memory & State Patterns: how context window management and episodic memory change the economics of every pattern above.
Part 1 — Field Guide overview: the cost-vs-latency framework used throughout this series.
If you found this useful, I write regularly on enterprise AI, agentic architectures, and applied GenAI adoption. My other recent posts:
Prompt, Context, Harness, Loop: The Four Layers of Engineering Reliable AI Agents
AI Agents, Multi-Agent Systems & LLM Council: A Practitioner's Guide to Enterprise Agentic AI
Want help choosing the right agent architecture for your use case? Book a discovery call →
Further reading
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR. arxiv.org/abs/2210.03629
Jimenez, C. E., et al. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR. arxiv.org/abs/2310.06770
Shinn, N., et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arxiv.org/abs/2303.11366
Agentic Protocols — Pattern Catalog — cross-reference catalog covering similar architectural ground
About the Author
Siddhesh Prabhugaonkar is a Generative AI & Agentic AI Enablement and Adoption Specialist with two decades as an Architect, Consultant, and Trainer across IT, Cloud, and Generative AI. He is a Microsoft Certified Trainer, a Pluralsight Instructor, and helps enterprises move from GenAI curiosity to production adoption at scale.
His consulting and training practice spans GenAI, Azure, Microsoft Foundry, Anthropic Claude, GitHub Copilot, Google Gemini, OpenAI Codex, Devin, and modern full‑stack engineering (.NET, MEAN, MERN). Notable engagements include GenAI enablement for ADP, IoT platform consulting for IIT Bombay's E‑Yantra program, and early work on Microsoft's Repository platform (which later became Entity Framework).
Empowering organizations and individuals to adopt, build, and scale with Generative AI, Cloud, and Modern Software Engineering.
Connect & explore:
💼 LinkedIn — linkedin.com/in/siddheshprabhugaonkar
📝 Blog — azureauthority.in
📬 Newsletter — cloud-authority.com
🎥 YouTube — youtube.com/c/SiddheshPrabhugaonkar




