# Single-Agent Design Patterns: Five Ways One Model Reasons and Acts

*Author: Siddhesh Prabhugaonkar — Enterprise AI & Cloud Consultant*

*Part 2 of the* [*Agentic Design Patterns series*](/a-field-guide-to-agentic-design-patterns-the-17-building-blocks-of-modern-ai-agents)*. Part 1 covered the four-family map and the cost-vs-latency framework used throughout.*

Before any agent talks to another agent, it has to decide how it talks to *itself* — how it interleaves thinking with doing, how far ahead it plans before it acts, and what it does when it's wrong. That's the single-agent layer, and almost every multi-agent system in Part 3 is really just several instances of one of these five patterns wired together.

> **Building a tool-using agent and unsure which control loop fits your latency budget?** [Book a discovery call →](https://topmate.io/siddheshp)

* * *

## 1\. Reason & Act (ReAct) Loop

**Intent:** Interleave an explicit reasoning trace with a concrete tool call, one step at a time, so every action is conditioned on the freshest observation.

![](https://cdn.hashnode.com/uploads/covers/651bff05e4455a8ac9ec7688/0a3eab33-f648-4d92-bf8c-6aaaeffa45ee.png align="center")

This is the pattern most people mean when they say "agent." Originally described in the ReAct paper (Yao et al., 2023), it's the substrate underneath AutoGPT-style loops and most production coding agents evaluated on benchmarks like SWE-bench, where the agent reasons its way through locating, editing, and verifying a fix across a large codebase.

|  |  |
| --- | --- |
| **Strength** | Maximally adaptive — the agent reacts to real tool output at every step, recovering from errors mid-flight |
| **Cost** | Grows with trajectory length; the entire history accumulates in context, so token cost climbs step over step |
| **Failure mode** | Error compounding — an early hallucinated observation-parse can drift the whole trajectory; without step budgets and stop conditions, loops can wander indefinitely |
| **Use it when** | The task is exploratory, the environment is only partially known up front, and step count is naturally small (single digits to low tens) |
| **Avoid it when** | You already know the full task decomposition — paying to re-derive it at every step is wasted spend; see Plan-Then-Execute below |

* * *

## 2\. Plan-Then-Execute

**Intent:** Decouple planning from execution. Generate the full plan once, then work through it — re-planning only when something breaks — instead of re-reasoning about the whole trajectory at every single step.

![](https://cdn.hashnode.com/uploads/covers/651bff05e4455a8ac9ec7688/1b6d70d0-dd52-44d9-9971-650a124c250c.png align="center")

Where ReAct re-derives "what now?" at every single step, Plan-Then-Execute pays that reasoning cost once and then runs a cheaper execution loop. This matters for long-horizon coherence: a plan made once, up front, doesn't drift the way a step-by-step trajectory can.

|  |  |
| --- | --- |
| **Strength** | Long-horizon coherence — the agent doesn't lose the thread over dozens of steps, because the thread was written down once |
| **Cost** | Lower than ReAct on long trajectories, since execution steps don't re-run full reasoning |
| **Failure mode** | Brittle to a changing environment — a plan made against stale assumptions can execute confidently into a wall; needs a re-planning trigger |
| **Use it when** | The task decomposition is knowable up front and the environment is stable enough that step 7 won't invalidate the assumptions behind step 2 |
| **Avoid it when** | The environment is highly dynamic or adversarial — you'll spend more on re-planning than you saved on execution |

* * *

## 3\. Search over Actions

**Intent:** Instead of committing to one action per step, generate multiple candidate next actions, evaluate the resulting states, and backtrack when a branch dead-ends — tree or graph search over the action space.

![](https://cdn.hashnode.com/uploads/covers/651bff05e4455a8ac9ec7688/f1adf8e3-3be2-408d-8d55-ecb0dc0b4be1.png align="center")

This pattern trades tokens for solution quality on tasks that look like search problems — puzzle-solving, code generation with multiple candidate solutions, planning under uncertainty. Because branches are largely independent, this is one of the rare "expensive" patterns that parallelizes well: you can evaluate several candidates concurrently rather than paying for them sequentially.

|  |  |
| --- | --- |
| **Strength** | Explores alternatives a single greedy trajectory would never generate; recovers from local dead ends by construction |
| **Cost** | 3–10×+ baseline, scaling with branching factor and depth — the most expensive pattern in the single-agent family |
| **Failure mode** | Combinatorial blow-up without a good evaluator/pruning function; a weak state-evaluator makes the search worse than a single greedy pass |
| **Use it when** | The task has a genuine search structure — multiple plausible paths, a way to score partial progress, and the value of a better answer justifies the spend |
| **Avoid it when** | The problem is largely sequential with one obviously correct next step — search adds cost without adding quality |

* * *

## 4\. Reflection & Self-Correction

**Intent:** Generate a candidate output, critique it against the goal (or against a verifier), and retry conditioned on that critique — a verbal reinforcement loop instead of a gradient update.

![](https://cdn.hashnode.com/uploads/covers/651bff05e4455a8ac9ec7688/bd613a9a-99b8-468d-8a3e-eff04b4753dc.png align="center")

This is the pattern behind Reflexion-style agents and most "self-healing" coding agents: run the tests, read the failure, retry with the failure as grounding context. The critical detail that separates this from a naive retry loop is that the critique gets *stored*, not just used once — so the second attempt is grounded in a specific, named mistake rather than a vague "try again."

|  |  |
| --- | --- |
| **Strength** | Catches its own mistakes without external supervision beyond a verifier; noticeably improves output quality on tasks with checkable success criteria |
| **Cost** | Roughly 2–4× baseline — you're paying for at least one extra generate-critique-retry cycle |
| **Failure mode** | Self-critique without an external grounding signal (a test suite, a validator) can hallucinate confidence in a still-wrong answer — reflection needs *something* objective to bounce off |
| **Use it when** | You have a cheap, reliable verifier — unit tests, a schema validator, a grader model — to ground the critique |
| **Avoid it when** | The only critic available is the same model marking its own homework with no external signal; you'll pay for a retry loop that doesn't actually improve accuracy |

* * *

## 5\. Code-as-Action

**Intent:** Instead of emitting one structured tool call per step, let the model write and execute a script — collapsing what would be many round-trips into a single composable action.

![](https://cdn.hashnode.com/uploads/covers/651bff05e4455a8ac9ec7688/3a44dffd-f628-42e6-85d1-c49496c7a2cd.png align="center")

Instead of "call tool A, wait, call tool B, wait, call tool C," the model writes `for item in fetch_all(): process(item)` once and the interpreter runs the whole loop. This is the pattern behind Open Interpreter, Code Interpreter-style tools, and increasingly, production coding agents that treat the shell itself as the action space. Because a for-loop over 50 files is one action instead of 50 round-trips, this collapses path depth dramatically for anything with repetitive structure.

|  |  |
| --- | --- |
| **Strength** | Collapses many sequential tool round-trips into one action — the single biggest latency win in this list for repetitive, structured tasks |
| **Cost** | Comparable to or cheaper than ReAct, since fewer round-trips means less repeated context re-transmission |
| **Failure mode** | The composability that makes this powerful also makes it the highest-risk pattern to run un-sandboxed — arbitrary code execution needs real isolation, not a system-prompt warning |
| **Use it when** | The task has repetitive or compositional structure (batch processing, data transformation, multi-file edits) and you have a real sandbox to run in |
| **Avoid it when** | You don't yet have execution isolation in place — see Part 5 on [Sandboxing & Permissioning](/agentic-design-patterns-interaction-integration) before shipping this pattern to production |

* * *

## Choosing between the five

| Pattern | Best fit | Relative cost | Path depth |
| --- | --- | --- | --- |
| ReAct Loop | Exploratory tasks, partially known environment | 1–3× | Deep |
| Plan-Then-Execute | Known decomposition, stable environment | 1–2× | Medium |
| Search over Actions | Genuine search structure, scorable partial progress | 3–10×+ | Deep, parallelizable |
| Reflection & Self-Correction | Checkable success criteria available | 2–4× | Medium |
| Code-as-Action | Repetitive/compositional structure, sandboxed execution | 1–2× | Shallow |

A useful rule of thumb: **start with ReAct because it's the simplest to reason about, then graduate to a more specialized pattern only once you can point to the specific failure mode ReAct is producing.** Plan-Then-Execute fixes drift on long trajectories. Reflection fixes silent wrongness on checkable tasks. Search fixes greedy dead-ends. Code-as-Action fixes round-trip latency on repetitive work. Don't reach for the fancier pattern until the plain loop has actually shown you the problem it doesn't solve.

* * *

## Next in the series

*   **Part 3 —** [**Multi-Agent Patterns**](/agentic-design-patterns-multi-agent)**:** what happens when you wire several of these single-agent loops together — orchestration, debate, role-based teams, and handoff.
    
*   **Part 4 —** [**Memory & State Patterns**](/agentic-design-patterns-memory-state)**:** how context window management and episodic memory change the economics of every pattern above.
    
*   **Part 1 —** [**Field Guide overview**](/a-field-guide-to-agentic-design-patterns-the-17-building-blocks-of-modern-ai-agents)**:** the cost-vs-latency framework used throughout this series.
    

* * *

*If you found this useful, I write regularly on enterprise AI, agentic architectures, and applied GenAI adoption. My other recent posts:*

*   [*Prompt, Context, Harness, Loop: The Four Layers of Engineering Reliable AI Agents*](https://azureauthority.in/prompt-context-harness-loop-the-four-layers-of-engineering-reliable-ai-agents)
    
*   [*AI Agents, Multi-Agent Systems & LLM Council: A Practitioner's Guide to Enterprise Agentic AI*](https://cloud-authority.com/ai-agents-multi-agents-llm-council-enterprise-guide)
    

**Want help choosing the right agent architecture for your use case?** [**Book a discovery call →**](https://topmate.io/siddheshp)

* * *

### Further reading

*   Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models.* ICLR. [arxiv.org/abs/2210.03629](https://arxiv.org/abs/2210.03629)
    
*   Jimenez, C. E., et al. (2024). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* ICLR. [arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770)
    
*   Shinn, N., et al. (2023). *Reflexion: Language Agents with Verbal Reinforcement Learning.* [arxiv.org/abs/2303.11366](https://arxiv.org/abs/2303.11366)
    
*   [Agentic Protocols — Pattern Catalog](https://www.agenticprotocols.dev/patterns.html) — cross-reference catalog covering similar architectural ground
    

* * *

## About the Author

**Siddhesh Prabhugaonkar** is a **Generative AI & Agentic AI Enablement and Adoption Specialist** with two decades as an Architect, Consultant, and Trainer across IT, Cloud, and Generative AI. He is a **Microsoft Certified Trainer**, a **Pluralsight Instructor**, and helps enterprises move from GenAI curiosity to production adoption at scale.

His consulting and training practice spans **GenAI, Azure, Microsoft Foundry, Anthropic Claude, GitHub Copilot, Google Gemini, OpenAI Codex, Devin**, and modern full‑stack engineering (.NET, MEAN, MERN). Notable engagements include GenAI enablement for **ADP**, IoT platform consulting for **IIT Bombay's E‑Yantra** program, and early work on Microsoft's Repository platform (which later became **Entity Framework**).

> *Empowering organizations and individuals to adopt, build, and scale with Generative AI, Cloud, and Modern Software Engineering.*

**Connect & explore:**

*   💼 LinkedIn — [linkedin.com/in/siddheshprabhugaonkar](https://www.linkedin.com/in/siddheshprabhugaonkar)
    
*   📝 Blog — [azureauthority.in](https://azureauthority.in/)
    
*   📬 Newsletter — [cloud-authority.com](https://cloud-authority.com/)
    
*   🎥 YouTube — [youtube.com/c/SiddheshPrabhugaonkar](https://www.youtube.com/c/SiddheshPrabhugaonkar)
