Functional AI is coming soon. Join the waitlist for early access.

// blog

6 min read

Agent Loop: Architecture and Failure Modes

An agent loop is the perceive-plan-act cycle at the core of every AI agent. Architecture, two dominant patterns, and four production failure modes.

An agent loop is the repeating cycle an AI agent executes to complete a task: perceive the current state, reason about what to do next, call a tool, observe the result, and decide whether to continue or stop. This cycle — not the model, not the system prompt — is where most production failures originate, and where reliability is won or lost.

Every major autonomous AI system converges on the same core pattern: an LLM inside a while loop, calling tools until a termination condition fires. The architecture is simple. The engineering around it is not.

What is inside an agent loop?

A minimal TypeScript implementation:

async function runAgent(goal: string): Promise<string> {
  const messages: Message[] = [
    { role: "system", content: SYSTEM_PROMPT },
    { role: "user", content: goal },
  ];
  let steps = 0;
  const MAX_STEPS = 12;

  while (steps < MAX_STEPS) {
    const response = await llm.chat(messages);
    messages.push(response);
    steps++;

    if (!response.toolCalls?.length) {
      return response.content; // final answer
    }

    for (const tc of response.toolCalls) {
      const result = await executeTool(tc.name, tc.args);
      messages.push({ role: "tool", content: result, toolCallId: tc.id });
    }
  }

  throw new Error(`Agent did not converge within ${MAX_STEPS} steps`);
}

Three things to notice before moving on:

  • MAX_STEPS is hard-coded, not optional.
  • The loop throws when it hits the ceiling — so your error monitoring actually catches it.
  • The function has one success path (no tool calls = final answer) and one failure path (ceiling reached).

That throw matters. A loop that returns an empty string or null on ceiling-hit masquerades as a success in your metrics while delivering nothing to the user.

The two dominant loop patterns

ReAct (Reason + Act) interleaves reasoning with action at every step. The model reasons about what to do, picks a tool, observes the result, and reasons again. Highly adaptive — the right default for most use cases because it can course-correct mid-run when a tool returns something unexpected.

Plan-and-Execute separates planning from execution. A planner model (often larger) generates a full task list upfront; an executor model (often smaller and cheaper) works through each step. Better for well-defined sequential workflows — data pipelines, report generation, anything you could write a checklist for. Brittle when step N returns something the planner did not anticipate. The fix is a re-plan gate: after every K steps, or on an unexpected tool output, ask the planner whether to revise the remaining steps before continuing.

For most product engineers shipping a first agent into production: start with ReAct. Graduate to Plan-and-Execute when you need a clear, auditable task plan, or when your tasks are predictable enough to sequence upfront.

Why production agent loops fail

The while loop itself is a solved problem. The engineering around it is where teams get burned. Four failure modes account for the majority of real-world incidents — and most guides skip three of them.

Failure mode 1: The infinite loop and the no-progress trap

A step ceiling stops a loop. It does not stop an agent from appearing to make progress while actually spinning. An agent calling search("auth errors"), then search("authentication errors"), then search("login failures") is looping semantically while producing different tool arguments each step — so the ceiling never fires early enough.

Research cataloging 63 confirmed production budget-overrun incidents across 21 orchestration frameworks found hallucination-driven loops — agents re-issuing tool calls without measurable progress — as one of the top recurring failure classes, with documented dollar losses from a single retry loop reaching thousands of dollars before an operator noticed.

The fix is action deduplication. Before executing any tool call, hash the (tool, args) pair and check whether you have seen it before:

const seen = new Set<string>();

function deduplicateCall(tool: string, args: unknown): string | null {
  const key = JSON.stringify({ tool, args });
  if (seen.has(key)) {
    return `Already called ${tool} with these exact arguments. Try a different approach.`;
  }
  seen.add(key);
  return null;
}

Return a synthetic observation if you have seen the call before. Exact-match deduplication catches roughly 80% of real-world loops without any additional latency. Semantic deduplication (embedding similarity) catches the rest, at the cost of added complexity.

Failure mode 2: Quadratic context accumulation

Naive agent loops re-serialize the entire message history on every step. Token usage does not grow linearly — it grows as the triangular number N(N+1)/2. A 20-step loop where each step adds 1,000 tokens produces 210,000 cumulative input tokens, not 20,000.

By step 30, a session that started at 5,000 tokens can be sending 80,000+ input tokens per call. Cost aside, accuracy degrades: most frontier models show measurable quality drops well before their advertised context limits in long-document tasks. The effective window is not the advertised window.

Mitigate with mid-loop summarization: once accumulated context exceeds roughly 60% of the model's effective window, summarize tool history into a compressed representation. Keep the system prompt and the last two or three turns verbatim; compress everything older. Pin any safety-critical instructions so they survive the summarization pass.

Failure mode 3: Silent tool failures

Tool-calling fails 3–15% of the time in production, depending on model size and task complexity. The most damaging failures are silent: the tool returns HTTP 200 with an empty payload, the model reads it as a valid no-op, and the agent "succeeds" at doing nothing. Standard infrastructure monitoring — uptime alerts, latency p99s, error rates — will not surface this.

Return a typed envelope from every tool and validate it before the model sees the result:

type ToolResult =
  | { status: "ok"; data: string }
  | { status: "empty"; reason: string }
  | { status: "error"; message: string };

function validateResult(raw: unknown): ToolResult {
  if (!raw || (typeof raw === "string" && raw.trim() === "")) {
    return { status: "empty", reason: "Tool returned no data" };
  }
  if (typeof raw === "object" && "error" in (raw as object)) {
    return { status: "error", message: String((raw as { error: unknown }).error) };
  }
  return { status: "ok", data: String(raw) };
}

When the model sees { status: "empty", reason: "Search returned no results" }, it makes a meaningfully different next-step decision than when it sees a blank string it may interpret as success.

Failure mode 4: Behavioral drift over long runs

By the twentieth tool call, an agent may be optimizing for something adjacent to the original goal — and there is no structural signal that anything is wrong. The output is well-formatted. The task completed. The status is 200. The result is subtly incorrect. Standard error-rate monitoring misses this failure class entirely because the diagnostic signal lives in output quality against original intent, not in latency or status codes.

The fix is a goal-satisfaction check: an independent model call (or a fast eval function) that scores the current output against the original task criteria and flags drift before the loop returns. This is the pattern Functional AI runs as a built-in on every Agentic Function execution — a configurable eval harness at the runtime level so that behavioral drift triggers a version alert before users see degraded behavior.

The four termination conditions you actually need

A step ceiling is one. All four are required for a production-grade loop:

ConditionWhat it catches
Max step countHard ceiling; stops absolute runaway
Token / budget ceilingCatches context ballooning before cost spirals
No-progress detectionCatches semantic loops the step counter misses
Goal-satisfaction checkStops the loop correctly, not just forcefully

Only the last one produces a successful result. The first three are circuit breakers. If circuit breakers fire on more than roughly 5% of production runs, the issue is tool design or goal ambiguity — not the ceiling values. Raise the ceiling and the problem follows you.

Tool design and context engineering

The most overlooked truth about agent loops: the bulk of what an agent sees is not your system prompt — it is tool outputs. Tool responses make up the majority of total token consumption across a typical agent conversation; the system prompt is a rounding error.

Design each tool schema for the agent's mental model, not the underlying API's surface. Return structured, human-readable output. Filter out metadata the agent does not need. Never dump raw JSON into the context when a two-sentence summary would do.

One pattern that pays for itself immediately: dynamic tool registration. A full roster of 20 tool schemas adds roughly 4,000 tokens to every call. A routing classifier that injects only the 3–4 schemas relevant to the current step adds about 400 tokens. That is 3,600 tokens recovered per call, compounding across every step in every run.

Where Functional AI fits

An agent loop running inside a while block on your own infrastructure is yours to own entirely: model selection, fallback on rate limits, context management, eval gating, version management. When GPT-4o returns a 429, your loop needs to route to Claude Sonnet without losing state. When a new model cuts cost by 40%, you need to eval it before promoting. When a prompt change breaks step 7 in production, you need to know immediately — not after users file support tickets.

Functional AI is an Agentic Function Runtime that manages these concerns at the runtime level. Define your loop as an Agentic Function — goal, tools, eval threshold, fallback chain — and run fn eval classify --threshold 0.95 to score a candidate version against your eval suite before promoting it. Automatic model fallback, quality-gated deploys, and per-version rollback are part of the runtime, not your codebase. Trust your agents in production without owning the entire harness.

FAQ

What is an agent loop? An agent loop is the repeating perceive-plan-act-observe cycle an AI agent runs until it meets a goal or hits a stop condition. It is the core architecture behind every autonomous AI system.

What causes an agent loop to run forever? An agent loop runs forever when it lacks a valid stopping condition: no step ceiling, no no-progress detection, no token budget, and a goal the agent cannot verifiably check. A step ceiling alone is necessary but not sufficient — semantic loops can repeat tool calls with varied arguments while the step count climbs toward the ceiling.

What is the difference between ReAct and Plan-and-Execute agent loops? ReAct interleaves reasoning and action at every step, making it adaptive but prone to drift on long runs. Plan-and-Execute generates a full task plan upfront and executes each step with a simpler model — efficient for well-defined sequential workflows, brittle when early steps return unexpected results.

How does context accumulation affect agent loop cost? Naive loops re-send the full message history on every step. Token usage grows as N(N+1)/2: a 20-step loop generating 1,000 tokens per step produces 210,000 cumulative input tokens, not 20,000. Context summarization at step boundaries is the standard mitigation.

How do I detect silent tool failures in an agent loop? Return a typed result envelope from every tool — { status: "ok" | "empty" | "error" } — and validate it before the model sees the result. Silent failures (HTTP 200 with empty payload) are the most damaging because no error surfaces to standard infrastructure monitoring.

What metrics should I track for a production agent loop? Track: steps per trace (p99), token cost per trace, no-progress terminations (loop hits ceiling before converging), tool error rate by tool name, and eval score per version. A step count trending up without a corresponding improvement in goal completion rate signals loop degradation.

// private beta

Ship agents you can trust.

Build your prompt, agent, or multi-agent workflow as an Agentic Function. Functional AI hosts it, certifies each version before it ships, switches to a backup model when the primary fails, and lets you release without a redeploy. You just call it from your code.

No spam. One email when Functional AI launches.

Shipping agents today? Become a design partner — design partners shape the roadmap and get early access.