Deconstructing Multi-Agent Consensus: Evaluator-Optimizer Loops vs Linear Chains

Why monolithic prompt chains degrade on complex tasks, how iterative evaluator-optimizer topologies converge, and practical patterns for production multi-agent systems.

Most commercial attempts to deploy AI agents stumble when moving beyond simple question-answering into multi-step knowledge work. The prevailing failure mode is the Linear Prompt Chain: connecting five LLM calls in series, passing the output of step $N$ directly as context to step $N+1$.

In a linear chain of length $k$, if each step has an independent success probability $p = 0.90$, the cumulative pipeline reliability collapses exponentially:

$$P(\text{Pipeline Success}) = p^k = (0.90)^5 \approx 0.59$$

A 41% pipeline error rate is unacceptable for mission-critical software. To build resilient agentic systems, we must abandon naive linear pipelines and adopt Evaluator-Optimizer Loops governed by explicit consensus constraints.


1. Topologies: Linear vs. Evaluator-Optimizer

Let’s contrast the structural differences between these two patterns:

[ Linear Chain: Cascading Degradation ]
Input ──► Step 1 ──► Step 2 (Accumulated Error) ──► Step 3 ──► Fragile Output

[ Evaluator-Optimizer Loop with Termination Guard ]
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β–Ό                             β”‚
Input ──► [ Generator Agent ] ──► [ Evaluator Agent ]
                 β–²                        β”‚
                 β”‚      (Feedback & Score)β”‚
                 └───────── Score < Threshold ?
                                  β”‚
                                  β–Ό (Score >= 0.92 or Max Retries)
                           Validated Output

The Roles in the Topology

  1. Generator / Worker: Produces an artifact (code, document, analysis, SQL query) conditioned on instructions and prior evaluation critiques.
  2. Evaluator / Judge: Evaluates the artifact against explicit, non-overlapping criteria (schema adherence, factual grounding, safety constraints, syntax correctness).
  3. Arbiter / Harness: An unyielding runtime harness (written in deterministic code, not LLM prompts) that controls loop bounds, enforces maximum iteration counts, calculates convergence deltas, and prevents infinite cycles.

2. Why Consensus Requires Deterministic Assertions

A common pitfall is using an LLM to evaluate another LLM without grounding in deterministic facts. This produces echo chambers where both the generator and evaluator hallucinate together.

To break this failure mode, evaluation must be hybrid:

                       Evaluation Pipeline
                               β”‚
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β–Ό                                     β–Ό
   Deterministic Tests                   Heuristic LLM Scoring
   β€’ JSON schema validation              β€’ Stylistic cohesion
   β€’ Type checking / Linter              β€’ Nuanced domain critique
   β€’ Unit test execution in sandbox      β€’ Ambiguity detection
   β€’ Regex / boundary constraints
            β”‚                                     β”‚
            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β–Ό
                    Combined Quality Metric

If deterministic tests fail (e.g., syntax error, invalid JSON, broken constraint), the evaluator immediately rejects the candidate without wasting tokens on heuristic analysis.


3. Minimal Implementation in TypeScript

Here is a concrete implementation of an Evaluator-Optimizer loop with deterministic fallback and bounded retries:

interface TaskResult {
  content: string;
  iteration: number;
  finalScore: number;
  history: string[];
}

interface EvaluationResult {
  score: number; // 0.0 to 1.0
  passed: boolean;
  critique: string;
  syntaxValid: boolean;
}

export async function runEvaluatorOptimizerLoop(
  taskPrompt: string,
  generator: (prompt: string, feedback?: string) => Promise<string>,
  evaluator: (candidate: string) => Promise<EvaluationResult>,
  maxIterations = 4,
  threshold = 0.90
): Promise<TaskResult> {
  let feedback: string | undefined = undefined;
  const history: string[] = [];

  for (let iter = 1; iter <= maxIterations; iter++) {
    // 1. Generate candidate
    const candidate = await generator(taskPrompt, feedback);
    history.push(`[Iteration ${iter}]: Candidate generated (${candidate.length} chars)`);

    // 2. Evaluate candidate
    const evalResult = await evaluator(candidate);

    if (evalResult.passed && evalResult.score >= threshold) {
      return {
        content: candidate,
        iteration: iter,
        finalScore: evalResult.score,
        history,
      };
    }

    // 3. Formulate targeted feedback for next iteration
    feedback = `Prior attempt scored ${evalResult.score.toFixed(2)}. Deficiencies: ${evalResult.critique}`;
    history.push(`[Iteration ${iter}]: Rejected (Score: ${evalResult.score.toFixed(2)})`);
  }

  throw new Error(`Evaluator-optimizer failed to converge within ${maxIterations} iterations.`);
}

4. Production Principles for Agent Engineers

When deploying multi-agent topologies into production, enforce the following five principles:

  1. State Isolation: Never allow agents to mutate a shared global context dictionary. Every message and revision must be an immutable event recorded in an append-only trace log.
  2. Cost & Latency Budgets: Each loop must declare an upper financial cost ceiling (e.g., maximum $0.15 per task) and execution timeout.
  3. Structured Reflection: Do not prompt the evaluator with generic requests like β€œCheck this work”. Prompt it with specific rubrics:
    • Does this meet constraint A? (Score 0-1)
    • Are there unreferenced entities? (True/False)
    • Specific remediation step if rejected.
  4. Compile Prompts with DSPy: Instead of manual prompt tweaking, use optimizer harnesses (like DSPy or self-hosted Bayesian optimizers) that programmatically explore few-shot examples against your golden test suite.
  5. Observability via Distributed Tracing: Export every turn, tool call, critique, and token count using OpenTelemetry or Cloudflare Workers Traces. Without granular tracing, debugging distributed agent swarms is impossible.