Deconstructing Multi-Agent Consensus: Evaluator-Optimizer Loops vs Linear Chains
Why monolithic prompt chains degrade on complex tasks, how iterative evaluator-optimizer topologies converge, and practical patterns for production multi-agent systems.
Most commercial attempts to deploy AI agents stumble when moving beyond simple question-answering into multi-step knowledge work. The prevailing failure mode is the Linear Prompt Chain: connecting five LLM calls in series, passing the output of step $N$ directly as context to step $N+1$.
In a linear chain of length $k$, if each step has an independent success probability $p = 0.90$, the cumulative pipeline reliability collapses exponentially:
$$P(\text{Pipeline Success}) = p^k = (0.90)^5 \approx 0.59$$
A 41% pipeline error rate is unacceptable for mission-critical software. To build resilient agentic systems, we must abandon naive linear pipelines and adopt Evaluator-Optimizer Loops governed by explicit consensus constraints.
1. Topologies: Linear vs. Evaluator-Optimizer
Letβs contrast the structural differences between these two patterns:
[ Linear Chain: Cascading Degradation ]
Input βββΊ Step 1 βββΊ Step 2 (Accumulated Error) βββΊ Step 3 βββΊ Fragile Output
[ Evaluator-Optimizer Loop with Termination Guard ]
βββββββββββββββββββββββββββββββ
βΌ β
Input βββΊ [ Generator Agent ] βββΊ [ Evaluator Agent ]
β² β
β (Feedback & Score)β
ββββββββββ Score < Threshold ?
β
βΌ (Score >= 0.92 or Max Retries)
Validated Output
The Roles in the Topology
- Generator / Worker: Produces an artifact (code, document, analysis, SQL query) conditioned on instructions and prior evaluation critiques.
- Evaluator / Judge: Evaluates the artifact against explicit, non-overlapping criteria (schema adherence, factual grounding, safety constraints, syntax correctness).
- Arbiter / Harness: An unyielding runtime harness (written in deterministic code, not LLM prompts) that controls loop bounds, enforces maximum iteration counts, calculates convergence deltas, and prevents infinite cycles.
2. Why Consensus Requires Deterministic Assertions
A common pitfall is using an LLM to evaluate another LLM without grounding in deterministic facts. This produces echo chambers where both the generator and evaluator hallucinate together.
To break this failure mode, evaluation must be hybrid:
Evaluation Pipeline
β
ββββββββββββββββββββ΄βββββββββββββββββββ
βΌ βΌ
Deterministic Tests Heuristic LLM Scoring
β’ JSON schema validation β’ Stylistic cohesion
β’ Type checking / Linter β’ Nuanced domain critique
β’ Unit test execution in sandbox β’ Ambiguity detection
β’ Regex / boundary constraints
β β
ββββββββββββββββββββ¬βββββββββββββββββββ
βΌ
Combined Quality Metric
If deterministic tests fail (e.g., syntax error, invalid JSON, broken constraint), the evaluator immediately rejects the candidate without wasting tokens on heuristic analysis.
3. Minimal Implementation in TypeScript
Here is a concrete implementation of an Evaluator-Optimizer loop with deterministic fallback and bounded retries:
interface TaskResult {
content: string;
iteration: number;
finalScore: number;
history: string[];
}
interface EvaluationResult {
score: number; // 0.0 to 1.0
passed: boolean;
critique: string;
syntaxValid: boolean;
}
export async function runEvaluatorOptimizerLoop(
taskPrompt: string,
generator: (prompt: string, feedback?: string) => Promise<string>,
evaluator: (candidate: string) => Promise<EvaluationResult>,
maxIterations = 4,
threshold = 0.90
): Promise<TaskResult> {
let feedback: string | undefined = undefined;
const history: string[] = [];
for (let iter = 1; iter <= maxIterations; iter++) {
// 1. Generate candidate
const candidate = await generator(taskPrompt, feedback);
history.push(`[Iteration ${iter}]: Candidate generated (${candidate.length} chars)`);
// 2. Evaluate candidate
const evalResult = await evaluator(candidate);
if (evalResult.passed && evalResult.score >= threshold) {
return {
content: candidate,
iteration: iter,
finalScore: evalResult.score,
history,
};
}
// 3. Formulate targeted feedback for next iteration
feedback = `Prior attempt scored ${evalResult.score.toFixed(2)}. Deficiencies: ${evalResult.critique}`;
history.push(`[Iteration ${iter}]: Rejected (Score: ${evalResult.score.toFixed(2)})`);
}
throw new Error(`Evaluator-optimizer failed to converge within ${maxIterations} iterations.`);
}
4. Production Principles for Agent Engineers
When deploying multi-agent topologies into production, enforce the following five principles:
- State Isolation: Never allow agents to mutate a shared global context dictionary. Every message and revision must be an immutable event recorded in an append-only trace log.
- Cost & Latency Budgets: Each loop must declare an upper financial cost ceiling (e.g., maximum $0.15 per task) and execution timeout.
- Structured Reflection: Do not prompt the evaluator with generic requests like βCheck this workβ. Prompt it with specific rubrics:
- Does this meet constraint A? (Score 0-1)
- Are there unreferenced entities? (True/False)
- Specific remediation step if rejected.
- Compile Prompts with DSPy: Instead of manual prompt tweaking, use optimizer harnesses (like DSPy or self-hosted Bayesian optimizers) that programmatically explore few-shot examples against your golden test suite.
- Observability via Distributed Tracing: Export every turn, tool call, critique, and token count using OpenTelemetry or Cloudflare Workers Traces. Without granular tracing, debugging distributed agent swarms is impossible.