By now, we’ve built a solid foundation.
In Part 1, we talked about Augmented LLMs and Prompt Chaining—simple yet powerful building blocks. In Part 2, we scaled things up with Routing, Parallelization, and the mighty Orchestrator.
But there’s still one thing missing.
Until now, our agents have been reactive. They generate, delegate, and execute—but they don’t evaluate their own work. They don’t reflect. And more importantly—they don’t learn from that reflection.
In this final part, we’ll explore two critical ideas:
- Evaluator–Optimizer Loops: where agents self-assess and improve
- Truly Autonomous Agents: systems that act independently with a sense of goal, memory, and planning
Let’s start with the first one.
Evaluator–Optimizer: When Agents Learn to Reflect
Most AI agents generate an answer and move on. But what if the answer isn’t good enough? What if the tone is off, or the logic is weak, or the summary just doesn’t hit the mark?
That’s where the Evaluator–Optimizer pattern shines.
This workflow introduces a feedback loop between generation, evaluation, and refinement. It's how we move from just outputting to actively improving—one iteration at a time.
In this enhanced setup, we introduce three roles:
- Generator Agent: Produces the initial output based on the task
- Evaluator Agent: Scores the output and gives feedback
- Refiner Agent: Improves the result based on that feedback
This continues until we either reach a high-enough quality score or hit a max iteration limit. It’s simple, powerful, and feels surprisingly human.
Here’s a real-world use case: Let’s say I want an AI to write a concise summary about renewable energy. I don’t just want the first draft—I want the best version it can produce, after feedback and refinement.
Here’s what that looks like in TypeScript using the AI SDK:
import { openai } from '@ai-sdk/openai';
import { generateText, generateObject } from 'ai';
import { z } from 'zod';
async function evaluatorOptimizerLoop(taskDescription: string) {
let currentOutput = '';
let iteration = 0;
const maxIterations = 5;
const qualityThreshold = 8; // Define your quality threshold
while (iteration < maxIterations) {
iteration += 1;
// Step 1: Generator Agent - Produce initial output
const generation = await generateText({
model: openai('gpt-4'),
system: 'You are a content generator.',
prompt: `Task: ${taskDescription}\n\nPlease provide the output.`,
});
currentOutput = generation.text;
// Step 2: Evaluator Agent - Assess the output
const evaluationSchema = z.object({
qualityScore: z.number().min(1).max(10),
feedback: z.string(),
});
const { object: evaluation } = await generateObject({
model: openai('gpt-4'),
schema: evaluationSchema,
system: 'You are an evaluator assessing the quality of the output.',
prompt: `Evaluate the following output:\n\n${currentOutput}\n\nProvide a quality score (1-10) and feedback for improvement.`,
});
const { qualityScore, feedback } = evaluation;
// Check if the quality meets the threshold
if (qualityScore >= qualityThreshold) {
console.log(`Iteration ${iteration}: Output meets quality standards.`);
break;
}
// Step 3: Refiner Agent - Improve the output based on feedback
const refinement = await generateText({
model: openai('gpt-4'),
system: 'You are a refiner improving content based on feedback.',
prompt: `Refine the following output based on the feedback:\n\nOutput:\n${currentOutput}\n\nFeedback:\n${feedback}\n\nProvide the improved output.`,
});
currentOutput = refinement.text;
}
return currentOutput;
}
// Example usage
evaluatorOptimizerLoop(
'Write a concise summary of the benefits of renewable energy.'
).then(console.log);This loop evaluates the generated output with a quality score (1–10). If the score is below our desired threshold, the refiner takes over, and we try again.
And here’s how the flow looks visually:

This kind of pattern brings several benefits:
- It enforces quality control
- Keeps roles modular (generation, evaluation, refinement)
- Enables future scaling with domain-specific evaluators or custom scoring logic
With this, you’re no longer building a tool that just says “yes” to everything. You’re building an agent that thinks, reflects, and adapts—a system that holds itself to a higher standard.
Truly Autonomous Agents: When AI Thinks for Itself
This is the final leap.
So far, we’ve seen how agents can generate, refine, route, parallelize, and even orchestrate complex workflows. But what if we removed ourselves from the loop almost entirely? What if we built agents that not only act—but plan, search, learn, and adapt without hand-holding?
That’s the vision behind Truly Autonomous Agents.
These systems go beyond “responding to a prompt.” They:
- Set intermediate goals
- Decide how to search or retrieve knowledge
- Evaluate relevance
- Generate follow-up questions
- And repeat the loop… until a meaningful outcome is reached
I recently tried to build one—an agent that could do deep web research, accumulate knowledge, adapt based on new questions, and then generate a final report—all while tracking usage cost, token consumption, and search relevance.
Here’s what the core looked like:
- A search agent using
exa-jsto pull real-time data - A query planner that came up with new angles to explore
- A learning extractor that summarized key takeaways and posed follow-up questions
- And a recursive loop that repeated the whole process to refine depth
Each loop became a cycle of:
Think → Search → Evaluate → Learn → Repeat
With the final product being a rich, AI-generated research report tailored to the user’s original goal.
Here’s the flow visualized in a simplified structure:

This kind of architecture introduces something rare in AI workflows:
- Long-term memory (accumulated learnings)
- Adaptive planning (depth + breadth control)
- Autonomous recursion (goal-aware querying)
In my case, the input was simple:
“Research the latest projects at Meta for a software engineer planning to apply.”
The agent broke it into sub-queries. It crawled results. It discarded irrelevant ones. It generated follow-up questions. And it kept digging deeper—until it had enough to generate a final summary.
All of this, without me guiding every step.
That’s what makes an agent feel autonomous—not just smart, but self-directed.
And while we’re still early in the world of true autonomy, the tools and patterns we’ve built so far—evaluators, orchestrators, tool integrations, and now recursive planners—are already paving the way.
Closing Thoughts on the Series
From Augmented LLMs to Autonomous Systems, we’ve traveled the full arc of what it means to build AI agents.
It’s no longer just about plugging in GPT. It’s about composing systems:
- That reason
- That coordinate
- That learn
- And soon… systems that understand goals and operate independently
If you’ve followed this far—you’re already thinking like a systems architect, not just a prompt engineer.
I hope this series demystified the buzz around AI agents—and gave you real tools, real code, and real workflows to experiment with.
👉 Go build. Go experiment. And if you do create something cool—send it my way.