All writing
Part 2 of 3Demystifying AI Agents

Demystifying AI Agents Part 2: Routing, Parallelization & Orchestrators

In Part 2, we dive into smart decision-making with Routing, faster thinking with Parallelization, and a glimpse into Orchestrators.

8 min read

In Part 1, we explored the basics—Augmented LLMs and Prompt Chaining. Those patterns laid the groundwork for how agents can access better information and reason in steps.

But now comes the fun part—giving our agents the ability to decide and scale. That’s what we’re unlocking with Routing, Parallelization, and eventually, Orchestration.

Routing

Once I had a few working agents under my belt, I realized something: not every prompt should be treated the same. If a user asks for a tweet, a blog post, or a customer support reply, the tone, format, and even the structure of the response need to change. Trying to handle all of that in one generic prompt? It gets messy fast.

That’s when I stumbled onto Routing—and everything started to make more sense.

Routing is the idea of detecting what the user wants, and then sending the request to the right prompt or agent based on that intent. Instead of forcing a single model to be a generalist, you let it be a specialist. Or better yet, you build multiple specialists and just direct traffic intelligently.

Think of it like a smart receptionist in a co-working space: “Oh, you're here for legal advice? Room 3, down the hall. Looking for marketing help? Head over to Room 7.”

This kind of logic unlocks modularity. You can build smaller, cleaner prompts, re-use them across flows, and scale up easily without bloating your code.

Some real-world examples that use routing:

  • An AI content generator that figures out if you want a tweet, blog post, or email—and responds accordingly
  • A personal assistant bot that chooses whether to search Google, look up your calendar, or summarize a doc
  • A multi-agent platform that delegates math queries to one model, code generation to another, and general Q&A to a third

Here’s a TypeScript example using the AI SDK, where the system classifies intent and picks the right response flow:

typescript
import { generateText } from 'ai';
import { openai } from '@ai-sdk/openai';
 
// Step 1: Intent detection
const router = await generateText({
  model: openai('gpt-4'),
  prompt: 'Decide the type: "Write me a tweet about AI agents." Options: tweet, blog, email',
});
 
const route = router.text.trim();
 
let finalResult: string;
 
if (route === 'tweet') {
  const result = await generateText({
    model: openai('gpt-4'),
    prompt: 'Write a tweet about AI agents evolving toward autonomy.',
  });
  finalResult = result.text;
} else if (route === 'blog') {
  const result = await generateText({
    model: openai('gpt-4'),
    prompt: 'Write a blog introduction about AI agents.',
  });
  finalResult = result.text;
} else {
  finalResult = 'Unknown route';
}
 
console.log(finalResult);

And here’s a quick visual of how that logic branches out internally:

What I love about routing is that it gives your agent a sense of judgment. It doesn't just answer questions—it first thinks about what kind of question it is and then picks the best tool for the job.

It’s one of the first steps toward building systems that feel less like bots, and more like assistants who understand context.

Parallelization

After setting up Routing, I found myself craving more speed. Some tasks didn’t need step-by-step logic. They didn’t even need decisions. They just needed to happen—all at once.

That’s where Parallelization became my go-to pattern—especially when using the Vercel AI SDK.

The SDK allows you to run multiple tools at the same time during a single generation step. That means if your tools don’t rely on each other, you don’t have to wait around. You can execute them concurrently and stitch the results together in a single, fast response.

Let me give you a real-world example.

Say a user asks:

“What’s the weather in San Francisco and what attractions should I visit?” Here, the model needs to fetch:

  • The current weather
  • A list of local attractions

These tasks don’t depend on each other, so running them sequentially would be unnecessary overhead. Instead, we call both tools in parallel.

Here’s how that looks using the AI SDK:

typescript
import { generateText, tool } from 'ai';
import { openai } from '@ai-sdk/openai';
import { z } from 'zod';
 
const result = await generateText({
  model: openai('gpt-4-turbo'),
  tools: {
    weather: tool({
      description: 'Get the weather in a location',
      parameters: z.object({
        location: z.string().describe('The location to get the weather for'),
      }),
      execute: async ({ location }: { location: string }) => ({
        location,
        temperature: 72 + Math.floor(Math.random() * 21) - 10,
      }),
    }),
    cityAttractions: tool({
      parameters: z.object({ city: z.string() }),
      execute: async ({ city }: { city: string }) => {
        if (city === 'San Francisco') {
          return {
            attractions: [
              'Golden Gate Bridge',
              'Alcatraz Island',
              'Fisherman\'s Wharf',
            ],
          };
        } else {
          return { attractions: [] };
        }
      },
    }),
  },
  prompt:
    'What is the weather in San Francisco and what attractions should I visit?',
});
 
console.log(result);

Under the hood, the agent runs the weather tool and the cityAttractions tool simultaneously—fetches both results—and then responds with a complete answer.

And here’s a visual that shows this concurrent logic flow in action:

This pattern is incredibly useful when:

  • You want to reduce latency and improve UX
  • Tasks don’t depend on one another
  • You’re combining data from multiple sources in one go

Parallelization with tools makes your AI agent feel sharp and responsive—like it knows how to multitask properly.

So now, instead of waiting for one thing to finish before starting the next, your agents can think in multiple directions at once—and deliver answers faster than ever.

Orchestration

By this point, I had multiple workflows running—some were routed smartly, others executed in parallel. But there was still something missing: coordination.

I didn’t just want intelligent pieces—I wanted them to work together like a team.

Enter the Orchestrator.

Think of the Orchestrator as the project manager in a multi-agent system. It doesn’t solve the problem directly. Instead, it knows what steps need to happen, which agent or tool is best suited for each, and in what order everything should run.

It often blends:

  • Routing to delegate tasks
  • Prompt chaining to break things down
  • Planning or memory to keep track of what’s been done and what’s next

The Orchestrator is what makes agents feel intentional. Instead of isolated prompts, you get full workflows—like a machine that actually plans and executes.

Some real-world examples:

  • An AI writing assistant that outlines, drafts, revises, and formats an article
  • A coding assistant that scans code, finds bugs, suggests tests, and offers fixes
  • A support bot that classifies the issue, fetches relevant data, and calls the right function

Here’s a simple orchestrated workflow using the AI SDK. It starts with generating a blog outline, expands each point into paragraphs, and stitches everything into a final draft:

typescript
import { openai } from '@ai-sdk/openai';
import { generateObject } from 'ai';
import { z } from 'zod';
 
async function implementFeature(featureRequest: string) {
  // Orchestrator: Plan the implementation
  const { object: implementationPlan } = await generateObject({
    model: openai('o3-mini'),
    schema: z.object({
      files: z.array(
        z.object({
          purpose: z.string(),
          filePath: z.string(),
          changeType: z.enum(['create', 'modify', 'delete']),
        }),
      ),
      estimatedComplexity: z.enum(['low', 'medium', 'high']),
    }),
    system:
      'You are a senior software architect planning feature implementations.',
    prompt: `Analyse this feature request and create an implementation plan:
${featureRequest}`,
  });
 
  // Workers: Execute the planned changes
  const fileChanges = await Promise.all(
    implementationPlan.files.map(async (file) => {
      // Each worker is specialised for the type of change
      const workerSystemPrompt = {
        create:
          'You are an expert at implementing new files following best practices and project patterns.',
        modify:
          'You are an expert at modifying existing code while maintaining consistency and avoiding regressions.',
        delete:
          'You are an expert at safely removing code while ensuring no breaking changes.',
      }[file.changeType];
 
      const { object: change } = await generateObject({
        model: openai('gpt-4o'),
        schema: z.object({
          explanation: z.string(),
          code: z.string(),
        }),
        system: workerSystemPrompt,
        prompt: `Implement the changes for ${file.filePath} to support:
${file.purpose}
 
Consider the overall feature context:
${featureRequest}`,
      });
 
      return {
        file,
        implementation: change,
      };
    }),
  );
 
  return {
    plan: implementationPlan,
    changes: fileChanges,
  };
}

Here’s a simple diagram showing how the orchestration flows from one step to another:

While this version is manually orchestrated, the same structure can scale to much more dynamic agents with memory, feedback loops, and even self-evaluation.

The takeaway? An Orchestrator isn’t about complexity—it’s about clarity. It helps your agents think like a team, not a bunch of disconnected parts.

That brings us to the end of Part 2.

So far, we’ve seen agents grow from simple prompt wrappers to intelligent systems that can decide, parallelize, and coordinate. With Routing, Parallelization, and Orchestration, you're no longer just building tools—you’re designing AI workflows.

But we’re not done yet.

In Part 3, we’ll explore the final leap:

  • Evaluators and Optimizers — agents that reflect on output and improve it
  • Truly Autonomous Agents — agents that set goals, plan, and act without hand-holding

It’s where things start to feel alive.

👉 Stay tuned. Part 3 is where it all comes together.