PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Building a Multi-Agent System With CrewAI
Multi-Agent Systems

Building a Multi-Agent System With CrewAI

CrewAI promises to make multi-agent coordination feel like assembling a team. This crewai tutorial shows you what that means in practice — the parts that work cleanly, the parts that fight you, and when to reach for a different tool.

September 21, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
from crewai import Agent, Task, Crew, Process
from crewai_tools import SerperDevTool

search_tool = SerperDevTool()

# Define the agents
researcher = Agent(
    role="Senior Research Analyst",
    goal="Find the most relevant and accurate information on the given topic",
    backstory="""You are an experienced research analyst who specializes in
    synthesizing information from multiple sources into clear, factual summaries.
    You cite specific data points and avoid vague generalities.""",
    tools=[search_tool],
    verbose=True
)

analyst = Agent(
    role="Data Analyst",
    goal="Identify key patterns and implications in research findings",
    backstory="""You are a critical thinker who evaluates research quality,
    identifies gaps in evidence, and surfaces the most significant insights.
    You are skeptical of claims that lack supporting data.""",
    verbose=True
)

writer = Agent(
    role="Technical Writer",
    goal="Produce clear, structured reports from analyzed findings",
    backstory="""You write clear, organized content with explicit structure.
    You never include content you cannot support from the provided research.""",
    verbose=True
)

# Define tasks
research_task = Task(
    description="Research the current state of {topic}. Find three to five key developments from the past year with specific data points, not general trends.",
    expected_output="A structured list of specific findings, each with source attribution and a key statistic or data point.",
    agent=researcher
)

analysis_task = Task(
    description="Analyze the research findings. Identify the two most significant implications and flag any claims that appear weakly supported.",
    expected_output="An analysis section with two key implications and a credibility assessment of the research.",
    agent=analyst,
    context=[research_task]
)

write_task = Task(
    description="Write a 400-word structured report on {topic} using the research and analysis provided.",
    expected_output="A report with an executive summary, three key findings, and a conclusion. Markdown formatted.",
    agent=writer,
    context=[research_task, analysis_task]
)

# Assemble and run the crew
crew = Crew(
    agents=[researcher, analyst, writer],
    tasks=[research_task, analysis_task, write_task],
    process=Process.sequential,
    verbose=True
)

result = crew.kickoff(inputs={"topic": "large language model reasoning capabilities"})
print(result.raw)

A startup's engineering team spent three days debugging a CrewAI pipeline that kept producing empty output from the final agent. They checked the agent prompts, the task definitions, the tool integrations. Everything looked correct. Eventually they found the bug: the final agent's task had a context parameter pointing to an intermediate task — but that intermediate task's output was a large nested dict, and CrewAI's internal serialization was silently truncating it. The final agent was reasoning from a blank slate while the team debugged the wrong layer entirely.

This is the kind of failure that a crewai tutorial doesn't usually cover. Framework documentation shows the happy path. Real usage shows you the edges.

This post covers both. We'll build a working three-agent research pipeline as a crewai tutorial, and then we'll dissect exactly what works, what fights you, and what you should use instead when CrewAI's design assumptions don't fit your task.

What CrewAI Is Actually Designed For

CrewAI models multi-agent collaboration as a crew of workers with defined roles executing a series of tasks in a defined order. The metaphor is deliberate: you're configuring a team, not writing an orchestrator.

Each agent has a role, a goal, a backstory, and optionally a set of tools. Each task has a description, an expected output, and an agent assigned to it. The Crew binds agents and tasks together, executes them sequentially (or hierarchically), and passes outputs from one task to the next.

This design is extremely well-suited to pipelines where the task order is fixed, the handoffs between agents are predictable, and role-based specialization provides real value. Content production, research summarization, structured report generation, and multi-step data enrichment all map naturally to CrewAI's model.

The CrewAI Tutorial: A Three-Agent Research Pipeline

Here's a minimal but representative crewai tutorial: a pipeline that researches a topic, analyzes the findings, and writes a structured report.

python
[object Object], crewai ,[object Object], Agent, Task, Crew, Process
,[object Object], crewai_tools ,[object Object], SerperDevTool

search_tool = SerperDevTool()

,[object Object],
researcher = Agent(
    role=,[object Object],,
    goal=,[object Object],,
    backstory=,[object Object],,
    tools=[search_tool],
    verbose=,[object Object],
)

analyst = Agent(
    role=,[object Object],,
    goal=,[object Object],,
    backstory=,[object Object],,
    verbose=,[object Object],
)

writer = Agent(
    role=,[object Object],,
    goal=,[object Object],,
    backstory=,[object Object],,
    verbose=,[object Object],
)

,[object Object],
research_task = Task(
    description=,[object Object],,
    expected_output=,[object Object],,
    agent=researcher
)

analysis_task = Task(
    description=,[object Object],,
    expected_output=,[object Object],,
    agent=analyst,
    context=[research_task]
)

write_task = Task(
    description=,[object Object],,
    expected_output=,[object Object],,
    agent=writer,
    context=[research_task, analysis_task]
)

,[object Object],
crew = Crew(
    agents=[researcher, analyst, writer],
    tasks=[research_task, analysis_task, write_task],
    process=Process.sequential,
    verbose=,[object Object],
)

result = crew.kickoff(inputs={,[object Object],: ,[object Object],})
,[object Object],(result.raw)

What this does: The crew runs three agents in sequence. The researcher uses a web search tool to gather information, producing structured findings. The analyst receives those findings via context=[research_task] and produces an evaluation. The writer receives both prior task outputs and produces the final report. The {topic} placeholder is interpolated from inputs at kickoff time — you can reuse the same crew definition for different topics without re-instantiation.

⚡ Pro tip: The backstory field does more work than its name implies. In testing across multiple crewai tutorial implementations, agents with backstories emphasizing skepticism, specificity, or a particular professional discipline produce measurably more focused output than agents with generic descriptions. Treat the backstory as a constraint specification, not a personality trait.

What Works Well in CrewAI

Role-based context scoping. When you assign a task to a specific agent, that agent's role and backstory frame how it interprets and responds to the task. This means you don't have to cram all behavioral context into the task description — the agent definition carries it. A researcher agent behaves like a researcher without being told to in every task prompt.

Sequential pipeline clarity. CrewAI's sequential process is straightforward to reason about. You read the task list top to bottom and know exactly what happens in what order. Debugging sequential pipelines is significantly easier than debugging dynamic routing systems — when something goes wrong, you know which task produced which output.

Task context passing. The context parameter lets you explicitly declare which prior task outputs a task needs. This is cleaner than global shared state: each task declares its dependencies, and CrewAI handles the serialization. For clean, predictable handoffs, this works well.

Low boilerplate for common patterns. A three-agent research pipeline in CrewAI takes 50 lines of configuration. The equivalent custom orchestration would require more infrastructure. For teams that need a working prototype quickly, CrewAI's abstractions save real time.

Where the Framework Works Against You

Output size and serialization. CrewAI serializes task outputs as strings passed to subsequent agents. Large outputs — extensive code, long structured documents, complex nested data — can be truncated, and the failure is often silent. The downstream agent receives truncated context but doesn't know it, so it proceeds with incomplete information and produces subtly wrong output.

No native parallelism. CrewAI's sequential process runs tasks one at a time. If you need two agents to work simultaneously — analyzing different aspects of the same document, for example — you're outside the framework's native model. The hierarchical process adds a manager agent to delegate tasks, but it adds latency and an extra LLM call; it's not true parallelism.

Opaque intermediate state. In sequential mode, intermediate task outputs are stored internally. You can set verbose=True to see them printed, but you can't easily access them programmatically during execution for logging, validation, or conditional branching based on output quality.

Dynamic routing is underdeveloped. If you want to route to different agents based on what a prior agent found — a classification step that sends inputs to different specialists — you need to implement this outside the framework. CrewAI's conditional logic support is limited; LangGraph is better suited for workflows with dynamic routing requirements.

⚡ Pro tip: For any crewai tutorial implementation going into production, add a post-task output validator. After each task completes, check that the output meets minimum length and structural requirements before passing it to the next agent. CrewAI doesn't do this for you, and silent output failures are the framework's biggest practical liability.

Handling the Context Truncation Problem

The startup's debugging problem has a practical fix. Instead of passing large structured objects as task context, serialize important outputs to a lightweight format before they're consumed by downstream tasks:

python
[object Object], pydantic ,[object Object], BaseModel

,[object Object], ,[object Object],(,[object Object],):
    findings: ,[object Object],[,[object Object],]
    sources: ,[object Object],[,[object Object],]
    key_statistic: ,[object Object],

research_task = Task(
    description=,[object Object],,
    expected_output=,[object Object],,
    agent=researcher,
    output_json=ResearchFindings
)

What this does: By specifying output_json, CrewAI validates the researcher's output against the Pydantic schema and stores it as a structured object rather than a raw string. Downstream tasks receive consistent, bounded output rather than a freeform text blob. This eliminates most truncation and schema-drift failures in practice.

⚡ Pro tip: Use output_json for any task whose output will be consumed by multiple downstream agents. Unstructured string outputs work for terminal tasks (those producing human-readable reports), but intermediate outputs that feed other agents benefit significantly from schema enforcement.

When to Use Something Else

CrewAI is the right choice when your pipeline is sequential, roles add genuine interpretive value, and the tasks map cleanly to a fixed order. It's the wrong choice when:

  • Tasks need to run in parallel for performance reasons
  • Routing between agents depends on the output of a classification step
  • You need fine-grained control over what context each agent receives
  • Your pipeline is simple enough that three lines of sequential API calls would work

For dynamic routing and stateful conditional logic, LangGraph is the right tool. For conversational multi-agent deliberation, AutoGen is better designed. For pipelines simple enough to fit in a single function, plain Python with direct API calls is easier to maintain than any framework.

⚠️ Common mistake: Treating CrewAI's backstory as optional flavor text. Teams that skip backstory definitions or use generic placeholders see significantly more role drift — agents that don't stay in their lane, or that produce output that mirrors the task description rather than applying domain-specific judgment. The backstory is the agent's behavioral frame. Invest the same care in it that you'd invest in a system prompt.

Running Your First CrewAI Pipeline

The fastest path through this crewai tutorial to a working system: start with two agents, not three. A researcher and a writer, sequential, no tools. Get the output format right first. Add tools, add a third agent, and move to structured output schemas only after the two-agent pipeline produces output you trust.

Debugging CrewAI Pipelines

When a CrewAI pipeline produces wrong output, the debugging path is more constrained than in a custom orchestrator. You have access to verbose output (set verbose=True on agents and the crew), task output attributes, and CrewAI's built-in usage_metrics. What you don't have is easy programmatic access to intermediate state during execution.

The most effective debugging approach for CrewAI pipelines is isolation testing: run each agent as a standalone agent (not inside a crew) with test inputs representing what it would receive in the pipeline. Verify each agent's output independently before assembling them into a crew. When agents work correctly in isolation but not in the crew, the problem is usually the context parameter — what's being passed between tasks.

Check these four things when a crew produces wrong output: first, whether the task descriptions contain {topic} or similar placeholders that weren't populated by kickoff(inputs=...). Second, whether context=[task] is set for tasks that depend on prior outputs. Third, whether the final task's expected_output description is specific enough — vague expected outputs produce vague outputs. Fourth, whether any agent's output was truncated in the context window.

For tasks where the final output format matters — structured JSON, specific markdown sections, particular word counts — specify the format in both the task's expected_output field and the agent's goal. Specification in one place is not sufficient.

Well-crafted agent definitions — roles, goals, and backstories — are reusable across projects. Once you've found formulations that produce consistent behavior, save them to PromptABCD and pull them into new crew configurations without starting from scratch each time.

multi-agent-systemscrewaicrewai-tutorialagent-frameworksllm-engineering

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousShared Memory in Multi-Agent SystemsNext →Building a Multi-Agent System With AutoGen
Share this post:
ShareShare