CrewAI in Depth: Roles, Tasks, and Process Orchestration
A practical tour of CrewAI's core abstractions — agents, tasks, processes, and Flows — with guidance on when this role-shaped framework beats a graph-based one.
Published on • October 4, 2026
AI Assistant

Introduction
Most multi-agent frameworks ask you to think like a graph engineer: nodes, edges, state machines. CrewAI asks you to think like a team lead. You define people (agents), jobs (tasks), and how the team runs its standup (process). That mental model is why CrewAI took off — it maps directly onto how humans already coordinate work, and it lets product engineers ship agentic workflows without first mastering a workflow DSL.
The adoption numbers back this up. CrewAI’s own site reports agentic workflows running at a scale of 450M+ per month, 4,000+ signups per week, and usage by 65% of the Fortune 500, with customers like DocuSign, Experian, PepsiCo, and IBM. Documented outcomes include Docusign achieving 75% faster first contact with leads, General Assembly cutting development time by 90%, and a food-ordering QA operation dropping from 74 hours to 3.
This post walks through the four layers of the framework — Agent, Task, Crew, Flow — and ends with a decision rule against LangGraph plus the production caveats nobody warns you about.
Agents: A Human Team, Not a Function Call
An agent in CrewAI is defined by three prompt fields that read like a job description:
- role — the title. “Senior Market Research Analyst.”
- goal — what success looks like for this person specifically.
- backstory — the resume that primes the model’s tone, depth, and domain vocabulary.
from crewai import Agent
from crewai_tools import SerperDevTool
analyst = Agent(
role="Senior Market Research Analyst",
goal="Identify credible market trends with sourced evidence",
backstory=(
"You have 12 years in B2B SaaS research. You distrust unsourced "
"claims and always separate observed data from speculation."
),
tools=[SerperDevTool()],
verbose=True,
)
The backstory is not decoration. It is the cheapest lever you have for controlling output quality: a skeptical researcher produces different text than an enthusiastic copywriter, even with the same task. Agents also accept llm (per-agent model routing), memory flags, max_iter (the internal loop cap before the agent gives up on a tool result), and delegation (whether it can hand subtasks to teammates).
The pattern that works: narrow roles, not generalists. Four agents with sharp boundaries beat one agent with four skills, because the prompt context each one carries stays small and relevant.
Tasks: The Unit of Work
If agents are people, tasks are tickets. A task specifies:
- description — what to do, written as an instruction to a specialist, not a vague wish.
- expected_output — the format contract. This is the single highest-leverage field in CrewAI; it is what the downstream task (or your code) actually parses.
- agent — who owns it.
- tools — per-task tool overrides, so a writing agent doesn’t get the database credentials for a research task.
- context — which prior tasks’ outputs get injected. Explicit, not implicit.
- output_pydantic / output_json — enforce structured output so you get a typed object instead of prose.
from crewai import Task
from pydantic import BaseModel
class TrendReport(BaseModel):
trend: str
evidence: list[str]
confidence: float
summarize = Task(
description="Summarize the top 3 trends for {topic} with evidence.",
expected_output="A TrendReport with trend, evidence, and confidence.",
output_pydantic=TrendReport,
agent=analyst,
)
Note the {topic} placeholder — task descriptions are templated against the inputs you pass to kickoff(). Treat expected_output as an API contract: if you cannot write an assertion against it, it is not specific enough.
Crew and Processes: Sequential vs Hierarchical
A Crew bundles agents and tasks and chooses a process — the scheduling strategy:
- sequential: tasks run in order, each seeing the previous output. Cheap, predictable, ideal for pipelines (research → draft → edit).
- ** hierarchical**: agents dynamically delegate through a manager agent. You must supply a
manager_llm. More flexible, more tokens, harder to predict.
from crewai import Crew, Process
crew = Crew(
agents=[analyst, writer],
tasks=[summarize, draft],
process=Process.sequential,
memory=True,
output_log_file=True,
)
result = crew.kickoff(inputs={"topic": "AI agent orchestration"})
Use hierarchical only when the task decomposition itself is genuinely unknown at authoring time. For everything else, sequential with an explicit context chain is easier to debug, cheaper, and reproducible.
Flows: Deterministic Control Around Crews
Here is the honest limitation of pure crews: an LLM decides execution order. Sometimes you need code to decide — validation gates, branching, retries, human approval. That is what Flows are for: a deterministic, event-driven shell that runs Python code and invokes crews as steps.
from crewai.flow.flow import Flow, listen, router, start
from pydantic import BaseModel
class ReportState(BaseModel):
topic: str = ""
report: str = ""
approved: bool = False
class ReportFlow(Flow[ReportState]):
@start()
def prepare(self):
self.state.topic = self.state.topic.strip()
@listen(prepare)
def run_crew(self):
result = Crew(...).kickoff(inputs={"topic": self.state.topic})
self.state.report = result.raw
@router(run_crew)
def check(self):
return "ok" if len(self.state.report) > 500 else "retry"
@listen("ok")
def publish(self):
publish_to_cms(self.state.report)
@start marks entry points, @listen reacts to a method’s output, and @router branches on a returned label that other @listen("label") methods consume. Flows also give you typed state (Pydantic), or_/and_ combinators, @persist for SQLite-backed state recovery across restarts, @human_feedback (CrewAI ≥ 1.8.0) for approval gates, and flow.usage_metrics — a provider-neutral token rollup across every crew, tool, and bare LLM call in the run. That last one matters for cost attribution.
The architecture that scales: Flow owns control flow; Crews own judgment. Deterministic code decides when; agents decide what.
Tooling and Memory
Tools are ordinary Python functions or classes decorated for LLM use; CrewAI ships a toolbox (Serper search, file readers, PDF scanners) and accepts anything from crewai-tools or MCP servers. Keep tools narrow with typed parameters — agents fail less when the argument schema is obvious.
Memory is layered: short-term per-run context, long-term persistence (LanceDB by default), and entity memory. On a crew, memory=True lets later tasks recall earlier findings; in Flows, self.remember(...) / self.recall(...) accumulate knowledge across runs. Useful for recurring workflows (weekly competitive scans), actively harmful for one-shot tasks where stale context pollutes the answer.
When CrewAI Shines vs When to Pick LangGraph
Choose CrewAI for role-shaped, collaborative workflows: research and reporting, content pipelines, lead enrichment, support triage — work that decomposes naturally into “specialist A hands to specialist B,” where you want readable prompts over precise state control. Gelato’s setup (3,000+ leads enriched monthly) is the archetype.
Choose LangGraph when correctness depends on exact state: transactional flows with compensation, long-running executions requiring checkpointed resumption at specific nodes, or anywhere you need to assert “the graph is at node X with state Y.” LangGraph gives you a explicit state machine; CrewAI gives you a team. Forcing a team to behave like a state machine is where CrewAI projects get painful.
Production Caveats
- Cost: hierarchical processes and memory multiply token spend. Log
usage_metrics(Flow) ortoken_usage(crew output) per run and alert on deltas. - Observability:
verbose=Trueis not observability. Ship traces (OpenTelemetry GenAI semantic conventions) capturing each agent’s inputs, tool calls, and outputs — otherwise a bad run is unreproducible. - Guardrails: use
output_pydanticvalidation, tool allowlists, and per-agentmax_iter. Add human approval via@human_feedbackbefore any side-effecting step. - Nondeterminism: same inputs, different outputs. Pin models, version prompts, and evaluate changes offline before production.