Artificial intelligence

How to Make LangGraph and CrewAI Agents Crash-Proof Without Rewriting Them

How to Make LangGraph and CrewAI Agents Crash-Proof Without Rewriting Them

LangGraph and CrewAI have become two of the most popular ways to build AI agents in Python. LangGraph gives you explicit control over agent state as a graph of nodes and edges. CrewAI gives you a higher-level abstraction of role-based agents working through tasks with tools. Both are excellent for building agents. Neither was designed to be the thing that keeps your agent alive when a Kubernetes pod is evicted halfway through a twenty-step run.

This tutorial shows how to add durable execution to an existing LangGraph graph or CrewAI agent so that every node or tool call becomes a recoverable, recorded step, without restructuring your agent code.

What goes wrong today

Take a LangGraph agent with a straightforward flow: research, draft, review, publish. You compile it, invoke it, and it runs as a single call in one Python process.

LangGraph does offer checkpointers, such as a Postgres-backed saver that persists graph state after each node. That’s valuable, but it’s important to understand what it doesn’t do. If the process crashes, the checkpoint is safely in the database, and then nothing happens. Nothing detects that the run died. Nothing schedules a resume on a healthy instance. If you run several replicas, nothing coordinates which one picks the run up. You end up writing a supervisor, and making that supervisor reliable is its own distributed-systems project.

CrewAI has a similar gap. If the process dies while the agent is iterating through tool calls, the in-flight task is lost, and a naive retry repeats every tool call from the beginning, including any that had side effects.

What you want is:

  1. Each node (LangGraph) or tool call (CrewAI) runs as a durable activity whose result is persisted.
  2. A crashed run is automatically resumed on any healthy instance.
  3. Completed steps are never re-executed. Their recorded results are replayed.
  4. Failed steps are retried with a policy, not surfaced as a fatal error.

That’s the definition of a durable workflow engine, and Dapr ships one.

The approach: wrap, don’t rewrite

The idea is to leave your graph or agent definition untouched and swap the runner. Instead of calling compiled.invoke(…) directly, you hand the compiled graph to a runner that executes it on Dapr Workflow. Each node execution becomes a workflow activity, and the workflow’s event history becomes the source of truth for where the run is.

Diagrid maintains a source-available Python integration package (diagrid on PyPI) that implements this for LangGraph, CrewAI, Google ADK, Strands, Pydantic AI, the OpenAI Agents SDK and others. The examples below use it.

Prerequisites

  • Python 3.10 or later
  • The Dapr CLI installed and initialized locally (dapr init), which also starts a local Redis container
  • An LLM API key for the model your agent uses

Install the extra for your framework:

pip install “diagrid[langgraph]”

# or

pip install “diagrid[crewai]”

Dapr Workflow needs a state store that supports actors. Create resources/statestore.yaml:

apiVersion: dapr.io/v1alpha1

kind: Component

metadata:

  name: statestore

spec:

  type: state.redis

  version: v1

  metadata:

  – name: redisHost

    value: localhost:6379

  – name: actorStateStore

    value: “true”

In production you’d point this at a managed PostgreSQL, Cosmos DB or another supported store. Your application code doesn’t change.

Part 1: A durable LangGraph agent

Here’s a minimal three-node graph. The graph definition is plain LangGraph:

from typing import List, TypedDict

from langgraph.graph import StateGraph, START, END

class State(TypedDict):

    messages: List[str]

    counter: int

def process_node(state: State) -> dict:

    return {“messages”: state[“messages”] + [“processed”],

            “counter”: state[“counter”] + 1}

def validate_node(state: State) -> dict:

    return {“messages”: state[“messages”] + [“validated”],

            “counter”: state[“counter”] + 1}

def finalize_node(state: State) -> dict:

    return {“messages”: state[“messages”] + [“finalized”],

            “counter”: state[“counter”] + 1}

graph = StateGraph(State)

graph.add_node(“process”, process_node)

graph.add_node(“validate”, validate_node)

graph.add_node(“finalize”, finalize_node)

graph.add_edge(START, “process”)

graph.add_edge(“process”, “validate”)

graph.add_edge(“validate”, “finalize”)

graph.add_edge(“finalize”, END)

compiled = graph.compile()

The only new code is the runner:

import asyncio

from diagrid.agent.langgraph import DaprWorkflowGraphRunner

async def main():

    runner = DaprWorkflowGraphRunner(graph=compiled, max_steps=50,

                                     name=”simple_graph”)

    runner.start()

    try:

        async for event in runner.run_async(

            input={“messages”: [“hello”], “counter”: 0},

            thread_id=”order-4711″,

        ):

            print(event[“type”], event.get(“status”, “”))

    finally:

        runner.shutdown()

asyncio.run(main())

Run it with a Dapr sidecar:

dapr run –app-id langgraph-agent –resources-path ./resources — python3 agent.py

Two details matter. First, thread_id is the workflow’s identity: if the process restarts and the same thread is resumed, Dapr continues the existing run instead of starting a new one. Second, the runner streams lifecycle events (workflow_started, workflow_status_changed, workflow_completed, workflow_failed), which you can forward to your UI or logs.

In a real agent, the nodes would call an LLM and tools. Because each node runs as an activity, a node that finished before a crash is never re-run. Its output is replayed from the workflow history, so you don’t pay for the same LLM call twice or trigger the same side effect twice. That is the core idea behind durable execution for LangGraph agents: the graph describes the logic, and the runtime guarantees it finishes.

Part 2: A durable CrewAI agent

For CrewAI, the unit of durability is the tool call. Define your agent, tools and task as you normally would:

import os

from crewai import Agent, Task

from crewai.tools import tool

def get_order(order_id: str) -> str:

    “””Return order status from the order service.”””

    return f”Order {order_id}: shipped”

agent = Agent(

    role=”Support Assistant”,

    goal=”Resolve customer order questions accurately”,

    backstory=”You help customers track and fix their orders.”,

    tools=[get_order],

    llm=os.getenv(“CREWAI_LLM”, “openai/gpt-4o-mini”),

)

task = Task(

    description=”Customer asks where order 4711 is. Find out and reply.”,

    expected_output=”A short, friendly status update.”,

    agent=agent,

)

Then run it through the workflow runner:

import asyncio

from diagrid.agent.crewai import DaprWorkflowAgentRunner

async def main():

    runner = DaprWorkflowAgentRunner(agent=agent, name=”support-agent”,

                                     max_iterations=10)

    runner.start()

    try:

        async for event in runner.run_async(task=task,

                                            session_id=”ticket-981″):

            print(event[“type”])

    finally:

        runner.shutdown()

asyncio.run(main())

Each tool invocation is now a workflow activity. If the process dies after get_order returns, the resumed run already has that result recorded and continues from the next reasoning step.

Testing crash recovery (do this before you ship)

Durability claims only count once you’ve tested them. A simple, repeatable test:

  1. Add a time.sleep(10) inside one of the middle nodes or tools so you have a window to act.
  2. Start the agent with dapr run and wait until the logs show that the first step has completed.
  3. Kill the process with Ctrl+C or kill -9.
  4. Start it again with the same thread_id or session_id.
  5. Confirm from the logs that the completed steps weren’t re-executed and the run finished.

Repeat the test with a tool that raises a transient exception on its first two calls, to confirm the retry behaviour matches your expectations. Put both tests in CI. They’re cheap and they catch regressions that unit tests miss.

Where the durability boundary sits

It’s worth being precise about what “durable” means here, because it shapes how you design nodes and tools.

For LangGraph, the unit of durability is the node. If a node makes three LLM calls and two tool calls internally, and the process crashes during the fourth call, the whole node runs again on recovery. That’s a good reason to keep nodes small and focused: one LLM call, or one tool call with its surrounding logic. Smaller nodes mean less repeated work and cheaper recoveries.

For CrewAI, the unit is the tool call, and the LLM reasoning between tool calls is replayed from recorded results. Keep tools narrow and side effects explicit: a tool that “creates a ticket and emails the customer” is harder to make safe than two separate tools.

In both cases, anything that must not happen twice belongs in its own step, with an idempotency key passed to the downstream API.

Production considerations

Once the local test passes, a few things change in production:

  • State store: Use a highly available database for workflow state, not a local Redis container.
  • Multiple replicas: Run more than one instance of the agent service. Dapr’s placement service spreads workflow execution across healthy replicas and moves it when one dies.
  • Idempotency at the edges: Durable execution guarantees a recorded step won’t be repeated, but a step that crashed mid-call may run again. Pass idempotency keys to payment, email and ticketing APIs.
  • Observability: Dapr emits OpenTelemetry traces and metrics for workflow and activity execution. Send them to your existing tracing backend so each agent run shows up as a trace, one span per step.
  • Security: In Kubernetes, sidecar-to-sidecar traffic is protected by mTLS with SPIFFE-based identities by default, which matters when agents call internal tools and other agents.

If you’d rather not run the workflow infrastructure yourself, Diagrid offers a managed option through Catalyst, which adds hosted durable execution and signed, verifiable execution history on top of the same integration.

Wrapping up

The frameworks you use to design agents and the infrastructure you use to run them solve different problems, and they don’t need to be the same tool. LangGraph and CrewAI are great at expressing agent logic. Durable workflow engines are great at making sure long-running, failure-prone processes finish. Keeping them separate, with a thin runner in between, gives you both without a rewrite: your graph stays readable, and your agents survive the restarts, deploys and flaky APIs that production will inevitably throw at them.

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This