Guide

LangGraph Checkpointer in Production: Setup and Fixes

A LangGraph checkpointer is what gives a graph memory. It saves the graph's state at every step so a conversation can continue, a run can pause for approval and resume, and a failed run can be inspected afterwards.

In development the default in-memory checkpointer just works. In production it is the part of LangGraph persistence most likely to fail quietly, because a wrong setting does not crash anything: state goes missing, or the database fills up.

This guide covers running a LangGraph checkpointer in production: which backend to choose, how to set up the Postgres checkpointer correctly, and how to fix the problems teams hit most often. If you want to use checkpoints to debug a misbehaving agent, see how to debug LangGraph agents in production.

What a LangGraph checkpointer stores

When you compile a graph with a checkpointer, LangGraph saves a checkpoint after each super-step, meaning each round of node execution. A checkpoint holds the state values, which nodes run next, and metadata about the step.

Checkpoints are grouped into threads. Every call names its thread with a thread_id in the run config, and everything the graph remembers about that conversation or job lives under that id:

config = {"configurable": {"thread_id": "ticket-4812"}}
graph.invoke(inputs, config)

Call the graph again with the same thread_id and it picks up where it left off. Call it with a new one and it starts from nothing.

Choosing a checkpointer for production

LangGraph ships several checkpointers, and the documentation is clear about which are meant for production:

CheckpointerPackageUse it for
InMemorySaverlanggraph-checkpointTests and local development only
SqliteSaver / AsyncSqliteSaverlanggraph-checkpoint-sqliteExperiments and local workflows
PostgresSaver / AsyncPostgresSaverlanggraph-checkpoint-postgresProduction
MongoDBSaver / AsyncMongoDBSaverlanggraph-checkpoint-mongodbProduction
CosmosDBSaverlangchain-azure-cosmosdbProduction on Azure

If your graph runs with ainvoke or astream, use the async variant of whichever backend you pick. Most teams already run Postgres, which makes the Postgres checkpointer the usual choice, and it is the one the rest of this guide uses.

Setting up the LangGraph Postgres checkpointer

Three details decide whether the Postgres checkpointer works, and none of them raise an error at the point where you get them wrong.

from_conn_string is a context manager. It opens a connection and closes it when the block ends, so use it with with:

from langgraph.checkpoint.postgres import PostgresSaver

with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
    checkpointer.setup()   # first time only
    graph = builder.compile(checkpointer=checkpointer)
    graph.invoke(inputs, config)

setup() must be called by you, once. It creates the checkpoint tables and runs migrations, and the library never calls it on its own, so a fresh database has no tables until you do. Run it in a deploy or migration step, not inside every request.

A connection you create yourself needs two settings. A long-running service usually passes a connection pool rather than a connection string, and the pool's connections must be created with autocommit=True and row_factory=dict_row:

from psycopg.rows import dict_row
from psycopg_pool import ConnectionPool
from langgraph.checkpoint.postgres import PostgresSaver

with ConnectionPool(
    conninfo=DB_URI,
    kwargs={"autocommit": True, "row_factory": dict_row},
) as pool:
    checkpointer = PostgresSaver(pool)
    graph = builder.compile(checkpointer=checkpointer)

Without autocommit=True, the tables setup() creates may never be committed. Without dict_row, the checkpointer's first read fails, because it reads columns by name and psycopg returns tuples by default.

Common LangGraph checkpointer problems and fixes

State disappears after a restart or deploy

The graph is using InMemorySaver (called MemorySaver in older code), which keeps checkpoints in the process's RAM, so every restart, deploy or crash wipes every thread. It often goes unnoticed in development because nothing restarts mid-conversation. Switch to a persistent checkpointer.

TypeError: tuple indices must be integers or slices, not str

The connection passed to PostgresSaver was created without row_factory=dict_row. Add it to the connection or pool settings as shown above.

The checkpoint tables do not exist

Either setup() was never called against this database, or it ran on a connection without autocommit=True and the table creation was never committed. Fix the connection settings, then run setup() once.

Errors on long thread IDs

PostgresSaver stores thread_id in a column of limited length, so keep ids under 255 characters. If you build ids from user, session and task names, hash them or use a UUID instead of joining raw strings.

Two users see each other's conversation

The checkpointer does exactly what it is told: every call with the same thread_id shares one state. A fixed id, a default copied from an example, or an id built from something that is not unique per conversation will merge conversations. Give every conversation or job its own id, and never fall back to a constant.

The checkpoint tables keep growing

Every super-step writes a checkpoint, so long-running threads and busy agents accumulate rows indefinitely, which raises storage costs and query latency. The LangGraph persistence documentation recommends pruning old checkpoints on a schedule, for example a job that deletes threads older than your retention window.

delete_thread(thread_id) removes a thread's checkpoints and pending writes together. For append-heavy state such as long chat histories, the documentation also describes DeltaChannel, which stores only what changed at each step instead of the full value.

An object in your state cannot be saved

The default serializer handles common types but not everything, a pandas DataFrame for example. JsonPlusSerializer(pickle_fallback=True) lets it pickle what it cannot otherwise encode.

Treat that as a trade-off. The Postgres checkpointer's own README advises restricting deserialization to known-safe types, with LANGGRAPH_STRICT_MSGPACK=true or an allowed_msgpack_modules list, so a compromised database cannot be used to run code. Where you can, keep state to plain data and store large objects elsewhere by reference.

A subgraph cannot see the parent graph's state

Each subgraph keeps its own checkpoint namespace. Data that has to cross a graph boundary belongs in a store, not in checkpointed state.

Checkpointer vs store in LangGraph

The two are easy to confuse because both persist data:

If something should be remembered after the thread ends, or read by a different thread, it belongs in a store. If it only matters within the current run, leave it in the graph state and let the checkpointer handle it.

A production checklist for your LangGraph checkpointer

It is also worth alerting on the checkpoint tables' growth and on checkpointer errors, alongside the signals in our guide to AI agent monitoring.

Where FleetHelp fits

A misconfigured checkpointer rarely announces itself. An agent keeps answering while it forgets context, mixes up conversations, or slows down as its tables grow, and the fix is usually one setting that is easy to find once you know where to look.

With FleetHelp's managed support, your agents message our support fleet over Telegram when something breaks, and get a diagnosis and a tested fix back, usually in under 60 seconds. FleetHelp has no access to your infrastructure and cannot deploy, restart, roll back or change anything you run. Plans are $99 a month per organisation; see pricing and subscribe.

Frequently Asked Questions

What does a LangGraph checkpointer do?

A LangGraph checkpointer saves a snapshot of your graph's state at every step and groups those snapshots into threads, identified by a thread_id. That is what lets a graph remember a conversation between calls, pause for human approval, resume after a failure, and be inspected or replayed later.

Which LangGraph checkpointer should I use in production?

Use a durable one: PostgresSaver or MongoDBSaver, or CosmosDBSaver on Azure, with the async variant if your graph runs async. InMemorySaver loses everything when the process restarts, and SqliteSaver is meant for local work and experiments.

Why does my LangGraph agent forget everything after a restart?

Because it is using InMemorySaver or MemorySaver, which keep checkpoints in RAM. Switch to a persistent checkpointer such as PostgresSaver and call setup() once so its tables exist.

What is the difference between a store and a checkpointer in LangGraph?

A checkpointer holds short-term memory for one thread: the graph state as a run moves through it. A store holds long-term memory your application defines, shared across threads, such as facts about a user. Use a store for anything that must survive beyond a single thread or cross into a subgraph.