Pi Durable Makes Crash Recovery a Promise Your Tools Have to Keep
Pi 1.0 shipped with a durable runtime that checkpoints every model request and tool call. The one-word replay flag on each tool decides whether resuming after a crash is safe or a second deploy.
Agent frameworks love to promise that your run survives a crash. Very few tell you what happens to the tool call that was halfway through when the process died. Pi Durable, which Earendil shipped on October 1 next to Pi 1.0, does tell you, in one sentence: "A tool call that was cut off reruns if it is safe to; otherwise the model is told it was interrupted."
That sentence is the whole design. And the word "safe" in it is not something the framework figures out. You declare it.
What landed in earendil-works/pi
Pi is an MIT-licensed coding-agent harness, and the earendil-works/pi monorepo sat at about 112k stars on a cache-busted shields read this morning. The Pi 1.0 post brings codemode with native MCP support, deferred tool loading, cache warming for Anthropic models, and mid-conversation system messages. The coding agent now publishes as @earendil-works/pi-coding-agent at 1.0.0.
The more interesting package sits in packages/durable. Pi Durable is "a framework for building any agentic application, coding agents included," published to npm as @earendil-works/pi-durable 1.0.0 under MIT. Earendil is careful to say it does not replace the Pi coding agent, and that it is experimental: "the API might still change."
The pitch is the stuff a terminal coding agent never needed and a product does. Runs survive process crashes. Any number of clients "can attach to any conversation in the harness," and any client "can steer a running conversation or queue a follow-up." One harness runs many conversations concurrently. That is the shape of an agent behind a web app, a Slack bot, or an internal tool many people share.
How the checkpointing works
Everything in a Pi Durable run is a task. The harness runs "one for each model request, one for each tool call, and one for compaction," and "every step of a run is a task that stores a checkpoint before it moves on." Timers survive restarts. Tasks can wait on other tasks.
State that is not the transcript, like a todo list, a plan, or "the sandbox a conversation runs in," lives in documents: typed JSON stored next to the transcript and "changed in the same atomic commits, so the state never disagrees with the transcript that produced it." That last guarantee is the real win of the design. Anyone who has debugged an agent whose plan file says step four while its transcript says step six knows why it matters.
Storage is pluggable. Pi Durable ships memory, SQLite, and JSONL backends, "plus a conformance suite and benchmarks for your own backend," and Earendil says the interface is small enough to put on top of "a key-value store or Postgres." One rule constrains all of it: "One process owns a storage at a time, and other clients attach to that process."
Duplicate submissions get the usual fix. Pass a requestId and a client that retries after a crash "gets the original submission back instead of asking twice":
const submission = await root.submit(
{ type: "input", content: "Hello", requestId: "greeting-1" },
context
);
The flag that carries the weight
Here is where crash recovery stops being a storage problem. When a process dies mid-tool-call, Pi Durable has two choices, and the package README makes the default explicit: the replay field accepts "safe" or nothing, and nothing means unsafe. A simplified version of the shape:
defineTool({
name: "example",
replay: "safe",
execute: async (args, api) => {
// implementation
},
})
A safe tool reruns on reopen. An unsafe one does not; the model gets an interrupted error result containing whatever output was committed before the crash, and has to decide what to do next.
This is the right default, and the reason is familiar to anyone who has used a workflow engine. Temporal's activity docs warn that an activity "may be executed multiple times and may even partially complete more than once," and tell you to design for it with idempotency keys. Pi Durable chose the opposite posture for agents: do not rerun unless told it is fine. Given that the next actor after an interruption is a language model that will happily try again, refusing to replay by default is the conservative call.
But look at what that means for you. The framework guarantees the transcript and the documents. It cannot guarantee the world outside them. Whether rerunning send_invoice, deploy_service, or post_to_slack is safe is a fact about your tool, and replay: "safe" is a sentence you write about it. Mark the wrong one and the crash-recovery feature becomes a double-charge feature.
The unsafe path has its own cost. Handing the model an interrupted result is honest, and it moves the judgment call from your code to the model. A model that sees "deploy interrupted" with partial output will often try the deploy again. The framework refused to replay; the agent may replay anyway.
Put this into practice
The lowest-friction way to try it:
npm install @earendil-works/pi-durable @earendil-works/pi-ai @earendil-works/chord
Start on the SQLite backend for a single process, and then do the part that actually matters, which takes a pen, not a terminal.
1. Sort every tool into three buckets. Pure reads (fetch a page, search a repo) are safe to replay. Writes keyed on something stable (upsert a row by ID, write a file to a fixed path) are usually safe. Anything that creates a new external effect each call (send, charge, deploy, open a PR) is unsafe until proven otherwise.
2. Make the middle bucket safe on purpose. If a tool can take an idempotency key, derive it from the run and the tool call, the same way Temporal suggests combining run and activity IDs, and pass it to the downstream API. Then the flag is true by construction, not by hope.
3. Write the interrupted path into your prompt. For unsafe tools, tell the model what an interrupted result means and what to check before retrying ("look up whether the deploy exists before calling deploy again"). That instruction is the only thing standing between an interruption and a duplicate.
4. Use requestId on every client submission. It is one field, and it removes a whole class of "the user clicked twice during a restart" bugs.
5. Point your agent at the source. Earendil says the entire runtime is about 15,000 lines, roughly "150,000 tokens with GPT and about 250,000 with Claude," and that "the storage backends alone are 3,000 lines it can usually skip." A codebase sized for an agent to read is a feature you should use when you get stuck.
Honest limitations
It is experimental. Earendil says so, and an API that may change is a real cost for anything you plan to run in production next quarter.
JavaScript only. Pi Durable runs on JavaScript runtimes. A Python agent stack cannot adopt it without a rewrite.
One owner per storage. A single process owns a storage at a time. That is simple to reason about and it is not horizontal scaling. If that process's machine dies, recovery waits for a new owner.
The SQLite default trades some durability for speed. The package README says the SQLite backend uses "WAL mode with synchronous = NORMAL: commits survive process crashes; the newest may be lost on power or host failure." A process crash is covered. A power cut can drop the latest step.
Pi itself has no permission layer. The Pi README states that it "does not include a built-in permission system for restricting filesystem, process, network, or credential access." Pi Durable makes runs survive. It does nothing to limit what those runs are allowed to touch, so pair it with whatever isolation your deployment already uses.
No numbers yet. Earendil ships benchmarks for backend authors, but the posts do not publish recovery-time or throughput figures for real workloads, and this article does not invent any.
The audit is the feature
The easy story about Pi Durable is "agents that survive crashes." The more useful one is that it made the hard question impossible to skip. Every tool you register now carries a one-word answer to "what happens if this runs twice?", and the default answer is "I don't know, so don't."
That is a good place for a framework to stand. It also means durability in agents is not something you install. You earn it, tool by tool, by knowing which of your side effects can repeat. Open your tool list today and mark each one. The ones you hesitate on are where your next incident lives.
Sources: Pi 1.0 announcement, Pi Durable announcement, pi-durable package README, earendil-works/pi repository, Temporal activity definition docs, npm registry metadata for @earendil-works/pi-durable 1.0.0 and @earendil-works/pi-coding-agent 1.0.0.
Medium metadata
- Title: Pi Durable Makes Crash Recovery a Promise Your Tools Have to Keep
- Subtitle: Pi 1.0 shipped with a durable runtime that checkpoints every model request and tool call. The replay flag on each tool decides whether resuming is safe.
- Tags: AI Agents, Software Architecture, Open Source, TypeScript, Distributed Systems
- Canonical URL: fervorai.dev (import from the published post)
- Reading time: about 8 minutes