Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · October 4, 2026 · repo

claude-memClaude CodeOperator Memoryagent-memoryclaude-codemcp

claude-mem 13.29.0 Adds the One Kind of Agent Memory You Can Audit

The memory plugin for Claude Code just shipped a keyed to-do ledger next to its vector recall. The split tells you which memories to trust.

On the same Saturday that one essay declared the whole agent memory category a mistake, one of its most popular open-source projects shipped a feature that half agrees with it.

On October 3, Kevin Liao published "Agents don't need memory, they need documentation". His line was blunt: "The entire memory plugin ecosystem is solving the wrong problem." The same day, claude-mem, the Apache 2.0 memory plugin for Claude Code with about 96,000 stars on GitHub, released v13.29.0. Its headline feature has nothing to do with better recall. It is a to-do list the agent writes on purpose, keyed and append-only, that loads every time a session starts.

That release draws a line most memory tools blur. On one side sits state an agent records deliberately and you can read back. On the other sit observations it retrieves because they look similar to the current prompt. Those two deserve very different levels of trust, and claude-mem now ships both in one box.

What claude-mem was before this release

The README describes a capture-and-replay system. Five lifecycle hooks (SessionStart, UserPromptSubmit, PostToolUse, Stop and SessionEnd) watch what the agent does. An observer model turns tool activity into compressed observations. Those land in SQLite with full-text search, plus a Chroma vector store for semantic search.

Retrieval runs through MCP tools in three layers. A search call returns a compact index of IDs, a timeline call adds chronological context around hits, and get_observations pulls full detail only for the IDs you pick. The project claims this saves about ten times the tokens of dumping everything, which is its own figure. You can wrap sensitive text in <private> tags to keep it out of storage.

The observer is a model call, so it costs something. The README lists a free claude-mem observer with a 14-day trial, OpenRouter, Gemini, or your Anthropic plan, and notes that memory "automatically falls back to your Anthropic plan unless you subscribe." Version 13.29.0 widens that with an OpenAI-compatible provider (presets include Ollama, LM Studio and vLLM) and a Codex subscription provider.

All of that is recall. It answers the question "what have I seen before that looks like this?"

What 13.29.0 adds

The new tools are work_state_write and work_state_read. Per the changelog, a write "appends one entry to a named list," with fields capped at 2,000 characters as JSON. An entry with a task field is a to-do item whose status is todo, doing, done or dropped. "The latest value of each key wins," and null clears a key.

At session start, the context opens with the work-state rules and every open item, capped at 3,000 characters. The data lives in an append-only table in claude-mem's database rather than in git, which the changelog says avoids merge conflicts.

Look at what this is not. It is not ranked by similarity. It is not summarized by an observer model. The agent decides to write it, it lands under a key, and the newest value replaces the old one. If you ask what the agent thinks is still open, there is exactly one answer, and you can read it.

That answers a different question: "what did I commit to, and where did I leave it?"

Why the split matters

Liao's essay lists five failures of retrieval memory. Two of them explain why work state is a different animal.

The first: "Memories are surfaced by similarity. Similarity search ranks how close two snippets are in embedding space. That's it." A vector match tells you two pieces of text look alike. It does not tell you which one is current. An observation from three weeks ago saying "we use the old auth flow" can rank above yesterday's decision to replace it, because it shares more words with today's prompt.

The second: "The store is unauditable." Thousands of embeddings in a local database are hard to inspect, so a wrong memory can sit there indefinitely, resurfacing whenever the wording lines up.

Keyed work state sidesteps both. Last write wins, so a stale entry gets overwritten instead of competing. The open list is short and capped, so a human can read the whole thing at session start. A task marked dropped stays dropped. That is the property Liao wants from documentation, delivered inside a memory plugin.

Picture a migration that spans four sessions. With recall alone, session three depends on the observer having summarized the right moment in session two, and on the new prompt sharing enough words with that summary to pull it back. With work state, session two ends by writing migrate-auth: doing, users table done, sessions table next, and session three opens with that line in front of it. One is a guess about relevance. The other is a note.

The essay itself deserves one caveat. Liao maintains Operator Memory, a BSD 3-Clause tool that stores agent knowledge in plain folders with, in its README's words, "no embeddings, no vector database." His evidence is more than a year of his own use, not a benchmark. Read the argument as a sharp position from a competitor, which is also why it is useful. The claude-mem release shows that a project with every reason to defend recall still found a job recall could not do.

Put this into practice

If you already run claude-mem, the upgrade path is short. If you do not, the install is npx claude-mem install on Node 20 or later, with Claude Code already set up.

  1. Give work state the job recall was doing badly. Open tasks, half-finished refactors and "do not touch this file until the migration lands" belong in work_state_write, not in the hope that an observation resurfaces.
  2. Read the open list at session start. It is capped at 3,000 characters for a reason. Treat it like a standup note. If something there is wrong, overwrite the key or mark it dropped.
  3. Treat recalled observations as hints. When the agent cites something it found through search, ask where it came from and how old it is. The timeline layer exists to answer exactly that.
  4. Turn on secret redaction. Version 13.29.0 adds opt-in redaction with 11 built-in patterns, including AWS, GitHub, OpenAI and Anthropic keys, JWTs and PEM private keys. It is off unless you enable it.
  5. Check the upgrade notes before you upgrade. CLAUDE_MEM_CONTEXT_MAIN_AGENT_ONLY now defaults to true, so subagent observations no longer appear at session start. The UserPromptSubmit hook timeout dropped from 60 to 15 seconds. The worker now refuses unknown hostnames as a DNS-rebinding guard, so a custom setup may need CLAUDE_MEM_ALLOWED_ORIGINS.
  6. Keep durable decisions in the repo. Work state lives in a local database. A design decision your teammates need still belongs in a file under version control.

Honest limitations

Work state is local. The changelog chose a database table over git to avoid merge conflicts, and that choice also means your teammates cannot review the agent's to-do list in a pull request. For a solo developer that is fine. For a team, the ledger is invisible in code review, because it never lands in the repository.

The 3,000-character cap at session start is a real ceiling. Long backlogs get cut, and the changelog does not say which open items lose out when the list overflows.

Nothing stops the agent from writing a wrong entry. Keyed state is easier to audit than embeddings, but "easier to audit" only helps if someone audits it. An agent that marks a task done when it is not has created a clean, authoritative, wrong record.

Recall has not gone away, and neither have its failure modes. The release also adds opt-in near-duplicate deduplication and ACT-R ranking for observations, which are improvements to recall, not replacements for it. If your agent's trouble is acting on stale context, work state helps only with the context you move into it.

And the observer still costs tokens. Every captured observation runs through a model, whether that is the trial observer, a paid provider, or your own Anthropic plan by fallback.

Decide what your agent is allowed to remember

The useful takeaway from this release has little to do with claude-mem specifically. Any memory system you give an agent holds two kinds of things: commitments it wrote down and associations it noticed. They fail differently. Commitments go stale when nobody updates them. Associations go stale while still looking relevant.

claude-mem now lets you keep them apart. Use that. Put what must be true into keyed state you read every morning, let recall suggest the rest, and never let a similarity score settle a question that a written record should answer.

Sources: claude-mem changelog, claude-mem README, Kevin Liao, Operator Memory.


Medium metadata

  • Title: claude-mem 13.29.0 Adds the One Kind of Agent Memory You Can Audit
  • Subtitle: The memory plugin for Claude Code just shipped a keyed to-do ledger next to its vector recall. The split tells you which memories to trust.
  • Tags: Claude Code, AI Agents, Software Development, Open Source, Developer Tools
  • Reading time: about 8 minutes
  • Canonical: import from the fervorai.dev URL