
Yurumi: the memory that had to forget
Four memory layers with different lifetimes, a TTL that never actually expired, and the decision that fixed it: sleeping is not deleting. How the memory lifecycle stopped being a loose field in the payload.
8 posts

Four memory layers with different lifetimes, a TTL that never actually expired, and the decision that fixed it: sleeping is not deleting. How the memory lifecycle stopped being a loose field in the payload.

Capivara's chat started answering without sources: HTTP 200, a polished sentence, zero memory. Restarting the embedding service looked like the obvious fix — and was precisely what kept it from coming back. Underneath, two defects from the same family: a process that only opens its port after waking up, and a fallback that doesn't know how to undo itself.

The embedder was not broken, it was loading. A one-way latch turned a cold load of minutes into a system that stayed broken until the next restart.

Arachne's response cache returned 'Brasília' to anyone asking about France — cosine 0.8531 against a 0.85 threshold. The story of a bug born in a CI investigation, the math that hid it, and the entity guard that fixed it without sacrificing recall.

Yurumi started as a local replacement for mem0 — storing conversation memories. In two weeks it gained intent, actions, learning, its own identity and a knowledge graph. The journey from a simple store to a complete agentic system.

Capivara kept its memories in a local ChromaDB, with its own embedder and L1-L4 layers. Yurumi kept the ecosystem's memory in a Qdrant. Two separate memories meant a brain that never talked to the rest of the body. The unification on 08/28 migrated 182/182 memories without losing a single one: hash dedup, backup before deleting, and an --apply gate that separates the plan from the execution.

Every project's AGENTS.md was the official instruction set of the ecosystem, but it lived outside the Yurumi's memory — the ingest only ran when someone remembered to run it. The fix was a signature watcher (mtime+size) that re-ingests only what changed. Along the way: a first_run that ran a full ingest every day, an os.stat that hung when WSL wedged, and the lesson of never trusting a junction.

Yurumi stored its memory with a single embedding model trained in Portuguese. It worked — until the first question in English. The fix was not swapping the model; it was running two of them side by side and fusing the rankings with RRF. After a full LoCoMo run, recall went from 0.139 to 0.1614.