
The restart button that was the problem
The embedder was not broken, it was loading. A one-way latch turned a cold load of minutes into a system that stayed broken until the next restart.
7 posts

The embedder was not broken, it was loading. A one-way latch turned a cold load of minutes into a system that stayed broken until the next restart.

Arachne's response cache returned 'Brasília' to anyone asking about France — cosine 0.8531 against a 0.85 threshold. The story of a bug born in a CI investigation, the math that hid it, and the entity guard that fixed it without sacrificing recall.

After the bilingual jump to 0.1614 recall@1, there was a number left on the bench: how much was the cross-encoder contributing on every search? The week I measured that latency and recall are a bill you pay on every query — and the answer ended up written as a permanent flag in the config.

A study about cron jobs: why they accumulate, fail silently, and how a cron that 'never sleeps' can take down an ecosystem. From Yurumi's consolidation to 289 orphan jobs — the lessons from a night of investigation.

Three embedding models, two search systems, one homegrown benchmark: the study that started with a question about semantic similarity and ended up revealing the tradeoffs between colibri, bge and openai — and how each one found its place in the ecosystem.

Yurumi stored its memory with a single embedding model trained in Portuguese. It worked — until the first question in English. The fix was not swapping the model; it was running two of them side by side and fusing the rankings with RRF. After a full LoCoMo run, recall went from 0.139 to 0.1614.

The simple question 'what does my agent actually remember?' turned into a full audit of Yurumi: 32 knowledge bases, 551 documents, a Qdrant collection with over 8 thousand vectors, and a LoCoMo baseline that handed back reality in percentage form.