
Arachne Knowledge OS: The Knowledge Operating System that united RAG, LLM and Obsidian
How Arachne evolved from a scraper into a personal knowledge platform with zero‑port infra, hybrid search and bot generation.
11 posts

How Arachne evolved from a scraper into a personal knowledge platform with zero‑port infra, hybrid search and bot generation.

Capivara's chat started answering without sources: HTTP 200, a polished sentence, zero memory. Restarting the embedding service looked like the obvious fix — and was precisely what kept it from coming back. Underneath, two defects from the same family: a process that only opens its port after waking up, and a fallback that doesn't know how to undo itself.

Before swapping Arachne's search engine for Yurumi, I had to prove the new one was at least as good as the old. Shadow dual-write, real traffic, overlap@10 — and the number nobody expects: 0.511.

After the bilingual jump to 0.1614 recall@1, there was a number left on the bench: how much was the cross-encoder contributing on every search? The week I measured that latency and recall are a bill you pay on every query — and the answer ended up written as a permanent flag in the config.

Three embedding models, two search systems, one homegrown benchmark: the study that started with a question about semantic similarity and ended up revealing the tradeoffs between colibri, bge and openai — and how each one found its place in the ecosystem.

Capivara kept its memories in a local ChromaDB, with its own embedder and L1-L4 layers. Yurumi kept the ecosystem's memory in a Qdrant. Two separate memories meant a brain that never talked to the rest of the body. The unification on 08/28 migrated 182/182 memories without losing a single one: hash dedup, backup before deleting, and an --apply gate that separates the plan from the execution.

Every project's AGENTS.md was the official instruction set of the ecosystem, but it lived outside the Yurumi's memory — the ingest only ran when someone remembered to run it. The fix was a signature watcher (mtime+size) that re-ingests only what changed. Along the way: a first_run that ran a full ingest every day, an os.stat that hung when WSL wedged, and the lesson of never trusting a junction.

Yurumi stored its memory with a single embedding model trained in Portuguese. It worked — until the first question in English. The fix was not swapping the model; it was running two of them side by side and fusing the rankings with RRF. After a full LoCoMo run, recall went from 0.139 to 0.1614.

The simple question 'what does my agent actually remember?' turned into a full audit of Yurumi: 32 knowledge bases, 551 documents, a Qdrant collection with over 8 thousand vectors, and a LoCoMo baseline that handed back reality in percentage form.

107 commits, 131 backend tests + 248 frontend, 2FA TOTP, local RAG with ChromaDB, R2 backup and D1 disaster recovery — the saga of turning a personal dashboard into reliable infrastructure.

777 commits, 2818 tests, hybrid RAG, ReAct Agent with 42 tools, MCP server, and the evolution from a simple scraper to a knowledge platform.