
Arachne Knowledge OS: The Knowledge Operating System that united RAG, LLM and Obsidian
How Arachne evolved from a scraper into a personal knowledge platform with zero‑port infra, hybrid search and bot generation.
8 posts

How Arachne evolved from a scraper into a personal knowledge platform with zero‑port infra, hybrid search and bot generation.

My portfolio's next feature isn't live yet: an AI-powered generator that rewrites my résumé around a job description. Before it becomes a button, the hard part was already solved — and it isn't generating text. It's keeping the AI from describing experiences that never happened.

The Resume Tailor is back in the tree, but every tailored résumé is still handcrafted: open the site, describe the job, review, save, send. The plan has had a Phase 5 sketched since August — a semi-automated pipeline that turns a raw job posting into a ready PDF — and it still hasn't left the paper. On the difference between shipping a tool and shipping a habit.

An LLM call ends in 200 OK with empty content and the pipeline moves on like nothing happened. Diagnosis: reasoning models burn the token budget thinking and nothing is left for the answer. The fixes: a token ceiling with real margin, and a parser that tolerates a data: [DONE] glued to the JSON.

The new stage in my transcription pipeline uses Gemini TEXT over 120s windows — no audio, cheap and fast. What keeps the model from making things up isn't a polite prompt: it's a numeric ±35% character invariant per window, with retry and flag when the math doesn't close.

300 training steps, a loss stuck at 10.8125, and a model that insisted on coming out blank. The cause wasn't a low learning rate — it was exploded weights in the starting checkpoint. The story of how a stalled run became a full diagnosis.

I actually read the BitNet b1.58 paper and got stuck on a math problem that refused to add up: how do three values (-1, 0, +1) fit in 1.58 bits? A study week that became a didactic clone of ternary quantization — and wrecked my mental model of 'precision equals quality'.

Hermes kept entering a continuation loop 4 times before giving up. The cause: Ollama, by default, only allocates 4096 context tokens — regardless of the model supporting 262K. The hard way: num_ctx of the model ≠ num_ctx of the inference slot.