
The gate that accepted 200 without reading the content
A health check that only looked at the HTTP status let truncated and empty responses through for days. The probe that came after measures the body, not the envelope.
20 posts

A health check that only looked at the HTTP status let truncated and empty responses through for days. The probe that came after measures the body, not the envelope.

I lost 810 training steps because a 2 GB checkpoint was copied straight onto its final name and the machine rebooted mid-copy. I fixed the write on the primary volume. Two days later the same bug was living in the second-drive mirror — and it had gotten worse: the truncated file passed the resume check.

A catatonic BitMamba 1B — broken tokenizer, 10^18 gradients, a midnight ENOMEM and a CUDA wedge that faked a dead GPU — closed 500 steps and the validation reproduced the eval loss to the fourth decimal. The diary of Phase 4B on a 12 GB RTX 3060.

I resumed a run from the last checkpoint and it looked healthy — steps advancing, gradient in range, zero errors. Except the learning rate had quietly shrunk by three orders of magnitude. The cause: the LR scheduler restores more than it should when loading old state, and the fix was two lines.

The Phase 4 run dropped mid-flight with no traceback, no OOM, no error log: the process simply stopped existing. Chasing the cause turned into a lesson about watching the wrong metrics — and about the invisible cost of letting your GPU driver sweep up your garbage.

300 training steps, a loss stuck at 10.8125, and a model that insisted on coming out blank. The cause wasn't a low learning rate — it was exploded weights in the starting checkpoint. The story of how a stalled run became a full diagnosis.

The codec post ended with a dangling issue: the lm_head diverging in validation. Two weeks later, a methodical audit rescaled the model's head, generated a clean checkpoint, and became the foundation of Phase 4 — the Resurrection.

A repo that accumulates build artifacts and syncs doesn't tell the clean story you want to show. TatuEngine went through versioning hygiene: build artifacts out of git, auto-sync running, and a repository ready for outsiders to look at without handing over the internal recipe.

The BitMamba-2 1B's 1:1 BVH took 27.26GB — unworkable. A compact ternary block codec delivered 114× compression, with the lm_head diverging mid-validation.

TatuEngine had 5.7GB of AI models committed to git since day one. GitHub rejects objects above 2GB with a 422 error — that push would never succeed. The fix wasn't 'pushing through': it was an orphan branch that preserved code, tests, and docs and rewrote the repository's history from scratch.

TatuEngine got a punishment layer v3: progressive fines for repeat offenses, partial cold restart, automatic tool denial and output reprimand. The idea: instead of only rewarding good behavior, punish bad behavior with growing intensity — like an immune system that learns from every mistake.

The autopoietic agent with ToolUse, an MCP server and a hybrid sandbox is the biggest attack surface in my ecosystem. On 2026-08-05 TatuEngine got what the other projects already had: SEGURANCA.md v1.0, a 24h watchdog, and a rule that changes how the project is developed.

After learning to infer (252× GPU speedup), TatuEngine went for the Master-Apprentice cycle: a Qwen2.5-3B teacher with LoRA distilling structured reasoning ([THOUGHT]/[ANSWER]) into the BitMamba-2 1B student. Six versions in one day (v0.17→v0.22.2), 381 examples, and the convergence that wouldn't come until the system message fix.

Born from the absurd question 'what if we used RT Cores to process neural networks?', killed the dream on WSL2, and became a hybrid CPU/GPU engine with BitMamba-2 1B, radical compression, an agent with 8 subsystems and a security sandbox — 302 commits later.

When your engine runs AI-generated code, you need a sandbox that doesn't hinder development but won't let path traversal slip through.

298 commits, zero new features in 30 days, and an SSM inference engine that runs unsupervised. The TatuEngine doesn't need changes — it just works.

Exploring the current state of TatuEngine — lessons, challenges, and discoveries.

The saga of making Student BitMamba-1B learn without exploding: Adafactor vs AdamW, saving 4 GB of VRAM, and the 4 bugs that produced NaN during bf16 training.

The journey (and the craters) of training a Mamba-2 with 1.4B parameters from scratch on an RTX 3060 — warmup, SFT, catastrophic eval, and Phase 4 of the resurrection.

Exploring the beginnings of TatuEngine — lessons, challenges, and discoveries.