8,343lines of Python
4dependencies
240exam questions across 8 domains
5ablation flags
A 365-day simulation of three agents designed to generate fine-tuning data. Bounded three-tier memory with utility-scored eviction, hand-rolled TF-IDF retrieval, an adversarial axiom pipeline, a 240-question exam with a raw-model control.
Designed and dry-run end to end. The full year has not been executed yet, and the page says so.