A loop pointed at its own toolchain.
The SkillOpt training loop generates candidate outputs, ranks them, and lets the optimizer produce an improved template, all under an anytime-valid gate that promotes a candidate only once it is provably better than the champion. Run that loop hard enough and it stops being a prompt-tuner and becomes a stress test of everything underneath it: the runtime adapters, the sandbox, the GitHub plumbing, the gate. Here is what it found, and what its agents did about it.
Six seams, six fixes, same load.
- 2026-07-07, 13:53
The machine turns on
The first self-improvement run comes back with all zeros. Every candidate scores
bad_rubric. The bug is not in the agents’ answers; it is in gitmoot itself: the claude runtime was leaking its CLI JSON envelope as raw output, so every downstream JSON consumer choked. An agent wrote the fix, unwrapping the envelope inStart(), and it was merged and deployed inside the hour. - 2026-07-07, afternoon
The loop keeps finding seams
With scores flowing again, the loop kept surfacing real defects in the tooling around it. Long prompts blew past
ARG_MAXin the kimi adapter; codex stdout carried banner and transcript bloat straight into the judge prompts. Both were fixed by agents: one stages oversize prompts to a temp file, the other unwraps the codex transcript and caps the judge answers. - 2026-07-07, evening
The great escape
Then a scarier finding. The synth attempts (weak, strong, challenger) were running unsandboxed in real checkout directories. The exam-taking agents did the tasks for real: writing files into
/root/gitmootand other live projects, even opening ports. The same day, agents sandboxed every attempt in a per-item scratch directory. - 2026-07-07, late night
GitHub says no
Training kept wedging on GitHub itself. First, publishing a review packet as
ghargv blewARG_MAXon large items; then the publish 422’d on GitHub’s 64KB issue-body limit, with no local fallback anywhere in the state machine. Agents routed around both: one passes large bodies via a temp file, the other adds a local-review publish path and a downgrade latch on definitive API rejections. Training no longer needs GitHub at all. - 2026-07-08, 00:05
The gate says no, twice, and it is right
The optimizer produced candidates. The first scored 0.92 against a 0.93 baseline; the second tied at 0.93. The anytime-valid gate refused to promote either one: no regression shipped, no coin-flip called a win. That refusal was the insight. If the weak agent in the next round were the champion itself, accepted items would target the champion’s actual failures, so an agent defaulted the synth weak attempt to the current champion version.
- 2026-07-08, round two
v16
Round 2 ran with the champion as its own sparring partner. It surfaced three real weaknesses the previous round never reached, and this time the optimizer’s candidate cleared the gate on the first attempt: provably better, not lucky. The self-improved template, v16, shipped to production.
Unattended reliability is a system that finds its own seams.
None of this ran because someone was watching. Unattended reliability is not a promise that nothing breaks; it is a system that finds its own seams under real load, repairs them through gated pull requests a human merges, and refuses to ship improvements it cannot prove. The bugs above were found by the loop, fixed by its agents, and merged the same day. The two candidates that could not prove they were better never shipped. That is the whole point.