← Use Cases

self-fixing agents

The week gitmoot’s agents fixed their own tools.

Over two days, gitmoot’s own self-improvement loop kept finding real bugs, not in the code it was asked to write, but in gitmoot itself. Each time, an agent shipped the fix through a gated pull request a human merged. This is that log, and every issue and PR in it is public.

✓ A true story. Every issue and PR below is public.
the setup

A loop pointed at its own toolchain.

The SkillOpt training loop generates candidate outputs, ranks them, and lets the optimizer produce an improved template, all under an anytime-valid gate that promotes a candidate only once it is provably better than the champion. Run that loop hard enough and it stops being a prompt-tuner and becomes a stress test of everything underneath it: the runtime adapters, the sandbox, the GitHub plumbing, the gate. Here is what it found, and what its agents did about it.

the log

Six seams, six fixes, same load.

  1. 2026-07-07, 13:53

    The machine turns on

    The first self-improvement run comes back with all zeros. Every candidate scores bad_rubric. The bug is not in the agents’ answers; it is in gitmoot itself: the claude runtime was leaking its CLI JSON envelope as raw output, so every downstream JSON consumer choked. An agent wrote the fix, unwrapping the envelope in Start(), and it was merged and deployed inside the hour.

  2. 2026-07-07, afternoon

    The loop keeps finding seams

    With scores flowing again, the loop kept surfacing real defects in the tooling around it. Long prompts blew past ARG_MAX in the kimi adapter; codex stdout carried banner and transcript bloat straight into the judge prompts. Both were fixed by agents: one stages oversize prompts to a temp file, the other unwraps the codex transcript and caps the judge answers.

  3. 2026-07-07, evening

    The great escape

    Then a scarier finding. The synth attempts (weak, strong, challenger) were running unsandboxed in real checkout directories. The exam-taking agents did the tasks for real: writing files into /root/gitmoot and other live projects, even opening ports. The same day, agents sandboxed every attempt in a per-item scratch directory.

  4. 2026-07-07, late night

    GitHub says no

    Training kept wedging on GitHub itself. First, publishing a review packet as gh argv blew ARG_MAX on large items; then the publish 422’d on GitHub’s 64KB issue-body limit, with no local fallback anywhere in the state machine. Agents routed around both: one passes large bodies via a temp file, the other adds a local-review publish path and a downgrade latch on definitive API rejections. Training no longer needs GitHub at all.

  5. 2026-07-08, 00:05

    The gate says no, twice, and it is right

    The optimizer produced candidates. The first scored 0.92 against a 0.93 baseline; the second tied at 0.93. The anytime-valid gate refused to promote either one: no regression shipped, no coin-flip called a win. That refusal was the insight. If the weak agent in the next round were the champion itself, accepted items would target the champion’s actual failures, so an agent defaulted the synth weak attempt to the current champion version.

  6. 2026-07-08, round two

    v16

    Round 2 ran with the champion as its own sparring partner. It surfaced three real weaknesses the previous round never reached, and this time the optimizer’s candidate cleared the gate on the first attempt: provably better, not lucky. The self-improved template, v16, shipped to production.

why this matters

Unattended reliability is a system that finds its own seams.

None of this ran because someone was watching. Unattended reliability is not a promise that nothing breaks; it is a system that finds its own seams under real load, repairs them through gated pull requests a human merges, and refuses to ship improvements it cannot prove. The bugs above were found by the loop, fixed by its agents, and merged the same day. The two candidates that could not prove they were better never shipped. That is the whole point.