luxifare self-improving loop run 41
The gauntlet said the new genome was better and the aggregate was never wrong. It was just incomplete. Split the same 9,600 games by position class and six of seven classes improve, one regresses, and the regression is large enough that its confidence interval never touches zero. This is the loop catching itself.
per-class breakdownordered worst first
◂ swipe the grid ▸ fill = share of that column’s own maximum score and Elo are genome 41’s, from 38’s side of the pairing length in moves
Each class contributes its Elo delta weighted by its share of the book. The seven contributions sum to the aggregate, which is exactly the problem: one negative term of −3.3 sits inside a total of +21.6 and never surfaces.
Genome 41 came from an LLM-proposed edit to four concepts, then 900 iterations of numeric tuning on the surviving parameters. Two of the seven changes act almost only when the pawn structure is locked, which is the definition of the class that regressed.
pawns.locked_chain_penaltyIntroduced at −18 cp, from nothing. The proposal was that a fixed chain restricts the side that owns it, and in open and semi-open games it does. In a genuinely closed position the chain is the asset, and the engine now pays to avoid building one.
Evidence: of the 1,320 closed games, 812 reached a position where this term fired for more than 30 consecutive plies, and genome 41 scored 43.1% in that subset against 52.2% in the rest.
mobility.knight_blockedDeepened from −5 to −9 cp. Blocked knights are worth less in most structures and worth more in closed ones, where they are the only piece that travels. The tuner had no closed-position weight in its objective, so nothing pushed back.
| Parameter | Concept | G38 | G41 | Origin |
|---|---|---|---|---|
pawns.locked_chain_penalty | Pawn structure | 0 | −18 | LLM proposal |
mobility.knight_blocked | Mobility | −5 | −9 | tuner |
mobility.bishop_open_diag | Mobility | +21 | +35 | LLM proposal |
king.shelter_storm_scale | King safety | 1.00 | 1.12 | tuner |
threats.hanging_minor | Threats | +28 | +39 | tuner |
space.behind_pawn_bonus | Space | +4 | +10 | tuner |
endgame.rook_activity | Endgame | +52 | +59 | tuner |
Genome 41 still ships. A net +21.6 Elo with a known, bounded, single-class regression is a better position to be in than a net +21.6 Elo with an unknown one, and the whole reason the evaluation is 125 named parameters rather than a weight matrix is that a finding like this can be pointed at a line.
The next generation inherits three constraints: closed positions are lifted to a
fixed 25% of the book so the tuner cannot ignore them again,
pawns.locked_chain_penalty is pinned as a candidate for reversion,
and every future gauntlet reports per class before it reports an aggregate. The
aggregate stops being the acceptance test.