luxifare self-improving loop run 41

Genome 41 is +22 Elo overall and −24 in closed positions.

The gauntlet said the new genome was better and the aggregate was never wrong. It was just incomplete. Split the same 9,600 games by position class and six of seven classes improve, one regresses, and the regression is large enough that its confidence interval never touches zero. This is the loop catching itself.

per-class breakdownordered worst first

Position class
Games
Score %
Elo Δ
Draws %
Length
Closed regressed
1,320
46.6
−24±13
52
78.2
Rook endgames
1,540
52.4
+17±12
46
88.3
Pawn storms
720
53.1
+22±18
24
58.9
Semi-open
2,140
53.9
+27±10
34
64.8
Open
1,860
54.8
+33±11
31
61.4
Minor-pieceendgames
1,040
55.0
+35±14
39
81.6
King attacks
980
56.2
+43±15
19
52.7

◂ swipe the grid ▸ fill = share of that column’s own maximum score and Elo are genome 41’s, from 38’s side of the pairing length in moves

Net over all 9,600 games 53.1% score +21.6 Elo ±4.1
Classes improved
6of 7
Every class except closed gained, by between 17 and 43 Elo.
Closed positions
−24Elo
46.6% over 1,320 games. The ±13 interval does not reach zero, so this is not noise.
Closed share of book
13.8%
Small enough that the loss is worth 3.3 Elo of the aggregate, and invisible inside it.
Params changed
7of 125
Four concepts touched. Two of the seven plausibly explain the closed-position loss.

Why the aggregate hid it

contribution to the +21.6

Each class contributes its Elo delta weighted by its share of the book. The seven contributions sum to the aggregate, which is exactly the problem: one negative term of −3.3 sits inside a total of +21.6 and never surfaces.

Closed −3.3
Pawn storms +1.6
Rook endgames +2.7
Minor-piece endgames +3.8
King attacks +4.4
Semi-open +6.0
Open +6.4
Sum weighted by share of book +21.6 Elo

What moved in the genome

7 of 125 params

Genome 41 came from an LLM-proposed edit to four concepts, then 900 iterations of numeric tuning on the surviving parameters. Two of the seven changes act almost only when the pawn structure is locked, which is the definition of the class that regressed.

pawns.locked_chain_penalty

prime suspect

Introduced at −18 cp, from nothing. The proposal was that a fixed chain restricts the side that owns it, and in open and semi-open games it does. In a genuinely closed position the chain is the asset, and the engine now pays to avoid building one.

Evidence: of the 1,320 closed games, 812 reached a position where this term fired for more than 30 consecutive plies, and genome 41 scored 43.1% in that subset against 52.2% in the rest.

mobility.knight_blocked

second suspect

Deepened from −5 to −9 cp. Blocked knights are worth less in most structures and worth more in closed ones, where they are the only piece that travels. The tuner had no closed-position weight in its objective, so nothing pushed back.

All seven parameter changes, genome 38 → 41
Parameter Concept G38 G41 Origin
pawns.locked_chain_penaltyPawn structure0−18LLM proposal
mobility.knight_blockedMobility−5−9tuner
mobility.bishop_open_diagMobility+21+35LLM proposal
king.shelter_storm_scaleKing safety1.001.12tuner
threats.hanging_minorThreats+28+39tuner
space.behind_pawn_bonusSpace+4+10tuner
endgame.rook_activityEndgame+52+59tuner

What the loop does with this

Genome 41 still ships. A net +21.6 Elo with a known, bounded, single-class regression is a better position to be in than a net +21.6 Elo with an unknown one, and the whole reason the evaluation is 125 named parameters rather than a weight matrix is that a finding like this can be pointed at a line.

The next generation inherits three constraints: closed positions are lifted to a fixed 25% of the book so the tuner cannot ignore them again, pawns.locked_chain_penalty is pinned as a candidate for reversion, and every future gauntlet reports per class before it reports an aggregate. The aggregate stops being the acceptance test.