Go-Trader · Issue #1159 · Bake-off Episode 9

Twenty-five entries.
Four judges.
One winner?

The teaser

This episode was finished. Then four more entries showed up. Twenty-five tries at the same dangerous job inside my live trading bot: teach it to hedge its own bets. Four entries landed after grading had already closed, from two models that had never competed here, so the old scores went in a drawer and everything got re-graded from scratch: every entry renamed to a blind label, read by four AI judges, scored to the same rubric. No coding background needed; I'll explain everything as we go.

The result is at the bottom. The re-grade did more than pick a winner.

scroll ↓
i.The job

Teach a trading bot to hedge.

In plain English: my bot buys and sells on its own. This episode's task was to teach it to hedge. When the bot bets on one coin going up, it should automatically place a smaller opposite bet on a closely related coin. The two coins tend to move together, so if the whole market dives, the second bet cushions the first. The hedge has to open when the main bet opens, grow when it grows, shrink when it shrinks, and close when it closes, including the emergency cases where the bot slams everything shut.

My own write-up rates this 90 out of 100 in difficulty. It touches live orders, shared account balances, and the emergency shutoffs. Get it wrong and the "protection" becomes a second way to lose money. Twenty-five entries tried it. Four of them were not supposed to be here at all.

ii.The field, and the four that came late

Twenty-five entries. Nine models. Four labs.

Each entry is a "PR," short for pull request: a proposed batch of code changes. Most models entered several times at different effort settings, the dial for how much thinking time a model gets (low, medium, high, xhigh). Then the twist: four entries landed after grading had already closed. Three from Opus 5 and one from Kimi, two models that had never appeared in this series. Grafting late entries onto finished scores would mix grading conditions, so the old round went in a drawer and all twenty-five were graded fresh, together, under identical conditions. Nothing in this episode uses the old numbers.

Opus 5

Anthropic · debut · late arrival
3 entries · xhigh, high, medium

Kimi

Moonshot AI · debut · late arrival
1 entry · high

Fable 5

Anthropic
3 entries · high, medium, low

GPT-5.6 Sol

OpenAI
3 entries · xhigh, high, medium

GPT-5.6 Terra

OpenAI
2 entries · xhigh, high

GPT-5.6 Luna

OpenAI
2 entries · high, xhigh

Grok 4.5

xAI
3 entries · high, medium, low

Sonnet 5

Anthropic
4 entries · 2× xhigh, 2× high

Opus 4.8

Anthropic
4 entries · 2× xhigh, 2× high

A finished contest plus four uninvited entries is a problem with exactly one honest fix: grade everyone again, blind, at the same time. So that's what happened. How, exactly, matters more this episode than any other.

iii.How it was judged

Four judges. Blind labels. One round.

In plain English: four AI judges (Fable 5, GPT-5.6 Sol, Grok 4.5, and Opus 5), every one at high effort, in a single round. Every entry was renamed to a blind label, A through Y, so no judge knew whose code it was reading. Each judge read the actual code of all twenty-five entries and scored it to a fixed rubric. The published score for each entry is the average of its four scores. Simple. Two things about that average are anything but simple, and you should read both before you trust a single number below.

⚖️ Honesty block 1 of 2: all four judges are also contestants Fable 5, Sol, Grok 4.5, and Opus 5 all have entries in the field they graded. Twelve of the hundred grades below are a judge scoring its own model's work. Those grades stay in at full weight, marked with a star, and a no-self column beside the average strips them back out. Any fairness claim on this page quotes that column. This time it changed the podium; the how and who comes after the reveal.
⚖️ Honesty block 2 of 2: nobody checked the judges No verification pass ran this round. Every score is exactly what the judge submitted, averaged. What you are reading is four opinions, never audited, never fact-checked against the code by a fifth party. Where the judges split wildly, the table says so instead of pretending the number is precise.
⚑ The low-confidence flags Three entries carry a low-confidence flag panel split: their four blind scores were 34.5 to 41 points apart. One judge read a near-perfect entry; another read a broken one; it was the same code. Read a flagged score as a shrug, never as a measurement.
iv.The full table

One hundred grades, nothing hidden.

Every judge's score for every entry, as submitted. Star = a judge grading its own model's entry. Amber rows = the low-confidence flags. The no-self column is the average with self-scores removed: fifteen of the twenty-five entries shift at least one place between the two columns. The biggest movers: Sol's high entry climbs five places, Fable's low entry drops five, Grok's high entry drops three. The rest of this page counts the ranking down in order, so the table sits behind a click.

The full table · spoilers insideall 25 rows, all four judges, ranked. Skip it to keep the countdown.
RankPRModel (effort)FableSolGrokOpus 5MeanNo-selfFlag
1#1404Opus 5 xhigh969710095★97.0097.67
2#1333Fable 5 high97★92.5948892.8891.50
3#1405Kimi high10073979190.2590.25
4#1367GPT-5.6 Sol xhigh94.575.5★97.591.589.7594.50
5#1406Opus 5 high9574.510088★89.3889.83
6#1364GPT-5.6 Terra xhigh9381958588.5088.50
7#1407Opus 5 medium95.563.510087.5★86.6386.33panel split
8#1332GPT-5.6 Sol high93.569.5★97.585.586.5092.17
9#1371Grok 4.5 medium89.577.592.5★83.585.7583.50
10#1363GPT-5.6 Luna high89.570.584.586.582.7582.75
11#1378Sonnet 5 high807492.583.582.5082.50
12#1361Opus 4.8 xhigh92.57286.577.582.1382.13
13#1376Fable 5 medium87.5★73867580.3878.00
14#1331Grok 4.5 high935890★7879.7576.33
15#1375Fable 5 low89★52937878.0074.33panel split
16#1337Opus 4.8 xhigh87599372.577.8877.88
17#1379Sonnet 5 xhigh8759.586.57777.5077.50
18#1380Sonnet 5 xhigh86.559.580.581.577.0077.00
19#1336Opus 4.8 high85568673.575.1375.13
20#1373Grok 4.5 low83.562.573★7072.2572.00
21#1365GPT-5.6 Sol medium81.561★87.555.571.3874.83
22#1360Opus 4.8 high8855.58553.570.5070.50panel split
23#1377Sonnet 5 high885765.561.568.0068.00
24#1366GPT-5.6 Luna xhigh64.557.5865265.0065.00
25#1362GPT-5.6 Terra high45.536.533.534.537.5037.50
★ = self-score (a judge grading its own model's entry, kept at full weight). Mean = average of the four scores as submitted; no-self = the average with self-scores removed. Amber rows carry the panel-split flag: four-judge ranges of 34.5 to 41 points, so treat those means as rough. Fifteen of the twenty-five ranks shift between the two columns.
v.The overachievers

Two underdogs punched above their weight.

The quiet story in the table is two entries from GPT-5.6's lesser known siblings. Terra at xhigh scored 88.50 and took sixth, ahead of Sol's high entry and every Opus 4.8, Sonnet 5 and Grok entry in the field. Luna at plain high scored 82.75 and took tenth, ahead of both Opus 4.8 entries running at xhigh. Neither model came in with a reputation, and both out-placed bigger names running harder settings.

The catch: neither result survives a different effort setting, and they break in opposite directions. Terra needed the extra effort; its high entry is the 37.50 dead-last card whose hedge can never switch on. Luna choked on it; its xhigh entry fell to 65.00, second from the bottom. Same family, same dial, opposite reactions. If this table hands you one practical rule, it is this: test your model at more than one effort level before you trust it.

The reckoning

So who
actually won?

Twenty-five entries, one hundred blind grades, averaged. The published order below uses the mean of four as submitted; the fairness column waits on the other side of the reveal. First, the field in one breath.

Counting down, from twenty-fifth place…

25
GPT-5.6 Terra high
37.50
24
GPT-5.6 Luna xhigh
65.00
23
Sonnet 5 high
68.00
22
Opus 4.8 high panel split · 34.5 pts
70.50
21
GPT-5.6 Sol medium
71.38
20
Grok 4.5 low
72.25
19
Opus 4.8 high
75.13
18
Sonnet 5 xhigh
77.00
17
Sonnet 5 xhigh
77.50
16
Opus 4.8 xhigh
77.88
15
Fable 5 low panel split · 41 pts
78.00
14
Grok 4.5 high
79.75
13
Fable 5 medium
80.38
12
Opus 4.8 xhigh
82.13
11
Sonnet 5 high
82.50
10
GPT-5.6 Luna high
82.75
9
Grok 4.5 medium
85.75
8
GPT-5.6 Sol high
86.50
7
Opus 5 medium panel split · 36.5 pts
86.63
6
GPT-5.6 Terra xhigh
88.50
5
Opus 5 high
89.38

Last place earned it the hard way: all four judges independently found the same defect in #1362. Its hedge machinery exists, but nothing in the live bot ever calls it, so the safety feature can never switch on. That is worse than skipping the job, because it looks done. The top four get their own screens.

· Fourth place ·
GPT-5.6 Solxhigh · PR #1367 Three judges scored this entry 94.5, 97.5, and 91.5. The fourth score, a 75.5, came from Sol itself, the panel's harshest grader anywhere. Strip the self-scores and this entry's average jumps to 94.50. Hold that thought for the other side of the reveal.
89.75/100
· Third place ·
Kimihigh · PR #1405 · debut One of the four late arrivals, on its first ever entry in this series, straight onto the podium. Also the panel's loudest argument: one judge gave it a 100, another a 73. Same code. Its podium spot survives the fairness column untouched, because Kimi wasn't judging anyone.
90.25/100
· Second place ·
Fable 5high · PR #1333 Last episode's champion, one step short of repeating. One asterisk, literally: the highest of its four scores, a 97, came from its own judge. Remove the self-scores and it sits at 91.50. Whether that 97 was favoritism gets answered after the reveal, and the answer surprised me.
92.88/100
Before first place: three things that shouldn't both be true
📏

Biggest build vs. smallest build

5,507 vs. 465 lines

The largest rewrite in the field, 5,507 lines added across 27 files, is still ahead of you on this page. The smallest, 465 lines across 14 files, is already behind you: last place, and its hedge can never actually run. Size is never scored. This episode it still tells the story.

🪙

The nine-point coin flip

Same code, 9 points apart

Two of the twenty-five entries are byte-for-byte identical code under different PR numbers. One judge scored that identical code twice: 86.5 and 77.5. Nine points is this table's noise floor. Two mid-table entries a few points apart are effectively tied.

🏅

The podium that moves

15 of 25 shift

Remove the twelve starred self-scores and fifteen of the twenty-five entries change rank, including two of the three podium spots. Who moves, which judge moved them, and why it isn't the bias you'd guess: after the reveal.

🏆
· First place ·
The late arrival took it.

Four entries crashed a finished contest and forced a re-grade. One of them won the whole thing, on its first attempt in this series.

Opus 5
Anthropic · xhigh · series debut
97.00 /100

The panel's calmest verdict. Its four blind scores (96, 97, 100, 95) sit within five points of each other, the tightest agreement in the field (SD 1.87), and it was the only entry every judge put in its own personal top three. It wins by 4.12 points, and by more (97.67) when self-scores come out. It also shipped the largest change in the field: +5,507 / −50 across 27 files.

First under the published average. First under the no-self average. First under every alternative way I sliced the scores. On a table where the judges split by up to 41 points elsewhere, they agreed hardest about this one.

The re-grade exists because this entry and three others arrived late. Redoing the whole round blind was the only honest way to let them in, and the newcomer beat the reigning champion by 4.12 points.

Now the promised part

The podium moves when the self-scores leave.

Under the published average, the podium is Opus 5, Fable 5, Kimi. Throw out every judge's score of its own work and it becomes Opus 5, then Sol's xhigh entry (94.50), then Sol's high entry (92.17). Fifteen of the twenty-five entries shift at least one place between the two columns; the big movers are Sol's high entry (up five places), Fable's low entry (down five), and Grok's high entry (down three).

And the mechanism is the opposite of what you'd guess. Nobody inflated their own grade. Fable 5 gave its own entry a 97, the highest score that entry got from anyone, but Fable graded the whole field generously, 7.91 points above the panel average, and that habit reached its own paper too. Sol graded its own two best entries at 75.5 and 69.5, two of its most brutal verdicts anywhere; Sol is the panel's harshest judge, 12.67 points below the panel average, and it never once scored any entry above the group's mean, its own included. Its harshness hit its own work hardest and knocked both entries off the podium.

I ran the check across all four judges: none shows a systematic bias toward its own entries. Fable is slightly stingier with its own work than its lenient norm, Sol even harsher on its own than usual, Grok less generous on its own, and Opus 5 close to neutral. The flip comes from judges with strong personal grading scales sitting inside the race they're grading. The winner is the one entry the flip can't touch: first in both columns.

Summary

The top eight, and the floor.

Anthropic · xhigh · debut
Opus 5
Winner
Mean of 4
97.00 /100
No-self
97.67
Verdict
Only entry in every judge's top three · tightest agreement (SD 1.87) · +5,507 / −50 across 27 files, the field's largest
Anthropic · high
Fable 5
2nd
Mean of 4
92.88 /100
No-self
91.50
Verdict
Last episode's champion · its own judge's 97 was its highest grade · drops to 3rd in the no-self column
Moonshot AI · high · debut
Kimi
3rd
Mean of 4
90.25 /100
No-self
90.25
Verdict
Podium on its first entry ever · the panel's widest single argument: a 100 from one judge, a 73 from another
OpenAI · xhigh
GPT-5.6 Sol
4th
Mean of 4
89.75 /100
No-self
94.50
Verdict
2nd in the no-self column · kept off the published podium by its own 75.5 self-grade, one of Sol's harshest anywhere
Anthropic · high · debut
Opus 5
5th
Mean of 4
89.38 /100
No-self
89.83
Verdict
Two perfect-adjacent grades (95, 100) against Sol's 74.5 · the debut model's second entry in the top five
OpenAI · xhigh
GPT-5.6 Terra
6th
Mean of 4
88.50 /100
No-self
88.50
Verdict
No self-scores, no drama: identical in both columns · the same model's other entry finished 25th
Anthropic · medium · debut
Opus 5 panel split
7th
Mean of 4
86.63 /100
No-self
86.33
Verdict
Low-confidence: judges split 36.5 points (a 100 and a 63.5 on the same code) · read the mean as rough
OpenAI · high
GPT-5.6 Sol
8th
Mean of 4
86.50 /100
No-self
92.17
Verdict
The no-self column's biggest climber: up five places to 3rd once Sol's own 69.5 comes out
OpenAI · high
GPT-5.6 Terra
25th
Mean of 4
37.50 /100
No-self
37.50
Verdict
All four judges found the same defect: hedge code the live bot never calls · +465 / −14 across 14 files, the field's smallest
Epilogue: what shipped

Nothing, yet.

The hedge issue is still open and none of the twenty-five entries has merged. When one does, or when I blend the best of them, that's a future episode's story. For now the verdict stands on four blind opinions, averaged, with every star and every split disclosed above.