Teach a trading bot to hedge.
In plain English: my bot buys and sells on its own. This episode's task was to teach it to hedge. When the bot bets on one coin going up, it should automatically place a smaller opposite bet on a closely related coin. The two coins tend to move together, so if the whole market dives, the second bet cushions the first. The hedge has to open when the main bet opens, grow when it grows, shrink when it shrinks, and close when it closes, including the emergency cases where the bot slams everything shut.
My own write-up rates this 90 out of 100 in difficulty. It touches live orders, shared account balances, and the emergency shutoffs. Get it wrong and the "protection" becomes a second way to lose money. Twenty-five entries tried it. Four of them were not supposed to be here at all.
Twenty-five entries. Nine models. Four labs.
Each entry is a "PR," short for pull request: a proposed batch of code changes. Most models entered several times at different effort settings, the dial for how much thinking time a model gets (low, medium, high, xhigh). Then the twist: four entries landed after grading had already closed. Three from Opus 5 and one from Kimi, two models that had never appeared in this series. Grafting late entries onto finished scores would mix grading conditions, so the old round went in a drawer and all twenty-five were graded fresh, together, under identical conditions. Nothing in this episode uses the old numbers.
Opus 5
Kimi
Fable 5
GPT-5.6 Sol
GPT-5.6 Terra
GPT-5.6 Luna
Grok 4.5
Sonnet 5
Opus 4.8
A finished contest plus four uninvited entries is a problem with exactly one honest fix: grade everyone again, blind, at the same time. So that's what happened. How, exactly, matters more this episode than any other.
Four judges. Blind labels. One round.
In plain English: four AI judges (Fable 5, GPT-5.6 Sol, Grok 4.5, and Opus 5), every one at high effort, in a single round. Every entry was renamed to a blind label, A through Y, so no judge knew whose code it was reading. Each judge read the actual code of all twenty-five entries and scored it to a fixed rubric. The published score for each entry is the average of its four scores. Simple. Two things about that average are anything but simple, and you should read both before you trust a single number below.
One hundred grades, nothing hidden.
Every judge's score for every entry, as submitted. Star = a judge grading its own model's entry. Amber rows = the low-confidence flags. The no-self column is the average with self-scores removed: fifteen of the twenty-five entries shift at least one place between the two columns. The biggest movers: Sol's high entry climbs five places, Fable's low entry drops five, Grok's high entry drops three. The rest of this page counts the ranking down in order, so the table sits behind a click.
▸The full table · spoilers insideall 25 rows, all four judges, ranked. Skip it to keep the countdown.
| Rank | PR | Model (effort) | Fable | Sol | Grok | Opus 5 | Mean | No-self | Flag |
|---|---|---|---|---|---|---|---|---|---|
| 1 | #1404 | Opus 5 xhigh | 96 | 97 | 100 | 95★ | 97.00 | 97.67 | |
| 2 | #1333 | Fable 5 high | 97★ | 92.5 | 94 | 88 | 92.88 | 91.50 | |
| 3 | #1405 | Kimi high | 100 | 73 | 97 | 91 | 90.25 | 90.25 | |
| 4 | #1367 | GPT-5.6 Sol xhigh | 94.5 | 75.5★ | 97.5 | 91.5 | 89.75 | 94.50 | |
| 5 | #1406 | Opus 5 high | 95 | 74.5 | 100 | 88★ | 89.38 | 89.83 | |
| 6 | #1364 | GPT-5.6 Terra xhigh | 93 | 81 | 95 | 85 | 88.50 | 88.50 | |
| 7 | #1407 | Opus 5 medium | 95.5 | 63.5 | 100 | 87.5★ | 86.63 | 86.33 | panel split |
| 8 | #1332 | GPT-5.6 Sol high | 93.5 | 69.5★ | 97.5 | 85.5 | 86.50 | 92.17 | |
| 9 | #1371 | Grok 4.5 medium | 89.5 | 77.5 | 92.5★ | 83.5 | 85.75 | 83.50 | |
| 10 | #1363 | GPT-5.6 Luna high | 89.5 | 70.5 | 84.5 | 86.5 | 82.75 | 82.75 | |
| 11 | #1378 | Sonnet 5 high | 80 | 74 | 92.5 | 83.5 | 82.50 | 82.50 | |
| 12 | #1361 | Opus 4.8 xhigh | 92.5 | 72 | 86.5 | 77.5 | 82.13 | 82.13 | |
| 13 | #1376 | Fable 5 medium | 87.5★ | 73 | 86 | 75 | 80.38 | 78.00 | |
| 14 | #1331 | Grok 4.5 high | 93 | 58 | 90★ | 78 | 79.75 | 76.33 | |
| 15 | #1375 | Fable 5 low | 89★ | 52 | 93 | 78 | 78.00 | 74.33 | panel split |
| 16 | #1337 | Opus 4.8 xhigh | 87 | 59 | 93 | 72.5 | 77.88 | 77.88 | |
| 17 | #1379 | Sonnet 5 xhigh | 87 | 59.5 | 86.5 | 77 | 77.50 | 77.50 | |
| 18 | #1380 | Sonnet 5 xhigh | 86.5 | 59.5 | 80.5 | 81.5 | 77.00 | 77.00 | |
| 19 | #1336 | Opus 4.8 high | 85 | 56 | 86 | 73.5 | 75.13 | 75.13 | |
| 20 | #1373 | Grok 4.5 low | 83.5 | 62.5 | 73★ | 70 | 72.25 | 72.00 | |
| 21 | #1365 | GPT-5.6 Sol medium | 81.5 | 61★ | 87.5 | 55.5 | 71.38 | 74.83 | |
| 22 | #1360 | Opus 4.8 high | 88 | 55.5 | 85 | 53.5 | 70.50 | 70.50 | panel split |
| 23 | #1377 | Sonnet 5 high | 88 | 57 | 65.5 | 61.5 | 68.00 | 68.00 | |
| 24 | #1366 | GPT-5.6 Luna xhigh | 64.5 | 57.5 | 86 | 52 | 65.00 | 65.00 | |
| 25 | #1362 | GPT-5.6 Terra high | 45.5 | 36.5 | 33.5 | 34.5 | 37.50 | 37.50 |
Two underdogs punched above their weight.
The quiet story in the table is two entries from GPT-5.6's lesser known siblings. Terra at xhigh scored 88.50 and took sixth, ahead of Sol's high entry and every Opus 4.8, Sonnet 5 and Grok entry in the field. Luna at plain high scored 82.75 and took tenth, ahead of both Opus 4.8 entries running at xhigh. Neither model came in with a reputation, and both out-placed bigger names running harder settings.
The catch: neither result survives a different effort setting, and they break in opposite directions. Terra needed the extra effort; its high entry is the 37.50 dead-last card whose hedge can never switch on. Luna choked on it; its xhigh entry fell to 65.00, second from the bottom. Same family, same dial, opposite reactions. If this table hands you one practical rule, it is this: test your model at more than one effort level before you trust it.
So who
actually won?
Twenty-five entries, one hundred blind grades, averaged. The published order below uses the mean of four as submitted; the fairness column waits on the other side of the reveal. First, the field in one breath.
Counting down, from twenty-fifth place…
Last place earned it the hard way: all four judges independently found the same defect in #1362. Its hedge machinery exists, but nothing in the live bot ever calls it, so the safety feature can never switch on. That is worse than skipping the job, because it looks done. The top four get their own screens.
Biggest build vs. smallest build
The largest rewrite in the field, 5,507 lines added across 27 files, is still ahead of you on this page. The smallest, 465 lines across 14 files, is already behind you: last place, and its hedge can never actually run. Size is never scored. This episode it still tells the story.
The nine-point coin flip
Two of the twenty-five entries are byte-for-byte identical code under different PR numbers. One judge scored that identical code twice: 86.5 and 77.5. Nine points is this table's noise floor. Two mid-table entries a few points apart are effectively tied.
The podium that moves
Remove the twelve starred self-scores and fifteen of the twenty-five entries change rank, including two of the three podium spots. Who moves, which judge moved them, and why it isn't the bias you'd guess: after the reveal.
Four entries crashed a finished contest and forced a re-grade. One of them won the whole thing, on its first attempt in this series.
The panel's calmest verdict. Its four blind scores (96, 97, 100, 95) sit within five points of each other, the tightest agreement in the field (SD 1.87), and it was the only entry every judge put in its own personal top three. It wins by 4.12 points, and by more (97.67) when self-scores come out. It also shipped the largest change in the field: +5,507 / −50 across 27 files.
First under the published average. First under the no-self average. First under every alternative way I sliced the scores. On a table where the judges split by up to 41 points elsewhere, they agreed hardest about this one.
The re-grade exists because this entry and three others arrived late. Redoing the whole round blind was the only honest way to let them in, and the newcomer beat the reigning champion by 4.12 points.
The podium moves when the self-scores leave.
Under the published average, the podium is Opus 5, Fable 5, Kimi. Throw out every judge's score of its own work and it becomes Opus 5, then Sol's xhigh entry (94.50), then Sol's high entry (92.17). Fifteen of the twenty-five entries shift at least one place between the two columns; the big movers are Sol's high entry (up five places), Fable's low entry (down five), and Grok's high entry (down three).
And the mechanism is the opposite of what you'd guess. Nobody inflated their own grade. Fable 5 gave its own entry a 97, the highest score that entry got from anyone, but Fable graded the whole field generously, 7.91 points above the panel average, and that habit reached its own paper too. Sol graded its own two best entries at 75.5 and 69.5, two of its most brutal verdicts anywhere; Sol is the panel's harshest judge, 12.67 points below the panel average, and it never once scored any entry above the group's mean, its own included. Its harshness hit its own work hardest and knocked both entries off the podium.
I ran the check across all four judges: none shows a systematic bias toward its own entries. Fable is slightly stingier with its own work than its lenient norm, Sol even harsher on its own than usual, Grok less generous on its own, and Opus 5 close to neutral. The flip comes from judges with strong personal grading scales sitting inside the race they're grading. The winner is the one entry the flip can't touch: first in both columns.
The top eight, and the floor.
Nothing, yet.
The hedge issue is still open and none of the twenty-five entries has merged. When one does, or when I blend the best of them, that's a future episode's story. For now the verdict stands on four blind opinions, averaged, with every star and every split disclosed above.