Kaggle Playground
Two Playground competitions, run the same way: one validation contract frozen early, every candidate measured against it, and the public leaderboard treated as too noisy to tune on. Both times the private standing came out well above the public one — 92nd to 13th on S6E9, 183rd to 88th on S6E8.
// evidence
- S6E9 private
- 13 / 3,576
- S6E8 private
- 88 / 3,532
- best
- top 0.4%
- S6E9 AUC
- 0.94573
- S6E9 public
- 92nd
- S6E8 public
- 183rd
// S6E9 — Predicting EV purchases
Binary classification scored on ROC AUC, in a team of two. Private 0.94573, 13th of 3,576 teams; public 0.94680, 92nd.
The final submission stacks fourteen logit columns with a converged L2 logistic regression, every member on one shared five-fold contract. The code and a full write-up, including what the teams above us did differently, are public.
// What the data turned out to be
- Income as a row ID
- The data was synthetic, generated from a 10,000-row original. 97.9 percent of incomes were exact original values, about sixty synthetic rows each, so income behaved like an ID of the original row, and the models that estimated that effect well did best.
- Income as GPT-2 tokens
- An argument on the forum, turned into features by another competitor, was that an LLM generated the rows and wrote numbers as tokens. Ported onto our folds, that view was the largest single gain of the last week — credited in the repository, not claimed as ours.
- No leak
- No exact duplicate rows, and the row id alone scored an AUC of 0.49999.
// Rules the stack had to pass
- The 1/k rule
- Averaging a model over k fold splits measured as about 2.1e-4 / k below its limit. So a one-split column was never compared with a multi-split one; every candidate was re-measured at the incumbent’s split count.
- Admission
- A new member had to lift the stack’s out-of-fold AUC by more than 1e-5, hold across folds, and survive a paired bootstrap. Three late additions fell below that bar and went in as a judgement call, and the write-up says so.
- Not the public board
- At 57,000 rows the public leaderboard cannot resolve differences below about 5e-5. Private moved with out-of-fold almost one to one, about 1.0e-3 lower for every stack.
// Where it fell short
Every out-of-fold member trained on 80 percent of the rows, where several teams above used ten to twenty folds. There was no language model trained on the original data, and TabPFN 3.5 was left out on purpose because its licence is non-commercial and every member was meant to be OSI-licensed.
Most gradient-boosted members early-stop on the scored fold. That is common Playground practice, and it makes their solo scores slightly optimistic.
// S6E8 — public to private
A 192-member stack on a frozen five-fold split, with an honest out-of-fold score of 0.9700980. The standing moved from 183rd on the public leaderboard to 88th of 3,532 on the private one.
The offset between out-of-fold and leaderboard settled at +0.00112, which made it possible to predict a score before spending a submission on finding out.
// S6E8 negative results
- max_bin
- +0.0020 on its own, 0.0000002 once inside the ensemble.
- Groupby aggregates
- +0.0012 alone, 0.0000034 inside the ensemble.
- Dead ends
- Marginalisation failed outright. Row order, duplicate detection and parent-row lookup all went nowhere.