KaggleBench
A benchmark testing whether AI agents can repair tabular preprocessing pipelines, not just write them. 1,700 scored runs across six systems.
Origin
A four-person project from a course on AI agents (CS 498). I authored 5 of the 20 tasks and 6 human baselines, built and co-built agent implementations and the evaluation tooling, and ran the campaign for my tasks.
The problem
Most agent benchmarks test whether an AI can write code. We tested something harder: whether agents can repair a preprocessing pipeline, deciding both what to add and which already-present harmful steps to remove.
The design
Twenty tasks grounded in real Kaggle competitions, with ground truth distilled from top-scoring public solutions rather than generated by a model. Agents pick from a fixed bank of candidate actions, which turns scoring into a deterministic set operation instead of one model judging another.
Remove recall shares the headline metric with add F1, because missing a harmful step is the costly error. The classic traps live in the banks as harmful actions: leakage, post-outcome features, imputing across the train and test boundary.
The results
A well-engineered single model call beat every agentic scaffold on both cost and accuracy, at an estimated $0.057 per run. In our benchmark, the extra agent steps did not improve on a single call with precomputed context.
Across our tasks, systems were better at adding useful steps than removing harmful ones.
My teammate Jonghu wrote an extended write-up of the project