Sean Kraemer
← Projects

KaggleBench

A benchmark testing whether AI agents can repair tabular preprocessing pipelines, not just write them. 1,700 scored runs across six systems.

Python · Kaggle APIRepo

Origin
A four-person project from a course on AI agents (CS 498). I authored 5 of the 20 tasks and 6 human baselines, built and co-built agent implementations and the evaluation tooling, and ran the campaign for my tasks.

The problem

Most agent benchmarks test whether an AI can write code. We tested something harder: whether agents can repair a preprocessing pipeline, deciding both what to add and which already-present harmful steps to remove.

The design

Twenty tasks grounded in real Kaggle competitions, with ground truth distilled from top-scoring public solutions rather than generated by a model. Agents pick from a fixed bank of candidate actions, which turns scoring into a deterministic set operation instead of one model judging another.

Remove recall shares the headline metric with add F1, because missing a harmful step is the costly error. The classic traps live in the banks as harmful actions: leakage, post-outcome features, imputing across the train and test boundary.

The results

A well-engineered single model call beat every agentic scaffold on both cost and accuracy, at an estimated $0.057 per run. In our benchmark, the extra agent steps did not improve on a single call with precomputed context.

Across our tasks, systems were better at adding useful steps than removing harmful ones.

My teammate Jonghu wrote an extended write-up of the project