Sean Kraemer
← Projects

AI Agent

A coding-agent harness with structured tools, session memory, and an evaluation runner. The entire loop is testable offline against a deterministic model.

Python · LLM tool useRepo

Origin
Adapted and extended from the upstream mini-swe-agent project. The agent loop design is theirs.

The argument

The upstream project this builds on argues that bash is all you need for a coding agent. I built the counter-argument on top of it. Structured tools buy reliability that a shell cannot: schema validation, edits that fail rather than guess, a syntax gate before the expensive test run.

What I added

A 15-tool registry where each tool is a dataclass with JSON-schema parameters. Edits require a unique match or they refuse, which pushes disambiguation onto the model that has the context to resolve it. A lint gate after each edit kills broken patches early.

Opt-in session memory with recency truncation and deterministic replay. A SWE-bench runner that generates predictions and leaves scoring to the external harness, so generation is my code and scoring is the benchmark’s.

Making an agent testable

Agents are miserable to test, so the keystone is a deterministic offline model. Scripted outputs make the entire agent loop unit-testable, so the full suite runs in CI without a single live API call.

Memory truncates rather than summarizes, and that is deliberate. Summarization would add a model dependency to the memory path and make the tests nondeterministic. It is the simplest solution that works, which for a v1 is the right call.

The README leads with a recording of the agent fixing a real failing test suite end to end, every tool call visible, for eight cents. The script that records it is committed, so the demo is reproducible rather than curated.