Sean Kraemer
← Projects

Code-LLM Prompting Study

We tested prompting strategies on a small code model. In our code-generation experiments, cleaning up the output improved results more than changing the prompt.

Python · PyTorch · transformersRepo

Origin
A four-person project in a graduate machine-learning-for-software-engineering seminar (CS 598). I set up the experiments, wrote and revised the write-ups, and presented our findings to the class.

The question

How far can prompting alone push a small quantized code model? We ran three controlled studies across HumanEval-family tasks: code generation, output prediction, test generation, Python to Java translation, and bug classification.

Fixed model, greedy decoding, and a deterministic 20-problem subset shared across every study so the comparisons stay paired. Everything ran on Google Colab GPUs, which is what the 4-bit quantization was for: fitting a real code model into free-tier VRAM.

What we found

The most consequential variable was not the prompt. Raw model output is nearly unusable at 5 percent pass@1, and a simple post-processor that keeps only the first function definition is worth up to 40 points. Evaluation hygiene dominates before prompting enters the picture.

Prompting does pay on reasoning tasks: output prediction went from 35 to 60 percent, translation from 45 to 75. But asking for more is not free. A coverage-obsessed test-generation prompt produced more tests with a lower pass rate.

These are point estimates on 20 problems, not significance claims. Every number in the README regenerates from committed run artifacts, and CI fails on drift.

My role

I set up the experiments, wrote and revised the write-ups, and presented the team’s findings to the class. Through that semester I also presented AI research papers, including the DeepSeek-Coder paper, which is the same model family the study ran on.

After the course, I prepared the repository for publication: cleaned up its history, added a script to regenerate results from saved run artifacts, and added tests and a CI check for changes to the reported results.