aisa-group/PostTrainBench

581 stars · Last commit 2026-10-02

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

README preview

# PostTrainBench: Can LLM Agents Automate LLM Post-Training?

[![Website](https://img.shields.io/badge/Website-posttrainbench.com-c17d5a)](http://posttrainbench.com/)

We introduce PostTrainBench, a benchmark that measures the ability of CLI agents to post-train pre-trained large language models (LLMs). In PostTrainBench, the agent's task is to improve the performance of a base LLM on a given benchmark. The agent is given access to an evaluation script and 10 hours on an H100 GPU. Performance is measured by the benchmark score of the post-trained LLM. This setup naturally evaluates an agent's ability to conduct AI R&D.

> [!IMPORTANT]
> **Run it on cloud GPUs via [Harbor](https://github.com/harbor-framework/harbor).** This repository's reference pipeline targets our HPC cluster (HTCondor), but `src/harbor_adapter/` runs the full benchmark (same prompt, judges and evaluation) on Modal, with no cluster needed. See [its README](src/harbor_adapter/README.md).

## Scaffolds

Agents are run through one of 4 CLI scaffolds: Claude Code, Codex CLI, Gemini CLI, and OpenCode. 


## Evaluation Tasks

PostTrainBench includes 6 benchmarks spanning reasoning, knowledge, math, health, and code:

1. **AIME 2025** - Math competition problems
2. **Arena Hard Writing** - Creative writing benchmark adapted from ArenaHard v2

View full repository on GitHub →