aisa-group/PostTrainBench
581 stars · Last commit 2026-10-02
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
README preview
# PostTrainBench: Can LLM Agents Automate LLM Post-Training? [](http://posttrainbench.com/) We introduce PostTrainBench, a benchmark that measures the ability of CLI agents to post-train pre-trained large language models (LLMs). In PostTrainBench, the agent's task is to improve the performance of a base LLM on a given benchmark. The agent is given access to an evaluation script and 10 hours on an H100 GPU. Performance is measured by the benchmark score of the post-trained LLM. This setup naturally evaluates an agent's ability to conduct AI R&D. > [!IMPORTANT] > **Run it on cloud GPUs via [Harbor](https://github.com/harbor-framework/harbor).** This repository's reference pipeline targets our HPC cluster (HTCondor), but `src/harbor_adapter/` runs the full benchmark (same prompt, judges and evaluation) on Modal, with no cluster needed. See [its README](src/harbor_adapter/README.md). ## Scaffolds Agents are run through one of 4 CLI scaffolds: Claude Code, Codex CLI, Gemini CLI, and OpenCode. ## Evaluation Tasks PostTrainBench includes 6 benchmarks spanning reasoning, knowledge, math, health, and code: 1. **AIME 2025** - Math competition problems 2. **Arena Hard Writing** - Creative writing benchmark adapted from ArenaHard v2