AMD-AGI/AgentKernelArena
107 stars · Last commit 2026-07-29
AgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, and GEAK—can be evaluated side-by-side on the same GPU kernel tasks, using objective and reproducible metrics.
README preview
# AgentKernelArena: An A/B Testing and RL-Ready Environment for GPU Kernel Agents AgentKernelArena is a controlled experimentation platform for developing AI agents on real GPU kernel optimization tasks. It enables reproducible A/B testing across models, prompts, tools, and agent policies, while providing objective compilation, correctness, and performance signals that can serve as rewards for agent reinforcement learning. ## Overview AgentKernelArena makes changes to an agent measurable. Run the same task set with a baseline and a treatment—such as a different model, prompt, MCP server, skill, tool, memory strategy, or policy—and compare the outcomes under the same execution and scoring pipeline. The platform provides: - **Controlled A/B experiments**: Label and compare repeated runs while holding tasks, hardware, environment, and evaluation rules constant. - **RL-ready feedback**: Produce per-task compilation, correctness, runtime, speedup, and score signals that can be consumed as rewards by an external reinforcement-learning system. - **Multiple agent integrations**: Run Cursor Agent, Claude Code, Codex, GEAK-based agents, mini-swe-agent-based flows, or custom agents through a shared interface. - **Real GPU task environments**: Work with HIP, Triton, FlyDSL, PyTorch-to-kernel conversion, instruction-to-kernel generation, and repository-level optimization tasks. - **Isolated and reproducible execution**: Give every task its own timestamped workspace and preserve logs, modified sources, and structured results. - **Centralized evaluation**: Measure compilation, correctness, and GPU performance independently of the optimizing agent. - **Multi-GPU scheduling**: Start one isolated Docker worker per GPU and dynamically claim tasks from a shared queue. - **Resumable experiments**: Resume a run without repeating tasks that already produced a completion report. - **Held-out evaluation**: Test optimized kernels on unseen shapes and measure the generalization gap. - **Task validation and visualization**: Validate task quality with a dedicated agent and compare local run reports in a dashboard.