prime-radiant-inc/superpowers-evals

124 stars · Last commit 2026-09-12

Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and deterministic post-checks.

README preview

# Superpowers Evals

Behavioral eval lab for [superpowers](https://github.com/obra/superpowers).
**Quorum** drives and grades real coding-agent CLIs (Claude, Codex,
Antigravity, Gemini, Hermes, Kimi, OpenCode, Pi, and Copilot). New scenarios use
a user-like conversation followed by a fresh evidence-based assessment against
private criteria plus deterministic post-checks.

Code, CLI, paths, and inline prose all use lowercase `quorum`; the capitalized
form `Quorum` appears in headings and the actor table.

This is not a generic benchmark suite. It is an eval lab for workflow
compliance: skill triggering, worktree behavior, subagent coordination,
verification reflexes, review quality, and cost-shaping patterns.

## Safety Model

quorum has two very different execution modes:

- **Static/unit checks** are safe for public CI. They run `biome`, `tsc`, and

View full repository on GitHub →