edonadei/caliper

203 stars · Last commit 2026-10-04

Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.

README preview

<h1 align="center">
  <picture>
    <source media="(prefers-color-scheme: dark)" srcset="docs/assets/icon-dark.svg">
    <img alt="" src="docs/assets/icon-light.svg" width="64" align="center">
  </picture>
  <br>Caliper
</h1>

<h3 align="center">Know if your agent skill actually works.</h3>

<p align="center">
Your skill worked when you tried it. Will it work the next nine times?<br>
After the next model update? When another skill competes for the same prompt?
</p>

<p align="center">
  <a href="https://pypi.org/project/caliper-eval/"><img src="https://img.shields.io/pypi/v/caliper-eval.svg" alt="PyPI"></a>
  <a href="https://pypi.org/project/caliper-eval/"><img src="https://img.shields.io/pypi/pyversions/caliper-eval.svg" alt="Python"></a>
  <a href="https://skills.sh/edonadei/caliper"><img src="https://skills.sh/b/edonadei/caliper" alt="Skills"></a>
  <img src="https://img.shields.io/badge/license-MIT-blue" alt="License: MIT">

View full repository on GitHub →