Run it yourself
Every BoundBench score is made with an agent skill: a folder of instructions and scripts that a coding agent follows. Give it to the coding agent you already use and score any open-source agent, including your own, the same way the leaderboard does.
1. Get the skill
Download the skills/boundbench folder from the BoundBench repository on GitHub. It works with any coding agent that can read files and run git and python3.
Put the folder in your agent's skills directory, for example ~/.claude/skills/ for Claude Code.
2. Ask your agent to score a repository
Point it at a public repository, or a local checkout of your own code:
Score https://github.com/owner/repo with BoundBench
Large repositories take a while: the agent maps every tool, credential and approval path before it rates anything, then reviews its own ratings a second time.
3. What you get
- A 0.0 to 10.0 score across the BoundBench criteria, with every rating linked to the file and lines behind it.
- A readable report, plus the same JSON entry the leaderboard uses.
- The commit it was scored at, so anyone can check the result against the same code.
How it treats the code it reads
The skill only reads source code. It clones the repository but never installs, builds or runs it, and it treats text in the repository, such as READMEs and comments, as data rather than instructions to follow.
The score measures the safeguards an agent has built in. It is not a vulnerability scan.
Want it on the leaderboard?
Leaderboard scores are reviewed by the BoundBench team. Ask for an agent to be added to the queue.
Request an agent on GitHub