writes your code.Canary tests it.
Canary is the independent tester for code your agents write. It runs your app, tries to break every change, and reports what would have reached production.
Get started in under 2 minutes.
Works with Claude Code, Cursor, Codex and any of your agentsInstall the Canary CLI with `npm i -g @runcanary/cli`, then run `canary skills` and follow its instructions to onboard this repository.The blast radius is the unknown unknown.
- 01
Your agents raise more PRs than your team can review. Review turned into a skim.
- 02
The tests pass. The same agent wrote them, and nobody on your team reads them. Nor should they.
- 03
The bug shows up as a Slack message from a customer, or a page that wakes your on-call at 2am.
- Teams on Canary have seen bug reports on new code go down, and on-call pages with them.
Release the flock.
One loop. From your coding agent, and on every pull request.
Reviewers read the diff. Canary runs the app, tries to break the diff, and reports the blind spots you missed.
Canary reads the diff and your codebase and maps every flow the change can reach.
Runs where you write code.
Two triggers. One loop. Every run ends in a report.
Before your agent calls a task done. Canary snapshots the uncommitted tree, boots it in a clean sandbox and tests it. No commit, no PR.
Works where you work.
Canary reads the context you already have, and files every break back to your coding agent.
Every failure arrives with proof.
Four failures Canary caught on real customer pull requests.
Field notes from the frontier.
We publish what we learn measuring AI against real, messy codebases.
Find it before your users do.
Put the flock on your next pull request. It runs on every one after.