Make your
tests prove it.
A passing test is a claim. Change the code and see if that claim holds.
Mutineer breaks your Ruby code one change at a time. Your tests should catch it. When they don’t, you get the exact line they missed.
Open source. MIT licensed. Prism + stdlib only.
return 0 if qty >= 10return 0 if qty > 10Still passes? Your tests missed qty = 10.
Show the test that catches it
assert_equal 0, calculator.discount(10)The original returns 0 at 10. If the changed method returns a different value, this assertion fails and kills the mutant.
Illustrative example; no Ruby runs in this page.
8,170spec examples, down from 17,702 (−54%)
Stephen Margheim had Opus 5.5 use Mutineer to check that every spec in his team's suite did real work, and refactor the suite. Lines of spec code fell 27%.
Read Stephen's post on XMutate. Run. Report.
A surviving mutant shows a change your tests did not catch. Inspect it, strengthen the test, and run again. Some changes are equivalent and need no new test.
Mutate
Mutineer parses your source with Prism and applies one change per mutant — flip a < to <=, swap && for ||, delete a statement. Each mutant is re-parsed to confirm it's still valid Ruby.
Run
Every mutant runs your suite in a forked, isolated process — across all your cores. A coverage map means each mutant only runs the test files that actually reach it.
Report
Killed = your tests caught it. Survived = they didn't. You get a mutation score and a unified diff for every survivor, pinned to the exact line. Set a --threshold to fail CI, or --format html for a shareable report. Inspect example report data → in the website theme.
96% mutation score. 24 ÷ (24 + 1). An uncovered mutant is excluded from this score.
Inspect a sample reportGive your pipeline evidence.
Line coverage tells you which code ran under test. It can't tell you whether a test would notice if that code broke — exactly where AI-generated tests are weakest. Mutineer answers that, programmatically: versioned JSON, mutant ids that survive unrelated edits, structured exit codes, and diff-scoped runs.
Close the loop on test quality
An agent writes code and tests, then runs Mutineer on just the diff. Each survivor ships a ready-made diff to feed back: "write a test that fails under this change." Stop when survivors hit zero.
--since origin/main --format json
Fail only when a PR makes tests worse
Diff the run against a stored baseline by mutant id. Exit 1 on any new survivor or a score drop — adopt mutation testing on a legacy suite without fixing everything first.
--baseline prior.json --threshold 90
One step in your workflow
A composite action wraps the CLI: point it at your sources, a baseline, and a threshold. Reads the diff, posts the report, sets the exit code.
uses: davidteren/mutineer@v1
Small tool. Exacting work.
Prism and the standard library provide the runtime. Test-framework changes stay in disposable child processes; Rails apps can use the optional --rails or --boot entry point.
Zero runtime deps
Prism ships with Ruby ≥ 3.4 and everything else is stdlib. Nothing to pin, nothing to break on upgrade.
Valid mutants only
Every mutation is re-parsed before it runs. Syntactically broken mutants are skipped, not counted against your score.
Fork isolation
Each mutant runs in its own forked process on Linux & macOS — no state bleed between runs, parallel by default.
Coverage-mapped
A per-mutant coverage map runs only the test files that reach the mutated line. Digest-keyed cache skips unchanged work.
CI & agent ready
Stable, sorted --format json with a versioned schema and mutant ids that survive unrelated edits. --threshold and --baseline gate the build.
Rails-aware
--rails boots config/environment once, then forks per mutant with fork-safe ActiveRecord — dogfooded against a real Rails app.
Twenty ways to break your code.
Five run by default. Fifteen Tier-2 operators are off until you ask for them with --operators or .mutineer.yml — they're noisier, but they catch deeper gaps.
| Operator | Mutation rule | Tier |
|---|---|---|
| comparison | < ↔ <= , > ↔ >= , == ↔ != | default |
| arithmetic | + ↔ − , * ↔ / , % → * , ** → * | default |
| boolean_connector | && ↔ || | default |
| boolean_literal | true ↔ false , nil → true | default |
| statement_removal | replace a non-final statement with nil | default |
| return_nil | replace a return / final expression with nil | tier 2 |
| literal_mutation | integer → 0, 1, n+1 ; string → empty | tier 2 |
| condition_negation | wrap if/unless/ternary condition in !( … ) | tier 2 |
| string_literal | non-empty string → "" ; "" → "mutineer" | tier 2 |
| regex | drop ^ / $ anchors ; + ↔ * | tier 2 |
| collection_method | map ↔ each , all? ↔ any? , first ↔ last , min ↔ max , select ↔ reject | tier 2 |
| safe_navigation | &. → . | tier 2 |
| range | .. ↔ ... | tier 2 |
| negation_removal | !x , not x → x | tier 2 |
| chain_link | drop one call from a chain: a.b.c → a.c | tier 2 |
| operand_removal | a && b → a , b | tier 2 |
| array_literal | [a, b] → [] | tier 2 |
| condition_true | if / elsif / unless / ternary / modifier / case-in guard condition → true | tier 2 |
| condition_false | if / elsif / unless / ternary / modifier / case-in guard condition → false | tier 2 |
| operator_assignment | += ↔ −= , *= ↔ /= , %= → *= , **= → *= (not ||= , &&=) | tier 2 |
Run mutineer --list-operators to see the live set for your installed version.
Put your tests to the test.
Ruby ≥ 3.4 is the only requirement. Point Mutineer at your source and the tests that cover it.
Install
Run against a file and its test
Gate a PR against a baseline
Preview mutations without running tests
Configure once in .mutineer.yml
Select JSON output on the command line:
Key flags: --operators, --threshold, --baseline, --since, --rails, --only, --jobs, --format human|json|html, --output FILE, --dry-run, --fail-fast. Typed flags override .mutineer.yml.