Mutineer
A clean-room mutation-testing tool for Ruby. Mutineer mutates your source one change at a time, runs your test suite (Minitest or RSpec) against each mutant, and reports the ones your tests failed to catch — the gaps where your suite isn't actually testing anything.
- Prism + stdlib only — zero runtime dependencies (Ruby ≥ 3.4).
- One mutation per mutant, validity-checked by re-parsing.
- Fork-isolated, parallel execution (Linux + macOS).
- Coverage-guided — each mutant runs only the test files that cover its line.
- Stops at the first failing test — in-process runs (not
--daemonor--test-command) stop a mutant's test run at the first failure.
📖 mutineer.github.io → — overview, operators, and usage.
Install
gem install mutineer
Or in a Gemfile:
gem "mutineer", group: :test
Usage
mutineer run <source...> --test <test...> [options]
Mutate lib/calculator.rb, checking it against its test, and fail CI if the
mutation score drops below 90%:
mutineer run lib/calculator.rb --test test/calculator_test.rb --threshold 90
Options
| Flag | Meaning | ||
|---|---|---|---|
--test FILE |
Test file covering the sources (repeatable) | ||
--operators LIST |
Comma-separated operator names (default: the Tier-1 set) | ||
--threshold FLOAT |
Exit 1 when the score is below FLOAT, or when nothing could be scored and something broke, or more than one mutant produced no verdict and they exceed 10% of those attempted (default: 0 = off) | ||
--only NAME |
Restrict to one fully-qualified subject, e.g. Calculator#add |
||
--framework NAME |
minitest (default) or rspec; auto-detected as rspec when most --test files end in _spec.rb |
||
--since REF |
Only mutate lines changed since git REF (e.g. origin/main) — ideal for PR CI |
||
--no-since |
Disable diff scoping; a typed no beats a .mutineer.yml since: key |
||
--baseline FILE |
Compare against a prior --format json run; exit 1 on new survivors / score drop (score drop is skipped under --since, whose score covers a different denominator; see CI) |
||
--baseline-epsilon FLOAT |
Score-drop tolerance for --baseline (default: 0) |
||
--jobs N |
Parallel worker count (default: processor count); forced to 1 by --test-command, --fail-fast, or --rails without --daemon |
||
--boot FILE |
Require an app entry point once before forking; select at least one test file | ||
--rails |
Boot config/environment and reconnect ActiveRecord per fork; without --daemon, defaults to redefine and runs serially |
||
--verbose |
Surface the real error when a fork capture fails (alias --debug) |
||
--strategy NAME |
Mutation application: reload whole-file (default) or redefine surgical (7a/7b accepted as deprecated aliases) |
||
--test-command CMD |
Run the suite as a subprocess in the app's own runtime (for apps on Ruby < 3.4); CMD must contain %{files}. See Apps on Ruby < 3.4 |
||
--daemon |
Boot the app once in a persistent daemon and fork per mutant, with per-worker DB isolation so --jobs N is safe under Rails (needs --rails/--boot; not with --test-command). See the daemon backend |
||
| `--format human\ | json\ | html` | Report format (default: human; html is a self-contained file) |
--output FILE |
Write the report to FILE instead of stdout | ||
--dry-run |
List candidate mutations without executing (honors suppression) | ||
--fail-fast |
Stop at the first surviving mutant | ||
--list-operators |
List available operators (default vs optional) and exit | ||
--version, --help |
Print version / usage and exit |
Exit codes
| Code | Meaning |
|---|---|
0 |
Score ≥ threshold (or no gate) and no baseline regression. |
1 |
Score below --threshold, OR nothing could be scored and something broke, or more than one mutant produced no verdict and they exceed 10% of those attempted, OR a --baseline regression, OR a runtime error. |
2 |
Usage / invalid-flag error (mistyped flag, bad path, unreadable baseline). |
Operators
Run mutineer --list-operators to see them. Default (Tier 1): arithmetic,
comparison, boolean_connector, boolean_literal, statement_removal.
Available but off by default (Tier 2, enable via --operators): return_nil,
literal_mutation, condition_negation, string_literal, regex,
collection_method, safe_navigation, range, negation_removal, chain_link,
operand_removal, array_literal, condition_true,
condition_false, operator_assignment.
Rails apps
Rails code needs its environment booted before the suite runs, so point Mutineer
at your app with --rails and run it inside the project's bundle:
RAILS_ENV=test bundle exec mutineer run \
app/models/order.rb --test test/models/order_test.rb --rails
--rails boots config/environment once in the parent process (every mutant
then forks and inherits it), defaults --strategy to redefine without --daemon
(surgical — it avoids reloading files into the app tree; --daemon keeps reload), and reconnects ActiveRecord in each
fork so the database connection is fork-safe. Use --boot FILE to boot a
different entry point. Boot mode requires at least one --test file and is
coverage-guided — each mutant runs only the test files that exercise its line
(coverage is captured by forking the booted app, then cached).
Add Mutineer to your Gemfile's test group:
gem "mutineer", group: :test, require: false
Faster, parallel-safe Rails (the --daemon backend)
--rails boots your app once but runs mutants serially — parallel --jobs
under Rails is unsafe, because every worker shares one test database and clobbers
the others' fixtures. --daemon fixes both: it boots the app once in a persistent
helper and forks per mutant, and gives each parallel worker its own database,
so --jobs N is safe and its verdicts are proven identical to a serial run.
RAILS_ENV=test bundle exec mutineer run \
app/models/order.rb --test test/models/order_test.rb \
--rails --daemon --jobs 4
- One boot, forked per mutant — restores the shared-boot speed.
- Coverage-guided — each mutant runs only its covering tests (like
--rails); a mutant on an uncovered line isno_coverage, so the score stays comparable to the in-process--railsscore. - Safe
--jobs N— each worker routes to its own copy of the test database, so parallel verdicts equal serial (no fixture cross-talk). - One backend at a time —
--daemoncan't be combined with--test-command(choose one), and it needs an app to boot (--railsor--boot).
Status: SQLite today (hermetic, CI-proven). Postgres per-worker
provisioning is in progress (#34/#35); until it lands, use --daemon with a
SQLite test database, or drop --jobs to run serially on other adapters.
Apps on Ruby < 3.4
Mutineer's own process needs Ruby ≥ 3.4 (it parses with stdlib Prism), and the
--rails path above boots your app inside Mutineer's process — so it can't run
against an app pinned to an older Ruby (ruby "3.1.6" in the Gemfile), where the
bundle rejects 3.4.
--test-command decouples the two: Mutineer stays on ≥ 3.4, but your suite runs
as a subprocess in your app's own runtime (whatever Ruby its bundle resolves
to). Run Mutineer with a 3.4+ Ruby and hand it the command that runs your tests:
RAILS_ENV=test mutineer run app/models/order.rb \
--test test/models/order_test.rb \
--test-command "bundle exec rails test %{files}"
%{files}is required; it expands to the--testpaths as separate arguments (a path with a space stays one argument — there is no shell).- Environment: vars like
RAILS_ENV/DATABASE_URLset on the Mutineer command are inherited. Mutineer unsetsBUNDLE_*,GEM_*,RUBY*,RBENV_VERSION,ASDF_RUBY_VERSION, andRBENV_DIRin the child (so Mutineer's own Ruby cannot pin the suite), and drops version-manager version bins (e.g.~/.rbenv/versions/3.4.x/bin) fromPATH, then prepends rbenv/asdf shims when a pin was scrubbed. Do not rely onRBENV_VERSION=…on the Mutineer command for the suite; use.ruby-versionor a wrapper. Don't putKEY=valprefixes inside--test-command(no shell; that would be treated as the program name).
Under a version manager (rbenv / asdf / chruby)
Automatic scrub targets rbenv and asdf (shims + version bins). chruby
has no shims: Mutineer still strips …/rubies/…/bin so it cannot leave Mutineer's
Ruby pinned, but you need a wrapper that sources chruby and selects the app
version. If the smoke check still reports a Ruby version mismatch, wrap the
suite. Example rbenv wrapper (bin/mutineer-test in the app):
#!/usr/bin/env bash
set -euo pipefail
cd "$(dirname "$0")/.."
unset GEM_HOME GEM_PATH RUBYLIB RUBYOPT BUNDLE_GEMFILE BUNDLE_BIN_PATH BUNDLER_VERSION
export RBENV_VERSION="$(cat .ruby-version 2>/dev/null || true)"
export RAILS_ENV="${RAILS_ENV:-test}"
export PATH="${HOME}/.rbenv/shims:${PATH}"
exec bundle exec rails test "$@"
# Mutineer on 3.4+; suite on the app's Ruby via the wrapper
mutineer run app/models/order.rb \
--test test/models/order_test.rb \
--test-command "bin/mutineer-test %{files}"
Mutineer also surfaces a targeted smoke-check message when Bundler prints
RubyVersionMismatch, instead of only blaming DB/migrations.
Tradeoffs — this path is correct but not free:
- Slower: your app re-boots for every mutant (no shared boot yet).
- No coverage narrowing: every mutant runs the full
--testset, so the score is an upper bound and not comparable to an in-process (--rails) score — uncovered mutants count as survivors, and an infrastructure failure is scored as a kill. Mutineer prints this caveat on every run and aborts up front (a "smoke check") if your unmutated suite isn't green. - Reload strategy only (
--strategy redefineis rejected on this path) and serial (--jobsis forced to 1). For apps on Ruby ≥ 3.4,--daemongives safe parallelism instead (see the daemon backend).
Suppressing equivalent mutants
Some mutants are equivalent (behaviour-identical) and survive forever — keeping a
file off 100%. Suppress them so the score and --threshold gate stay meaningful:
- Inline:
some_line # mutineer:disable-line(or scope it:# mutineer:disable-line comparison). Put a reason after--:# mutineer:disable-line comparison -- the test checks only 20. - Config: a
.mutineer.ymlignore:list of mutant ids. Each survivor'sidis printed in the JSON report, so copy it straight intoignore:.
Suppressed mutants are excluded from the score (so 100% becomes reachable).
Mutant ids
A mutant id is 12 hex characters. It hashes the file path (relative to the
project root), the method's qualified name, the operator, the mutated code, and
the mutant's position among identical mutants in that method. When one file has
two methods with the same qualified name (for example two top-level def index
in two DSL blocks), the second and later ones also hash their position among
those methods, so their ids differ. The first one's id does not change. An edit
outside the method does not change the id. Moving or renaming the file,
renaming the method or its class, or adding an identical mutant earlier in the
method does. Adding a method with the same name earlier in the same file also
does.
- The project root is the directory mutineer runs from (in the Action, the
working-directory). Run from the same root to get the same ids. - A source outside the project root uses its absolute path, so its ids differ between machines.
Migrating from ids without the file path. Before 1.3, ids did not include the file path, so two files could share an id (#126). Old-format ids keep working until 2.0, with a warning:
ignore:An old entry still suppresses its mutants. The run prints the new ids for each old entry, each with its file and method. When it names one mutant, replace the entry with that id. When it names several (in different files, or same-named methods in one file), the old entry over-matched: it also hid mutants you did not mean to ignore. The warning says so. Keep only the ids for the mutant you meant to ignore, not all of them. The list covers only the sources and operators in that run, so run over every source with every operator set you use (for example your Tier-2--operators) for the full list.--baselineAn old baseline still matches: a survivor matches a stored one with the same old id in the same file. A stored file that is an absolute path outside the project root (a baseline written on another machine) matches on the old id alone. The run tells you to regenerate it. Regenerate it with--format json, but only after every gate that reads it runs 1.3 or later (the Action'sversion:pin, your CIGemfile.lock). An older version treats every new-format survivor as new.
The JSON report's summary.id_format is 2 for the new format.
summary.legacy_id_matches.ignore counts the old-format ignore entries a run
matched, and summary.legacy_id_matches.baseline counts the survivors matched
only through an old baseline id.
Ids are relative to the directory you run mutineer from. mutineer finds
.mutineer.yml by walking up. When the file it loads is in a parent directory
(other than your home directory), it warns that the ignore ids will not match
and tells you which directory to run from.
CI gating
Store a JSON run as a baseline, then fail the build only when a PR makes things worse:
mutineer run app/ --baseline .mutineer/baseline.json # exit 1 on NEW survivors or a score drop
--baseline reports which survivors are new (by mutant id) and any score drop. It
combines with --threshold (the worse of the two sets the exit code). Pass a
directory (or several sources) to audit a whole layer in one boot — tests are
auto-paired by convention and the report breaks down per source.
GitHub Action
This repo ships a composite action (action.yml) that wraps the CLI for CI:
- uses: actions/checkout@v4
- uses: ruby/setup-ruby@v1
with: { ruby-version: "3.4", bundler-cache: true }
- uses: davidteren/mutineer@v1
with:
sources: app/
baseline: .mutineer/baseline.json
threshold: "90"
Default change: on pull_request events (not pull_request_target) the
action scopes the run to the PR's changed lines, diffing against the PR's exact
base commit (fetched by the action itself when the checkout is shallow; falls
back to the base branch tip). Pass since: none for a full scan, or an
explicit since: ref (which needs fetch-depth: 0 on checkout).
With the default JSON format the action also:
- writes a score summary to the job's step summary;
- annotates surviving mutants on the PR diff, up to 50 (
errorlevel when the gate failed,warningwhen it passed); - exposes the report path via the
reportoutput for later steps (withformat: human/htmlthis needs theoutputinput).
The CLI prints a progress line to the log at every 10% of the run, whatever the format.
For AI agents & pipelines
Mutineer is built for programmatic use — versioned JSON, mutant ids that survive unrelated edits, structured exit codes, and diff-scoped runs. See:
- AI agents & CI recipes — the agent inner-loop and CI-gate recipes (and how to avoid infinite loops on equivalent mutants): rendered · source
- JSON schema reference — the
--format jsonschema and its versioning contract: rendered · source - Ruby API (YARD) — class reference for the shipped gem: https://davidteren.github.io/mutineer/api/
Configuration
Mutineer reads an optional .mutineer.yml from the project root (nearest one,
walking up). CLI flags override config; config overrides defaults.
Sources are positional CLI arguments and test files come from --test. The
config file accepts these keys:
| Key | Value and purpose |
|---|---|
operators |
An operator name or list of names; defaults to the Tier-1 set |
threshold |
A number from 0 to 100; 0 turns the score gate off |
jobs |
A positive integer; the default is the processor count. test_command, fail_fast, or --rails without --daemon forces 1. |
only |
A fully-qualified subject name, such as Calculator#add |
require |
A path or list of extra files to load before mutating |
boot |
The app entry point to require once before forking |
rails |
true or false; enables the Rails boot defaults |
since |
A nonblank git ref, or false to disable diff scoping |
framework |
minitest or rspec; an explicit value is kept during test pairing |
verbose |
true or false; shows capture diagnostics |
ignore |
A mutant id or list of ids to suppress |
baseline |
The path to a prior JSON report |
fail_fast |
true or false; stops scheduling after the first survivor |
test_command |
The external-runtime suite command, including %{files}; see Apps on Ruby < 3.4 |
daemon |
true or false; uses the persistent app daemon with worker DB isolation |
In 1.4, invalid values for known scalar keys exit 2 with a message naming the
file and key. The list keys (operators, require, ignore) are not checked
this way: an unknown operator name warns and is skipped, so operators: [bogus]
runs no mutants and exits 0 (#167). Boolean keys take true or false (quoted forms also work), not "yes".
jobs must be positive; a string value contains digits only. String values for
threshold and the CLI-only --baseline-epsilon use plain decimals such as 90
or 0.5, not +2, 1e2, or 1_0. String options such as only and baseline
cannot be null or boolean. A blank since is invalid; use since: false to turn
scoping off. Unknown keys and operator names warn and are ignored.
format, strategy, output, baseline_epsilon, and dry_run are CLI-only.
For JSON output, use --format json, not a format: config key. To select RSpec
in the file, add framework: rspec.
# .mutineer.yml
operators: [arithmetic, comparison, boolean_connector, boolean_literal, statement_removal]
threshold: 90
jobs: 4
require:
- config/environment
Coverage results are cached in .mutineer/coverage.json (digest-keyed; rebuilt
automatically when sources change). Add .mutineer/ to your .gitignore.
License
MIT — see LICENSE.