Module: Mutineer::MutantId
- Defined in:
- lib/mutineer/mutant_id.rb
Overview
Content-based id for a mutant — NOT byte offsets. Pure function, reused
by the Runner (matching the ignore list), the Reporter (emitting a copy-
pasteable id per survivor), and #13 baseline gating (diffing id-sets run to
run). digest is stdlib, so zero new deps.
Offset-free by design: keyed on the subject's project-relative file path + qualified_name (a method, not a byte position) + operator + the normalized mutated token + an occurrence ordinal among same-(operator, token) twins WITHIN the subject + (only when positive) the subject's ordinal among same-named subjects in its file. Unrelated edits that only shift byte offsets can preserve the id. File moves, renames, project-root changes, and changes to repeated-name or repeated-mutation order can change it.
Class Method Summary collapse
-
.digest(parts) ⇒ String
NUL-joins the id parts and returns the first 12 hex chars of their SHA256.
-
.for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) ⇒ String
Computes the stable id for a single mutant.
-
.for_subject(subject, source, mutations, path:, subject_ordinal: 0) ⇒ Array<String>
Computes ids for a subject's full mutation list, in input order, assigning each its 0-based occurrence among twins sharing the same (operator, token).
-
.legacy_for(subject, mutation, source, occurrence = 0) ⇒ String
The pre-1.3 id: the MutantId.for formula without the path, so it collides across files.
-
.legacy_for_subject(subject, source, mutations) ⇒ Array<String>
MutantId.for_subject for the pre-1.3 id format (see MutantId.legacy_for).
-
.normalized_token(mutation, source) ⇒ String
Extracts the exact code being mutated, whitespace-collapsed — the same normalization the Reporter's
diff_foruses for its token label. -
.with_occurrences(mutations, source) {|mutation, occurrence| ... } ⇒ Array
Maps each mutation to the block's result, passing its 0-based occurrence among earlier mutations with the same (operator, token).
Instance Method Summary collapse
-
#digest(parts) ⇒ String
private
NUL-joins the id parts and returns the first 12 hex chars of their SHA256.
-
#for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) ⇒ String
private
Computes the stable id for a single mutant.
-
#for_subject(subject, source, mutations, path:, subject_ordinal: 0) ⇒ Array<String>
private
Computes ids for a subject's full mutation list, in input order, assigning each its 0-based occurrence among twins sharing the same (operator, token).
-
#legacy_for(subject, mutation, source, occurrence = 0) ⇒ String
private
The pre-1.3 id: the MutantId.for formula without the path, so it collides across files.
-
#legacy_for_subject(subject, source, mutations) ⇒ Array<String>
private
MutantId.for_subject for the pre-1.3 id format (see MutantId.legacy_for).
-
#normalized_token(mutation, source) ⇒ String
private
Extracts the exact code being mutated, whitespace-collapsed — the same normalization the Reporter's
diff_foruses for its token label. -
#with_occurrences(mutations, source) {|mutation, occurrence| ... } ⇒ Array
private
Maps each mutation to the block's result, passing its 0-based occurrence among earlier mutations with the same (operator, token).
Class Method Details
.digest(parts) ⇒ String
NUL-joins the id parts and returns the first 12 hex chars of their SHA256.
112 113 114 |
# File 'lib/mutineer/mutant_id.rb', line 112 def digest(parts) Digest::SHA256.hexdigest(parts.join("\x00"))[0, 12] end |
.for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) ⇒ String
Computes the stable id for a single mutant.
NUL-joined so token delimiters (||=, spaces, ::, #) can never collide
with the separator; SHA256[0,12] gives a fixed-length, copy-pasteable key.
The file path is part of the key, so the same qualified name in two files
(an owner-less def, or a class reopened elsewhere) cannot collide. path
is required so no caller can get the colliding legacy id by accident.
43 44 45 46 47 |
# File 'lib/mutineer/mutant_id.rb', line 43 def for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) parts = [path, subject.qualified_name, mutation.operator, normalized_token(mutation, source), occurrence] parts << subject_ordinal if subject_ordinal.positive? digest(parts) end |
.for_subject(subject, source, mutations, path:, subject_ordinal: 0) ⇒ Array<String>
Computes ids for a subject's full mutation list, in input order, assigning
each its 0-based occurrence among twins sharing the same (operator, token).
This is what disambiguates a + b + c's two + mutants without an offset.
60 61 62 63 64 |
# File 'lib/mutineer/mutant_id.rb', line 60 def for_subject(subject, source, mutations, path:, subject_ordinal: 0) with_occurrences(mutations, source) do |m, occ| self.for(subject, m, source, occ, path: path, subject_ordinal: subject_ordinal) end end |
.legacy_for(subject, mutation, source, occurrence = 0) ⇒ String
The pre-1.3 id: the for formula without the path, so it collides across files. Kept only to match ignore entries and baselines stored in the old format; removed in 2.0.
75 76 77 78 |
# File 'lib/mutineer/mutant_id.rb', line 75 def legacy_for(subject, mutation, source, occurrence = 0) digest([subject.qualified_name, mutation.operator, normalized_token(mutation, source), occurrence]) end |
.legacy_for_subject(subject, source, mutations) ⇒ Array<String>
for_subject for the pre-1.3 id format (see legacy_for).
86 87 88 |
# File 'lib/mutineer/mutant_id.rb', line 86 def legacy_for_subject(subject, source, mutations) with_occurrences(mutations, source) { |m, occ| legacy_for(subject, m, source, occ) } end |
.normalized_token(mutation, source) ⇒ String
Extracts the exact code being mutated, whitespace-collapsed — the same
normalization the Reporter's diff_for uses for its token label.
122 123 124 |
# File 'lib/mutineer/mutant_id.rb', line 122 def normalized_token(mutation, source) source.byteslice(mutation.start_offset...mutation.end_offset).gsub(/\s+/, " ").strip end |
.with_occurrences(mutations, source) {|mutation, occurrence| ... } ⇒ Array
Maps each mutation to the block's result, passing its 0-based occurrence among earlier mutations with the same (operator, token).
98 99 100 101 102 103 104 105 106 |
# File 'lib/mutineer/mutant_id.rb', line 98 def with_occurrences(mutations, source) seen = Hash.new(0) mutations.map do |m| key = [m.operator, normalized_token(m, source)] occ = seen[key] seen[key] += 1 yield m, occ end end |
Instance Method Details
#digest(parts) ⇒ String (private)
NUL-joins the id parts and returns the first 12 hex chars of their SHA256.
112 113 114 |
# File 'lib/mutineer/mutant_id.rb', line 112 def digest(parts) Digest::SHA256.hexdigest(parts.join("\x00"))[0, 12] end |
#for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) ⇒ String (private)
Computes the stable id for a single mutant.
NUL-joined so token delimiters (||=, spaces, ::, #) can never collide
with the separator; SHA256[0,12] gives a fixed-length, copy-pasteable key.
The file path is part of the key, so the same qualified name in two files
(an owner-less def, or a class reopened elsewhere) cannot collide. path
is required so no caller can get the colliding legacy id by accident.
43 44 45 46 47 |
# File 'lib/mutineer/mutant_id.rb', line 43 def for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) parts = [path, subject.qualified_name, mutation.operator, normalized_token(mutation, source), occurrence] parts << subject_ordinal if subject_ordinal.positive? digest(parts) end |
#for_subject(subject, source, mutations, path:, subject_ordinal: 0) ⇒ Array<String> (private)
Computes ids for a subject's full mutation list, in input order, assigning
each its 0-based occurrence among twins sharing the same (operator, token).
This is what disambiguates a + b + c's two + mutants without an offset.
60 61 62 63 64 |
# File 'lib/mutineer/mutant_id.rb', line 60 def for_subject(subject, source, mutations, path:, subject_ordinal: 0) with_occurrences(mutations, source) do |m, occ| self.for(subject, m, source, occ, path: path, subject_ordinal: subject_ordinal) end end |
#legacy_for(subject, mutation, source, occurrence = 0) ⇒ String (private)
The pre-1.3 id: the for formula without the path, so it collides across files. Kept only to match ignore entries and baselines stored in the old format; removed in 2.0.
75 76 77 78 |
# File 'lib/mutineer/mutant_id.rb', line 75 def legacy_for(subject, mutation, source, occurrence = 0) digest([subject.qualified_name, mutation.operator, normalized_token(mutation, source), occurrence]) end |
#legacy_for_subject(subject, source, mutations) ⇒ Array<String> (private)
for_subject for the pre-1.3 id format (see legacy_for).
86 87 88 |
# File 'lib/mutineer/mutant_id.rb', line 86 def legacy_for_subject(subject, source, mutations) with_occurrences(mutations, source) { |m, occ| legacy_for(subject, m, source, occ) } end |
#normalized_token(mutation, source) ⇒ String (private)
Extracts the exact code being mutated, whitespace-collapsed — the same
normalization the Reporter's diff_for uses for its token label.
122 123 124 |
# File 'lib/mutineer/mutant_id.rb', line 122 def normalized_token(mutation, source) source.byteslice(mutation.start_offset...mutation.end_offset).gsub(/\s+/, " ").strip end |
#with_occurrences(mutations, source) {|mutation, occurrence| ... } ⇒ Array (private)
Maps each mutation to the block's result, passing its 0-based occurrence among earlier mutations with the same (operator, token).
98 99 100 101 102 103 104 105 106 |
# File 'lib/mutineer/mutant_id.rb', line 98 def with_occurrences(mutations, source) seen = Hash.new(0) mutations.map do |m| key = [m.operator, normalized_token(m, source)] occ = seen[key] seen[key] += 1 yield m, occ end end |