Module: Mutineer::MutantId

Defined in:
lib/mutineer/mutant_id.rb

Overview

Content-based id for a mutant — NOT byte offsets. Pure function, reused by the Runner (matching the ignore list), the Reporter (emitting a copy- pasteable id per survivor), and #13 baseline gating (diffing id-sets run to run). digest is stdlib, so zero new deps.

Offset-free by design: keyed on the subject's project-relative file path + qualified_name (a method, not a byte position) + operator + the normalized mutated token + an occurrence ordinal among same-(operator, token) twins WITHIN the subject + (only when positive) the subject's ordinal among same-named subjects in its file. Unrelated edits that only shift byte offsets can preserve the id. File moves, renames, project-root changes, and changes to repeated-name or repeated-mutation order can change it.

Class Method Summary collapse

Instance Method Summary collapse

Class Method Details

.digest(parts) ⇒ String

NUL-joins the id parts and returns the first 12 hex chars of their SHA256.

Parameters:

  • parts (Array) —

    the values that make up the id.

Returns:

  • (String) —

    a 12-character hex id.



112
113
114
# File 'lib/mutineer/mutant_id.rb', line 112

def digest(parts)
  Digest::SHA256.hexdigest(parts.join("\x00"))[0, 12]
end

.for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) ⇒ String

Computes the stable id for a single mutant.

NUL-joined so token delimiters (||=, spaces, ::, #) can never collide with the separator; SHA256[0,12] gives a fixed-length, copy-pasteable key. The file path is part of the key, so the same qualified name in two files (an owner-less def, or a class reopened elsewhere) cannot collide. path is required so no caller can get the colliding legacy id by accident.

Parameters:

  • subject (Mutineer::Subject) —

    the subject (method) the mutant lives in; its qualified_name anchors the id to a method rather than a byte position.

  • mutation (Mutineer::Mutation) —

    the atomic edit whose operator is hashed.

  • source (String) —

    the full, unmutated source the mutation indexes into.

  • occurrence (Integer) (defaults to: 0) —

    0-based ordinal among twins sharing the same (operator, token) within the subject, disambiguating otherwise-identical mutants.

  • path (String) —

    the subject's file, normalized with ProjectPath.relative against the project root (an absolute real path when outside the root).

  • subject_ordinal (Integer) (defaults to: 0) —

    0-based ordinal among subjects in the same file sharing this qualified name (two owner-less def index in two DSL blocks). Hashed only when positive, so a subject whose name is unique in its file keeps the id it had without it.

Returns:

  • (String) —

    a 12-character hex id, stable while its keyed identity and occurrence order stay unchanged.



43
44
45
46
47
# File 'lib/mutineer/mutant_id.rb', line 43

def for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0)
  parts = [path, subject.qualified_name, mutation.operator, normalized_token(mutation, source), occurrence]
  parts << subject_ordinal if subject_ordinal.positive?
  digest(parts)
end

.for_subject(subject, source, mutations, path:, subject_ordinal: 0) ⇒ Array<String>

Computes ids for a subject's full mutation list, in input order, assigning each its 0-based occurrence among twins sharing the same (operator, token). This is what disambiguates a + b + c's two + mutants without an offset.

Parameters:

  • subject (Mutineer::Subject) —

    the subject the mutations belong to.

  • source (String) —

    the full, unmutated source for token normalization.

  • mutations (Array<Mutineer::Mutation>) —

    the subject's mutations, in order.

  • path (String) —

    the subject's normalized file path (see for).

  • subject_ordinal (Integer) (defaults to: 0) —

    the subject's ordinal among same-named subjects in its file (see for).

Returns:

  • (Array<String>) —

    one 12-character id per mutation, positionally aligned.



60
61
62
63
64
# File 'lib/mutineer/mutant_id.rb', line 60

def for_subject(subject, source, mutations, path:, subject_ordinal: 0)
  with_occurrences(mutations, source) do |m, occ|
    self.for(subject, m, source, occ, path: path, subject_ordinal: subject_ordinal)
  end
end

.legacy_for(subject, mutation, source, occurrence = 0) ⇒ String

The pre-1.3 id: the for formula without the path, so it collides across files. Kept only to match ignore entries and baselines stored in the old format; removed in 2.0.

Parameters:

  • subject (Mutineer::Subject) —

    the subject (method) the mutant lives in.

  • mutation (Mutineer::Mutation) —

    the atomic edit whose operator is hashed.

  • source (String) —

    the full, unmutated source the mutation indexes into.

  • occurrence (Integer) (defaults to: 0) —

    0-based ordinal among same-(operator, token) twins.

Returns:

  • (String) —

    a 12-character hex id in the old format.



75
76
77
78
# File 'lib/mutineer/mutant_id.rb', line 75

def legacy_for(subject, mutation, source, occurrence = 0)
  digest([subject.qualified_name, mutation.operator,
          normalized_token(mutation, source), occurrence])
end

.legacy_for_subject(subject, source, mutations) ⇒ Array<String>

for_subject for the pre-1.3 id format (see legacy_for).

Parameters:

  • subject (Mutineer::Subject) —

    the subject the mutations belong to.

  • source (String) —

    the full, unmutated source for token normalization.

  • mutations (Array<Mutineer::Mutation>) —

    the subject's mutations, in order.

Returns:

  • (Array<String>) —

    one old-format id per mutation, positionally aligned.



86
87
88
# File 'lib/mutineer/mutant_id.rb', line 86

def legacy_for_subject(subject, source, mutations)
  with_occurrences(mutations, source) { |m, occ| legacy_for(subject, m, source, occ) }
end

.normalized_token(mutation, source) ⇒ String

Extracts the exact code being mutated, whitespace-collapsed — the same normalization the Reporter's diff_for uses for its token label.

Parameters:

  • mutation (Mutineer::Mutation) —

    supplies the byte range to slice.

  • source (String) —

    the source to byteslice (byte offsets, never char).

Returns:

  • (String) —

    the mutated token with runs of whitespace collapsed to one space.



122
123
124
# File 'lib/mutineer/mutant_id.rb', line 122

def normalized_token(mutation, source)
  source.byteslice(mutation.start_offset...mutation.end_offset).gsub(/\s+/, " ").strip
end

.with_occurrences(mutations, source) {|mutation, occurrence| ... } ⇒ Array

Maps each mutation to the block's result, passing its 0-based occurrence among earlier mutations with the same (operator, token).

Parameters:

  • mutations (Array<Mutineer::Mutation>) —

    the subject's mutations, in order.

  • source (String) —

    the full, unmutated source for token normalization.

Yield Parameters:

  • mutation (Mutineer::Mutation) —

    the current mutation.

  • occurrence (Integer) —

    its ordinal among same-(operator, token) twins.

Returns:

  • (Array) —

    the block's results, positionally aligned.



98
99
100
101
102
103
104
105
106
# File 'lib/mutineer/mutant_id.rb', line 98

def with_occurrences(mutations, source)
  seen = Hash.new(0)
  mutations.map do |m|
    key = [m.operator, normalized_token(m, source)]
    occ = seen[key]
    seen[key] += 1
    yield m, occ
  end
end

Instance Method Details

#digest(parts) ⇒ String (private)

NUL-joins the id parts and returns the first 12 hex chars of their SHA256.

Parameters:

  • parts (Array) —

    the values that make up the id.

Returns:

  • (String) —

    a 12-character hex id.



112
113
114
# File 'lib/mutineer/mutant_id.rb', line 112

def digest(parts)
  Digest::SHA256.hexdigest(parts.join("\x00"))[0, 12]
end

#for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0) ⇒ String (private)

Computes the stable id for a single mutant.

NUL-joined so token delimiters (||=, spaces, ::, #) can never collide with the separator; SHA256[0,12] gives a fixed-length, copy-pasteable key. The file path is part of the key, so the same qualified name in two files (an owner-less def, or a class reopened elsewhere) cannot collide. path is required so no caller can get the colliding legacy id by accident.

Parameters:

  • subject (Mutineer::Subject) —

    the subject (method) the mutant lives in; its qualified_name anchors the id to a method rather than a byte position.

  • mutation (Mutineer::Mutation) —

    the atomic edit whose operator is hashed.

  • source (String) —

    the full, unmutated source the mutation indexes into.

  • occurrence (Integer) (defaults to: 0) —

    0-based ordinal among twins sharing the same (operator, token) within the subject, disambiguating otherwise-identical mutants.

  • path (String) —

    the subject's file, normalized with ProjectPath.relative against the project root (an absolute real path when outside the root).

  • subject_ordinal (Integer) (defaults to: 0) —

    0-based ordinal among subjects in the same file sharing this qualified name (two owner-less def index in two DSL blocks). Hashed only when positive, so a subject whose name is unique in its file keeps the id it had without it.

Returns:

  • (String) —

    a 12-character hex id, stable while its keyed identity and occurrence order stay unchanged.



43
44
45
46
47
# File 'lib/mutineer/mutant_id.rb', line 43

def for(subject, mutation, source, occurrence = 0, path:, subject_ordinal: 0)
  parts = [path, subject.qualified_name, mutation.operator, normalized_token(mutation, source), occurrence]
  parts << subject_ordinal if subject_ordinal.positive?
  digest(parts)
end

#for_subject(subject, source, mutations, path:, subject_ordinal: 0) ⇒ Array<String> (private)

Computes ids for a subject's full mutation list, in input order, assigning each its 0-based occurrence among twins sharing the same (operator, token). This is what disambiguates a + b + c's two + mutants without an offset.

Parameters:

  • subject (Mutineer::Subject) —

    the subject the mutations belong to.

  • source (String) —

    the full, unmutated source for token normalization.

  • mutations (Array<Mutineer::Mutation>) —

    the subject's mutations, in order.

  • path (String) —

    the subject's normalized file path (see for).

  • subject_ordinal (Integer) (defaults to: 0) —

    the subject's ordinal among same-named subjects in its file (see for).

Returns:

  • (Array<String>) —

    one 12-character id per mutation, positionally aligned.



60
61
62
63
64
# File 'lib/mutineer/mutant_id.rb', line 60

def for_subject(subject, source, mutations, path:, subject_ordinal: 0)
  with_occurrences(mutations, source) do |m, occ|
    self.for(subject, m, source, occ, path: path, subject_ordinal: subject_ordinal)
  end
end

#legacy_for(subject, mutation, source, occurrence = 0) ⇒ String (private)

The pre-1.3 id: the for formula without the path, so it collides across files. Kept only to match ignore entries and baselines stored in the old format; removed in 2.0.

Parameters:

  • subject (Mutineer::Subject) —

    the subject (method) the mutant lives in.

  • mutation (Mutineer::Mutation) —

    the atomic edit whose operator is hashed.

  • source (String) —

    the full, unmutated source the mutation indexes into.

  • occurrence (Integer) (defaults to: 0) —

    0-based ordinal among same-(operator, token) twins.

Returns:

  • (String) —

    a 12-character hex id in the old format.



75
76
77
78
# File 'lib/mutineer/mutant_id.rb', line 75

def legacy_for(subject, mutation, source, occurrence = 0)
  digest([subject.qualified_name, mutation.operator,
          normalized_token(mutation, source), occurrence])
end

#legacy_for_subject(subject, source, mutations) ⇒ Array<String> (private)

for_subject for the pre-1.3 id format (see legacy_for).

Parameters:

  • subject (Mutineer::Subject) —

    the subject the mutations belong to.

  • source (String) —

    the full, unmutated source for token normalization.

  • mutations (Array<Mutineer::Mutation>) —

    the subject's mutations, in order.

Returns:

  • (Array<String>) —

    one old-format id per mutation, positionally aligned.



86
87
88
# File 'lib/mutineer/mutant_id.rb', line 86

def legacy_for_subject(subject, source, mutations)
  with_occurrences(mutations, source) { |m, occ| legacy_for(subject, m, source, occ) }
end

#normalized_token(mutation, source) ⇒ String (private)

Extracts the exact code being mutated, whitespace-collapsed — the same normalization the Reporter's diff_for uses for its token label.

Parameters:

  • mutation (Mutineer::Mutation) —

    supplies the byte range to slice.

  • source (String) —

    the source to byteslice (byte offsets, never char).

Returns:

  • (String) —

    the mutated token with runs of whitespace collapsed to one space.



122
123
124
# File 'lib/mutineer/mutant_id.rb', line 122

def normalized_token(mutation, source)
  source.byteslice(mutation.start_offset...mutation.end_offset).gsub(/\s+/, " ").strip
end

#with_occurrences(mutations, source) {|mutation, occurrence| ... } ⇒ Array (private)

Maps each mutation to the block's result, passing its 0-based occurrence among earlier mutations with the same (operator, token).

Parameters:

  • mutations (Array<Mutineer::Mutation>) —

    the subject's mutations, in order.

  • source (String) —

    the full, unmutated source for token normalization.

Yield Parameters:

  • mutation (Mutineer::Mutation) —

    the current mutation.

  • occurrence (Integer) —

    its ordinal among same-(operator, token) twins.

Returns:

  • (Array) —

    the block's results, positionally aligned.



98
99
100
101
102
103
104
105
106
# File 'lib/mutineer/mutant_id.rb', line 98

def with_occurrences(mutations, source)
  seen = Hash.new(0)
  mutations.map do |m|
    key = [m.operator, normalized_token(m, source)]
    occ = seen[key]
    seen[key] += 1
    yield m, occ
  end
end