Skip to content

Two stages hand-copy each other's trailing-word logic after #601/#602; share one predicate each #610

Description

@derek73

Why. The /simplify pass on #609 found two places where one stage re-derives a decision another stage owns. Each pair agrees today. Nothing pins the agreement, which is the drift class mechanisms.md#ONE-PREDICATE-PER-QUESTION names: each site keeps passing its own tests while they disagree about an input neither covers. #609 left both alone because fixing them reaches outside its diff.

1. P6's particle tail, computed in assign and in post_rules.

  • _assign.py's family-comma given-part run (Should a name word after a credential join the suffix? John Smith PhD Jones reads middle Smith PhD #602) works out p6_lo:p6_hi, the particles ending the part behind any post-nominals. It leaves those to P6, so the run doesn't absorb a word P6 will attach to the family (Smith, John PhD de).
  • _post_rules.py's P6 attachment finds the same tail its own way. It steps back past pieces with no name role, then back over particle pieces, stopping at a lone ambiguous member whose role is SUFFIX.
  • The two stop differently. Assign tests tags, post_rules tests roles plus AMBIGUOUS_ACRONYM_TAG. They agree only because the run has already given every non-particle word and every class member a SUFFIX or TITLE role. A change to either walk, or to the run's roles, could split them silently.
  • Proposed: one _pieces predicate for "the P6 tail of this part", called by both stages, each passing its own test for a word that holds no name. Cost: probably one frame on every family-comma parse, so _CALL_BASELINE would be re-recorded. Measure before deciding.

2. The maiden take's candidate filter copies the suffix peel's admissions.

  • _group._maiden_take decides which words at the end of a clause might belong to the trailing run (_clause_tail_word plus an inline numeral-shape clause). Then it lets tail_reading decide which of them it actually takes.
  • That filter is a hand copy of peel_trailing's per-piece admissions (suffix piece, end-of-walk numeral against the word before it, ambiguous member) and of H5's word test. Its contract is "a superset of what tail_reading can take". Nothing states or tests that contract. A new admission in the peel would be silently missing from the clause's view.
  • Rethink the maiden-marker rule (M2): a marker stands after the current name and takes words up to the trailing suffix run #601's own implementation already hit this once: the prototype left out the numeral fork, so John née Jones Smith VI kept VI in the maiden name.
  • Proposed: one may_trail predicate in _pieces, written beside peel_trailing and owning the numeral arm, plus a sweep test pinning that every word tail_reading takes is one may_trail admits. It touches maiden-clause parses only, so it costs no frames on ordinary names.

Done when: no behavior change. The case table, all five differential gates and an equivalence fuzz like #609's simplify check (every field, ambiguity and initials string over the corpora plus a random grid) are identical. Frame counts on ordinary names don't rise.

Related: #601, #602, #609.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions