Files
hermes-skills/skills/legal/contract-reviewer/references/pre-existing-tracked-changes-handling.md

4.8 KiB

Handling Versioned Files with Pre-Existing Tracked Changes

Problem

Contract files versioned as v1.1, v1.2, etc. typically contain tracked changes (w:ins, w:del) from other authors (e.g. 祺帆, 社区, 俊玮 吴). These must be preserved per review-rules ("已有他人的修订模式保持原样不动").

Critical Issue: python-docx .text vs Actual Content

python-docx paragraph.text does NOT include text from w:ins elements.

This means reading the contract via python-docx shows the original text (before others' modifications), not the accepted state (what the document actually says after accepting all changes). This leads to:

  • Identifying issues that have already been fixed
  • Missing the actual current state of clauses
  • Wrong tracked_replace targets (searching for text that no longer exists in rendered form)

Correct Approach: lxml XML Parsing for Accepted State

def get_para_full_text(p):
    """Get paragraph text in 'accepted' state (includes w:ins, excludes w:del)"""
    ns_w = 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'
    text = ''
    for t in p.iter(f'{{{ns_w}}}t'):
        if t.text:
            # Check if inside a w:del - skip deleted text
            parent = t.getparent()
            in_del = False
            while parent is not None:
                if parent.tag in (f'{{{ns_w}}}del', f'{{{ns_w}}}delText'):
                    in_del = True
                    break
                parent = parent.getparent()
            if not in_del:
                text += t.text
    return text

First step when receiving a versioned file: Always use this function to read the contract's actual state before starting review.

tracked_replace Failures Near Other Authors' Changes

Symptoms

  • tracked_replace returns False (text not found)
  • ValueError: Element is not a child of this node when the target text spans or touches a w:ins from another author

Root Cause

tracked_replace searches for text in regular w:r runs. Text inside w:ins from other authors is in a different element hierarchy — the runs are children of w:ins, not direct children of w:p.

Workarounds

  1. Text not found: The accepted-state text differs from what tracked_replace searches. Re-check what the actual run text says (without ins content) and target that.

  2. Double-period pattern (2026-07-06 赵巷健康云):

    • Original: "调解不成则。" (run ends with period)
    • Other author's ins: "向甲方所在地法院提起诉讼。" (also ends with period)
    • Rendered: "调解不成则向甲方所在地法院提起诉讼。。" (double period)
    • Fix: Manually create w:del for the orphaned original period run:
import copy
from lxml import etree
W = 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'

# Find the trailing period run (last regular w:r in paragraph)
runs = [c for c in paragraph if c.tag == f'{{{W}}}r']
last_run = runs[-1]  # verify it's "。"

# Create w:del tracked change
rev_id = str(editor._next_id())
del_el = etree.Element(f'{{{W}}}del')
del_el.set(f'{{{W}}}id', rev_id)
del_el.set(f'{{{W}}}author', 'WB')
del_el.set(f'{{{W}}}date', editor._revision_date)

del_run = etree.SubElement(del_el, f'{{{W}}}r')
orig_rpr = last_run.find(f'{{{W}}}rPr')
if orig_rpr is not None:
    del_run.append(copy.deepcopy(orig_rpr))

del_text = etree.SubElement(del_run, f'{{{W}}}delText')
del_text.set('{http://www.w3.org/XML/1998/namespace}space', 'preserve')
del_text.text = '。'

# Replace run with del element at same position
idx = list(paragraph).index(last_run)
paragraph.remove(last_run)
paragraph.insert(idx, del_el)

Font Verification: Mixed-Attribute Paragraphs

When a paragraph has runs with mixed rFonts attributes (some eastAsia=None, some eastAsia=宋体), the WB INS inherits from the adjacent run. If that adjacent run has ea=宋体 but the paragraph's first run has ea=None, the font-verify script reports a false-positive EASTASIA MISMATCH.

Fix: Set the WB INS rFonts to match the immediately preceding original run (the one tracked_replace copied from). If the preceding run has ea=None, remove the eastAsia attribute from the WB INS:

rfonts = ins_run_rpr.find(f'{{{W}}}rFonts')
if f'{{{W}}}eastAsia' in rfonts.attrib:
    del rfonts.attrib[f'{{{W}}}eastAsia']

Workflow: Review Checklist for Versioned Files

  1. Read accepted state via lxml (not python-docx .text)
  2. Identify existing change authors — know what's already been modified
  3. Re-assess issues — many may already be resolved by prior revisions
  4. Target tracked_replace carefully — use the raw run text, not accepted-state text
  5. Handle edge cases manually — double periods, text spanning ins boundaries
  6. Verify font — expect false positives in mixed-attribute paragraphs