Files
hermes-skills/skills/legal/contract-editor/references/high-density-tracked-changes-layering.md
T

5.1 KiB

Layering WB Revisions on High-Density Tracked Changes Documents

Problem

When a document already has extensive tracked changes from another author (e.g., 华诚-Z with 170+ INS and 90+ DEL), ContractEditor's tracked_replace frequently fails with ValueError: Element is not a child of this node because the paragraph structure is heavily fragmented with interleaved w:ins/w:del/w:r elements.

Solution: Direct lxml Operations

Strategy

Use zipfile + lxml to directly manipulate the XML instead of ContractEditor library. Three operation types:

1. Append text to existing paragraph end

Find the paragraph, locate the last content element, and append a w:ins after it.

# Find the last non-pPr child element in the paragraph
last_content = None
for child in p:
    if child.tag != f'{WNS}pPr':
        last_content = child

# Create INS element
ins = etree.SubElement(p, f'{WNS}ins')
ins.set(f'{WNS}id', str(next_id))
ins.set(f'{WNS}author', 'WB')
ins.set(f'{WNS}date', '2026-07-02T00:00:00Z')

r = etree.SubElement(ins, f'{WNS}r')
# Clone rPr from nearby run
rpr = get_reference_rpr(p)  # see below
if rpr is not None:
    r.insert(0, copy.deepcopy(rpr))

t = etree.SubElement(r, f'{WNS}t')
t.set('{http://www.w3.org/XML/1998/namespace}space', 'preserve')
t.text = "追加的文字内容"

2. Insert new paragraph (全段INS)

Clone neighboring paragraph's pPr, create a new w:p with all content inside w:ins.

# Clone pPr from reference paragraph
ref_p = paras[target_idx]  # the paragraph after which to insert
new_p = etree.Element(f'{WNS}p')

# Clone pPr
ref_ppr = ref_p.find(f'{WNS}pPr')
if ref_ppr is not None:
    new_p.append(copy.deepcopy(ref_ppr))

# Create INS wrapping all content
ins = etree.SubElement(new_p, f'{WNS}ins')
ins.set(f'{WNS}id', str(next_id))
ins.set(f'{WNS}author', 'WB')
ins.set(f'{WNS}date', '2026-07-02T00:00:00Z')

r = etree.SubElement(ins, f'{WNS}r')
rpr = get_reference_rpr(ref_p)
if rpr is not None:
    r.insert(0, copy.deepcopy(rpr))
t = etree.SubElement(r, f'{WNS}t')
t.set('{http://www.w3.org/XML/1998/namespace}space', 'preserve')
t.text = "新增条款全文"

# Insert after reference paragraph
ref_p.addnext(new_p)

3. Character-level replacement within high-density paragraph

When text to replace is inside an existing w:ins from another author (e.g., 华诚-Z), you need to split that ins element.

# Find the ins element containing target text
for ins_elem in p.findall(f'{WNS}ins'):
    for r in ins_elem.findall(f'{WNS}r'):
        t = r.find(f'{WNS}t')
        if t is not None and t.text and old_text in t.text:
            # Split: keep text before, add WB del+ins for changed part, keep text after
            pos = t.text.index(old_text)
            before = t.text[:pos]
            after = t.text[pos + len(old_text):]
            
            # Modify existing t to keep only 'before'
            t.text = before + after.replace(old_text, new_text)  # simplified
            # Or split into multiple elements...

Getting reference rPr

def get_reference_rpr(p):
    """Get rPr from first non-del run in paragraph, or from 华诚-Z ins"""
    # Try plain runs first
    for r in p.findall(f'{WNS}r'):
        rpr = r.find(f'{WNS}rPr')
        if rpr is not None:
            return rpr
    # Try non-WB ins elements
    for ins in p.findall(f'{WNS}ins'):
        if ins.get(f'{WNS}author') != 'WB':
            for r in ins.findall(f'{WNS}r'):
                rpr = r.find(f'{WNS}rPr')
                if rpr is not None:
                    return rpr
    # Try previous paragraph
    prev = p.getprevious()
    if prev is not None:
        return get_reference_rpr(prev)
    return None

Critical: Post-save sz fix

When INS runs clone rPr from paragraphs that lack explicit w:sz (relying on style inheritance), the INS will render at wrong size. Always run a post-save sweep:

# Determine dominant body sz from neighboring paragraphs
# Then fix all WB INS runs missing sz
for ins in body.iter(f'{WNS}ins'):
    if ins.get(f'{WNS}author') != 'WB':
        continue
    for r in ins.findall(f'{WNS}r'):
        rpr = r.find(f'{WNS}rPr')
        if rpr is not None:
            sz = rpr.find(f'{WNS}sz')
            if sz is None:
                sz = etree.SubElement(rpr, f'{WNS}sz')
                sz.set(f'{WNS}val', dominant_sz)  # e.g., '24' for 12pt
                szCs = etree.SubElement(rpr, f'{WNS}szCs')
                szCs.set(f'{WNS}val', dominant_sz)

Author Unification

After Doro reviews and confirms, unify all authors to WB:

python scripts/unify-author-wb.py input.docx [output.docx]

Lesson Learned (2026-07-02)

  • Doro will edit the files in OnlyOffice after upload. Always download Doro's version before doing further work.
  • "你自己要满意再给我" = self-verify before delivery, don't ask user to check.
  • "认真做" = thoroughness signal. Read full contract text, verify each modification landed correctly.
  • When Doro says "看看是否还有需要调整的" = compare your version vs Doro's, identify what Doro changed, assess if further work needed.
  • Unifying author is a standard final step — use the script, don't hand-code each time.