4.8 KiB
Handling Versioned Files with Pre-Existing Tracked Changes
Problem
Contract files versioned as v1.1, v1.2, etc. typically contain tracked changes (w:ins, w:del) from other authors (e.g. 祺帆, 社区, 俊玮 吴). These must be preserved per review-rules ("已有他人的修订模式保持原样不动").
Critical Issue: python-docx .text vs Actual Content
python-docx paragraph.text does NOT include text from w:ins elements.
This means reading the contract via python-docx shows the original text (before others' modifications), not the accepted state (what the document actually says after accepting all changes). This leads to:
- Identifying issues that have already been fixed
- Missing the actual current state of clauses
- Wrong
tracked_replacetargets (searching for text that no longer exists in rendered form)
Correct Approach: lxml XML Parsing for Accepted State
def get_para_full_text(p):
"""Get paragraph text in 'accepted' state (includes w:ins, excludes w:del)"""
ns_w = 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'
text = ''
for t in p.iter(f'{{{ns_w}}}t'):
if t.text:
# Check if inside a w:del - skip deleted text
parent = t.getparent()
in_del = False
while parent is not None:
if parent.tag in (f'{{{ns_w}}}del', f'{{{ns_w}}}delText'):
in_del = True
break
parent = parent.getparent()
if not in_del:
text += t.text
return text
First step when receiving a versioned file: Always use this function to read the contract's actual state before starting review.
tracked_replace Failures Near Other Authors' Changes
Symptoms
tracked_replacereturnsFalse(text not found)ValueError: Element is not a child of this nodewhen the target text spans or touches aw:insfrom another author
Root Cause
tracked_replace searches for text in regular w:r runs. Text inside w:ins from other authors is in a different element hierarchy — the runs are children of w:ins, not direct children of w:p.
Workarounds
-
Text not found: The accepted-state text differs from what
tracked_replacesearches. Re-check what the actual run text says (without ins content) and target that. -
Double-period pattern (2026-07-06 赵巷健康云):
- Original:
"调解不成则。"(run ends with period) - Other author's ins:
"向甲方所在地法院提起诉讼。"(also ends with period) - Rendered:
"调解不成则向甲方所在地法院提起诉讼。。"(double period) - Fix: Manually create
w:delfor the orphaned original period run:
- Original:
import copy
from lxml import etree
W = 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'
# Find the trailing period run (last regular w:r in paragraph)
runs = [c for c in paragraph if c.tag == f'{{{W}}}r']
last_run = runs[-1] # verify it's "。"
# Create w:del tracked change
rev_id = str(editor._next_id())
del_el = etree.Element(f'{{{W}}}del')
del_el.set(f'{{{W}}}id', rev_id)
del_el.set(f'{{{W}}}author', 'WB')
del_el.set(f'{{{W}}}date', editor._revision_date)
del_run = etree.SubElement(del_el, f'{{{W}}}r')
orig_rpr = last_run.find(f'{{{W}}}rPr')
if orig_rpr is not None:
del_run.append(copy.deepcopy(orig_rpr))
del_text = etree.SubElement(del_run, f'{{{W}}}delText')
del_text.set('{http://www.w3.org/XML/1998/namespace}space', 'preserve')
del_text.text = '。'
# Replace run with del element at same position
idx = list(paragraph).index(last_run)
paragraph.remove(last_run)
paragraph.insert(idx, del_el)
Font Verification: Mixed-Attribute Paragraphs
When a paragraph has runs with mixed rFonts attributes (some eastAsia=None, some eastAsia=宋体), the WB INS inherits from the adjacent run. If that adjacent run has ea=宋体 but the paragraph's first run has ea=None, the font-verify script reports a false-positive EASTASIA MISMATCH.
Fix: Set the WB INS rFonts to match the immediately preceding original run (the one tracked_replace copied from). If the preceding run has ea=None, remove the eastAsia attribute from the WB INS:
rfonts = ins_run_rpr.find(f'{{{W}}}rFonts')
if f'{{{W}}}eastAsia' in rfonts.attrib:
del rfonts.attrib[f'{{{W}}}eastAsia']
Workflow: Review Checklist for Versioned Files
- Read accepted state via lxml (not python-docx
.text) - Identify existing change authors — know what's already been modified
- Re-assess issues — many may already be resolved by prior revisions
- Target tracked_replace carefully — use the raw run text, not accepted-state text
- Handle edge cases manually — double periods, text spanning ins boundaries
- Verify font — expect false positives in mixed-attribute paragraphs