# Multi-Version Creation Pattern ## When to Use When creating multiple versions of the same contract (e.g., 版本1法定安排 vs 版本2反委托保护) or when needing to redo a version from scratch. ## Critical Rule **ALWAYS start from the original source file for each version. Never modify a previously modified version.** ## Step-by-Step Pattern ### 1. Preserve Original Source ```python # First time: copy original to safe location shutil.copy('/path/to/original.docx', '/tmp/original_backup.docx') ``` ### 2. For Each Version, Start Fresh ```python # Always reload from original with zipfile.ZipFile('/tmp/original_backup.docx', 'r') as zin: all_data = {n: zin.read(n) for n in zin.namelist()} doc_xml = all_data['word/document.xml'] root = etree.fromstring(doc_xml) body = root.find(f'{W}body') paras = body.findall(f'{W}p') # Get font template from original rpr_template = None for p in paras: for r in p.findall(f'{W}r'): t = r.find(f'{W}t') if t is not None and t.text and t.text.strip(): rpr_elem = r.find(f'{W}rPr') if rpr_elem is not None: rpr_template = copy.deepcopy(rpr_elem) break if rpr_template: break ``` ### 3. Apply All Modifications in One Pass ```python rev_id = 1000 # Start fresh revision ID counter # Batch all replacements replacements = [ (0, "old text", "new text"), (5, "old text", "new text"), # ... more replacements ] for idx, old_text, new_text in replacements: p = paras[idx] # Clear runs (but NOT comment anchors!) for child in list(p): tag = child.tag.split('}')[-1] if '}' in child.tag else child.tag if tag == 'r': # Only remove regular runs p.remove(child) d, i = make_tracked_replace(old_text, new_text, rpr_template, rev_id) rev_id += 2 p.append(d) p.append(i) # Insert new clauses insert_after = paras[10] for clause_text in new_clauses: new_p = make_ins_paragraph(clause_text, rpr_template, rev_id) rev_id += 1 insert_after.addnext(new_p) insert_after = new_p # Mark deletions (e.g., 承诺书) for idx in range(21, 30): p = paras[idx] text_parts = [] for r in p.findall(f'{W}r'): t = r.find(f'{W}t') if t is not None and t.text: text_parts.append(t.text) full_text = ''.join(text_parts) if not full_text.strip(): continue for child in list(p): tag = child.tag.split('}')[-1] if '}' in child.tag else child.tag if tag == 'r': p.remove(child) del_elem = make_tracked_delete(full_text, rpr_template, rev_id) rev_id += 1 p.append(del_elem) # Add comments LAST (after all structural changes) # Comment anchors are fragile - add them at the end ``` ### 4. Add Comments Carefully ```python # Check if paragraph already has comment anchors existing = p.find(f'{W}commentRangeStart') if existing is None: # Add new comment anchors comment_start = etree.Element(f'{W}commentRangeStart') comment_start.set(f'{W}id', str(comment_id)) p.insert(0, comment_start) comment_end = etree.Element(f'{W}commentRangeEnd') comment_end.set(f'{W}id', str(comment_id)) p.append(comment_end) comment_ref_run = etree.SubElement(p, f'{W}r') comment_ref = etree.SubElement(comment_ref_run, f'{W}commentReference') comment_ref.set(f'{W}id', str(comment_id)) # Update comments.xml if 'word/comments.xml' in all_data: croot = etree.fromstring(all_data['word/comments.xml']) else: croot = etree.Element(f'{W}comments', nsmap={'w': W_NS}) # Add or update comment new_comment = etree.SubElement(croot, f'{W}comment') new_comment.set(f'{W}id', str(comment_id)) new_comment.set(f'{W}author', author) new_comment.set(f'{W}date', datetime.now().isoformat()) p = etree.SubElement(new_comment, f'{W}p') r = etree.SubElement(p, f'{W}r') t = etree.SubElement(r, f'{W}t') t.set(XML_SPACE, 'preserve') t.text = comment_text ``` ### 5. Save and Verify ```python all_data['word/document.xml'] = etree.tostring(root, xml_declaration=True, encoding='UTF-8', standalone=True) all_data['word/comments.xml'] = etree.tostring(croot, xml_declaration=True, encoding='UTF-8', standalone=True) with zipfile.ZipFile('/tmp/version1.docx', 'w', zipfile.ZIP_DEFLATED) as zout: for name, data in all_data.items(): zout.writestr(name, data) # Verify immediately doc = Document('/tmp/version1.docx') print(f"OK: {len(doc.paragraphs)} paragraphs") ``` ## Common Pitfalls ### ❌ Don't Do This ```python # WRONG: Modifying v1 to create v2 shutil.copy('/tmp/v1.docx', '/tmp/v2.docx') with zipfile.ZipFile('/tmp/v2.docx', 'r') as zin: # ... load v1's modified structure # This will have v1's tracked changes, comments, etc. ``` ### ❌ Don't Clear Everything When Modifying ```python # WRONG: Clears comment anchors too! for child in list(p): if child.tag not in (f'{W}pPr',): p.remove(child) # Removes commentRangeStart/End! ``` ### ✅ Do This Instead ```python # RIGHT: Only clear regular runs, preserve comment anchors for child in list(p): tag = child.tag.split('}')[-1] if '}' in child.tag else child.tag if tag == 'r': # Only regular runs p.remove(child) # commentRangeStart, commentRangeEnd are preserved ``` ## Preserving Original Comments When the original document has comments (e.g., Alice, 法务, 杜律), the workflow must: 1. Read original comments.xml to get all comment IDs and content 2. Check which paragraphs have comment anchors (commentRangeStart/End) 3. When clearing runs, preserve comment anchors (they're not `w:r` elements) 4. Add new comments with NEW IDs (don't reuse original IDs) 5. Original comments remain unchanged in comments.xml ## Preserving Third-Party Tracked Changes (2026-07-01 华诚-Z案) When a contract file contains tracked changes from someone other than WB (e.g., 华诚-Z, Crystall, or any third-party reviewer), **those files must never be overwritten**. The tracked changes represent real editorial work that cannot be reconstructed from session notes alone. ### Backup Protocol ```python import shutil from datetime import datetime # BEFORE any modification to a file with third-party tracked changes: ts = datetime.now().strftime('%Y%m%d_%H%M%S') backup_path = f'/tmp/{filename}.bak_{ts}' shutil.copy(source_path, backup_path) print(f"Backed up to {backup_path}") ``` ### Detection: Does This File Have Third-Party Changes? ```python import zipfile, re with zipfile.ZipFile(filepath) as z: content = z.read('word/document.xml').decode('utf-8', errors='ignore') authors = set(re.findall(r'w:author="([^"]+)"', content)) third_party = authors - {'WB'} if third_party: print(f"⚠️ Third-party authors found: {third_party} — BACKUP REQUIRED") ``` ### Multi-Version with Third-Party Edits When creating v1 and v2 from a file that has both original content AND third-party edits: 1. **Backup the file with third-party edits** (e.g., `华诚-Z版.bak_20260701`) 2. **Backup the pristine original** (no tracked changes at all) 3. **For each version**: start from the pristine original, then layer on: - WB's own tracked changes - Third-party's tracked changes (with author renamed to WB) 4. **Never modify the backup files** — they are your insurance ### What Was Lost (华诚-Z案) - 华诚-Z made 3 tracked changes in OnlyOffice: 第六条 (removed specific legal citations), 第七条 (simplified correction process), 第八条 (added employee return placement clause) - These intermediate files in /tmp/ were overwritten during v1/v2 creation - Only session notes preserved the *content* of changes, not the actual tracked change markup (ids, timestamps, exact XML positions) - **Recovery was impossible** — Doro had to accept reconstructed versions ## Real Example from This Session - Original: 4 comments (Alice×2, 法务, 杜律) - Version 1: 5 comments (original 4 + WB legal risk) - Version 2: 5 comments (original 4 + WB legal risk) Both versions created independently from original, each with their own WB comment (different content for each version).