Files
hermes-skills/skills/legal/contract-editor/references/systematic-file-recovery-lost-authors.md

4.8 KiB

Systematic File Recovery for Lost Author Markers

When intermediate files have been overwritten during iterative editing (e.g., multiple versions of a contract revision), and you need to find a specific version that contains tracked changes by a particular author (e.g., "华诚-Z"), use this systematic scan approach.

Scenario

  • You made multiple intermediate files (v1, v2, v3...) in /tmp/ during contract editing
  • You overwrote files, losing the version with a specific author's tracked changes
  • You need to find ANY surviving file that still has that author's w:author attribute

Recovery Technique

Step 1: List all candidate files

Find all .docx files in the working directory that are newer than the original source file:

find /tmp -name '*.docx' -newer /tmp/original_file.docx 2>/dev/null | sort

Step 2: Check each file for the target author

import zipfile, re, os
from datetime import datetime

target_author = '华诚-Z'  # or whatever author you're looking for

files = [
    "/tmp/v1_clean.docx",
    "/tmp/v1_final.docx",
    # ... list all candidate files from Step 1
]

for f in files:
    if not os.path.exists(f):
        continue
    try:
        z = zipfile.ZipFile(f)
        content = z.read('word/document.xml').decode('utf-8', 'ignore')
        authors = set(re.findall(r'w:author="([^"]+)"', content))
        mt = datetime.fromtimestamp(os.path.getmtime(f)).strftime('%m-%d %H:%M')
        has_target = target_author in authors
        marker = '★' if has_target else ' '
        print(f"{marker} {os.path.basename(f):35s} {mt}  authors={sorted(authors)}")
        z.close()
    except Exception as e:
        print(f"  ERROR {f}: {e}")

Step 3: Extract the target author's changes

Once you find the file with the target author, extract their specific tracked changes:

WNS = '{http://schemas.openxmlformats.org/wordprocessingml/2006/main}'

z = zipfile.ZipFile('/tmp/file_with_target_author.docx')
with z.open('word/document.xml') as f:
    tree = etree.parse(f)
root = tree.getroot()
body = root.find(f'{WNS}body')
paras = body.findall(f'{WNS}p')

for i, p in enumerate(paras):
    has_target = False
    parts = []
    for child in p:
        tag = etree.QName(child.tag).localname
        if tag == 'r':
            t = child.find(f'{WNS}t')
            if t is not None and t.text:
                parts.append(('RUN', t.text, None))
        elif tag == 'ins':
            author = child.get(f'{WNS}author', '?')
            if target_author in author:
                has_target = True
            ins_texts = []
            for r in child.findall(f'{WNS}r'):
                t = r.find(f'{WNS}t')
                if t is not None and t.text:
                    ins_texts.append(t.text)
            if ins_texts:
                parts.append(('INS', ''.join(ins_texts), author))
        elif tag == 'del':
            author = child.get(f'{WNS}author', '?')
            if target_author in author:
                has_target = True
            del_texts = []
            for r in child.findall(f'{WNS}r'):
                t = r.find(f'{WNS}delText')
                if t is not None and t.text:
                    del_texts.append(t.text)
            if del_texts:
                parts.append(('DEL', ''.join(del_texts), author))
    
    if has_target:
        print(f"\n★ P{i}:")
        for kind, text, author in parts:
            if kind == 'RUN':
                print(f"  [原文] {repr(text)}")
            else:
                print(f"  [{kind} by {author}] {repr(text)}")

z.close()

Empirical Case (2026-07-01 反委托代发工资协议)

  • Made ~15 intermediate files in /tmp/ during iterative editing
  • Overwrote all files, changing all w:author attributes to "WB"
  • User (Doro) demanded recovery of 华诚-Z's tracked changes
  • Systematic scan found /tmp/v1_doro_updated.docx with authors=['WB', '华诚-Z']
  • Extracted 华诚-Z's 3 specific changes:
    • P6: INS "等" (between WB's "《劳务派遣暂行规定》" and "规定,")
    • P12: INS "退回派遣员工" + INS "由乙方依法自行安置处理,与甲方无涉。"

Key Pitfalls

  1. Don't assume the file is gone — check ALL intermediate files, not just the ones you expect
  2. Check timestamps — the file you need might be an early intermediate, not the latest
  3. Use w:author attribute — this is the definitive marker, not file content or naming
  4. Comments may also be lost — the recovered file might have lost some original comments (see references/comment-restoration-from-original.md)

Prevention (Better Than Recovery)

The existing skill already covers this, but worth repeating: 改前必备份 — before modifying any file with third-party tracked changes, save a timestamped backup:

cp file_with_third_party.docx file_with_third_party.bak_$(date +%Y%m%d_%H%M%S).docx