4.7 KiB
lxml XML Declaration Fix for docx Files
Problem (2026-07-01, 劳务派遣协议案)
When lxml serializes XML (via etree.tostring() or python-docx's Document.save()), it outputs:
- Single-quote XML declaration:
<?xml version='1.0' encoding='UTF-8' standalone='yes'?> - LF line endings (
\n)
Original docx files (created by Word/WPS/OnlyOffice) use:
- Double-quote XML declaration:
<?xml version="1.0" encoding="UTF-8" standalone="yes"?> - CRLF line endings (
\r\n)
OnlyOffice cannot open docx files with single-quote XML declarations. The file appears structurally valid (ZIP ok, XML parses fine, python-docx loads it, even x2t can convert it to PDF), but the OnlyOffice web editor refuses to open it.
Affected Files
Only XML files that were re-serialized by lxml are affected. In a typical ContractEditor workflow:
word/document.xml— always re-serialized (main editing target)word/settings.xml— re-serialized if trackRevisions was added/modified
Other XML files (styles.xml, fontTable.xml, theme1.xml, etc.) that were read and written back unchanged via zipfile retain their original format.
Diagnosis
import zipfile
def check_docx_xml_format(docx_path):
"""Check if any XML files have problematic single-quote declarations."""
issues = []
with zipfile.ZipFile(docx_path) as z:
for name in z.namelist():
if name.endswith('.xml') or name.endswith('.rels'):
data = z.read(name).decode('utf-8')
first_line = data.split('\n')[0]
has_single_quotes = "version='1.0'" in first_line
has_lf_only = '\r\n' not in data[:200]
if has_single_quotes or has_lf_only:
issues.append((name, has_single_quotes, has_lf_only))
return issues
Fix Script
import zipfile
import re
import os
import tempfile
def fix_xml_declarations(docx_path, output_path=None):
"""
Fix lxml-serialized XML files inside a docx:
1. Single quotes -> double quotes in XML declaration
2. LF -> CRLF line endings (only if file has no CRLF)
If output_path is None, fixes in-place (via temp file + rename).
"""
if output_path is None:
output_path = docx_path
tmp_fd, tmp_path = tempfile.mkstemp(suffix='.docx')
os.close(tmp_fd)
try:
with zipfile.ZipFile(docx_path, 'r') as zin:
with zipfile.ZipFile(tmp_path, 'w', zipfile.ZIP_DEFLATED) as zout:
for item in zin.infolist():
data = zin.read(item.filename)
if item.filename.endswith('.xml') or item.filename.endswith('.rels'):
text = data.decode('utf-8')
# Fix 1: Single quotes -> double quotes in XML declaration
text = re.sub(
r"<\?xml version='1\.0' encoding='UTF-8' standalone='yes'\?>",
'<?xml version="1.0" encoding="UTF-8" standalone="yes"?>',
text
)
# Fix 2: LF -> CRLF (only if no CRLF present)
if '\r\n' not in text and '\n' in text:
text = text.replace('\n', '\r\n')
data = text.encode('utf-8')
zout.writestr(item, data)
os.replace(tmp_path, output_path)
except:
if os.path.exists(tmp_path):
os.unlink(tmp_path)
raise
# Usage after ContractEditor.save() or manual zipfile write:
# fix_xml_declarations('/tmp/【修】contract.docx')
Integration Points
After ContractEditor.save()
ed = ContractEditor(src)
# ... edits ...
ed.save(output_path)
fix_xml_declarations(output_path) # Must run after every save
After manual zipfile+lxml write
with zipfile.ZipFile(output_path, 'w', zipfile.ZIP_DEFLATED) as zout:
for item in zin.infolist():
# ... write files ...
pass
fix_xml_declarations(output_path) # Must run after ZIP is closed
Key Insight
x2t(OnlyOffice converter CLI) tolerates single-quote declarations — it can convert the "broken" file to PDF successfully- The OnlyOffice web editor (WOPI-based document editing) does NOT tolerate single-quote declarations
python-docx Document()opens the file fine (lxml parses both formats)- Standard validation tools (zipfile.testzip(), etree.fromstring()) all pass
This makes the issue hard to diagnose — everything looks valid except OnlyOffice refuses to open it. The only reliable test is checking the raw bytes of the XML declaration in the ZIP.