feat: export core Hermes skills
This commit is contained in:
@@ -0,0 +1,598 @@
|
||||
---
|
||||
name: contract-review-general
|
||||
description: 合同法律审查通用工作流程。适用于各类商事合同的审查,包括主体背景调查、条款审查、审查意见输出。与Maggie共建的标准流程,持续优化迭代中。
|
||||
version: 0.1.0
|
||||
tags: [合同审查, 法律, 尽职调查, 背景调查]
|
||||
---
|
||||
|
||||
# 合同法律审查通用工作流程
|
||||
|
||||
> 与Maggie共建的标准化合同审查流程,适用于各类商事合同。
|
||||
> 流程文档同步存放在:~/.hermes/shared/合同审查交付件/合同审查工作流程.md
|
||||
|
||||
## 核心原则
|
||||
|
||||
1. **信息准确性**:所有检索结果必须来自官方或权威第三方来源,标明出处;无法获取权威来源的如实告知,绝不拼凑
|
||||
2. **先主体后条款**:开始审查条款之前,必须先完成交易双方的背景调查。原因:主体资格、关联关系、合规问题是地基,条款审查是在地基上盖的房子。如果交易本身存在合法性障碍(如关联交易未披露、同业竞争违规、主体不适格),条款再完美也没有意义。背景调查决定了审查的立场、侧重点和建议方向
|
||||
3. **整体审查**(铁律,2026-06-17 Maggie):每个条款当作整体审查——识别条款全部要件(适用主体/情形列举/权利义务/例外但书/兜底救济)后再定性,严禁摘单句断章取义;整个合同当作整体审查——条款间交叉比对、关联呼应、前后逻辑自洽、交叉引用核对指向、定义简称全文贯通;逐条逐款不跳不漏(含正文/附件/补充协议/签署页/表格);内容与逻辑并重。详见 contract-reviewer skill「整体审查原则」节
|
||||
4. **两遍独立审查**(手动审查铁律):法律条款审查和文字校对必须分两遍独立进行。第一遍按审查清单逐项检查法律条款(违约责任、管辖、保密、知识产权、转包等);第二遍逐段通读全文做文字校对(笔误、漏字、重复词语、年份错误、称谓不一致、模板残留如"药品"应为"医疗器械"等)。⚠️ 2026-06-12教训(数字健康城区运维V3):只做一遍遗漏5处文字问题,被Doro指出"认真点"后第二遍才全部发现。法律审查和文字校对是两个独立的认知任务,不能合并做
|
||||
5. **保密义务**:合同审查的所有信息严格保密,仅限授权人员知悉,未经委托方书面同意不得向第三方披露
|
||||
6. **逐步推进**:与Maggie一步一步确认,不跳步骤
|
||||
|
||||
## 第一阶段:交易主体背景调查
|
||||
|
||||
### Step 1: 提取合同中的主体信息
|
||||
- 从合同文本中提取各方的名称、统一社会信用代码、法定代表人、注册地址
|
||||
- 核对合同中载明的信息与工商登记是否一致
|
||||
|
||||
### Step 2: 工商基本信息检索
|
||||
对每一方主体检索:
|
||||
- 公司名称、统一社会信用代码
|
||||
- 法定代表人
|
||||
- 注册资本(认缴/实缴)
|
||||
- 成立日期
|
||||
- 经营范围(是否涵盖合同涉及的业务)
|
||||
- 证照情况
|
||||
|
||||
**权威来源**:国家企业信用信息公示系统(gsxt.gov.cn)、企查查(qcc.com)、天眼查(tianyancha.com)
|
||||
|
||||
### Step 3: 股权结构检索
|
||||
- 股东构成及持股比例
|
||||
- 股权变更历史(频繁变更需警惕)
|
||||
- 实际控制人(穿透股权层级)
|
||||
- 上市公司关联
|
||||
|
||||
### Step 4: 关联关系排查
|
||||
- 甲乙双方之间是否存在关联关系(交叉持股、共同股东、共同实控人)
|
||||
- 对外投资企业清单
|
||||
- 关联交易嫌疑排查
|
||||
|
||||
### Step 5: 法律风险排查
|
||||
- 诉讼/仲裁记录 → 中国裁判文书网(wenshu.court.gov.cn)
|
||||
- 被执行人/失信被执行人 → 中国执行信息公开网
|
||||
- 行政处罚 → 信用中国(creditchina.gov.cn)
|
||||
- 经营异常名录 → 国家企信系统
|
||||
|
||||
### Step 6: 履约能力评估
|
||||
- 注册资本 vs 合同金额
|
||||
- 经营范围 vs 合同内容
|
||||
- 企业存续时间
|
||||
- 财务状况(上市公司查年报)
|
||||
|
||||
## 第二阶段:合同条款审查
|
||||
|
||||
(待补充 — 随案例实践逐步完善)
|
||||
|
||||
## 手动审查流程(Workflow超时降级方案)
|
||||
|
||||
当uwf workflow的reviewer阶段连续超时(通常3次以上Timeout waiting for response),不要继续等待重试。按以下流程手动完成审查。
|
||||
|
||||
## 触发条件
|
||||
合同法律审查任务(非卫生中心系列),包括商事合同、劳动合同、服务协议等。
|
||||
|
||||
## 审查立场铁律(Doro 2026-07-01 严厉纠正)
|
||||
|
||||
### 1. 站客户立场,不站法律合规立场
|
||||
- 客户的商业安排是客户的决策,审查方无权否定
|
||||
- 例:客户选择"甲方代发工资"→保留代发安排,在代发框架下最大化保护客户
|
||||
- ❌ 错误:把"甲方代发"改成"乙方直接发"(=否定客户的商业决策)
|
||||
- ✅ 正确:保留代发,加批注提示风险,在代发框架下加保护条款
|
||||
|
||||
### 2. 不做法律价值判断
|
||||
- ❌ "强烈建议采用版本1"
|
||||
- ✅ 呈现两个版本的风险差异,由律师和客户决定
|
||||
|
||||
### 3. 不填写合同空白内容
|
||||
- 合同中空白处(金额、月数、日期等)是当事人商业条款
|
||||
- 审查方无权擅自填入任何数值
|
||||
- ❌ 质量保证期空白直接填"12个月"(无任何依据)
|
||||
- ✅ 批注提示"请根据招标文件/投标文件约定填写具体月数"
|
||||
|
||||
### 4. 不编造法律依据
|
||||
- 没有找到权威来源时如实说"未找到"
|
||||
- ❌ 编一个"司法实践中认定"糊弄
|
||||
- ✅ "未检索到直接裁判文书支持,以下为相关法条/学理分析,供参考"
|
||||
|
||||
### 5. 修订文本质量
|
||||
- 修订后的条款必须通读上下文,确保语句通顺、逻辑自洽
|
||||
- ❌ 机械替换文字后不管上下文是否衔接
|
||||
- ✅ 每次修订后完整读一遍该段落,确认作为独立语句读得通
|
||||
- Doro指示手动干预(如"你手动干吧")
|
||||
|
||||
### 操作步骤
|
||||
1. **杀掉卡住的进程**:`process(action='kill', session_id=...)`
|
||||
2. **读取合同**:用python-docx提取全文段落+表格,识别甲乙方和顾问单位
|
||||
3. **加载审查规则**:读取`/tmp/contract-review/rules/review-rules-root.md`(及目录特定rules)
|
||||
4. **按审查清单逐项检查**:合并reviewer+editor角色,一次性完成分析+修订——逐条对照review-rules.md中的审查清单(主体、违约责任、争议解决、保密/数据、知识产权、第三方侵权、转包、价款、服务成果使用权、条款逻辑),同时检查笔误和语法错误
|
||||
5. **执行修订**:修订模式(author=WB)。可用ContractEditor库或直接XML操作,二者均可。直接XML时注意:
|
||||
- `w:ins`/`w:del`必须有id、author、date属性
|
||||
- INS run的rPr至少包含`rFonts hint="eastAsia"`
|
||||
- 新增条款段落用对应样式(如style val='3'对应Heading 2),pPr中包含ins rPr标记
|
||||
- tracked_replace需正确处理跨run的文本匹配
|
||||
6. **终审验证**(不可跳过,同workflow终审标准):
|
||||
- XML级别字体检查(WB INS的rFonts + sz与原文一致)
|
||||
- **heading INS sz检查**:新增heading段落(style=2/3等)的INS run sz必须匹配styles.xml中对应样式的sz值,不能用body正文sz(2026-06-15教训:heading 2样式sz=24 vs INS实际sz=21导致标题偏小)
|
||||
- **交叉引用偏移检查**:用`第\s*\d+\s*条`搜索全文硬编码引用,核对新增heading条款后编号是否顺移但引用未更新(2026-06-15教训:新增2个heading条款后"第10条""第13条"引用未更新)\n - **手动编号与自动编号冲突检查(2026-06-15 端午节合同教训)**:新增条款若用手动键入的\"N、\"编号(如INS run里直接打\"6、转包限制\"),必须确认前一条不是用numbering.xml自动编号渲染的。陷阱:某条用`<w:numPr><w:numId w:val=\"3\"/>`且numbering.xml中该num定义`start=6 lvlText=%1、`,正文XML里看不到\"6、\"字样但渲染时自动显示\"6、\"。如果紧接着的INS新增条款手打\"6、\",就会出现两个6。检测方法:\n ```python\n # 1. 读numbering.xml确认每个numId的start值和lvlText格式\n # 2. 在document.xml中找出所有带numPr的段落,推算它们渲染出的实际编号\n # 3. 把自动编号段落和手动\"N、\"段落合并成一条完整序列,检查是否连续无重复\n ```\n 修复:手动编号顺延(6→7),且其后所有手动编号条款全部跟着+1(否则会出现两个7)。本案修复了7、转包→8、违约→9、争议三条。纯文本未跨run时可直接`xml.replace('6、转包限制:','7、转包限制:')`,替换前后用count断言各为1,确保精准。修订模式w:ins包裹要保持不变
|
||||
- 修订内容文本验证(接受修订后的文本是否正确)
|
||||
- 删除/插入计数确认
|
||||
7. **交付**:上传Nextcloud → scan → 清OnlyOffice缓存 → 更新tracker → 更新xlsx → 私信通知Doro
|
||||
|
||||
### 手动审查必须分两遍(2026-06-12教训)
|
||||
**第一遍**只做法律实质审查(按review-rules.md审查清单逐项),**第二遍**做文字校对(逐字通读,按contract-reviewer skill的「文字校对」10项必检清单)。两遍完成后合并issue清单,一次性修订。
|
||||
|
||||
❌ **错误做法**:边读边改、读完法律问题就动手 → 遗漏文字级错误(双重用词、笔误、称谓不一致)
|
||||
✅ **正确做法**:法律审查完 → 放下法律思维 → 逐字通读做校对 → 合并后再动手
|
||||
|
||||
2026-06-12数字健康城区运维合同教训:第一遍找到8个法律问题直接修订交付,Doro说"但你认真点",第二遍逐字通读补出5个文字问题(双"进行"、"2027月"、"由于因"、"买卖双方"、"如果或证实"),不得不从原文重做全部13处修订。
|
||||
|
||||
### 与workflow审查的区别
|
||||
- 手动审查是reviewer+editor合一,但**验证标准不降低**
|
||||
- 效率更高(一次性完成,无角色间传递开销),适合workflow反复超时的复杂合同
|
||||
- 缺点:没有独立reviewer复核环节,需要更仔细的自检——**必须分两遍**(法律实质+文字校对),不可合并
|
||||
|
||||
## Workflow `needs_clarification` 处理(乙方空白/无法识别顾问单位)
|
||||
|
||||
当classifier返回 `$status: needs_clarification`(通常因为合同某方名称空白、无法判断顾问单位):
|
||||
|
||||
### 第一步:报告Doro并询问
|
||||
直接告知Doro合同双方信息,请Doro确认乙方/顾问单位。
|
||||
|
||||
### 第二步:Doro不知道时——追溯文件来源
|
||||
如果Doro回复"合同名称无法确认该问谁"(2026-06-12教训),**不要再追问Doro**,自行追溯文件发送人:
|
||||
|
||||
1. **查企微缓存**:`ls ~/.hermes/cache/documents/*关键词*` → 确认文件到达时间
|
||||
2. **查session记录**:`session_search(query="文件名关键词", sort="newest")` → 找到文件接收的session,识别发送人
|
||||
3. **查Nextcloud activity**(如果上面找不到):
|
||||
```sql
|
||||
-- 容器: nextcloud-db-1, 用户: nextcloud, 密码见docker-compose.yml
|
||||
SELECT FROM_UNIXTIME(timestamp), user, subject, file
|
||||
FROM oc_activity WHERE file LIKE '%关键词%' ORDER BY timestamp DESC;
|
||||
```
|
||||
4. 确认发送人后,在群里@发送人询问乙方是哪个顾问单位
|
||||
|
||||
### 第三步:获得确认后继续workflow
|
||||
将乙方信息告知后,可以:
|
||||
- 取消当前thread(`uwf thread cancel <id>`)
|
||||
- 以明确prompt重新启动(`uwf thread start review-contract -p "审查合同...,站XX立场"`)
|
||||
- 或手动按已确认的顾问单位直接审查
|
||||
|
||||
## 社区卫生服务中心目录路由(手动交付必检)
|
||||
|
||||
部分顾问单位(朱家角镇社区卫生服务中心、练塘镇社区卫生服务中心、青浦区爱国卫生和健康促进指导中心等)在 Nextcloud 下有**独立的子目录**,而非放在通用的 `Doro合同审查任务/待审查/` 和 `Doro合同审查任务/任务交付/`:
|
||||
|
||||
```
|
||||
Doro合同审查任务/
|
||||
├── 待审查/ ← 通用待审查
|
||||
├── 任务交付/ ← 通用任务交付
|
||||
├── 朱家角镇社区卫生服务中心/
|
||||
│ ├── 待审查/ ← 朱家角专属
|
||||
│ ├── 任务交付/ ← 朱家角交付物
|
||||
│ └── review-rules.md ← 朱家角特殊规则(含审查意见模板要求)
|
||||
├── 练塘镇社区卫生服务中心/
|
||||
│ └── ...
|
||||
└── 青浦区爱国卫生和健康促进指导中心/
|
||||
└── ...
|
||||
```
|
||||
|
||||
**Workflow deliverer 只把文件产出到 `/tmp/contract-review/`,不会自动路由到这些子目录。** 交付时必须:
|
||||
1. 检查该顾问单位是否有独立子目录:`docker exec nextcloud-nextcloud-1 ls "/var/www/html/data/doro/files/Doro合同审查任务/" | grep 关键词`
|
||||
2. 有 → 复制到该子目录的 `任务交付/`(deliverer 通常投递到通用任务交付/,需手动路由):
|
||||
```bash
|
||||
docker exec nextcloud-nextcloud-1 cp "/var/www/html/data/doro/files/Doro合同审查任务/任务交付/【修】XX.docx" "/var/www/html/data/doro/files/Doro合同审查任务/朱家角镇社区卫生服务中心/任务交付/"
|
||||
docker exec nextcloud-nextcloud-1 cp "/var/www/html/data/doro/files/Doro合同审查任务/任务交付/【审】XX.docx" "/var/www/html/data/doro/files/Doro合同审查任务/朱家角镇社区卫生服务中心/任务交付/"
|
||||
docker exec nextcloud-nextcloud-1 chown www-data:www-data "/var/www/html/data/doro/files/Doro合同审查任务/朱家角镇社区卫生服务中心/任务交付/【修】XX.docx" "/var/www/html/data/doro/files/Doro合同审查任务/朱家角镇社区卫生服务中心/任务交付/【审】XX.docx"
|
||||
docker exec nextcloud-nextcloud-1 php occ files:scan doro
|
||||
```
|
||||
3. ⚠️ 2026-07-02 教训:恭兴合同和肃言合同都被 deliverer 投递到通用 `任务交付/`,需要手动复制到朱家角子目录
|
||||
|
||||
**特殊交付物:审查意见文档** — 部分卫生中心的 `review-rules.md` 要求生成独立的审查意见表格(条文|原文|修订后),使用指定模板(如 `~/.hermes/shared/模版库/朱家角 审查意见【模板】.docx`)。Workflow 的 deliverer 通常能正确生成,但交付路由仍需手动完成。审查意见文件是主合同的 companion,不需要单独的 tracker/xlsx 条目。
|
||||
|
||||
## .doc 文件转换(非 .docx 格式)
|
||||
|
||||
邱律师有时发送 `.doc` 格式文件(Word 97-2003),workflow 只能处理 `.docx`。
|
||||
|
||||
### 转换步骤
|
||||
```bash
|
||||
# 1. 复制到 /tmp/contract-review/ 并转换
|
||||
cp source.doc /tmp/contract-review/
|
||||
libreoffice --headless --convert-to docx /tmp/contract-review/source.doc --outdir /tmp/contract-review/
|
||||
# 输出: source.docx
|
||||
|
||||
# 2. 启动 workflow 时注明原始格式
|
||||
uwf thread start review-contract -p "审查合同:source.doc(已转换为source.docx),顾问单位:XXX"
|
||||
```
|
||||
|
||||
### 注意事项
|
||||
- LibreOffice 转换后文件名自动为 `原名.docx`(去掉 `.doc` 后缀加 `.docx`)
|
||||
- 转换后 .doc 和 .docx 都保留在 /tmp/contract-review/,workflow 会识别 .docx
|
||||
- 交付物文件名为 `【修】原名.docx`(不带 .doc 后缀)
|
||||
- 上传 Nextcloud 待审查目录时保留原始 .doc 格式(不转换)
|
||||
- 2026-07-02 恭兴合同.doc / 肃言合同.doc 均使用此流程成功审查
|
||||
|
||||
### 常见错误
|
||||
- `Error: source file could not be loaded`:文件名含特殊前缀(如 `doc_xxx_`)。确保复制到 /tmp/ 时用干净文件名
|
||||
- `failed to launch javaldx`:正常警告,不影响转换
|
||||
|
||||
## 大文件预处理(图片切割+缝合)
|
||||
|
||||
超大docx文件(常见于含招标公告截图、中标通知书等附件的合同)需要在审查前切割,审查后缝合:
|
||||
|
||||
### 切割(classifier阶段,审查前)
|
||||
```bash
|
||||
python3 ~/hc-contract-editor/scripts/contract_preprocess.py preprocess <合同文件> /tmp/contract-review/
|
||||
```
|
||||
- `action=none`:正常文件(文字为主),直接审查
|
||||
- `action=cut`:末尾有纯图片附件 → 输出`_stripped.docx`(去图片)+ `cutdata.json`(图片数据)
|
||||
- `action=ocr`:全文是扫描件图片,需OCR → 标记后人工处理
|
||||
|
||||
⚠️ 切割只移除末尾纯图片附件,不碰合同正文和文字附件。
|
||||
|
||||
### 缝合(deliverer阶段,交付前)
|
||||
```bash
|
||||
# 检查有没有cutdata
|
||||
ls /tmp/contract-review/*_cutdata.json
|
||||
# 有 → 还原
|
||||
python3 ~/hc-contract-editor/scripts/contract_preprocess.py restore <修订文件> <cutdata.json> <输出文件>
|
||||
# 质检:zipfile检查media文件数 vs cutdata记录一致
|
||||
```
|
||||
|
||||
### 教训(华新镇签约服务费合同,1.7MB→45KB)
|
||||
- 文件82%是图片(5张,1.6MB),全在末尾附件
|
||||
- 不切割直接审查:reviewer反复timeout(文件太大)
|
||||
- 切割后审查+还原:回注后媒体文件逐字节一致
|
||||
- 剥离脚本位置:`~/.hermes/scripts/strip_docx_media.py`
|
||||
|
||||
## 合同审查台账(合同审查清单.xlsx)
|
||||
|
||||
位置:Nextcloud `Doro合同审查任务/合同审查清单.xlsx`(根目录,非任务交付/子目录)
|
||||
- 5列:序号、交付日期、顾问单位名称(全称)、合同名称、文件名
|
||||
- 这是Doro可以随时打开看的正式工作记录
|
||||
|
||||
### 更新规则
|
||||
1. 终审通过后、通知Doro之前 → 新增一行到xlsx
|
||||
2. 顾问单位名称从classifier output的`our_party_name`字段获取,不自己解析合同文本
|
||||
3. 上传Nextcloud覆盖旧版 → 清OnlyOffice缓存
|
||||
4. xlsx记录永久保留,pass后删的只是合同文件本身
|
||||
5. ⚠️ Doro收到通知时xlsx必须已经是最新的
|
||||
6. Doro说pass/completed时,一次性更新所有当批合同到xlsx(不要一份一份更新上传)
|
||||
|
||||
### openpyxl样式复制注意事项
|
||||
复制已有行的样式到新行时,**不能直接赋值border**:
|
||||
```python
|
||||
# ❌ 报错 TypeError: unhashable type: 'StyleProxy'
|
||||
new_cell.border = ref_cell.border
|
||||
|
||||
# ✅ 正确:用copy()或重建Side对象
|
||||
from copy import copy
|
||||
new_cell.font = copy(ref_cell.font)
|
||||
new_cell.alignment = copy(ref_cell.alignment)
|
||||
new_cell.border = Border(
|
||||
left=Side(style=ref_border.left.style, color=ref_border.left.color),
|
||||
right=Side(style=ref_border.right.style, color=ref_border.right.color),
|
||||
top=Side(style=ref_border.top.style, color=ref_border.top.color),
|
||||
bottom=Side(style=ref_border.bottom.style, color=ref_border.bottom.color),
|
||||
)
|
||||
```
|
||||
|
||||
### 内部tracker
|
||||
位置:`~/.hermes/data/contract-tracker.json`
|
||||
- 记录每份合同的完整生命周期:queued→running→ready→passed→cleaned
|
||||
- 用于runner脚本流程控制,不是给Doro看的
|
||||
|
||||
## 第三阶段:审查意见输出
|
||||
|
||||
### 输出方式一:批注+修订模式(推荐用于正式交付)
|
||||
|
||||
当律师要求"在文件里批注审核意见+修订条款"时,直接操作docx文件:
|
||||
|
||||
**批注(Comments)**:
|
||||
- ⚠️ ContractEditor 库**没有** add_comment 方法。批注必须用独立的 zipfile+lxml 函数实现。
|
||||
- ContractEditor.tracked_replace() 也**不接受** author 参数(author 固定为 'WB',在 __init__ 中设置)。
|
||||
- 批注实现模式:先用 ContractEditor 做 tracked_replace 并 save,然后用独立函数 add_comments_to_docx() 在已保存的文件上追加批注。
|
||||
- 每条批注需要三个位置协调:comments.xml(定义)+ document.xml(锚点 commentRangeStart/End + commentReference)+ rels/content_types(注册)
|
||||
- comments.xml 根元素用 `etree.Element(qn('comments'), nsmap={'w': W, 'r': R_NS})`(不能用 .set('xmlns:w', ...),lxml 会报错)
|
||||
- Content_Types 需添加 Override: `/word/comments.xml` → `application/vnd.openxmlformats-officedocument.wordprocessingml.comments+xml`
|
||||
- document.xml.rels 需添加 Relationship: Type=`.../relationships/comments`, Target=`comments.xml`
|
||||
- 参考实现见 `/tmp/review_all.py` 的 add_comments_to_docx() 函数(经实测可用)
|
||||
- 技术细节参见 word-docx skill 的 "Adding Comments Programmatically" 章节
|
||||
|
||||
**修订痕迹(Tracked Changes)**:
|
||||
- 可修改的条款直接用 w:del + w:ins 修订
|
||||
- 仅做批注建议的条款只加批注不改正文
|
||||
- 技术细节参见 word-docx skill 的 "Surgical Tracked-Change Edits" 章节
|
||||
|
||||
**交付要求**:
|
||||
- 一份文件同时包含批注(解释原因/建议)和修订(具体修改),律师打开即可审阅
|
||||
- 批注中给出修改理由和法律依据,修订中直接改条款文字
|
||||
|
||||
### 输出方式二:独立审查意见书(适用于分析报告)
|
||||
|
||||
当不直接修改合同、仅出具审查意见时:
|
||||
- 按条款逐项列出问题、风险等级、修改建议
|
||||
- 适合先发给委托人讨论,确认后再改合同
|
||||
|
||||
### 批量独立审查(刘婷律师等批量交办模式)
|
||||
|
||||
当律师一次交办多份合同审查(非workflow),流程如下:
|
||||
|
||||
**工具链**:ContractEditor(tracked_replace) + add_comments_to_docx(zipfile+lxml手写comments.xml)
|
||||
|
||||
**ContractEditor注意事项**:
|
||||
- `tracked_replace(old_text, new_text)` — 不接受author参数,默认author='WB'
|
||||
- 不支持 `add_comment()` 方法——必须用独立的zipfile+lxml函数操作comments.xml
|
||||
- 先做tracked_replace → save → 再调add_comments函数(因为save会重写document.xml)
|
||||
|
||||
**add_comments_to_docx函数关键点**:
|
||||
- comments.xml的根元素用 `etree.Element(qn('comments'), nsmap={'w': W, 'r': R_NS})`,不能用set('xmlns:w', ...)
|
||||
- commentRangeStart插入位置用 `p.insert(0, crs)` 而非 `p.insert(list(p).index(runs[0]), crs)`——后者在tracked_replace修改过的段落中会报ValueError(runs[0]不在p的直接子元素中)
|
||||
- Content_Types.xml和document.xml.rels都需要注册comments关系
|
||||
- 批注author统一用"WB"
|
||||
|
||||
**交付方式**:
|
||||
- 上传Nextcloud对应目录(如 `Doro合同审查任务/刘婷律师合同审查-YYYYMMDD/`)
|
||||
- 通知Doro(企微私信或send_message)
|
||||
- wecom_dm.py不支持发文件(无--file参数),发文件超时时改用Nextcloud+通知
|
||||
|
||||
**政府采购合同模板常见问题**(朱家角社区医院系列):
|
||||
- 联系人和电话字段写反("电话:柳老师"/"联系人:021-xxx")
|
||||
- 质量保证金、履约保证金金额未填("/")
|
||||
- 调解机构空白("可以向___提请调解")
|
||||
- 付款方式一次性100%对甲方不利,建议分期
|
||||
- 仲裁条款需确认是否符合甲方意愿
|
||||
|
||||
### 独立审查(非workflow模式)
|
||||
|
||||
适用场景:非Doro的合同审查任务(如莎莎、Maggie、刘婷律师直接交办),不走uwf workflow。
|
||||
|
||||
**批量审查流程(刘婷律师模式)**:
|
||||
刘婷律师经常一次发送多份合同要求批量审查。处理方式:
|
||||
1. 收齐所有文件后确认审查立场和输出方式
|
||||
2. .doc 文件先用 libreoffice --headless --convert-to docx 转换
|
||||
3. 用 ContractEditor 做 tracked_replace(修订),然后用 add_comments_to_docx() 追加批注
|
||||
4. 逐份发送修订版文件 + 简要审查要点说明
|
||||
5. 不需要走 tracker/xlsx/Nextcloud 流程(非Doro体系)
|
||||
|
||||
**常见审查清单(中小型服务/施工/采购合同)**:
|
||||
- 法律引用是否过时(合同法→民法典)
|
||||
- 甲方名称是否完整一致(正文 vs 签约页)
|
||||
- 联系人/电话等空白项是否需要补充
|
||||
- 违约金比例是否对等/是否过高(千分之五/日=年化182.5%,远超司法保护上限)
|
||||
- 质保期是否合理(消防≥2年,固定设施≥1年)
|
||||
- 付款是否与验收挂钩
|
||||
- 是否缺少争议解决/不可抗力/保密条款
|
||||
- 知识产权等模板残留条款
|
||||
- 合同份数是否合理
|
||||
- 编号/条款标题是否完整
|
||||
- 错别字(效益→效力等)
|
||||
|
||||
**流程**:
|
||||
1. 确认审查立场(代表哪方?关注什么重点?)
|
||||
2. 读取合同全文,掌握交易结构
|
||||
3. 针对性法律研究(政策合规性、法条适用等)
|
||||
4. 按委托人要求的重点逐条审查
|
||||
5. 以批注+修订或独立意见书形式输出
|
||||
6. 邮件/企微发送交付
|
||||
|
||||
**与workflow审查的区别**:
|
||||
- 不走classifier→reviewer→editor→deliverer流程
|
||||
- 不需要查review-rules.md和tracker
|
||||
- 不上传Doro的Nextcloud目录
|
||||
- 审查规则和交付方式由委托律师当场指定
|
||||
|
||||
### 双语合同完善 / 模板残留清理(参见 references/bilingual-contract-completion.md)
|
||||
|
||||
莎莎/Maggie/客户直接交办"在现有草稿上完善"的**中英对照合同**(法律服务协议、委托代理协议等),草稿常是从所内旧模板/旧案改来、残留前案痕迹时,参见该专题文件。核心铁律:
|
||||
1. **模板残留是常态**——中文侧和英文侧可能各自残留*不同*旧案内容,且费率/金额/付款等商业条款中英文互相矛盾,必须逐条中英对照
|
||||
2. **商业条款绝不替客户猜**——费率/金额/工时/付款分期/第三方代付有矛盾时,列清单让委托律师给准数,确认后再动手
|
||||
3. **"以中文为准"→先定中文再对齐英文**(符合协议兜底条款,也符合用户预期)
|
||||
4. **vision 报"截断/多空格"先回源核实**——justified 渲染图把跨页断词、字距、自动换行误报为截断,本类任务一次就误报4次;查源 docx `p.text` + `word/settings.xml` 是否含 autoHyphenation 即可证伪
|
||||
5. **关键数字(账号/金额/信用代码)逐字符回源核对**,不信 vision 读数(曾把15位账号读成14位)
|
||||
6. 含**段落级 run 合并替换**函数、多页 OnlyOffice 渲染+逐页 vision 验收管线、交付清稿 vs tracked-changes 两版本的处理
|
||||
|
||||
### 光伏/能源合同审查要点(参见 references/solar-lease-review.md)
|
||||
|
||||
屋顶光伏租赁合同是典型的投资方格式合同,审查时参见专题参考文件。
|
||||
|
||||
## 注意事项
|
||||
|
||||
### 法律意见边界(2026-07-01 总结,多次纠正)
|
||||
|
||||
**角色定位**:执行审查规则、呈现发现。不是法律顾问。
|
||||
|
||||
1. **不做法律价值判断**:不说"强烈建议采用版本X"。呈现各方案的法律风险和后果,由律师决定。
|
||||
2. **不教学**:给客户的意见不用表格对比+教学式分析。律师风格=先列要点(递进排列),再给整体修改文本,不引法条、不用表格(参见 legal-advice-output-style skill)。
|
||||
3. **区分法律依据与行业惯例**:法条引用必须验证原文逐段数;裁判倾向/司法实践不能说成法律结论;律所文章≠官方,不得作确定性结论。
|
||||
4. **公益合同不加商事条款**:公益/捐赠/框架协议不适用对抗性违约金、严苛保证金等商事套路。按 contract_nature 分类审查。
|
||||
5. **批注只写方案不写理由**:批注格式="建议修改为……",不加【新增】【修改】标签,不写法律理由。
|
||||
6. **金额是商业条款绝对不改**。
|
||||
7. **法律检索铁律**:信息来源仅限官方资料(法律/法规/政策/裁判文书原文)。引用须标注并说明未经裁判文书验证;给出结论必须列出信息来源链接。
|
||||
|
||||
### 邱律师批量文件接收分流(Intake Triage)
|
||||
邱律师一次性发送多份合同时,**必须先交叉核对再决定是否启动workflow**。四类分流:今日已交付 / 正在审查 / 历史已审重发 / 全新未处理。详见 `references/intake-triage-crosscheck.md`。
|
||||
- ⚠️ 重发≠重审:邱律师经常重新发送已审查过的文件,必须报告Doro确认是否需要重新审查
|
||||
- ⚠️ auto_notify可能漏文件:不能假设自动化已处理所有新文件,手动核对xlsx是必要的
|
||||
- xlsx读取需 `sudo` + venv Python(文件属www-data)
|
||||
|
||||
### `uwf thread exec --background` 的 notify_on_complete 陷阱(2026-07-02教训)
|
||||
|
||||
通过 `terminal(background=true, notify_on_complete=true)` 运行 `uwf thread exec <id> --count 20 --background` 时:
|
||||
- **parent进程**立即退出(输出 "Step 1 xxx → running"),触发 terminal 的 notify_on_complete
|
||||
- **实际工作**在独立的 `--_background-worker` 子进程中继续(可通过 `ps aux | grep <thread_id>` 看到)
|
||||
- ⚠️ **不要**在收到通知后立即再次 exec —— 会报 "thread is already being executed by PID xxx"
|
||||
|
||||
**正确做法**:收到通知后先 `uwf thread show <id>` 检查状态:
|
||||
- `Status: running` + 有对应进程 → 正在执行,等待
|
||||
- `Status: idle` + 无进程 → 已完成当前步骤,可继续 exec
|
||||
- `Status: end` → workflow已结束
|
||||
|
||||
**推荐方式**:不要用 terminal 的 background 模式包裹 `uwf thread exec --background`(双层后台)。改用:
|
||||
```bash
|
||||
# 直接在前台运行uwf的后台模式,让uwf自己管理后台
|
||||
/home/maggie/.hermes/node/bin/uwf thread exec <id> --count 20 --background
|
||||
# 然后用 ps + uwf thread show 轮询状态
|
||||
```
|
||||
|
||||
### ⚠️ 收到新合同后:先检查queue-runner是否已启动thread(2026-07-03白鹤镇教训)
|
||||
|
||||
文件到达 `/tmp/contract-review/` 后,`contract-queue-runner.sh`(常驻后台进程)会**自动检测并启动workflow thread**。如果agent也手动 `uwf thread start`,就会产生两个thread竞争同一份合同。
|
||||
|
||||
**收到合同后的正确顺序**:
|
||||
```bash
|
||||
# 1. 确认文件已到 /tmp/contract-review/
|
||||
ls /tmp/contract-review/*.docx
|
||||
|
||||
# 2. 检查queue-runner是否在跑(在跑=可能已自动启动thread)
|
||||
ps aux | grep "[c]ontract-queue-runner"
|
||||
|
||||
# 3. 检查是否已有该合同的running thread
|
||||
/home/maggie/.hermes/node/bin/uwf thread list 2>&1 | grep running
|
||||
|
||||
# 4. 有running thread → 只监控,不重复启动
|
||||
# 没有running thread + queue-runner不在 → 才手动start
|
||||
```
|
||||
|
||||
**双线程症状**:`ps aux | grep uwf-hermes` 显示两个进程审查同一份文件(prompt中文件名相同但thread_id不同)。发现后立即kill较晚的那个。
|
||||
|
||||
### 批量合同启动前必做:清理僵尸线程+杀queue-runner(2026-07-01/02教训)
|
||||
多份合同串行处理前,**必须先清理 /tmp/contract-review 和僵尸线程**。详见 `references/workflow-systemic-issues-202607.md` §7-8。
|
||||
- **第一步:杀 queue-runner**(否则它会自动启动重复线程!2026-07-02恭兴合同教训——queue-runner和手动start同时对恭兴合同启动了两个thread,两个uwf-hermes进程竞争同一个【修】文件)
|
||||
```bash
|
||||
kill $(ps aux | grep "[c]ontract-queue-runner" | awk '{print $2}') 2>/dev/null
|
||||
```
|
||||
- **第二步:检查+取消已有running threads**
|
||||
```bash
|
||||
uwf thread list | grep running # 列出所有running
|
||||
# 逐个cancel非当前任务的
|
||||
uwf thread cancel <thread_id>
|
||||
```
|
||||
- **第三步:杀死所有 uwf-hermes 进程**(cancel thread 不够,进程不感知status变化)
|
||||
```bash
|
||||
ps aux | grep "[u]wf-hermes" | grep -v grep | awk '{print $2}' | xargs kill 2>/dev/null
|
||||
```
|
||||
- **第四步:清空 /tmp/contract-review**(保留 rules/)
|
||||
```bash
|
||||
cd /tmp/contract-review && rm -f *.docx *.doc *.py *.pdf *.png *.xml *.bak 2>/dev/null
|
||||
```
|
||||
- 然后逐份:复制→start→exec→等end→下一份
|
||||
- ⚠️ 多份合同绝对不能并行启动,共用/tmp/contract-review会互相覆盖
|
||||
- ⚠️ 双线程竞争症状:`thread is already being executed by PID xxx`错误、/tmp下出现不属于当前合同的文件
|
||||
|
||||
### Reviewer/Editor 超时卡死恢复(2026-07-02 夏阳合同教训)
|
||||
|
||||
当 thread 卡在某个 role(通常 reviewer)超过 20 分钟且 Head 不变:
|
||||
|
||||
**诊断**:
|
||||
```bash
|
||||
# 1. 确认 Head 未变化(连续两次 5min 间隔检查)
|
||||
uwf thread show <thread_id> # 记录 Head 值
|
||||
sleep 300
|
||||
uwf thread show <thread_id> # Head 相同 = 卡死
|
||||
|
||||
# 2. 找到对应的 uwf-hermes 进程
|
||||
ps aux | grep "<thread_id>" | grep "uwf-hermes" | grep -v grep
|
||||
```
|
||||
|
||||
**恢复**:
|
||||
```bash
|
||||
# 1. 杀死卡住的 uwf-hermes 进程
|
||||
kill <PID>
|
||||
|
||||
# 2. 等 10 秒,thread 变为 suspended 状态
|
||||
sleep 10
|
||||
uwf thread show <thread_id>
|
||||
# 输出: Status=suspended, Suspend="agent command failed (uwf-hermes)"
|
||||
|
||||
# 3. 重新执行(从当前 role 继续,不会丢失之前的修订进度)
|
||||
uwf thread exec <thread_id> --count 15 --background
|
||||
```
|
||||
|
||||
**注意**:
|
||||
- thread exec 会从 suspended 的当前 role 重新开始该步骤(reviewer 重审/editor 重修)
|
||||
- 之前 editor 产出的 【修】文件仍在 /tmp/contract-review/,不会丢失
|
||||
- 如果同一 role 连续卡死 3 次以上(同一 Head 值不变),升级处理:
|
||||
|
||||
**Final review 循环退回问题(2026-07-03 白鹤镇教训)**:
|
||||
|
||||
Workflow 的 final_review 发现问题后会退回 editor→reviewer→deliverer→final_review 循环。邱律师明确要求**"只审查一次就行"**——当 final_review 退回时,如果第一轮修订已通过 reviewer 第2轮复核(通常复核确认修订无误),应直接杀掉进程交付,不要进入第三轮以后的循环。
|
||||
|
||||
**判断时机**:从 deliverer 进程的 prompt 中确认"复核通过"→ 到 final_review 退回 editor → 立即终止整个 thread 进程链,取已有的【修】文件手动交付。
|
||||
|
||||
**3次卡死后的升级路径(2026-07-02 家庭医生签约合同教训)**:
|
||||
|
||||
当 reviewer 在同一 Head 反复卡死 3 次(kill→resume→再卡死),说明 LLM 对该 prompt 无法正常返回。此时【修】文件通常已经存在且完整(editor 已产出),可直接手动交付:
|
||||
|
||||
```python
|
||||
# 验证【修】文件质量(必做!)
|
||||
from zipfile import ZipFile
|
||||
import re
|
||||
path = '/tmp/contract-review/【修】XXX.docx'
|
||||
with ZipFile(path) as z:
|
||||
doc_xml = z.read('word/document.xml')
|
||||
ins_count = len(re.findall(b'w:ins ', doc_xml))
|
||||
del_count = len(re.findall(b'w:del ', doc_xml))
|
||||
wb_count = len(re.findall(b'w:author="WB"', doc_xml))
|
||||
print(f'INS={ins_count}, DEL={del_count}, WB={wb_count}')
|
||||
# 确认: ins_count > 0, wb_count == ins_count + del_count
|
||||
```
|
||||
|
||||
验证通过后直接手动上传:
|
||||
```bash
|
||||
docker cp "/tmp/contract-review/【修】XXX.docx" "nextcloud-nextcloud-1:/var/www/html/data/doro/files/Doro合同审查任务/任务交付/"
|
||||
docker exec nextcloud-nextcloud-1 chown www-data:www-data "/var/www/html/data/doro/files/Doro合同审查任务/任务交付/【修】XXX.docx"
|
||||
docker exec nextcloud-nextcloud-1 php occ files:scan doro
|
||||
```
|
||||
|
||||
⚠️ 手动交付跳过了 final_review 步骤,需要额外注意:
|
||||
- 检查 INS run 字体(rFonts + sz)是否与原文一致
|
||||
- 检查交叉引用是否需要更新
|
||||
- 仍需更新 tracker + xlsx + 清 OO 缓存
|
||||
|
||||
### 合同组合梳理(Portfolio Audit)
|
||||
对已有合同进行批量合规审查时(如客户要求梳理所有在履约合同),参见 `references/contract-portfolio-audit.md`。与逐份审查修订不同,产出物为批注PDF+Excel汇总表。
|
||||
- 文件命名:客户名称+文件名+批注+日期(非标准版本号格式)
|
||||
- Excel结构:总览sheet + 每个分组一个sheet(租赁按校区,其他按合同类型)
|
||||
- 先做2-3份样本验证表头,确认后再批量推进
|
||||
- ⚠️ 完成sheet后必须做文件数量核对(源文件数 - 已知重复 = 表格行数),并在第四项中添加【文件汇总说明】供客户对照原始文件与表格条目,详见 `references/contract-portfolio-audit.md` step 7
|
||||
- 租赁合同按校区聚合时,租赁+物业+补充协议打包在一起看,不分家
|
||||
- ⚠️ 风险点栏(K列)不仅要列常规风险,**必须包含合同变更/提前解除/减面积相关的责任义务分析**(客户最关心)——通知期、违约金、押金处理、已有退租先例。详见 `references/template-comparison-methodology.md` "Thematic risk additions to K column"
|
||||
- ⚠️ openpyxl生成的Excel在OnlyOffice中**所有多行内容cell的行高都会截断**,必须用脚本逐行计算并设置显式高度,详见 reference 中的 pitfall 说明
|
||||
|
||||
### 标准模版对比(新增L列)
|
||||
当客户有标准合同模版(常见于集团客户),需要对比签订的合同与模版的差异时,参见 `references/template-comparison-methodology.md`。产出物为表格新增一列"与标准模版差异"。
|
||||
|
||||
### 多方修订对账(Multi-Party Revision Reconciliation)
|
||||
当合同经多方修订(如WB、客户方律师、对方等不同修订人),需要核对叠加修订效果是否符合协商一致条件、对比模板差异、评估影响、统一修订人署名时,参见 `references/multi-party-revision-reconciliation.md`。典型场景:MCN合同经双方律师各自修订后需验证商业条件落实情况。
|
||||
|
||||
关键技术点(2026-07-03实证):
|
||||
- **嵌套修订处理**:B在A的`w:ins`内部做`w:del`→先接受嵌套删除→清除空壳→统一作者
|
||||
- **恢复已删内容**:要把WB之前的del恢复回来时,不能简单删除del元素(原文已消失),需要替换为ins
|
||||
- **追加模板修改**:在统一后的文件上继续做del+ins tracked changes,找到目标run→remove→insert(del_elem, ins_elem)
|
||||
|
||||
### 被问合同内容时必须先查文件(2026-07-09 Maggie纠正)
|
||||
|
||||
用户问"这个条款写了什么""保密信息包含什么""这条有没有问题"等**针对具体合同内容的问题**时,**立即找到文件并读取**,不得从一般法律知识或推理回答。
|
||||
|
||||
❌ 错误:用户问"保密信息包含什么" → 从一般施工合同常识回答"通常包含..."
|
||||
✅ 正确:用户问"保密信息包含什么" → 找到文件 → 读取保密条款原文 → 引用原文回答
|
||||
|
||||
Maggie原话:"这是修订的内容,我让你看让你查,你做了吗"——被问即查,不凭推测回答。这与memory中"Maggie核查口令"一致:被问即回工具核实。
|
||||
|
||||
### 信息检索限制
|
||||
- 企查查、天眼查、爱企查、国家企信系统等中国工商数据源对境外IP有访问限制
|
||||
- 遇到此限制时,请Maggie协助查询,或通过国内代理访问
|
||||
- 上市公司信息可通过巨潮资讯网(cninfo.com.cn)、新浪财经等渠道获取
|
||||
|
||||
### xlsx 读取的正确方法(权限问题)
|
||||
|
||||
xlsx文件属 www-data,maggie用户无法直接读取。正确流程:
|
||||
```bash
|
||||
# 方案1:复制到/tmp后读取
|
||||
sudo cp "/home/maggie/nextcloud/data/data/doro/files/Doro合同审查任务/合同审查清单.xlsx" /tmp/xlsx_check.xlsx
|
||||
sudo chmod 644 /tmp/xlsx_check.xlsx
|
||||
/home/maggie/.hermes/hermes-agent/venv/bin/python3 -c "import openpyxl; ..."
|
||||
```
|
||||
注意:`python3` 的 openpyxl 在 `/home/maggie/.hermes/hermes-agent/venv/bin/python3`,系统 python3 和 sudo python3 都没有 openpyxl。
|
||||
|
||||
### 工作文档
|
||||
- 每次合同审查的工作流程文档存放在:~/.hermes/shared/合同审查交付件/
|
||||
- 流程文档随审查进展同步更新
|
||||
@@ -0,0 +1,72 @@
|
||||
# 双语合同完善 / 模板残留清理(中英对照法律服务协议等)
|
||||
|
||||
## 适用场景
|
||||
莎莎 / Maggie / 客户直接交办一份"在现有草稿基础上完善"的中英对照合同(法律服务协议、委托代理协议等),不走 Doro 的 uwf workflow。草稿往往是从所内旧模板 / 旧案改来的,残留前案痕迹。
|
||||
|
||||
## 铁律1:模板残留是常态,必须逐条中英对照
|
||||
律所复用双语模板时,**中文侧和英文侧可能各自残留不同旧案的内容**,且核心商业条款中英文互相矛盾。
|
||||
普力克 Procore 案(2026-06-25)实例:
|
||||
- 中文侧整体是"社保调查"案,英文侧整体是"喆航买卖合同纠纷"案——**两个不同旧案混在一份草稿里**
|
||||
- 客户名残留旧客户(艾达 / AMOS),分布在抬头、正文、付款条款第三方代付人、签署页多处
|
||||
- 费率:中文 2,500 元/小时·2-3 小时一次性;英文 2,000 元/小时·每月 5-6 小时——三项数据互相打架
|
||||
- 付款:中文付全额;英文付 50% 即 72,500 元 + AMOS 新加坡公司代付
|
||||
|
||||
**做法**:python-docx 提全文(含 `doc.tables`),把每个条款的中文段和紧邻的英文段配对,逐条比对。重点盯:服务内容、服务范围、主要联系人、费率、工时、付款方式/金额、第三方付款人、签署日期。
|
||||
|
||||
## 铁律2:商业条款绝不替客户猜,先确认后动手
|
||||
费率、金额、工时、付款分期、第三方代付——这些是商业条款。草稿里中英文打架时**不能自己挑一个填进去**。把所有矛盾点列成清单,让委托律师给准数,确认后再出稿。
|
||||
本案先问了 7 点(真实业务事由 / 非诉 or 诉讼 / 联系人 / 费率 / 工时 / 收款方式 / 签署日),拿到 3 点核心确认(服务内容=王晓玲社保调查、费率 2,500×2-3 小时一次性全额、签署日 6/25)才动手。
|
||||
|
||||
## 铁律3:以中文为准 → 先定中文,再对齐英文
|
||||
协议若含"如有出入以中文为准"条款(多数中英对照协议都有),完善顺序必须是:**先把中文条款定稿,再逐条把英文改成与中文一致**。用户本案也明确指示此顺序。这样既符合协议约定,也符合用户预期。该兜底条款审完保留——对己方有利,不要删。
|
||||
|
||||
## 技术:段落级 run 合并替换(保格式整段替换)
|
||||
python-docx 里一个段落常被切成几十个碎 run(中文逐字符切、英文逐词切,bold/size 各异)。整段替换文本又要保留段落样式,用这个函数:
|
||||
```python
|
||||
def set_para_text(p, text):
|
||||
"""整段文本写入第一个 run,清空其余 run,保留首 run 格式。"""
|
||||
if not p.runs:
|
||||
p.add_run(text); return
|
||||
p.runs[0].text = text
|
||||
for r in p.runs[1:]:
|
||||
r.text = ""
|
||||
```
|
||||
- 表格单元格同理:`set_para_text(cell.paragraphs[idx], text)`
|
||||
- 删除整个空段落(如清掉被替空的旧英文残段,避免留空行):`el = p._element; el.getparent().remove(el)`
|
||||
- 替换后务必复查:遍历全文(段落 + 表格 cell)拼成大字符串,断言旧内容(艾达/AMOS/旧金额/旧费率等)`not in full`,再断言新数据出现次数正确。
|
||||
|
||||
## 验收:多页 OnlyOffice 渲染 → 逐页 vision 核对
|
||||
单页固定版面用 `print-ready-pdf/scripts/oo_render_check.sh`。多页双语合同用以下管线(Maggie/Doro/莎莎都用 OnlyOffice 看文件,必须用 x2t 不用 LibreOffice):
|
||||
```bash
|
||||
docker cp file.docx nextcloud-onlyoffice-1:/tmp/_chk.docx
|
||||
# 容器内写 TaskQueueDataConvert xml,m_nFormatTo=513,跑 x2t(见 print-ready-pdf)
|
||||
docker cp nextcloud-onlyoffice-1:/tmp/_chk.pdf /tmp/out.pdf
|
||||
pdftoppm -png -r 110 /tmp/out.pdf /tmp/pg # 每页一张 PNG
|
||||
```
|
||||
然后 `vision_analyze` 逐页核对:客户名、服务内容、费率、付款方式金额、签署日期、签署页主体、中英一致性。
|
||||
|
||||
## 铁律4:vision 报"截断/截字/多空格"先回源核实,别急着改
|
||||
justified(两端对齐)渲染图里,vision 模型会把以下**正常排版现象**误报为"文字截断/排版错误"——本案一次任务里就误报了 4 次:
|
||||
- 跨页连字符断词(第1页底 `accor-` → 第2页头 `ding`)
|
||||
- 跨页句子接续(第4页底 `Negotiations shall` → 第5页头 `commence...`)
|
||||
- 两端对齐的字距假象(`with` 被读成 `wit h`)
|
||||
- 两端对齐的自动换行断词(`aforemen` / `tioned`,且无连字符)
|
||||
|
||||
判断方法(任一即可证伪 vision 的"截断"):
|
||||
1. 从源 docx 读该段 `p.text`,确认是完整单词 / 完整句子;
|
||||
2. 查 `word/settings.xml` 是否含 `autoHyphenation`(没开就不会真断字);
|
||||
3. 查文本不含软连字符 `\u00ad`。
|
||||
|
||||
**渲染层表象 ≠ 文档缺陷**,Word / OnlyOffice 编辑视图会正常回流。
|
||||
|
||||
## 铁律5:关键数据(账号/金额/信用代码)逐字符回源核对,不信 vision 读数
|
||||
vision 把收款账号 `121955296610001`(15 位)读成 `12195296610001`(漏 1 位)。账号、金额、统一社会信用代码这类关键数字,**必须从源 docx 逐字符提取核对**,绝不采信 vision 对渲染图的读数。本案账号与原稿逐位一致,确认后未改。
|
||||
|
||||
## 顺手订正范围("条款明确得当"要求)
|
||||
完善双语合同时,除主体/商业条款外,一并修掉:
|
||||
- 中文病句("在上述法律服务作为客户的代理人" → "就上述法律服务事项作为客户的代理人")、错别字("疑异"→"疑义")、异体字("⻓"→"长")
|
||||
- 英文逗号粘连句(comma splice,逗号改分号或加 and)、口语化用词("really" → "truthfully",与同条 "faithfully" 体系一致)
|
||||
|
||||
## 交付
|
||||
- 命名规则:当事人名称+文件名称+版本+修改人+日期(本案:`普力克贸易(上海)有限公司-法律服务协议-v1-20260625.docx`)
|
||||
- 本类任务默认交"用印前清稿"(直接定稿、非修订模式)。若需**带修订痕迹**版本供对方/同事看改了哪些,另出一版 tracked changes(用 ContractEditor 库,不裸写 XML)。交付时主动告知用户这两种版本的区别,问要哪种。
|
||||
@@ -0,0 +1,193 @@
|
||||
# Contract Portfolio Audit (合同梳理/合规审查)
|
||||
|
||||
Different from individual contract review (docx track changes). This is batch reading of an existing contract collection to produce a compliance overview for the client.
|
||||
|
||||
## When to use
|
||||
- Client wants to understand the status and risks of their existing contract portfolio
|
||||
- Goal is compliance oversight, not redlining individual contracts
|
||||
- Deliverables: annotated PDFs + Excel summary table
|
||||
|
||||
## Workflow
|
||||
|
||||
### 1. File inventory
|
||||
- Map the full directory structure and count files
|
||||
- Group by logical categories (the client's folder structure usually reflects this)
|
||||
- Record in MemPalace as a long-term project if ongoing
|
||||
|
||||
### 2. OCR scanned PDFs
|
||||
- Most contracts from clients are scanned PDFs (no text layer)
|
||||
- Check with `page.get_text().strip()` — empty = scanned
|
||||
- Use `deepseek-ocr` skill (page-by-page image fallback if API unstable)
|
||||
- See `deepseek-ocr` skill for retry strategy
|
||||
- For large batches: convert all PDFs to images first, then OCR in a single loop with `time.sleep(1-2)` between pages
|
||||
|
||||
### 3. Analyze and annotate PDFs
|
||||
- Add sticky note annotations via pymupdf (see `ocr-and-documents` skill)
|
||||
- Author field: use `annot.set_info(title="WB")` + `annot.update()`
|
||||
- Annotation content: risk points, missing clauses, suggestions — no risk level tags (【高风险】etc.)
|
||||
- Suggestions must be specific and actionable ("建议增加……""建议修改为……"), not vague descriptions
|
||||
|
||||
### 4. Excel summary table
|
||||
- Use openpyxl with proper styling (微软雅黑, header fill, borders, wrap_text, freeze panes)
|
||||
|
||||
#### File naming for deliverables (differs from standard versioned naming):
|
||||
```
|
||||
批注PDF: 客户名称-文件名称-批注-YYYYMMDD.pdf
|
||||
汇总表: 客户名称-XX合同汇总表-YYYYMMDD.xlsx
|
||||
```
|
||||
Example:
|
||||
```
|
||||
南通新东方-BOSS直聘服务合同-批注-20260609.pdf
|
||||
南通新东方-人力资源合同汇总表-20260609.xlsx
|
||||
南通新东方-租赁合同汇总表-20260609.xlsx
|
||||
```
|
||||
|
||||
#### Standard fields (adjustable per project):
|
||||
| Field | Notes |
|
||||
|-------|-------|
|
||||
| 序号 | Sequential |
|
||||
| 合同名称 | |
|
||||
| 合同编号 | "无" if not present |
|
||||
| 合同类型 | 采购服务/采购货物/租赁/etc. |
|
||||
| 甲方(我方主体) | May have multiple entities |
|
||||
| 乙方(对方主体) | |
|
||||
| 合同期限 | Start-end dates |
|
||||
| 合同金额 | With currency |
|
||||
| 服务内容 | Brief description |
|
||||
| 合同状态 | 履行中/即将到期/已到期 |
|
||||
| 主要风险点 | Numbered list |
|
||||
| 建议 | Numbered list, actionable |
|
||||
|
||||
### 5. Grouping logic (confirmed with Maggie for 南通新东方)
|
||||
- **房租物业合同**: Group by 校区/租赁物 — rental + property management together in one sheet per campus
|
||||
- **非模板其他合同**: Group by 合同类型 (采购货物/采购服务/租赁/etc.)
|
||||
- **人力资源**: Flat list
|
||||
|
||||
### 6. Excel structure for rental contract portfolios
|
||||
|
||||
**Multi-sheet workbook**: one Excel file for all campuses.
|
||||
|
||||
| Sheet | Content |
|
||||
|-------|---------|
|
||||
| **总览** | One row per campus: campus name, tenant entity, landlord, current area, term, quarterly rent, quarterly property fee, deposit total, key risks, notes |
|
||||
| **校区A** | All contracts for that campus in detail |
|
||||
| **校区B** | ... |
|
||||
| ... | ... |
|
||||
|
||||
**Per-campus sheet structure** (sections with merged header rows):
|
||||
|
||||
```
|
||||
一、原租赁合同系列(乙方:XXX)
|
||||
[header row]
|
||||
主合同 → 补充协议 → 退租 → 三方转让 → ...
|
||||
|
||||
二、扩租合同系列(乙方:YYY) [if applicable]
|
||||
[header row]
|
||||
主合同 → 补充协议 → ...
|
||||
|
||||
三、物业管理合同系列
|
||||
[header row]
|
||||
原租物业 → 补充协议 → 退租 → 扩租物业 → ...
|
||||
|
||||
四、校区整体风险分析与建议
|
||||
[merged A:K cell — 【整体风险分析】+ numbered risks + 【建议】]
|
||||
[merged A:K cell — 【文件汇总说明】+ source file ↔ table entry mapping]
|
||||
```
|
||||
|
||||
**Section 四 has TWO merged cells** (both A:K, wrap_text=True, vertical=top):
|
||||
1. **风险分析**: overall risks + recommendations (existing)
|
||||
2. **文件汇总说明**: client-facing reconciliation — lists every source file by original directory, maps each to its table entry number, and explains any discrepancies (duplicates removed, entries combined, files reclassified to different sections). This lets the client cross-reference their original files against the table without guessing.
|
||||
|
||||
**Per-campus detail columns**:
|
||||
序号, 文件名称, 合同类型, 合同当事人, 租赁标的/服务范围, 面积(㎡), 合同期限, 金额/费用, 核心内容, 当前状态, 风险点/备注
|
||||
|
||||
**Key points**:
|
||||
- Each section has its own header row (repeated column headers after each section divider)
|
||||
- Section dividers are merged cells with blue background
|
||||
- Within each section, documents are listed in logical/chronological order (main contract → supplements → amendments)
|
||||
- The risk analysis section at the bottom is a single large merged cell covering all columns
|
||||
- Use freeze_panes on the title area so data scrolls while title stays
|
||||
|
||||
### 7. File count reconciliation (MANDATORY before upload)
|
||||
|
||||
After completing all sheets, reconcile source files against table entries:
|
||||
|
||||
1. **Count source files**: list all PDFs in the source directory tree
|
||||
2. **Subtract known duplicates**: identical content files (same size + same date often = duplicate scan)
|
||||
3. **Count table entries**: sum all numbered rows across the sheet
|
||||
4. **Reconcile**: source_files - duplicates MUST equal table_entries. If not, identify missing files.
|
||||
|
||||
**Common miss patterns**:
|
||||
- Two files with similar names/content silently merged into one row (e.g., `扩租物业补充协议.pdf` and `扩租物业补充协议(电费教育科技)1.pdf` — same theme but different signing parties, should be separate entries or explicitly noted as combined)
|
||||
- A file in one subdirectory overlooked because a "similar" file in another subdirectory was already covered (e.g., `物业补充协议(电费培训学校).pdf` missed because `物业补充xiey.pdf` seemed to cover the same topic)
|
||||
- Files across subdirectories not cross-checked against the master list
|
||||
|
||||
**Rule**: Every unique file gets its own row unless explicitly combined (in which case the combined row must list all file names and the备注 must explain why they're grouped). One-to-one is the default; combining requires justification.
|
||||
|
||||
**Client-facing summary**: After reconciliation, add a 【文件汇总说明】section to the Excel (inside section 四, as a separate merged A:K cell below the risk analysis). Format:
|
||||
```
|
||||
【文件汇总说明】
|
||||
|
||||
客户提供XX校区相关合同文件共N份,分布在M个文件夹中。经核对整理,去重后为X份独立文件,汇总表中归纳为Y条记录。具体对照如下:
|
||||
|
||||
一、文件夹A/(N份)
|
||||
1. 文件名.pdf → 表第X项
|
||||
2. 文件名.pdf → 表第Y项
|
||||
...
|
||||
|
||||
二、文件夹B/(N份,含M份重复)
|
||||
...
|
||||
|
||||
说明:
|
||||
- 重复文件:XXX同时出现在A/和B/目录中,内容一致,仅计1份
|
||||
- 合并记录:XXX两份内容一致,合并为表第N项
|
||||
- 分类调整:A/目录下的XXX实为YYY,已归入ZZZ系列
|
||||
```
|
||||
This is NOT optional — every campus/category sheet must include this summary so the client can verify completeness without asking.
|
||||
|
||||
### 8. Upload to Nextcloud
|
||||
- Place annotated PDFs alongside originals in same directory
|
||||
- Place Excel in the parent category folder (e.g., 房租物业合同/ for the rental summary)
|
||||
- `docker cp` → `chown www-data` → `php occ files:scan`
|
||||
|
||||
### 9. Sample-first validation
|
||||
**Always start with a small sample** (2-3 contracts of different types):
|
||||
1. Pick contracts that represent different complexity levels
|
||||
2. Do full OCR → analyze → annotate → fill Excel
|
||||
3. Send to Maggie for review of format and content
|
||||
4. Adjust table structure based on feedback
|
||||
5. Then batch process remaining contracts
|
||||
|
||||
This avoids rework — the table structure almost always needs adjustment after the first review.
|
||||
|
||||
## Key differences from individual contract review
|
||||
| Aspect | Individual review | Portfolio audit |
|
||||
|--------|------------------|-----------------|
|
||||
| Input | Usually docx | Usually scanned PDF |
|
||||
| Output | Track changes docx | Annotated PDF + Excel summary |
|
||||
| Depth | Clause-by-clause redline | Risk overview and key terms extraction |
|
||||
| Audience | Lawyer (for negotiation) | Client management (for compliance) |
|
||||
| Volume | 1 contract at a time | Batch (tens to hundreds) |
|
||||
| Workflow | uwf reviewer→editor | OCR → analyze → annotate → summarize |
|
||||
| File naming | 当事人+文件+版本+修改人+日期 | 客户名+文件名+批注/汇总表+日期 |
|
||||
|
||||
## Related references
|
||||
- `references/template-comparison-methodology.md` — when client has a standard template, how to produce the "与标准模版差异" column
|
||||
|
||||
## Pitfalls
|
||||
- **Section 四 formatting**: The risk analysis section must use properly merged A:K cells. Common mistakes: (1) row heights set wrong (section headers should be ~30, not 80/350); (2) content cell not merged across all columns so text only shows in column A; (3) empty separator rows getting large fixed heights instead of auto. After building section 四, verify: `ws.merged_cells.ranges` includes both A:K rows, row heights are sane (None or ~30 for headers), and wrap_text=True on content cells.
|
||||
- **CRITICAL: OnlyOffice does NOT reliably auto-size row heights — for merged cells OR regular cells with wrap_text.** Setting `height=None` (auto) on ANY cell with multi-line content means text WILL be cut off. This is the #1 user complaint ("文字显示不出来"). **Always set explicit heights on EVERY row with multi-line content**:
|
||||
1. After populating all cells, run a height estimation pass over ALL data rows
|
||||
2. For each row, find the cell with the most visual lines (accounting for CJK chars at ~2 width units each, line wrapping at `col_width * 1.2` chars)
|
||||
3. Set height = max_visual_lines × 15pt + 30-50% buffer
|
||||
4. For merged cells spanning A:K (~231 width), one text line ≈ one visual line (no wrapping)
|
||||
5. For narrow columns (H=20, K/L=40), long CJK lines wrap heavily — a 60-char Chinese line in a width-20 column = ~4 visual lines
|
||||
6. Write a reusable height-check script rather than eyeballing — the number of rows that need fixing always exceeds expectations
|
||||
- **File count reconciliation**: after completing a campus/category sheet, always count source files and compare against table rows BEFORE uploading. The user should never be the one to discover a missing file. See step 7 above.
|
||||
- Start with a small sample (2-3 contracts) to validate table headers with Maggie before batch processing
|
||||
- Chinese text in Python strings with nested quotes (especially inside f-strings or dicts): use single quotes inside double or vice versa, or put long strings in variables
|
||||
- openpyxl SyntaxError: Chinese punctuation like 、()【】 inside nested Python string literals can confuse the parser — write the script to a .py file first, then run it, rather than using inline heredocs
|
||||
- For complex campuses (like 北翼玖玖 with 17 files across 3 subdirs): use delegate_task to read all OCR'd files in parallel, then synthesize the analysis yourself
|
||||
- When building a campus sheet, always trace the full document chain (main contract → what amended it → what superseded it) to understand the current state
|
||||
- Watch for different legal entities signing different contracts at the same location (e.g., 培训学校 vs 教育科技) — flag this as a management complexity risk
|
||||
- Watch for landlord entity changes mid-lease (common in commercial real estate) — trace the transfer chain and note deposit movements
|
||||
@@ -0,0 +1,101 @@
|
||||
# 合同接收分流:邱律师批量文件交叉核对
|
||||
|
||||
当邱律师(QiuTing)一次性发送多份合同时,必须先交叉核对分流,再决定是否启动workflow。
|
||||
|
||||
## 分流四类
|
||||
|
||||
| 类别 | 判定方法 | 处理 |
|
||||
|------|---------|------|
|
||||
| ✅ 今日已交付 | xlsx有今天日期的对应条目 | 跳过 |
|
||||
| 🔄 正在审查 | `uwf thread list`显示running | 等完成,不重复启动 |
|
||||
| 📋 历史已审、今日重发 | xlsx有该文件但日期更早 | **报告Doro**:可能是修改稿需重审,也可能是误发 |
|
||||
| ❌ 全新未处理 | xlsx无匹配条目 | 启动workflow |
|
||||
|
||||
## 核对步骤
|
||||
|
||||
### Step 1: 提取今日QiuTing文件清单
|
||||
```bash
|
||||
for f in ~/.hermes/cache/documents/*.meta; do
|
||||
sender=$(python3 -c "import json; d=json.load(open('$f')); print(d.get('sender_id',''))" 2>/dev/null)
|
||||
if [ "$sender" = "QiuTing" ]; then
|
||||
ts=$(python3 -c "import json; d=json.load(open('$f')); print(int(d.get('timestamp',0)))" 2>/dev/null)
|
||||
# 比对今日时间戳范围
|
||||
filename=$(basename "$f" .meta | sed 's/^doc_[a-f0-9]*_//')
|
||||
echo "$filename"
|
||||
fi
|
||||
done
|
||||
```
|
||||
按北京时间筛选当天00:00~24:00范围内的文件。
|
||||
|
||||
### Step 2: 查xlsx台账匹配
|
||||
```bash
|
||||
sudo /home/maggie/.hermes/hermes-agent/venv/bin/python3 << 'PYEOF'
|
||||
import openpyxl
|
||||
xlsx = "/home/maggie/nextcloud/data/data/doro/files/Doro合同审查任务/合同审查清单.xlsx"
|
||||
wb = openpyxl.load_workbook(xlsx, read_only=True, data_only=True)
|
||||
ws = wb.active
|
||||
# 用关键词模糊匹配文件名(去掉前缀/后缀后取核心词)
|
||||
keywords = ["关键词1", "关键词2"] # 从文件名提取
|
||||
for row in ws.iter_rows(min_row=2, values_only=True):
|
||||
row_str = str(row)
|
||||
for kw in keywords:
|
||||
if kw in row_str:
|
||||
print(f"#{row[0]} | {row[1]} | {row[2]} | {row[3]} | {row[4]}")
|
||||
break
|
||||
wb.close()
|
||||
PYEOF
|
||||
```
|
||||
|
||||
**匹配要点**:文件名常有微小差异(空格、下划线、版本号),用核心关键词模糊匹配,不要精确匹配全名。
|
||||
|
||||
### Step 3: 查uwf运行状态
|
||||
```bash
|
||||
/home/maggie/.hermes/node/bin/uwf thread list | grep running
|
||||
```
|
||||
|
||||
### Step 4: 查/tmp/contract-review/和progress目录
|
||||
确认是否有文件正在队列中等待。
|
||||
|
||||
### Step 5: 汇总报告Doro
|
||||
按四类分类呈现,对「历史已审重发」类必须询问Doro:是修改稿需重审还是误发。
|
||||
|
||||
## 合同发送量核查协议(Doro问"发了几份/审了几份"时)
|
||||
|
||||
⚠️ **Doro问邱律师发了几份合同时,必须多源交叉核对,不能只查一个来源。** 2026-06-29/30教训:Doro问了三遍"邱律师发了几份",每次agent都漏东西。根因:只查了gateway.log的inbound消息,漏了document cache的meta文件(文件消息不产生inbound text log)。
|
||||
|
||||
### 必查5个源(缺一不可)
|
||||
|
||||
| # | 数据源 | 命令 | 提取信息 |
|
||||
|---|--------|------|----------|
|
||||
| 1 | gateway.log inbound消息 | `grep "2026-MM-DD" gateway.log \| grep "QiuTing" \| grep "inbound"` | 文字消息+指令(合同审查等) |
|
||||
| 2 | document cache meta文件 | `for f in ~/.hermes/cache/documents/*.meta; do` 读sender_id+timestamp | 实际收到的文件(⚠️文件消息在gateway.log中msg为空) |
|
||||
| 3 | 待审查目录 | `sudo find .../待审查/ -type f` | 已上传但未审完的合同 |
|
||||
| 4 | 任务交付目录 | `sudo find .../任务交付/ -type f -name "【修】*"` | 已审查交付的合同 |
|
||||
| 5 | xlsx台账 | `sudo python3 -c "openpyxl..."` | 已pass登记的合同 |
|
||||
|
||||
### meta文件时间戳转换
|
||||
```python
|
||||
from datetime import datetime, timezone, timedelta
|
||||
bj = timezone(timedelta(hours=8))
|
||||
# meta的timestamp是Unix秒,转北京时间
|
||||
dt = datetime.fromtimestamp(ts, tz=bj)
|
||||
print(f'{dt.strftime("%m-%d %H:%M")} Beijing - {filename}')
|
||||
```
|
||||
|
||||
### 汇总格式(给Doro的报告)
|
||||
按时间顺序列出每份合同:
|
||||
- 文件名 | 发送时间(北京时间) | 审查状态(已审/待审/未审) | 交付文件(有/无)
|
||||
- 去重:同文件名的重发标记"重发,与XX:XX相同"
|
||||
- 最终统计:N份独立合同,X份已审,Y份未审
|
||||
|
||||
### 关键陷阱
|
||||
1. **gateway.log里文件消息的msg字段为空** — QiuTing发送文件时,gateway记录`msg=''`,看不到文件名。必须查meta文件才能知道发了什么文件
|
||||
2. **服务器时区是UTC,meta时间戳是Unix秒** — 必须转UTC+8才是北京时间,不能直接显示
|
||||
3. **待审查目录可能有未清理的旧文件** — 不能只看待审查目录判断今天收了几份,要按meta时间戳筛选
|
||||
|
||||
## 关键陷阱
|
||||
|
||||
1. **重发≠重审**:邱律师有时重新发送已审查过的文件(今天5/12份是这种情况),不代表需要重新审查。必须报告Doro确认。
|
||||
2. **auto_notify可能漏文件**:`auto_notify_new_file.sh` 监控inotify有时不触发(进程挂掉、文件到达太快等),不能假设自动化已处理所有新文件。手动核对是必要的。
|
||||
3. **xlsx需要sudo+venv Python**:文件权限属于www-data,必须用 `sudo /home/maggie/.hermes/hermes-agent/venv/bin/python3` 读取。
|
||||
4. **同名不同版**:文件名相同但日期不同的可能是新版本(如 `购销合同 朱家角` 在6/27和6/29各有一份),需比较内容确认是否修改稿。
|
||||
+298
@@ -0,0 +1,298 @@
|
||||
# Multi-Party Revision Reconciliation (多方修订对账)
|
||||
|
||||
## 场景
|
||||
|
||||
合同经多方修订(如WB、华诚-Z、Adon hase等不同修订人),需要:
|
||||
1. 核对叠加修订后的最终效果是否符合协商一致的商业条件
|
||||
2. 对比最终版与模板的差异
|
||||
3. 评估差异对特定方权利义务的影响
|
||||
4. 统一修订人署名
|
||||
|
||||
## 技术方法
|
||||
|
||||
### 1. 提取带修订标记的全文(含作者归属)
|
||||
|
||||
```python
|
||||
from docx import Document
|
||||
from lxml import etree
|
||||
|
||||
ns = {'w': 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'}
|
||||
W = '{http://schemas.openxmlformats.org/wordprocessingml/2006/main}'
|
||||
doc = Document('contract.docx')
|
||||
body = doc.element.body
|
||||
|
||||
# 统计修订作者
|
||||
authors = set()
|
||||
for elem in body.iter():
|
||||
author = elem.get(f'{W}author')
|
||||
if author:
|
||||
authors.add(author)
|
||||
|
||||
# 分作者统计插入/删除
|
||||
ins_by_author = {}
|
||||
del_by_author = {}
|
||||
for ins in body.findall(f'.//{W}ins'):
|
||||
a = ins.get(f'{W}author', 'unknown')
|
||||
ins_by_author[a] = ins_by_author.get(a, 0) + 1
|
||||
for d in body.findall(f'.//{W}del'):
|
||||
a = d.get(f'{W}author', 'unknown')
|
||||
del_by_author[a] = del_by_author.get(a, 0) + 1
|
||||
```
|
||||
|
||||
### 2. 带修订标记的文本提取
|
||||
|
||||
格式:`[+作者: 插入文本]` / `[-作者: 删除文本]` / 普通文本
|
||||
|
||||
```python
|
||||
def get_text_with_revisions(body):
|
||||
result = []
|
||||
for para in body.findall(f'.//{W}p'):
|
||||
para_text = []
|
||||
for elem in para.iter():
|
||||
if elem.tag == f'{W}ins':
|
||||
author = elem.get(f'{W}author', '?')
|
||||
texts = [t.text for t in elem.findall(f'.//{W}t') if t.text]
|
||||
if texts:
|
||||
para_text.append(f"[+{author}: {''.join(texts)}]")
|
||||
elif elem.tag == f'{W}del':
|
||||
author = elem.get(f'{W}author', '?')
|
||||
texts = [t.text for t in elem.findall(f'.//{W}delText') if t.text]
|
||||
if texts:
|
||||
para_text.append(f"[-{author}: {''.join(texts)}]")
|
||||
elif elem.tag == f'{W}t':
|
||||
in_revision = False
|
||||
p = elem
|
||||
while p is not None:
|
||||
if p.tag in [f'{W}ins', f'{W}del']:
|
||||
in_revision = True
|
||||
break
|
||||
p = p.getparent()
|
||||
if not in_revision and elem.text:
|
||||
para_text.append(elem.text)
|
||||
if para_text:
|
||||
result.append(''.join(para_text))
|
||||
return result
|
||||
```
|
||||
|
||||
### 3. 三版对比分析
|
||||
|
||||
对比维度:
|
||||
- **模板** → 我方标准条款(基准线)
|
||||
- **对方修订版** → 对方修改后接受所有修订的版本(ins=0, del=0 说明已全部接受)
|
||||
- **当前叠加版** → 在WB修订基础上加入另一方修改
|
||||
|
||||
### 4. 协商一致核对表
|
||||
|
||||
按商业条件逐项核对:
|
||||
|
||||
| 协商条件 | 条款位置 | 当前版本内容 | 是否符合 |
|
||||
|---|---|---|---|
|
||||
| 费用40/60/60/80/80 | 3.3条 | 具体金额 | ✅/❌ |
|
||||
| 违约责任按我方 | 第4条 | ... | ✅/⚠️ |
|
||||
|
||||
### 5. 模板偏差影响评估
|
||||
|
||||
对每处偏差评级:
|
||||
- ✅ 对己方有利(扩大了己方权利/对方义务)
|
||||
- ⚠️ 中等影响(条件调整但不改变核心权利义务)
|
||||
- ❗ 重要(实质性削弱己方权利/扩大己方义务/给对方逃出合同的通道)
|
||||
|
||||
## 修订人统一规则
|
||||
|
||||
当需要将多个修订人统一为一个(如统一为WB):
|
||||
|
||||
### 处理优先级
|
||||
1. A修订了B的修订 → **以A为准**(接受A对B的修改)
|
||||
2. A和B各自独立修订 → 两者都保留,统一署名
|
||||
|
||||
### 核心技术挑战:嵌套修订
|
||||
|
||||
**关键结构:B(华诚-Z)在A(WB)的 `w:ins` 内部做了 `w:del`**
|
||||
|
||||
XML表现为:
|
||||
```xml
|
||||
<w:ins w:author="WB">
|
||||
<w:r><w:t>个月未给乙方安排工作的,</w:t></w:r>
|
||||
<w:del w:author="华诚-Z">
|
||||
<w:r><w:delText>拍摄</w:delText></w:r>
|
||||
</w:del>
|
||||
</w:ins>
|
||||
```
|
||||
含义:WB插入了"个月未给乙方安排拍摄工作的,",华诚-Z在WB的插入中删除了"拍摄",最终效果="个月未给乙方安排工作的,"。
|
||||
|
||||
### 三步统一流程(2026-07-03 MCN模特合同实证)
|
||||
|
||||
**Step 1: 接受嵌套删除**(B对A修订的修改)
|
||||
|
||||
```python
|
||||
def accept_nested_deletions(body, inner_author='华诚-Z', outer_author='WB'):
|
||||
"""接受inner_author对outer_author修订的修改(删除嵌套del元素)"""
|
||||
for ins_elem in body.findall(f'.//{W}ins'):
|
||||
if ins_elem.get(f'{W}author') != outer_author:
|
||||
continue
|
||||
# 找到outer_author的ins内部,inner_author做的del
|
||||
for del_elem in ins_elem.findall(f'.//{W}del'):
|
||||
if del_elem.get(f'{W}author') == inner_author:
|
||||
parent = del_elem.getparent()
|
||||
parent.remove(del_elem)
|
||||
|
||||
accept_nested_deletions(body)
|
||||
```
|
||||
|
||||
**Step 2: 清除空壳元素**
|
||||
|
||||
接受嵌套删除后,某些WB的ins可能变空(内容全被华诚-Z删了,华诚-Z在旁边插入了替代文本):
|
||||
|
||||
```python
|
||||
def remove_empty_ins(body):
|
||||
"""删除没有任何文本内容的ins元素"""
|
||||
for ins_elem in body.findall(f'.//{W}ins'):
|
||||
has_text = False
|
||||
for t in ins_elem.findall(f'.//{W}t'):
|
||||
if t.text and t.text.strip():
|
||||
has_text = True
|
||||
break
|
||||
if not has_text:
|
||||
parent = ins_elem.getparent()
|
||||
if parent is not None:
|
||||
parent.remove(ins_elem)
|
||||
|
||||
remove_empty_ins(body)
|
||||
```
|
||||
|
||||
**Step 3: 统一作者名**
|
||||
|
||||
```python
|
||||
def rename_author(body, old_author, new_author):
|
||||
"""修改所有修订元素的author属性"""
|
||||
count = 0
|
||||
for elem in body.iter():
|
||||
if elem.get(f'{W}author') == old_author:
|
||||
elem.set(f'{W}author', new_author)
|
||||
count += 1
|
||||
return count
|
||||
|
||||
count = rename_author(body, '华诚-Z', 'WB')
|
||||
```
|
||||
|
||||
### 在统一后的文件上追加模板修改
|
||||
|
||||
如果发现与模板有偏差需要修正(如恢复模板中的固定金额违约金选项、自动续约条件等),在已统一的文件上新增tracked changes:
|
||||
|
||||
```python
|
||||
from copy import deepcopy
|
||||
|
||||
def get_rPr_from_run(run):
|
||||
"""获取run的格式属性用于新增修订"""
|
||||
rpr = run.find(f'{W}rPr')
|
||||
return deepcopy(rpr) if rpr is not None else None
|
||||
|
||||
def make_ins(text, rPr=None, author='WB', date='2026-07-03T06:00:00Z'):
|
||||
"""创建tracked insertion"""
|
||||
ins = etree.Element(f'{W}ins')
|
||||
ins.set(f'{W}id', str(abs(hash(text)) % 100000))
|
||||
ins.set(f'{W}author', author)
|
||||
ins.set(f'{W}date', date)
|
||||
r = etree.SubElement(ins, f'{W}r')
|
||||
if rPr is not None:
|
||||
r.append(deepcopy(rPr))
|
||||
t = etree.SubElement(r, f'{W}t')
|
||||
t.set('{http://www.w3.org/XML/1998/namespace}space', 'preserve')
|
||||
t.text = text
|
||||
return ins
|
||||
|
||||
def make_del(text, rPr=None, author='WB', date='2026-07-03T06:00:00Z'):
|
||||
"""创建tracked deletion"""
|
||||
d = etree.Element(f'{W}del')
|
||||
d.set(f'{W}id', str(abs(hash(text + 'del')) % 100000))
|
||||
d.set(f'{W}author', author)
|
||||
d.set(f'{W}date', date)
|
||||
r = etree.SubElement(d, f'{W}r')
|
||||
if rPr is not None:
|
||||
r.append(deepcopy(rPr))
|
||||
dt = etree.SubElement(r, f'{W}delText')
|
||||
dt.set('{http://www.w3.org/XML/1998/namespace}space', 'preserve')
|
||||
dt.text = text
|
||||
return d
|
||||
|
||||
# 修改模式:找到目标run → 替换为del(旧) + ins(新)
|
||||
target_run = ... # 找到包含目标文本的w:r元素
|
||||
rPr = get_rPr_from_run(target_run)
|
||||
old_text = target_run.find(f'.//{W}t').text
|
||||
|
||||
del_elem = make_del(old_text, rPr)
|
||||
ins_elem = make_ins(new_text, rPr)
|
||||
|
||||
idx = list(para).index(target_run)
|
||||
para.remove(target_run)
|
||||
para.insert(idx, del_elem)
|
||||
para.insert(idx + 1, ins_elem)
|
||||
```
|
||||
|
||||
### 恢复已被删除的内容(撤回WB之前的del)
|
||||
|
||||
当需要把WB之前删除的内容恢复回来(因为模板中有这个内容):
|
||||
|
||||
```python
|
||||
# 找到WB的del元素
|
||||
for child in para:
|
||||
if child.tag == f'{W}del' and child.get(f'{W}author') == 'WB':
|
||||
del_texts = ''.join(dt.text for dt in child.findall(f'.//{W}delText') if dt.text)
|
||||
if '目标文本' in del_texts:
|
||||
# 策略:删除这个del,插入一个ins代替(因为原文已经"被删"了)
|
||||
inner_r = child.find(f'.//{W}r')
|
||||
inner_rPr = get_rPr_from_run(inner_r)
|
||||
ins_restore = make_ins(del_texts, inner_rPr)
|
||||
idx = list(para).index(child)
|
||||
para.remove(child)
|
||||
para.insert(idx, ins_restore)
|
||||
break
|
||||
```
|
||||
|
||||
## 验证步骤
|
||||
|
||||
统一修订人后必须验证:
|
||||
|
||||
```python
|
||||
# 1. 确认只剩一个作者
|
||||
authors = set()
|
||||
for elem in body.iter():
|
||||
a = elem.get(f'{W}author')
|
||||
if a:
|
||||
authors.add(a)
|
||||
assert authors == {'WB'}, f"Unexpected authors: {authors}"
|
||||
|
||||
# 2. 获取接受所有修订后的最终文本
|
||||
def get_accepted_text(para):
|
||||
texts = []
|
||||
for elem in para.iter():
|
||||
if elem.tag == f'{W}t':
|
||||
in_del = False
|
||||
p = elem
|
||||
while p is not None:
|
||||
if p.tag == f'{W}del':
|
||||
in_del = True
|
||||
break
|
||||
p = p.getparent()
|
||||
if not in_del and elem.text:
|
||||
texts.append(elem.text)
|
||||
return ''.join(texts)
|
||||
|
||||
# 3. 逐条核对关键条款的最终文本
|
||||
```
|
||||
|
||||
## 输出格式
|
||||
|
||||
报告分三部分:
|
||||
1. **协商一致核对** — 逐项确认商业条件是否落实(✅/⚠️/❌)
|
||||
2. **模板差异表** — 与标准模板的偏差点 + 对己方权利义务的影响评估
|
||||
3. **结论与建议** — 哪些已符合、哪些需关注、是否需要进一步协商
|
||||
|
||||
## 注意事项
|
||||
|
||||
- 对方修订版如果修订已全部接受(ins=0, del=0),说明是"接受所有修订后"的干净版本
|
||||
- 比较时需要分别提取:(1)当前版本接受所有修订后的最终文本;(2)带修订标记的过程文本
|
||||
- 对方修改了措辞但实质不变的情况(如"整容"→"形象变化"),需评估措辞变化是否改变法律效果——措辞更宽泛时分析对哪方有利
|
||||
- 费用条款的结构性改写(如合并年度/拆分年度)要验证数学正确性(年总额×期数=合同总额)
|
||||
- 处理段落29-31连读场景:华诚-Z可能将原来的多个年度段落(第三、四年/第五年分开写)合并为一个(第四、第五年),导致中间段落被del清空,需要把前后段落连起来读才能看到完整的费用结构
|
||||
- **修订人统一后的检查清单**:作者唯一性、ins/del总数合理、关键条款最终文本正确、文件可正常打开
|
||||
@@ -0,0 +1,76 @@
|
||||
# 屋顶光伏租赁合同审查要点(出租方/甲方视角)
|
||||
|
||||
> 来源:2026-06-15 莎莎委托审查"全额上网 租赁合同"的实践总结
|
||||
> 适用:代表屋顶业主(出租方)审查光伏投资方提供的格式合同
|
||||
|
||||
## 合同性质
|
||||
|
||||
- 司法实践认定为**综合性无名合同**(最高院2020年判决),包含:
|
||||
- 租赁关系 → 适用《民法典》合同编租赁合同章
|
||||
- 电力供应关系 → 适用供用电合同相关规定
|
||||
- **没有专门立法**,多部法律法规叠加适用
|
||||
|
||||
## 三种上网模式
|
||||
|
||||
| 模式 | 含义 | 政策限制 |
|
||||
|------|------|---------|
|
||||
| 全额上网 | 发电全部卖给电网 | 2025.5.1后工商业分布式**禁止** |
|
||||
| 自发自用余电上网 | 优先自用,余电卖电网 | 各省有自用比例要求(30%-80%) |
|
||||
| 全部自发自用 | 全部自用,不上网 | 需防逆流装置 |
|
||||
|
||||
**关键政策**:《分布式光伏发电开发建设管理办法》(国能发新能规〔2025〕7号),2025年1月23日发布
|
||||
|
||||
## 甲方视角核心审查清单
|
||||
|
||||
### 一、合同有效性
|
||||
- [ ] 上网模式是否符合现行政策(全额上网已禁止工商业分布式)
|
||||
- [ ] 租赁期限是否超过20年(《民法典》第705条上限)
|
||||
- [ ] "自动续签"条款是否可执行(≠重新签合同,超出部分可能无效)
|
||||
- [ ] 建筑是否有合法手续(无建设工程规划许可证→租赁合同无效,司法解释第2条)
|
||||
- [ ] 是否存在在先抵押/查封(先抵后租→买卖不破租赁不适用,司法解释第14条)
|
||||
|
||||
### 二、租金保障
|
||||
- [ ] 首次租金支付有无明确的并网截止期限
|
||||
- [ ] 是否以甲方先开票为付款前提(税务资质限制风险)
|
||||
- [ ] 有无租金递增机制(20年合同无递增→购买力大幅贬值)
|
||||
- [ ] 乙方逾期付租有无违约责任
|
||||
- [ ] 日期是否正确(常见错误:6月31日、2月30日等)
|
||||
|
||||
### 三、违约责任对等性
|
||||
- [ ] 甲方违约赔偿范围和计算方式(常见:剩余年限全部预期收益,金额巨大)
|
||||
- [ ] 乙方违约责任是否对等(典型格式合同:甲方3条详尽、乙方仅1条极简)
|
||||
- [ ] 赔偿有无上限
|
||||
- [ ] 甲方有无主动解除权(格式合同常常不给甲方任何解除权)
|
||||
|
||||
### 四、甲方义务合理性
|
||||
- [ ] 甲方能否自由进入自己的屋顶(常见:需乙方书面同意)
|
||||
- [ ] 甲方对第三方行为是否承担连带责任
|
||||
- [ ] 是否需免费提供大量场地设施(设备房、电缆通道等)
|
||||
- [ ] 乙方转让/质押资产是否仅需"通知"甲方
|
||||
- [ ] 有无开放式义务(如"甲方应积极支持")
|
||||
|
||||
### 五、其他风险
|
||||
- [ ] 拆迁补偿分配(格式合同常约定全归乙方)
|
||||
- [ ] 管辖法院(格式合同常约定乙方所在地)
|
||||
- [ ] 政策变化解除时甲方有无补偿
|
||||
- [ ] 有无保险要求
|
||||
- [ ] 合同到期后设备拆除和屋顶恢复义务
|
||||
|
||||
## 关键法律依据
|
||||
|
||||
| 法律法规 | 条文 | 内容 |
|
||||
|---------|------|------|
|
||||
| 《民法典》第705条 | 租赁期限 | 不得超过20年,超过部分无效 |
|
||||
| 《民法典》第725条 | 买卖不破租赁 | 所有权变动不影响租赁效力 |
|
||||
| 房屋租赁司法解释第2条 | 合同无效 | 无建设工程规划许可证→租赁合同无效 |
|
||||
| 房屋租赁司法解释第14条 | 不受保护 | 先抵押/查封后租赁→不受保护 |
|
||||
| 《企业破产法》第18条 | 管理人选择权 | 破产后管理人可解除未完合同 |
|
||||
| 分布式光伏管理办法(2025) | 上网模式 | 工商业禁止全额上网 |
|
||||
|
||||
## 司法实践要点
|
||||
|
||||
- **零租金合同风险**:法院可能不认定租赁关系→无法适用买卖不破租赁(浙06民终4075号)
|
||||
- **BAPV vs BIPV**:附加式光伏通常不构成添附(浙02民终4005号),一体化光伏可能构成添附
|
||||
- **无规划许可证**:直接认定租赁协议无效(京民申2129号)
|
||||
- **屋顶漏水**:法院严格审查因果关系,需第三方鉴定报告
|
||||
- **预期收益损失**:破产清算中难以获得支持,需合同中明确约定计算方式
|
||||
@@ -0,0 +1,120 @@
|
||||
# Template Comparison Methodology (标准模版对比)
|
||||
|
||||
## When to use
|
||||
When a client has a standard template and wants to know how their signed contracts deviate from it.
|
||||
This is an ADD-ON to the portfolio audit workflow — produces an extra column in the Excel summary.
|
||||
|
||||
## Setup
|
||||
1. Locate template file(s) in the client's 参考文件/ directory
|
||||
2. Read template with python-docx (if .docx) or OCR (if .pdf)
|
||||
3. Map template articles to a checklist
|
||||
|
||||
## Comparison Dimensions
|
||||
|
||||
### 1. Missing clauses (模版有但合同缺失)
|
||||
Check each template article against the actual contract. Flag:
|
||||
- Entire articles/sections missing
|
||||
- Sub-clauses within a section that were dropped
|
||||
- Protective language removed (e.g., "且不承担违约责任")
|
||||
- Quantitative requirements removed (e.g., time limits, penalty caps)
|
||||
|
||||
Risk levels:
|
||||
- 🔴 高:directly impacts party's core rights (押金退还、法定解除权、发票保护)
|
||||
- ⚠️ 中:reduces protection but workarounds exist
|
||||
- Low: cosmetic or minor scope differences
|
||||
|
||||
### 2. Added clauses (合同有但模版没有)
|
||||
Note whether additions are:
|
||||
- ✅ Favorable to the client
|
||||
- ❌ Unfavorable (e.g., additional burden of proof requirements)
|
||||
- Neutral
|
||||
|
||||
### 3. Substantive differences (实质性差异)
|
||||
Side-by-side of same clause with different wording/numbers:
|
||||
- Notice periods (e.g., 1 month → 3 months)
|
||||
- Penalty rates (e.g., 0.1‰ vs 0.1% — **watch for OCR errors on ‰**)
|
||||
- Scope qualifiers added/removed (e.g., "严重影响" vs "影响")
|
||||
- Rights downgraded (e.g., "征得同意" → "通知")
|
||||
- Obligation standards changed (e.g., "按现状交付" → "装修保持原状")
|
||||
|
||||
## Key articles to always check (rental contracts)
|
||||
|
||||
| Template Article | What to check |
|
||||
|-----------------|---------------|
|
||||
| 发票条款 | 甲方不开票→乙方可延付不违约?税务损失由谁担? |
|
||||
| 押金退还 | 仅"期满"还是"期满/解除/终止"均可退? |
|
||||
| 单方解除 | 有无"除法定或本合同约定外"限定? |
|
||||
| 甲方违约赔偿 | 是否含退押金+退预付款+装修损失+诉讼费律师费? |
|
||||
| 不可抗力范围 | 是否含"政府政策变更"+"行业治理"?还是仅限特定事件? |
|
||||
| 无法办证解除 | "房屋本身原因+政策原因"还是仅"政策原因"? |
|
||||
| 续租/优先权时限 | 通知期多长?模版通常1个月 |
|
||||
| 看房权 | "征得同意"还是"通知"? |
|
||||
| 迁离标准 | "按现状"还是"保持原状"(暗示恢复义务)? |
|
||||
| 甲方备案义务 | 有无时限?不配合有无违约后果? |
|
||||
| 竞业限制范围 | "大楼"还是"大楼、商圈等"? |
|
||||
| 补充条款 | 是否利用了模版预留的补充条款空间? |
|
||||
|
||||
## Output format for L column (与标准模版差异)
|
||||
|
||||
For contracts with MAJOR deviations:
|
||||
```
|
||||
⚠️ 与标准模版存在多处重大偏离:
|
||||
1. 【条款名】模版要求XX;本合同YY
|
||||
2. 【条款名】模版要求XX;本合同YY
|
||||
...
|
||||
✅ [any positive deviation noted]
|
||||
```
|
||||
|
||||
For contracts largely matching template:
|
||||
```
|
||||
✅ 与标准模版高度一致,仅以下小偏差:
|
||||
1. 【条款名】差异描述
|
||||
...
|
||||
其余主要条款均与模版一致
|
||||
```
|
||||
|
||||
For contracts SAME-SOURCE as template but with substantive deviations (本合同就是 07 模版填空而成、骨架同源,但个别关键条款被改动/删除 —— 人民中路即此类,最易被误判为"独立友好范本"):
|
||||
```
|
||||
✅ 本合同与 07 标准模版同源(系模版填空而成,条款体系/编号/措辞逐条对应),但有以下实质偏离:
|
||||
1. 🔴【条款名】模版为XX;本合同被改为YY(对乙方不利/缺失关键保护)
|
||||
2. 🔴【条款名】模版有XX条款;本合同整条删除
|
||||
...
|
||||
其余条款与模版一致(含0.1‰逾期、含疫情/行业治理、优先权等模版标配——属同源固有,非本合同特有优势)
|
||||
```
|
||||
> ⚠️ **此档对治的典型错误(人民中路实证)**:把模版本身就有的标配条款(0.1‰、含疫情、优先权)当成"本合同特别友好"列为优势,却漏掉真正被改动/删除的实质偏离(抵押"不得→可"、办学许可证免责款被删)。判别口诀:**先认"是不是模版填空而成"——是 → 用本档,差异只写"被改/被删"的地方,绝不把模版标配当本合同优势。**
|
||||
|
||||
For non-comparable documents:
|
||||
```
|
||||
补充协议,非模版对比范围
|
||||
```
|
||||
or
|
||||
```
|
||||
物业服务协议,无对应标准模版
|
||||
```
|
||||
|
||||
## Implementation pattern
|
||||
|
||||
Use `delegate_task` with two parallel subagents:
|
||||
1. **Structural comparison subagent**: reads template + actual contracts, produces clause-by-clause diff
|
||||
2. **Thematic extraction subagent**: extracts specific clause categories the client cares about (e.g., termination, modification, penalties)
|
||||
|
||||
Pass to each subagent:
|
||||
- Template full text
|
||||
- All contract .md files with their table entry numbers
|
||||
- Clear focus instructions
|
||||
|
||||
After delegation, verify outputs exist and are coherent before incorporating into Excel.
|
||||
|
||||
## Thematic risk additions to K column
|
||||
|
||||
When the client asks for specific risk categories (e.g., 合同变更与提前解除):
|
||||
- APPEND to existing K column content with a clear sub-header: `【合同变更与提前解除】`
|
||||
- Include: notice period, penalty formula, deposit handling, practical implications
|
||||
- Flag favorable precedents (e.g., a previously successful partial surrender with zero penalty)
|
||||
- In section 四, add a comprehensive summary with cost estimates for all scenarios
|
||||
|
||||
## OCR artifact awareness
|
||||
- ‰ (千分号) frequently OCR'd as % — flag for manual verification on originals
|
||||
- Table cells may have garbled column alignment
|
||||
- Stamps/seals over text cause character corruption
|
||||
- Article numbering may be inconsistent in OCR output
|
||||
@@ -0,0 +1,109 @@
|
||||
# Workflow 系统性问题清单(2026-06-27 至 2026-07-01)
|
||||
|
||||
## 1. auto_notify 静默失效(模式B)
|
||||
|
||||
**表现**:进程在跑但 inotifywait 没有捕获事件,日志完全为空。
|
||||
**根因**:inotifywait 文件描述符失效/内核事件丢失。
|
||||
**现状**:watchdog cron (`63bb31d4f050`) 只能重启进程,无法修复内核级别的事件丢失。
|
||||
**Fallback**:手动检查 `~/.hermes/cache/documents/` + relay 脚本串行启动 workflow。
|
||||
|
||||
## 2. 非合同文件误审
|
||||
|
||||
**表现**:香花桥家庭医生招标需求被当合同审查。
|
||||
**根因**:classifier 的 `not_a_contract` 逻辑虽然写了(第13-19行),但 LLM 执行时没走到——可能因标题含"服务""项目"等关键词触发了合同判断。
|
||||
**修复状态**:通知脚本已存在(`wecom_group_notify.py`),routing 已配置(`$END`)。需加强 classifier prompt 中对招标/技术需求的识别。
|
||||
|
||||
## 3. 审查意见不该生成
|
||||
|
||||
**表现**:非卫生中心合同(如生育友好-计生协会)也自动生成了审查意见。
|
||||
**根因**:workflow 对所有合同统一执行 editor 生成审查意见的逻辑,未按 ruleset_type 区分。
|
||||
**修复方向**:在 editor procedure 中加条件判断——只有 ruleset_type=health-centers 且 review-rules.md 中明确要求"审查意见"交付物时才生成。
|
||||
|
||||
## 4. 文件名 hash 前缀残留
|
||||
|
||||
**表现**:`【修】doc_0ad4ee63bff5_印刷品制作合同2026.6(1).docx`
|
||||
**根因**:deliverer 从 cache 路径取文件名时未剥离 `doc_[0-9a-f]{12}_` 前缀。
|
||||
**修复状态**:editor procedure step 7 已有清理逻辑(第147-149行),但 deliverer 步骤可能绕过了 editor 的命名。
|
||||
|
||||
## 5. 新增条款插入位置错乱
|
||||
|
||||
**表现**:健康积分合同新增"数据归属""转包分包"条款标题和内容分离。
|
||||
**根因**:editor 用 python-docx 插入段落时,对文档结构的理解不够(插入在错误的 anchor 位置)。
|
||||
**自愈**:reviewer 第2轮检出并返回 editor 修复,第3轮通过。但多花了2轮=多花10分钟+token。
|
||||
|
||||
## 6. idle thread 不自动继续
|
||||
|
||||
**表现**:反委托代发的 thread 卡在 idle classifier。
|
||||
**根因**:auto_notify 启动了 thread (`uwf thread start`) 但没有执行 (`uwf thread exec --count 20 --background`)。
|
||||
**修复**:需要确保 relay/queue 脚本在 start 后立即 exec。
|
||||
|
||||
## 7. /tmp/contract-review 被僵尸线程污染(2026-07-01)
|
||||
|
||||
**表现**:新启动的 classifier 在 /tmp/contract-review/ 中看到旧线程残留的文件(其他合同的【修】文件、.py脚本、.pdf等),导致分类器混乱或长时间卡住。
|
||||
**根因**:
|
||||
- queue-runner.sh 和手动启动并存,多个 thread exec 进程同时操作 /tmp/contract-review/
|
||||
- 僵尸线程(6月26日、6月30日的stuck threads)被queue-runner定期唤醒但无法推进,占用/tmp/contract-review/
|
||||
- `uwf thread show` 显示 "running" 但实际 uwf-hermes 子进程已死或卡住
|
||||
|
||||
### 7b. Queue-runner 自动启动重复线程(2026-07-02 恭兴合同教训)
|
||||
|
||||
**表现**:手动启动的 thread `06FJ1DY529R7C85BDPBQG1QBBW` 和 queue-runner 自动启动的 `06FJ1DVM3PCZYVDXAAXE1DTYBG` 同时处理恭兴合同.docx,两个 uwf-hermes reviewer 进程同时写 /tmp/contract-review/【修】恭兴合同.docx。
|
||||
**根因**:queue-runner.sh 检测到 Nextcloud 待审查目录中的新文件后自动 `uwf thread start + exec`,与手动操作撞车。
|
||||
**症状**:`ps aux | grep uwf-hermes` 可以看到**两个**不同 thread-id 的 reviewer/editor 进程,prompt 里有相同的文件名。
|
||||
**修复**:手动管理合同时,**必须先 kill queue-runner**:
|
||||
```bash
|
||||
kill $(ps aux | grep "[c]ontract-queue-runner" | awk '{print $2}') 2>/dev/null
|
||||
```
|
||||
确认只有自己的线程在跑后再继续。如果发现重复线程已经启动,cancel 它并 kill 其进程。
|
||||
**教训**:仅 cancel thread 不够——对应的 uwf-hermes 子进程可能仍在运行(进程不感知 thread status 变化),必须同时 kill 进程 PID。
|
||||
|
||||
**清理步骤(批量开工前必做)**:
|
||||
```bash
|
||||
# 1. 查看正在运行的uwf进程
|
||||
ps aux | grep "[u]wf-hermes" | grep -v grep
|
||||
ps aux | grep "[u]wf.*thread exec" | grep -v grep
|
||||
|
||||
# 2. 杀死所有uwf-hermes子进程
|
||||
kill $(ps aux | grep "[u]wf-hermes" | grep -v grep | awk '{print $2}') 2>/dev/null
|
||||
|
||||
# 3. 杀死queue-runner
|
||||
kill $(ps aux | grep "[c]ontract-queue-runner" | awk '{print $2}') 2>/dev/null
|
||||
|
||||
# 4. 取消所有非当前的stuck threads
|
||||
/home/maggie/.hermes/node/bin/uwf thread list | grep running
|
||||
# 对每个非当前任务的thread: uwf thread cancel <id>
|
||||
|
||||
# 5. 清理/tmp/contract-review(保留rules/)
|
||||
cd /tmp/contract-review
|
||||
rm -f *.docx *.docx.bak *.py *.pdf *.png *.yaml *.xml 【修】*.docx 2>/dev/null
|
||||
|
||||
# 6. 确认干净
|
||||
ls /tmp/contract-review/ | grep -v rules
|
||||
# 应该为空
|
||||
|
||||
# 7. 复制新合同文件进去,启动thread
|
||||
```
|
||||
|
||||
**预防**:串行处理合同(一份完成再启动下一份)。`uwf thread show` 显示 end 后再清理+启动下一份。
|
||||
|
||||
## 8. 批量合同串行处理实操(2026-07-01 四份合同)
|
||||
|
||||
**场景**:邱律师一次发2-4份合同,需要逐份串行审查。
|
||||
**正确流程**:
|
||||
1. 收到所有文件 → 上传Nextcloud待审查目录 → 复制到本地shared目录
|
||||
2. 第1份:清理/tmp → 复制文件 → `uwf thread start` → `uwf thread exec --count 20 --background` → 等`status=end`
|
||||
3. 第2份:清理/tmp → 复制文件 → 新thread → exec → 等end
|
||||
4. 全部完成后:统一更新tracker、不逐份通知(通知只在全部完成后发一次)
|
||||
|
||||
**等待策略**:
|
||||
- `sleep N && uwf thread show <id>` 轮询,每轮5-10分钟
|
||||
- 典型合同全程约30-60分钟(classifier 2-3min → reviewer 5-10min → editor 5-10min → reviewer复核 5-10min → 可能再editor → final_review 3-5min → deliverer 2-3min)
|
||||
- 如果同一个role卡超过15分钟没变化,检查 `ps aux | grep uwf-hermes` 进程是否存活
|
||||
|
||||
## 待修复优先级
|
||||
|
||||
1. 【高】加强 classifier 对非合同文件的识别(prompt 强化)
|
||||
2. 【中】审查意见按 ruleset_type 条件生成
|
||||
3. 【中】/tmp/contract-review 隔离(考虑按thread_id创建子目录,避免多线程冲突)
|
||||
4. 【低】auto_notify 根因(可能需要换 inotifywait 为 polling 方案)
|
||||
5. 【低】新增条款插入位置的准确性(需要更结构化的段落插入逻辑)
|
||||
@@ -0,0 +1,116 @@
|
||||
"""
|
||||
Standalone function to add comments to a .docx file using zipfile + lxml.
|
||||
Use AFTER ContractEditor.save() since ContractEditor rewrites document.xml.
|
||||
|
||||
Usage:
|
||||
from add_comments_to_docx import add_comments_to_docx
|
||||
# comments = [(anchor_text, comment_text), ...]
|
||||
placed = add_comments_to_docx('/tmp/【修】contract.docx', comments)
|
||||
"""
|
||||
import zipfile, io
|
||||
from datetime import datetime
|
||||
from lxml import etree
|
||||
|
||||
W = 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'
|
||||
R_NS = 'http://schemas.openxmlformats.org/officeDocument/2006/relationships'
|
||||
|
||||
def qn(tag):
|
||||
return f'{{{W}}}{tag}'
|
||||
|
||||
def add_comments_to_docx(filepath, comments, author='WB'):
|
||||
"""Add comments to a docx file.
|
||||
|
||||
Args:
|
||||
filepath: path to docx (modified in place)
|
||||
comments: list of (anchor_text, comment_text) tuples
|
||||
author: comment author name (default 'WB')
|
||||
|
||||
Returns:
|
||||
int: number of comments successfully placed
|
||||
"""
|
||||
with open(filepath, 'rb') as f:
|
||||
data = f.read()
|
||||
zin = zipfile.ZipFile(io.BytesIO(data))
|
||||
buf = io.BytesIO()
|
||||
zout = zipfile.ZipFile(buf, 'w', zipfile.ZIP_DEFLATED)
|
||||
doc_xml = zin.read('word/document.xml')
|
||||
doc_tree = etree.fromstring(doc_xml)
|
||||
body = doc_tree.find(qn('body'))
|
||||
|
||||
nsmap = {'w': W, 'r': R_NS}
|
||||
comments_xml = etree.Element(qn('comments'), nsmap=nsmap)
|
||||
|
||||
comment_id = 200
|
||||
placed = 0
|
||||
|
||||
for anchor_text, comment_text in comments:
|
||||
cid = str(comment_id)
|
||||
comment_id += 1
|
||||
|
||||
# Create comment element
|
||||
comment_el = etree.SubElement(comments_xml, qn('comment'))
|
||||
comment_el.set(qn('id'), cid)
|
||||
comment_el.set(qn('author'), author)
|
||||
comment_el.set(qn('date'), datetime.now().strftime('%Y-%m-%dT%H:%M:%SZ'))
|
||||
cp = etree.SubElement(comment_el, qn('p'))
|
||||
cr = etree.SubElement(cp, qn('r'))
|
||||
ct = etree.SubElement(cr, qn('t'))
|
||||
ct.set('{http://www.w3.org/XML/1998/namespace}space', 'preserve')
|
||||
ct.text = comment_text
|
||||
|
||||
# Find anchor in document
|
||||
for p in body.iter(qn('p')):
|
||||
runs = p.findall(f'.//{qn("r")}')
|
||||
full = ''.join(''.join(t.text or '' for t in r.findall(qn('t'))) for r in runs)
|
||||
if anchor_text in full:
|
||||
# Insert commentRangeStart at beginning of paragraph
|
||||
crs = etree.Element(qn('commentRangeStart'))
|
||||
crs.set(qn('id'), cid)
|
||||
p.insert(0, crs)
|
||||
|
||||
# Append commentRangeEnd + reference
|
||||
cre = etree.Element(qn('commentRangeEnd'))
|
||||
cre.set(qn('id'), cid)
|
||||
p.append(cre)
|
||||
|
||||
ref_run = etree.SubElement(p, qn('r'))
|
||||
ref_rpr = etree.SubElement(ref_run, qn('rPr'))
|
||||
ref_style = etree.SubElement(ref_rpr, qn('rStyle'))
|
||||
ref_style.set(qn('val'), 'CommentReference')
|
||||
ref_cr = etree.SubElement(ref_run, qn('commentReference'))
|
||||
ref_cr.set(qn('id'), cid)
|
||||
|
||||
placed += 1
|
||||
break
|
||||
|
||||
# Rewrite zip
|
||||
for item in zin.namelist():
|
||||
if item == 'word/document.xml':
|
||||
zout.writestr(item, etree.tostring(doc_tree, xml_declaration=True, encoding='UTF-8', standalone=True))
|
||||
elif item == '[Content_Types].xml':
|
||||
ct_xml = zin.read(item)
|
||||
ct_tree = etree.fromstring(ct_xml)
|
||||
if not any('comments.xml' in (el.get('PartName') or '') for el in ct_tree):
|
||||
override = etree.SubElement(ct_tree, 'Override')
|
||||
override.set('PartName', '/word/comments.xml')
|
||||
override.set('ContentType', 'application/vnd.openxmlformats-officedocument.wordprocessingml.comments+xml')
|
||||
zout.writestr(item, etree.tostring(ct_tree, xml_declaration=True, encoding='UTF-8', standalone=True))
|
||||
elif item == 'word/_rels/document.xml.rels':
|
||||
rels_xml = zin.read(item)
|
||||
rels_tree = etree.fromstring(rels_xml)
|
||||
if not any('comments.xml' in (el.get('Target') or '') for el in rels_tree):
|
||||
rel = etree.SubElement(rels_tree, 'Relationship')
|
||||
rel.set('Id', f'rId{len(rels_tree) + 10}')
|
||||
rel.set('Type', 'http://schemas.openxmlformats.org/officeDocument/2006/relationships/comments')
|
||||
rel.set('Target', 'comments.xml')
|
||||
zout.writestr(item, etree.tostring(rels_tree, xml_declaration=True, encoding='UTF-8', standalone=True))
|
||||
else:
|
||||
zout.writestr(item, zin.read(item))
|
||||
|
||||
zout.writestr('word/comments.xml', etree.tostring(comments_xml, xml_declaration=True, encoding='UTF-8', standalone=True))
|
||||
zout.close()
|
||||
zin.close()
|
||||
|
||||
with open(filepath, 'wb') as f:
|
||||
f.write(buf.getvalue())
|
||||
return placed
|
||||
Reference in New Issue
Block a user