Просмотр исходного кода

feat: skill-case-get v1.0.0 — 案例采集/拆解/归档技能包(可独立分发)

FmodeAgent 7 часов назад
Сommit
2cb6f5dc13

+ 4 - 0
.gitignore

@@ -0,0 +1,4 @@
+node_modules/
+outputs/
+*.log
+.DS_Store

+ 120 - 0
README.md

@@ -0,0 +1,120 @@
+# skill-case-get
+
+> 案例采集 / 拆解 / 归档技能包 —— 把各种素材变成**可入库的案例**。
+
+```bash
+npx skill-case-get@latest workspace          # 装进 ./.claude/skills/skill-case-get
+npx skill-case-get@latest selftest           # 本机自检(不联网)
+```
+
+```
+收到素材 → 拆解分析 → 能否进案例库 → 打标 → 上传归档(待审清单)
+```
+
+## 解决什么问题
+
+同事在微信里连发 9 张图,说「这是上次 IC 那个学生的案例」。
+这 **不是 9 个素材,是 1 个案例**,图要按九宫格阅读顺序排。以前只能手工数、手工起名、手工打标。
+
+本技能把它变成一条可重复的管线:
+
+1. **识别**来源:连续多图 / Word / PPT / 图文混合 / 裸图 / 视频
+2. **拆解**:图片走 OCR + 语义理解;`.docx`/`.pptx` 用 python3 标准库解包;视频抽帧 + 转写
+3. **分组排序**:尺寸 / 时间 / 内容三信号判同组,九宫格给显式 `order`(左上 → 右下)
+4. **分离**:「案例说明」进案例字段,「图片视频」进 `materialAssets[]`,不混
+5. **合规**:能不能进库?是客诉负面吗?有无授权?
+6. **打标**:静态属性 + 跟进状态七段 + 亮点/场景/异议 + 学校别名归一
+7. **归档**:`caseSubmit` 写待审清单(`pending`,不进公共素材库)
+8. **自进化**:新标签追加回 `references/tag-dictionary.json`
+
+## 30 秒上手
+
+```bash
+# 只看分组对不对(不调模型、不联网)
+npx skill-case-get@latest plan -- --image g1.jpg --image g2.jpg --image g9.jpg
+
+# 完整采集(推荐先 dry-run,确认后再 submit)
+npx skill-case-get@latest ingest -- \
+  --image g1.jpg --image g9.jpg \
+  --doc 案例稿.docx \
+  --authorization authorized \
+  --asset-base-url https://cdn.example.com/cases
+
+# 确认顺序无误 → 提交待审
+npx skill-case-get@latest submit -- \
+  --package case-get-outputs/<runId>/case-package.json
+```
+
+`plan` / `ingest` / `submit` 之后要加 `--`,参数透传给运行器。
+
+## 目录结构
+
+```
+skills/skill-case-get/
+├── SKILL.md                      技能正文(frontmatter 参照 skill-vision / skill-listen)
+├── README.md                     本文件
+├── package.json                  name/version/bin
+├── skill-package-manifest.json   分发清单
+├── bin/skill-case-get.js         安装器 + 运行器透传
+├── references/tag-dictionary.json  标签字典(维度/取值/别名,可增量更新)
+└── scripts/
+    ├── case-intake.mjs           主入口:plan / ingest / submit / reorder / dictionary / selftest
+    ├── office_extract.py         .docx/.pptx 解包(纯标准库 zipfile + ElementTree)
+    ├── vision-bridge.mjs         图片理解(复用 skill-vision)
+    ├── listen-bridge.mjs         视频抽帧 + 转写(ffmpeg + skill-listen)
+    ├── cloud-client.mjs          云函数 caseSubmit / casePendingList
+    ├── grouping.mjs              分组 + 九宫格顺序
+    ├── classify.mjs              案例/素材分离、打标、别名归一、合规判断
+    ├── image-size.mjs            零依赖图片尺寸探测(PNG/JPEG/GIF/WebP/BMP)
+    ├── lib.mjs                   共享工具(路径/字典/幂等/素材规范化)
+    ├── smoke.mjs                 冒烟检查
+    └── tests/                    node --test 单测(34 个)
+```
+
+## 关键设计
+
+**为什么用 python3 标准库解包 `.docx` / `.pptx`?**
+本机没有 python-docx / python-pptx,Node 也没有对应库,而 `.docx`/`.pptx` 本质就是 zip。
+`zipfile` + `xml.etree.ElementTree` 足够抽出文字段落、表格、内嵌媒体、超链接。
+**零依赖**意味着技能可以直接被拷进 `~/.claude/skills/` 而不用 `npm install`。
+
+**分组为什么要求「至少 2 种信号」?**
+只看内容相似度会把两份不同学生的相似截图并成一组;只看时间会把同一时刻收到的不同素材并成一组。
+尺寸(同源同排版)+ 时间(连发)+ 内容(共同话术)三者取二,才既不漏并也不误并。
+
+**为什么不自动上传二进制?**
+项目规则明确「不新建独立业务服务器处理普通 CRUD」。技能只做本机拆解与归档编排,
+媒体由既有 CDN / Parse Files 承载,用 `--asset-base-url` 派生 URL。没给就明确警告,而不是造一个假 URL。
+
+## 依赖
+
+| 依赖 | 必需 | 用途 |
+| --- | --- | --- |
+| Node ≥ 18 | ✅ | 运行器(含 `fetch`、`node --test`) |
+| python3 ≥ 3.8 | ✅ | `.docx` / `.pptx` 解包(仅标准库) |
+| skill-vision | 推荐 | 图片 OCR + 语义理解;宿主自带多模态时由宿主读图 |
+| skill-listen | 可选 | 视频音轨转写(走 Fmode 网关,凭据仅服务端) |
+| ffmpeg / ffprobe | 可选 | 视频抽帧、时长探测 |
+
+缺任何一个可选依赖都不会伪造结果——会在 `warnings` 里说清楚。
+
+## 安全红线
+
+- **只处理已获授权素材**:`authorizationStatus != authorized` 一律不入库。
+- **新案例不进公共素材库**:只写 `reviewStatus='pending'` + `readyForUse=false`,等 `caseApprove` 放行。
+- **PII 不落库**:`privacyFindings[].snippet` 是掩码后的(`138****5678`),真实 PII 只在内存里过一遍。
+- **凭据不落盘**:token 只从环境变量读取,仅内存持有。
+
+详见 `SKILL.md` 与仓库 `CLAUDE.md` / `.claude/rules/data-and-compliance.md`。
+
+## 开发
+
+```bash
+npm run smoke    # 冒烟(含单测)
+npm test         # node --test scripts/tests/
+node scripts/case-intake.mjs selftest
+```
+
+## License
+
+MPL-2.0 — Copyright (c) 2026 未来飞马 Fmode

+ 188 - 0
SKILL.md

@@ -0,0 +1,188 @@
+---
+name: skill-case-get
+description: "案例采集 / 拆解 / 归档。把各种素材(连续多图 / 九宫格、Word .docx、PPT .pptx、图文混合、裸图、视频)拆解为「案例信息」与「素材数组 materialAssets[]」,做合规判断(能否入库 / 客诉负面 / 隐私发现)与打标(学校别名归一 + 跟进状态七段 + 亮点/场景/异议),最后经云函数 caseSubmit 落到待审清单。适用场景:(1) 同事连发 6/9 张图其实是同一个素材,需要按九宫格阅读顺序编号归档;(2) 一份 Word/PPT 案例稿要抽出文字说明并拆出内嵌配图;(3) 判断这堆素材能不能进案例库、是不是客诉负面事件;(4) 按案例库标签体系打标并归一学校别名(IC → 帝国理工)。本 Skill 只处理已获授权素材。"
+description_en: "Case intake, decomposition and archiving. Turns raw material (image grids, .docx, .pptx, mixed text+image, bare images, video) into case fields plus an ordered materialAssets[] array, runs compliance triage and tagging, then submits to the review queue via the caseSubmit cloud function. Use for: (1) a run of 6/9 images that is ONE material needing 3x3 reading-order numbering, (2) extracting text and embedded media from Word/PPT, (3) deciding whether material may enter the case library or is a complaint/negative event, (4) tagging against the case-library tag system and normalising school aliases (IC -> Imperial College London). This skill only processes authorised material."
+version: 1.0.0
+author: Yuyang001 (FmodeAgent)
+license: MPL-2.0
+copyright: "Copyright (c) 2026 未来飞马 Fmode"
+tags: [未来飞马, 智能体技能, 超级技能, 服务级, 案例采集, 案例拆解, 素材归档, FmodeAgent, Hermes Agent, FmodeCode, Claude Code, case-library, case-intake, ocr, material-assets, 九宫格, 打标]
+---
+
+本 Skill 只处理已获授权素材。
+
+# Fmode Case Get — 案例采集 / 拆解 / 归档
+
+## Overview
+
+把素材变成**可入库的案例**:
+
+```
+各种素材 → 拆解分析 → 能否进案例库判断 → 打标 → 上传归档(待审清单)
+```
+
+| 来源类型 | 处理方式 |
+| --- | --- |
+| ① 连续多张图片(朋友圈九宫格) | skill-vision OCR + 语义理解;按尺寸/时间/内容相似度**判为同一组**,按九宫格阅读顺序给显式 `order` |
+| ② Word `.docx` | python3 标准库 `zipfile` + `ElementTree` 解包:文字段落、表格、内嵌媒体、超链接 |
+| ③ PPT `.pptx` | 同上:按放映顺序抽每页标题/正文/备注/图片/表格 |
+| ④ 图文混合 | 文档文字走「案例说明」,内嵌图并进素材组一起编号 |
+| ⑤ 裸图片无文字说明 | 靠视觉理解补 `label` / `description` / `usageSuggestion`;分组回落到尺寸 + 时间信号 |
+| ⑥ 视频(附加) | ffmpeg 抽帧 → 交给图片链路;音轨走 skill-listen 转写 |
+
+**关键区分(必须分开,不能混)**
+
+- 「案例说明」——什么时候用、大概情况、背景 → 案例字段 `title / summary / targetCustomer / usageSuggestion / resultEvidence / outcome`
+- 「图片 / 视频」→ `materialAssets[]`(`order` 1 起,同组内按左上 → 右下)
+
+## 红线(不可协商)
+
+1. **授权前置**:`authorizationStatus != authorized` 一律不入库、不产生案例对象。默认值是 `pending`,必须显式传 `--authorization authorized`。
+2. **不进公共素材库**:采集只写 `reviewStatus='pending'` + `readyForUse=false`,由负责人用 `caseApprove` 放行后才进检索结果。
+3. **不落 PII**:真实 PII 只在本次请求内存中处理。`privacyFindings[].snippet` **本身就是掩码后的**(`138****5678`),不打码不落库。
+4. **不伪造结果**:缺 ffmpeg / skill-vision / skill-listen 时,在 `warnings` 里明说,不编造 OCR 或转写内容。
+5. **不写凭据**:token 只从环境变量 / 本地配置读,仅内存持有,不落盘、不进日志、不进报告。
+
+## 鉴权与依赖
+
+- **云函数**:`POST https://server.lumistedu.com/api/functions`,请求体 `{"path":"/caseSubmit","params":{...},"token":"..."}`。
+  session token 来源:`CASE_PARSE_SESSION_TOKEN` → `PARSE_SESSION_TOKEN`。
+- **图片理解**:复用兄弟技能 `skill-vision` 的 `analyze()`(`~/.claude/skills/skill-vision/scripts/vision-client.mjs`)。
+  该函数在**宿主自带多模态**时(FmodeCode / Claude Code)会返回读图指令而非调 API,此时由宿主 Agent 用自己的 Read 工具读图后补全。
+  可用 `--vision-model <name>` 覆盖,或 `--skip-vision` 跳过。
+- **音视频**:复用兄弟技能 `skill-listen` 的 `listen-runner.mjs`(走 Fmode 网关,凭据仅服务端)。
+- **python3**:解包 `.docx` / `.pptx`,**只用标准库**(本机没有 python-docx / python-pptx)。
+- **ffmpeg / ffprobe**:视频抽帧、探测时长(可选)。
+
+> 缺 token 是「缺配置」,不是「用不了」——请勿点任何付费/充值弹窗。
+
+## 用法
+
+### 命令行
+
+```bash
+# 只有图片(九宫格):先看分组与顺序,不调模型
+npx skill-case-get@latest plan -- --image g1.jpg --image g2.jpg ... --image g9.jpg
+
+# 完整采集:拆解 + 图片理解 + 打标,产出案例包(不提交)
+npx skill-case-get@latest ingest -- \
+  --image g1.jpg --image g9.jpg \
+  --doc 案例稿.docx --deck 复盘.pptx \
+  --authorization authorized \
+  --source-ref wx-group-2026-10 \
+  --asset-base-url https://cdn.example.com/cases
+
+# 只给一个目录,自动按扩展名分流
+npx skill-case-get@latest plan -- --input ./素材目录
+
+# 调顺序(把第 3 张挪到第 1 位)
+npx skill-case-get@latest reorder -- --material-assets <out>/material-assets.json --move "3:1"
+
+# 提交到待审清单
+npx skill-case-get@latest submit -- --package <out>/case-package.json --submitted-by user123
+
+# 自检(含 office 解包、九宫格分组、打标、合规、顺序调整)
+npx skill-case-get@latest selftest
+```
+
+`ingest` / `plan` 之后必须加 `--`,其后参数透传给运行器。
+
+### 在 Node 脚本中调用
+
+技能目录被复制进 `.claude/skills/` 时**不含 node_modules**,运行器是零依赖 ESM:
+
+```js
+import { runPlan, runIngest, runSubmit } from './scripts/case-intake.mjs';
+
+const report = await runIngest({
+  inputs: ['g1.jpg', 'g2.jpg', '案例稿.docx'],
+  outDir: './case-get-outputs/run-001',
+  authorizationStatus: 'authorized',
+  assetBaseUrl: 'https://cdn.example.com/cases',
+});
+console.log(report.casePackage.materialAssets);   // order 已按九宫格排好
+console.log(report.casePackage.riskFlags);        // 合规判断
+if (report.canSubmit) await runSubmit({ package: report.artifacts.casePackage });
+```
+
+## 产物
+
+| 文件 | 内容 |
+| --- | --- |
+| `extracted.json` | 文档文字段落 / 表格 / 超链接 / 内嵌媒体 / 视频转写 |
+| `group-plan.json` | 分组(`groupId` / `layout` / `count` / `orderingRule`)与组内顺序 |
+| `vision.json` | 每张图的 `ocrText` / `description` / `label` / `usageSuggestion` / `piiHints` |
+| `material-assets.json` | `materialAssets[]`(可手工调顺序后回灌) |
+| `case-package.json` | 云函数 `caseSubmit` 入参,**逐字对齐共享契约** |
+| `report.md` | 人类可读报告(顺序表 / 标签 / 合规 / 下一步) |
+
+## 分组与顺序规则(能力 3)
+
+同事连发 6 张、9 张图**不是多个素材,而是同一个素材**(朋友圈素材按九宫格排)。
+
+三种信号判同组,**至少 2 种成立**才并组:
+
+| 信号 | 判据 | 阈值 |
+| --- | --- | --- |
+| 尺寸 | 宽高比相似度 | ≥ 0.88 |
+| 时间 | 文件时间差 | ≤ 15 分钟(窗口内线性衰减) |
+| 内容 | OCR/描述 字符二元组 Jaccard | ≥ 0.25 |
+
+- 只有 1 种信号成立 → **不并组**(避免把不同素材误并);
+- 完全拿不到尺寸/时间(纯转发图)时,会按现有信号降级判定,理由写进 `reasons`;
+- 顺序:`orderingRule = left-to-right,top-to-bottom`,`order` **1 起**连续编号;
+  原始排版位置丢失时用「收图时间升序 + 文件名自然序」近似发送顺序;
+- 要人工干预:`reorder --move "3:1"` 或 `--order "3,1,2"`,也可在输入 manifest 里写 `groups[].items[].order`。
+
+九宫格 = `layout: "3x3"`;6 张 = `3x2`;4 张 = `2x2`。
+
+## 打标(能力 6)
+
+- **硬性静态属性**:`schoolCanonical` / `country` / `major` / `stage` / `subject`
+- **学校别名归一**:`SchoolAlias` 表(`IC → 帝国理工`、`UCL → 伦敦大学学院`)。
+  别名表存在 `references/tag-dictionary.json` 的 `schoolAlias.entries`;命中的写法写入 `schoolAliases[]`。
+  英文别名按词边界匹配(`IC` 不会命中 `MAGIC`)。
+- **销售阶段七段**:`新进线 / 挖需中 / 方案推荐中 / 异议处理中 / 待决策 / 沉默待跟进 / 已成交`
+  按 `docs/prd/dashboard/10-案例推荐策略.md` 的策略表反推,写进 `fitStatus[]`。
+- **亮点 / 场景 / 异议**:`highlightTypes` / `scenarioTags` / `objectionTags`
+- **字典自进化**:采集到字典里没有的新取值时追加进 `references/tag-dictionary.json`
+  (`--no-learn` 可关;本技能**只改文件、不自动 push**)。
+
+## 合规判断(能力 5)
+
+| 判定 | 落到哪 |
+| --- | --- |
+| 授权缺失 / pending / denied | `riskFlags: UNAUTHORIZED` + blocker,**不入库** |
+| 客诉 / 负面事件(投诉、退费纠纷、挂科、维权…) | `riskFlags: COMPLAINT` / `NEGATIVE_EVENT` + blocker,**不得对外展示** |
+| 隐私片段(手机号 / 邮箱 / 微信号 / 姓名…) | `privacyFindings[{field, rule, snippet}]`,snippet **已掩码** |
+| 与留学/课程无关 | `riskFlags: OFF_TOPIC` |
+| 无任何素材 | blocker,无法构成案例 |
+
+`complianceBlockers` 非空 → `submit` **拒绝提交**(即使授权通过)。
+
+## 云函数
+
+| 函数 | 用途 |
+| --- | --- |
+| `caseSubmit` | 提交案例;写 `reviewStatus='pending'`、`readyForUse=false`;`tenantId + idempotencyKey` 幂等 |
+| `casePendingList` | 查待审清单 |
+| `caseApprove` | 负责人审核放行(`approve` → `readyForUse=true`) |
+| `caseSearch` | 检索(只返回 `reviewStatus='approved'` 且 `readyForUse=true`) |
+
+`tenantId` 由云函数从**服务端会话上下文**解析,客户端不传。
+
+## 安装为 FmodeCode / Claude Code 技能
+
+```bash
+npx --yes skill-case-get@latest workspace   # 项目级 → ./.claude/skills/skill-case-get
+npx --yes skill-case-get@latest install     # 用户级 → ~/.claude/skills/skill-case-get
+```
+
+安装后可直接提示:`这 9 张图是同一个案例素材,帮我拆解、打标,然后归档到待审清单。`
+
+## 注意事项
+
+- `.docx` / `.pptx` 解包走 python3 标准库,**不要**为此引入 python-docx / python-pptx / 任何 npm 包。
+- 本机不存储 / 不上传二进制:本地素材的 `url` 由 `--asset-base-url` 派生;没给就会在 `warnings` 里提示(不伪造 URL)。
+- 视频抽帧是均匀取点(默认 6 帧,`--frame-count` 调整);转写走网关按音频时长计费。
+- 枚举统一**小写**(对齐线上:`pending` / `authorized`),展示层再做中文化。

+ 271 - 0
bin/skill-case-get.js

@@ -0,0 +1,271 @@
+#!/usr/bin/env node
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * skill-case-get 安装器 + 运行器透传
+ *
+ *   skill-case-get workspace [--smoke]        安装到 ./.claude/skills/skill-case-get
+ *   skill-case-get install  [--smoke]         安装到 ~/.claude/skills/skill-case-get
+ *   skill-case-get install --target <dir> [--force]
+ *   skill-case-get check | path | smoke | selftest
+ *   skill-case-get ingest [参数...]           透传给 scripts/case-intake.mjs
+ *   skill-case-get plan|submit|reorder|dictionary [...]
+ */
+
+const fs = require('fs');
+const os = require('os');
+const path = require('path');
+const { spawnSync } = require('child_process');
+
+const SKILL_NAME = 'skill-case-get';
+// 本包就是技能本体:skills/skill-case-get/ 同时是 npm 包根与技能目录
+// (与 skill-listen 的「包根 / 内含 skills/<name>」布局不同,产物清单以任务书为准)。
+const SKILL_SOURCE = path.resolve(__dirname, '..');
+const SOURCE_ROOT = SKILL_SOURCE;
+const RUNNER = path.join(SKILL_SOURCE, 'scripts', 'case-intake.mjs');
+const WORKSPACE_ROOT = process.cwd();
+const GLOBAL_TARGET = path.join(os.homedir(), '.claude', 'skills', SKILL_NAME);
+const WORKSPACE_TARGET = path.join(WORKSPACE_ROOT, '.claude', 'skills', SKILL_NAME);
+const WORKSPACE_SKILLS_ROOT = path.join(WORKSPACE_ROOT, '.claude', 'skills');
+const GLOBAL_SKILLS_ROOT = path.join(os.homedir(), '.claude', 'skills');
+
+// 透传给 case-intake.mjs 的子命令
+const RUNNER_COMMANDS = new Set(['ingest', 'plan', 'submit', 'reorder', 'dictionary', 'selftest']);
+
+const REQUIRED_ENTRIES = [
+  'SKILL.md',
+  'README.md',
+  'package.json',
+  'skill-package-manifest.json',
+  'references/tag-dictionary.json',
+  'scripts/case-intake.mjs',
+  'scripts/office_extract.py',
+  'scripts/vision-bridge.mjs',
+  'scripts/listen-bridge.mjs',
+  'scripts/cloud-client.mjs',
+  'scripts/grouping.mjs',
+  'scripts/classify.mjs',
+  'scripts/image-size.mjs',
+  'scripts/lib.mjs',
+];
+
+function expandHome(value) {
+  return String(value || '').replace(/^~(?=$|[\\/])/, os.homedir());
+}
+
+// ---------------------------------------------------------------------------
+// 运行器透传
+// ---------------------------------------------------------------------------
+
+function runRunner(passthrough) {
+  const result = spawnSync(process.execPath, [RUNNER, ...passthrough], { stdio: 'inherit', shell: false });
+  if (result.error) {
+    console.error(`skill-case-get: 无法启动运行器:${result.error.message}`);
+    process.exit(1);
+  }
+  process.exit(result.status == null ? 1 : result.status);
+}
+
+function passthroughArgs(argv) {
+  const rest = argv.slice(1);
+  if (rest[0] === '--') return rest.slice(1);
+  return rest;
+}
+
+// ---------------------------------------------------------------------------
+// 子命令
+// ---------------------------------------------------------------------------
+
+function parseArgs(argv) {
+  const first = argv[0] && !argv[0].startsWith('--') ? argv[0] : 'install';
+  const args = { command: first, target: GLOBAL_TARGET, smoke: false, force: false, help: false };
+  if (first === 'workspace' || first === 'install-workspace') {
+    args.command = 'install';
+    args.target = WORKSPACE_TARGET;
+  }
+  for (let i = first === argv[0] ? 1 : 0; i < argv.length; i++) {
+    const token = argv[i];
+    if (token === '--target' && argv[i + 1]) args.target = argv[++i];
+    else if (token.startsWith('--target=')) args.target = token.slice('--target='.length);
+    else if (token === '--workspace') args.target = WORKSPACE_TARGET;
+    else if (token === '--global') args.target = GLOBAL_TARGET;
+    else if (token === '--smoke') args.smoke = true;
+    else if (token === '--force') args.force = true;
+    else if (token === '--help' || token === '-h') args.help = true;
+  }
+  args.target = path.resolve(expandHome(args.target));
+  return args;
+}
+
+function printHelp() {
+  console.log([
+    'skill-case-get — 案例采集 / 拆解 / 归档技能 + FmodeCode / Claude Code 安装器',
+    '',
+    '把素材(连续多图 / Word / PPT / 图文混合 / 裸图 / 视频)拆解成案例:',
+    '  npx skill-case-get@latest ingest -- --image a.jpg --image b.jpg --authorization authorized',
+    '  npx skill-case-get@latest plan   -- --input ./素材目录',
+    '  npx skill-case-get@latest submit -- --package case-get-outputs/<runId>/case-package.json',
+    '  npx skill-case-get@latest reorder -- --material-assets <file> --move "3:1"',
+    '  npx skill-case-get@latest selftest',
+    '',
+    '安装 Skill:',
+    '  npx skill-case-get@latest workspace [--smoke]   # 安装到 ./.claude/skills/skill-case-get',
+    '  npx skill-case-get@latest install [--smoke]     # 安装到 ~/.claude/skills/skill-case-get',
+    '  npx skill-case-get@latest install --target <dir> [--force]',
+    '  npx skill-case-get@latest check',
+    '  npx skill-case-get@latest path',
+    '',
+    'Options:',
+    '  --workspace      安装到 ./.claude/skills/skill-case-get',
+    '  --global         安装到 ~/.claude/skills/skill-case-get(默认)',
+    '  --target <dir>   安装到自定义目录',
+    '  --force          允许覆盖自定义目录',
+    '  --smoke          安装后运行冒烟检查',
+    '  --help, -h       显示帮助',
+    '',
+    '依赖:python3(.docx/.pptx 解包,仅标准库);可选 ffmpeg(视频抽帧)。',
+    '图片理解复用 skill-vision,音视频转写复用 skill-listen——未安装时会在 warnings 里提示,不伪造结果。',
+  ].join('\n'));
+}
+
+function ensureDir(dirPath) { fs.mkdirSync(dirPath, { recursive: true }); }
+
+function isInside(parentDir, childDir) {
+  const relative = path.relative(path.resolve(parentDir), path.resolve(childDir));
+  return relative === '' || (!!relative && !relative.startsWith('..') && !path.isAbsolute(relative));
+}
+
+function canOverwriteTarget(targetDir, force) {
+  return force
+    || path.resolve(targetDir) === path.resolve(GLOBAL_TARGET)
+    || isInside(WORKSPACE_SKILLS_ROOT, targetDir)
+    || isInside(GLOBAL_SKILLS_ROOT, targetDir);
+}
+
+const SKIP_DIR_NAMES = new Set(['node_modules', 'outputs', '.git', '.claude']);
+
+function copyDirRecursive(source, destination, guard) {
+  const stat = fs.statSync(source);
+  if (stat.isDirectory()) {
+    ensureDir(destination);
+    for (const child of fs.readdirSync(source)) {
+      // outputs 是运行产物;.claude 是安装目标所在;node_modules / .git 永远不拷
+      if (SKIP_DIR_NAMES.has(child)) continue;
+      const childSource = path.join(source, child);
+      // 目标目录恰好长在源目录里(--target 指到源内)时跳过,否则会自己拷自己直到路径超长
+      if (guard) {
+        let childReal = '';
+        try { childReal = fs.realpathSync(childSource); } catch { childReal = childSource; }
+        if (childReal === guard || isInside(childReal, guard)) continue;
+      }
+      copyDirRecursive(childSource, path.join(destination, child), guard);
+    }
+    return;
+  }
+  ensureDir(path.dirname(destination));
+  fs.copyFileSync(source, destination);
+  // 保留可执行位(.py / .mjs 直接跑)
+  try {
+    const mode = fs.statSync(source).mode;
+    if (mode & 0o111) fs.chmodSync(destination, mode);
+  } catch { /* 某些文件系统不支持,忽略 */ }
+}
+
+function installSkill(target, force) {
+  if (!fs.existsSync(SKILL_SOURCE)) {
+    throw new Error(`Skill 源目录不存在:${SKILL_SOURCE}`);
+  }
+  const resolvedTarget = path.resolve(target);
+  if (path.resolve(SKILL_SOURCE) === resolvedTarget) {
+    throw new Error('安装目标与源目录相同,无需安装');
+  }
+  if (fs.existsSync(target)) {
+    if (!canOverwriteTarget(target, force)) {
+      throw new Error(`拒绝覆盖自定义目录(需要 --force):${target}`);
+    }
+    fs.rmSync(target, { recursive: true, force: true });
+  }
+  ensureDir(target);
+  let guard = resolvedTarget;
+  try { guard = fs.realpathSync(resolvedTarget); } catch { guard = resolvedTarget; }
+  copyDirRecursive(SKILL_SOURCE, target, guard);
+}
+
+function checkSkill(target) {
+  const missing = REQUIRED_ENTRIES.filter((entry) => !fs.existsSync(path.join(target, entry)));
+  if (missing.length) {
+    throw new Error(`安装目标缺少必需文件:${missing.join(', ')}`);
+  }
+  // 标签字典必须是合法 JSON,否则技能跑不起来
+  const dictPath = path.join(target, 'references', 'tag-dictionary.json');
+  const dict = JSON.parse(fs.readFileSync(dictPath, 'utf8'));
+  if (!dict.dimensions) throw new Error('标签字典缺少 dimensions');
+  return {
+    status: 'ok',
+    skill: SKILL_NAME,
+    target,
+    required: REQUIRED_ENTRIES,
+    dimensions: Object.keys(dict.dimensions).length,
+  };
+}
+
+function runSmoke() {
+  const result = spawnSync(process.execPath, ['scripts/smoke.mjs'], { cwd: SOURCE_ROOT, stdio: 'inherit', shell: false });
+  if (result.status !== 0) throw new Error('smoke 未通过');
+}
+
+function printNextSteps(target) {
+  const workspaceMode = isInside(WORKSPACE_SKILLS_ROOT, target);
+  console.log('');
+  console.log('安装完成。');
+  console.log(`Skill 位置:${target}`);
+  console.log('');
+  if (workspaceMode) {
+    console.log('项目级技能已就绪。若当前 VSCode FmodeCode / Claude Code 会话是开着的,重启一次。');
+  } else {
+    console.log('用户级技能已就绪,对所有 FmodeCode / Claude Code 工作区生效。');
+  }
+  console.log('');
+  console.log('在 FmodeCode / Claude Code 里可以直接这样说:');
+  console.log('  这 9 张图是同一个案例素材,帮我拆解、打标,然后归档到待审清单。');
+  console.log('');
+  console.log('提醒:无授权(authorizationStatus != authorized)一律不入库;新案例先进待审清单,不进公共素材库。');
+}
+
+function main() {
+  const argv = process.argv.slice(2);
+  const command = argv[0] && !argv[0].startsWith('--') ? argv[0] : 'install';
+
+  // 运行器透传放在前面:这些子命令的参数要原样交给 case-intake.mjs
+  if (RUNNER_COMMANDS.has(command)) {
+    runRunner(passthroughArgs(argv));
+    return;
+  }
+
+  const args = parseArgs(argv);
+  if (args.help || args.command === 'help') { printHelp(); return; }
+  if (args.command === 'path') { console.log(args.target); return; }
+  if (args.command === 'install') {
+    installSkill(args.target, args.force);
+    console.log(JSON.stringify(checkSkill(args.target), null, 2));
+    if (args.smoke) runSmoke();
+    printNextSteps(args.target);
+    return;
+  }
+  if (args.command === 'check') { console.log(JSON.stringify(checkSkill(args.target), null, 2)); return; }
+  if (args.command === 'smoke') { runSmoke(); return; }
+  printHelp();
+  process.exitCode = 1;
+}
+
+try { main(); }
+catch (error) { console.error(`skill-case-get 失败:${error.message}`); process.exit(1); }

+ 64 - 0
package.json

@@ -0,0 +1,64 @@
+{
+  "name": "skill-case-get",
+  "version": "1.0.0",
+  "description": "案例采集 / 拆解 / 归档技能包:把连续多图、Word、PPT、图文混合、裸图等素材拆解为案例字段与素材数组(materialAssets),做合规判断与打标,经云函数 caseSubmit 入库待审。",
+  "type": "commonjs",
+  "main": "./scripts/case-intake.mjs",
+  "exports": {
+    ".": {
+      "import": "./scripts/case-intake.mjs",
+      "default": "./scripts/case-intake.mjs"
+    }
+  },
+  "bin": {
+    "skill-case-get": "bin/skill-case-get.js"
+  },
+  "scripts": {
+    "smoke": "node bin/skill-case-get.js smoke",
+    "test": "node --test scripts/tests/",
+    "vision": "node scripts/vision-bridge.mjs",
+    "extract": "python3 scripts/office_extract.py",
+    "listen": "node scripts/listen-bridge.mjs"
+  },
+  "files": [
+    "bin/",
+    "README.md",
+    "references/",
+    "scripts/",
+    "skill-package-manifest.json",
+    "SKILL.md"
+  ],
+  "keywords": [
+    "fmode",
+    "hermes",
+    "harness-loop",
+    "ai-agent",
+    "agent-skill",
+    "super-skill",
+    "case-library",
+    "case",
+    "intake",
+    "index",
+    "ocr",
+    "material-assets",
+    "未来飞马",
+    "智能体技能",
+    "超级技能",
+    "案例采集",
+    "案例拆解",
+    "素材归档",
+    "九宫格",
+    "打标"
+  ],
+  "author": "Yuyang001 (FmodeAgent)",
+  "license": "MPL-2.0",
+  "repository": {
+    "type": "git",
+    "url": "git+ssh://git@github.com/fmodecn/skill-case-get.git"
+  },
+  "homepage": "https://github.com/fmodecn/skill-case-get#readme",
+  "engines": {
+    "node": ">=18"
+  },
+  "dependencies": {}
+}

+ 152 - 0
references/tag-dictionary.json

@@ -0,0 +1,152 @@
+{
+  "schemaVersion": "1.0.0",
+  "updatedAt": "2026-10-10T00:00:00.000Z",
+  "note": "案例库标签字典。技能在采集过程中遇到字典里没有的取值时,会追加到对应维度(自进化),供下次复用。新增条目统一带 \"source\":\"learned\" 与 \"learnedAt\"。",
+  "enumCase": "lowercase",
+  "dimensions": {
+    "productLine": {
+      "label": "产品线",
+      "multiple": false,
+      "hardRule": "只能是 大学课程 / 国际课程 / 申诉 三者之一",
+      "values": ["大学课程", "国际课程", "申诉"]
+    },
+    "country": {
+      "label": "国家/地区",
+      "multiple": false,
+      "values": ["英国", "美国", "澳大利亚", "加拿大", "新西兰", "爱尔兰", "中国香港", "新加坡", "日本", "韩国", "荷兰", "德国"]
+    },
+    "schoolCanonical": {
+      "label": "学校(标准名)",
+      "multiple": true,
+      "values": [
+        "帝国理工", "伦敦大学学院", "伦敦政治经济学院", "伦敦国王学院", "曼彻斯特大学",
+        "爱丁堡大学", "华威大学", "布里斯托大学", "格拉斯哥大学", "南安普顿大学",
+        "利兹大学", "伯明翰大学", "诺丁汉大学", "谢菲尔德大学", "杜伦大学",
+        "多伦多大学", "英属哥伦比亚大学", "麦吉尔大学",
+        "悉尼大学", "新南威尔士大学", "墨尔本大学", "昆士兰大学", "莫纳什大学",
+        "纽约大学", "加州大学洛杉矶分校", "南加州大学", "东北大学", "伊利诺伊大学厄巴纳-香槟分校",
+        "香港大学", "香港中文大学", "香港科技大学",
+        "新加坡国立大学", "南洋理工大学"
+      ]
+    },
+    "schoolAlias": {
+      "label": "学校别名归一表",
+      "multiple": true,
+      "note": "对应 SchoolAlias 表(aliasText → canonicalName)。命中别名时写入 CaseAsset.schoolAliases,并把 schoolCanonical 归一为标准名。",
+      "entries": [
+        { "aliasText": "IC", "canonicalName": "帝国理工", "country": "英国" },
+        { "aliasText": "Imperial", "canonicalName": "帝国理工", "country": "英国" },
+        { "aliasText": "imperial college", "canonicalName": "帝国理工", "country": "英国" },
+        { "aliasText": "UCL", "canonicalName": "伦敦大学学院", "country": "英国" },
+        { "aliasText": "LSE", "canonicalName": "伦敦政治经济学院", "country": "英国" },
+        { "aliasText": "KCL", "canonicalName": "伦敦国王学院", "country": "英国" },
+        { "aliasText": "曼大", "canonicalName": "曼彻斯特大学", "country": "英国" },
+        { "aliasText": "Manchester", "canonicalName": "曼彻斯特大学", "country": "英国" },
+        { "aliasText": "爱大", "canonicalName": "爱丁堡大学", "country": "英国" },
+        { "aliasText": "Edinburgh", "canonicalName": "爱丁堡大学", "country": "英国" },
+        { "aliasText": "华威", "canonicalName": "华威大学", "country": "英国" },
+        { "aliasText": "Warwick", "canonicalName": "华威大学", "country": "英国" },
+        { "aliasText": "UofT", "canonicalName": "多伦多大学", "country": "加拿大" },
+        { "aliasText": "UBC", "canonicalName": "英属哥伦比亚大学", "country": "加拿大" },
+        { "aliasText": "USYD", "canonicalName": "悉尼大学", "country": "澳大利亚" },
+        { "aliasText": "UNSW", "canonicalName": "新南威尔士大学", "country": "澳大利亚" },
+        { "aliasText": "墨大", "canonicalName": "墨尔本大学", "country": "澳大利亚" },
+        { "aliasText": "NYU", "canonicalName": "纽约大学", "country": "美国" },
+        { "aliasText": "UCLA", "canonicalName": "加州大学洛杉矶分校", "country": "美国" },
+        { "aliasText": "USC", "canonicalName": "南加州大学", "country": "美国" },
+        { "aliasText": "HKU", "canonicalName": "香港大学", "country": "中国香港" },
+        { "aliasText": "CUHK", "canonicalName": "香港中文大学", "country": "中国香港" },
+        { "aliasText": "NUS", "canonicalName": "新加坡国立大学", "country": "新加坡" },
+        { "aliasText": "NTU", "canonicalName": "南洋理工大学", "country": "新加坡" }
+      ]
+    },
+    "major": {
+      "label": "专业",
+      "multiple": true,
+      "values": [
+        "计算机科学", "数据科学", "人工智能", "软件工程", "电子电气工程", "机械工程",
+        "数学", "统计学", "金融", "会计与金融", "经济学", "管理学", "市场营销",
+        "生物医学", "心理学", "传媒", "教育学", "法律", "建筑学", "化学工程"
+      ]
+    },
+    "stage": {
+      "label": "年级/阶段",
+      "multiple": false,
+      "values": ["预科", "本科一年级", "本科二年级", "本科三年级", "本科四年级", "硕士", "博士"]
+    },
+    "subject": {
+      "label": "科目",
+      "multiple": true,
+      "values": [
+        "高等数学", "线性代数", "概率论与数理统计", "微积分", "离散数学",
+        "数据结构与算法", "操作系统", "计算机网络", "数据库", "机器学习",
+        "宏观经济学", "微观经济学", "财务会计", "金融学", "统计学",
+        "大学物理", "电磁学", "力学", "电路分析", "信号与系统",
+        "学术英语", "学术写作", "生物化学", "有机化学"
+      ]
+    },
+    "fitStatus": {
+      "label": "适用跟进状态(销售阶段七段)",
+      "multiple": true,
+      "reference": "docs/prd/dashboard/10-案例推荐策略.md",
+      "values": ["新进线", "挖需中", "方案推荐中", "异议处理中", "待决策", "沉默待跟进", "已成交"]
+    },
+    "highlightTypes": {
+      "label": "亮点类型",
+      "multiple": true,
+      "values": [
+        "提分", "录取", "申诉成功", "补考通过", "稳分", "高分冲刺",
+        "挂科挽救", "退费保障", "家长认可", "出分反馈", "首课体验",
+        "导师匹配", "群内答疑", "课时反馈", "大课时囤课"
+      ]
+    },
+    "scenarioTags": {
+      "label": "场景标签",
+      "multiple": true,
+      "values": [
+        "考前冲刺", "开学季", "期中考试", "期末考试", "补考季", "论文季",
+        "选课指导", "转专业", "申诉期", "退费咨询", "家长陪同", "多地时差",
+        "时间紧张", "基础薄弱", "目标高分", "临近毕业"
+      ]
+    },
+    "objectionTags": {
+      "label": "异议类型",
+      "multiple": true,
+      "values": ["贵", "不放心老师", "犹豫", "怕没效果", "怕时间不够", "要和家里商量", "对比其他机构", "担心退费难"]
+    },
+    "riskFlags": {
+      "label": "风险标记",
+      "multiple": true,
+      "note": "命中任一即不得进入公共素材库;无授权一律不入库。",
+      "values": ["COMPLAINT", "NEGATIVE_EVENT", "PII_RISK", "UNAUTHORIZED", "OFF_TOPIC", "DUPLICATE"]
+    },
+    "sourceType": {
+      "label": "采集来源类型",
+      "multiple": false,
+      "note": "对齐 SHARED-CONTRACT.md,统一小写",
+      "values": ["image_group", "document", "presentation", "mixed", "single_image", "middleware"]
+    },
+    "reviewStatus": {
+      "label": "审核状态",
+      "multiple": false,
+      "values": ["pending", "approved", "rejected"]
+    },
+    "authorizationStatus": {
+      "label": "授权状态",
+      "multiple": false,
+      "values": ["authorized", "pending", "denied"]
+    }
+  },
+  "promotionStatusMapping": {
+    "authorizationStatus": {
+      "AUTHORIZED": "authorized",
+      "PENDING": "pending",
+      "DENIED": "denied"
+    },
+    "fitStatusNote": "对齐线上小写枚举;展示层再做中文化。"
+  },
+  "maskingMusts": {
+    "note": "打码是红线:识别并记录,未通过抽检不得进入检索结果。本期只记录 privacyFindings,不打码不落 PII。",
+    "fields": ["姓名", "头像", "手机号", "邮箱", "微信号", "QQ", "身份证号", "银行卡号", "具体住址"]
+  }
+}

+ 1127 - 0
scripts/case-intake.mjs

@@ -0,0 +1,1127 @@
+#!/usr/bin/env node
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * skill-case-get —— 案例采集 / 拆解 / 归档 主入口
+ *
+ * 管线:
+ *   ① 输入识别:连续多图 / .docx / .pptx / 图文混合 / 裸图 / 视频
+ *   ② 拆解提取:图片→skill-vision(OCR+语义);docx/pptx→python3 标准库解包;视频→抽帧+skill-listen 转写
+ *   ③ 分组与顺序:尺寸/时间/内容相似度判同组;九宫格按左上→右下给显式 order
+ *   ④ 案例信息 vs 素材 分离
+ *   ⑤ 合规判断:能否入库?客诉/负面?→ riskFlags / privacyFindings
+ *   ⑥ 打标:静态属性 + 跟进状态七段 + 亮点/场景/异议 + 学校别名归一
+ *   ⑦ 入库:caseSubmit 写待审清单(reviewStatus=pending, readyForUse=false),带幂等键
+ *   ⑧ 字典自进化:新标签追加进 references/tag-dictionary.json
+ *
+ * 子命令:
+ *   plan     只识别 + 提取 + 分组(不调任何模型,纯本机)
+ *   ingest   plan + 图片理解 + 分类打标 → 产出 case-package.json(默认不提交)
+ *   submit   读取 case-package.json(可先 reorder)→ 调云函数 caseSubmit
+ *   reorder  调整 materialAssets 顺序(支持数组 + 顺序调整)
+ *   dictionary  查看 / 追加标签字典
+ *   selftest 本机自检(office 解包 + 分组 + 打标 + 合规)
+ *
+ * 安全:绝不打印 token/PII;生成的产物默认落在 <cwd>/case-get-outputs/。
+ */
+
+import fs from 'node:fs';
+import path from 'node:path';
+import zlib from 'node:zlib';
+import { spawnSync } from 'node:child_process';
+import { parseArgs } from 'node:util';
+import {
+  TENANT_DEFAULT, SKILL_ROOT, buildIdempotencyKey, deriveAssetUrl, ensureDir, firstText, imageUrlsFrom,
+  learnAliases, learnTags, loadDictionary, newRunId, normalizeMaterialAsset, outputsDir,
+  readJson, reorderAssets, resolvePython, saveDictionary, sortAssets, truncate, uniq, writeJson,
+} from './lib.mjs';
+import { readImageInfo, IMAGE_EXTENSIONS, VIDEO_EXTENSIONS, DOC_EXTENSIONS, DECK_EXTENSIONS, TEXT_EXTENSIONS } from './image-size.mjs';
+import { applyGroupOverrides, groupMaterials, DEFAULT_ORDERING_RULE } from './grouping.mjs';
+import { classifyCase } from './classify.mjs';
+import { processVideo } from './listen-bridge.mjs';
+import { analyzeImages } from './vision-bridge.mjs';
+import { CloudError, submitCase } from './cloud-client.mjs';
+
+// ---------------------------------------------------------------------------
+// 输入识别
+// ---------------------------------------------------------------------------
+
+export function detectSourceType({ images = 0, documents = 0, decks = 0, videos = 0, texts = 0 }) {
+  const kinds = [images > 0, documents > 0, decks > 0, videos > 0, texts > 0].filter(Boolean).length;
+  if (kinds > 1) return 'mixed';
+  if (images > 1) return 'image_group';
+  if (images === 1 && videos === 0) return 'single_image';
+  if (videos > 0) return 'image_group'; // 视频素材组,按 group 处理
+  if (decks > 0) return 'presentation';
+  if (documents > 0) return 'document';
+  return 'mixed';
+}
+
+export function classifyInput(filePath) {
+  const ext = path.extname(filePath).toLowerCase();
+  if (IMAGE_EXTENSIONS.has(ext)) return 'image';
+  if (VIDEO_EXTENSIONS.has(ext)) return 'video';
+  if (DOC_EXTENSIONS.has(ext)) return 'docx';
+  if (DECK_EXTENSIONS.has(ext)) return 'pptx';
+  if (TEXT_EXTENSIONS.has(ext)) return 'text';
+  return 'unknown';
+}
+
+/** 展开目录(只展开一层里的文件,不递归太深,避免误扫)。 */
+export function expandInputs(inputs, maxDepth = 2) {
+  const out = [];
+  const walk = (target, depth) => {
+    let stat;
+    try {
+      stat = fs.statSync(target);
+    } catch {
+      out.push({ path: target, exists: false });
+      return;
+    }
+    if (stat.isDirectory()) {
+      if (depth > maxDepth) return;
+      for (const entry of fs.readdirSync(target).sort()) {
+        if (entry.startsWith('.')) continue;
+        walk(path.join(target, entry), depth + 1);
+      }
+      return;
+    }
+    out.push({ path: target, exists: true, size: stat.size, mtimeMs: stat.mtimeMs });
+  };
+  for (const input of inputs) walk(input, 0);
+  return out;
+}
+
+// ---------------------------------------------------------------------------
+// ② 拆解提取
+// ---------------------------------------------------------------------------
+
+function runOfficeExtract(files, mediaOut) {
+  if (!files.length) return { available: true, results: [], error: null };
+  const script = path.join(SKILL_ROOT, 'scripts', 'office_extract.py');
+  if (!fs.existsSync(script)) return { available: false, results: [], error: `缺少 ${script}` };
+  const python = resolvePython();
+  const args = [script, '--media-out', mediaOut];
+  for (const file of files) args.push('--input', file);
+  const run = spawnSync(python, args, { encoding: 'utf8', maxBuffer: 64 * 1024 * 1024 });
+  if (run.error) {
+    return { available: false, results: [], error: `${python} 启动失败:${run.error.message}(.docx/.pptx 解包需要 python3 标准库)` };
+  }
+  let parsed = null;
+  try {
+    parsed = JSON.parse(run.stdout || '{}');
+  } catch {
+    return { available: true, results: [], error: `office_extract 输出不是 JSON:${truncate(run.stderr || run.stdout, 300)}` };
+  }
+  return {
+    available: true,
+    results: parsed.results || [],
+    error: parsed.ok ? null : `部分文档解包失败:${(parsed.results || []).filter((r) => !r.ok).map((r) => r.error).join(', ')}`,
+  };
+}
+
+function sleep(ms) {
+  return new Promise((resolve) => { setTimeout(resolve, ms); });
+}
+
+// ---------------------------------------------------------------------------
+// plan:识别 + 提取 + 分组
+// ---------------------------------------------------------------------------
+
+export async function runPlan(options = {}) {
+  const outDir = ensureDir(options.outDir);
+  const mediaDir = ensureDir(path.join(outDir, 'media'));
+
+  const entries = expandInputs(options.inputs || []);
+  const missing = entries.filter((e) => !e.exists).map((e) => e.path);
+  const buckets = { image: [], video: [], docx: [], pptx: [], text: [], unknown: [] };
+
+  for (const entry of entries) {
+    if (!entry.exists) continue;
+    buckets[classifyInput(entry.path)].push(entry.path);
+  }
+
+  const warnings = [];
+  for (const file of missing) warnings.push(`输入不存在:${file}`);
+  if (buckets.unknown.length) warnings.push(`忽略无法识别的文件:${buckets.unknown.join(', ')}`);
+
+  // ---- docx / pptx 解包
+  const docxResults = runOfficeExtract(buckets.docx, path.join(mediaDir, 'docx'));
+  const pptxResults = runOfficeExtract(buckets.pptx, path.join(mediaDir, 'pptx'));
+  if (docxResults.error) warnings.push(docxResults.error);
+  if (pptxResults.error) warnings.push(pptxResults.error);
+
+  const textBundle = { paragraphs: [], notes: [], titles: [], transcript: '' };
+  const mediaFromDocs = [];
+
+  for (const result of [...docxResults.results, ...pptxResults.results]) {
+    if (!result.ok) {
+      warnings.push(`${result.fileName || result.input} 解包失败:${result.error}`);
+      continue;
+    }
+    for (const p of result.paragraphs || []) if (p.text) textBundle.paragraphs.push(p.text);
+    for (const h of result.headings || []) if (h) textBundle.titles.push(h);
+    for (const note of result.notes || []) {
+      if (typeof note === 'string') textBundle.notes.push(note);
+      else if (note && note.text) textBundle.notes.push(note.text);
+    }
+    for (const slide of result.slides || []) if (slide.title) textBundle.titles.push(slide.title);
+    for (const media of result.media || []) {
+      if (media.kind === 'image' && media.localPath) mediaFromDocs.push({ ...media, sourceDoc: result.fileName || result.input });
+    }
+  }
+
+  // ---- 纯文本输入(.md/.txt):并入说明语料
+  for (const file of buckets.text) {
+    try {
+      const text = fs.readFileSync(file, 'utf8');
+      textBundle.paragraphs.push(...text.split(/\r?\n/).map((l) => l.trim()).filter(Boolean));
+    } catch (error) {
+      warnings.push(`文本读取失败 ${file}:${error.message}`);
+    }
+  }
+
+  // ---- 图片:读尺寸/时间
+  const imageItems = [];
+  for (const file of buckets.image) {
+    const info = readImageInfo(file);
+    if (!info.ok) warnings.push(`图片信息读取失败 ${path.basename(file)}:${info.error}`);
+    imageItems.push({
+      kind: 'image',
+      localPath: path.resolve(file),
+      width: info.width || null,
+      height: info.height || null,
+      mtimeMs: info.mtimeMs || fs.statSync(file).mtimeMs,
+      ocrText: '',
+      description: '',
+      label: path.basename(file),
+      origin: 'direct',
+    });
+  }
+  // 文档内嵌图片也参与分组(图文混合)。sourceKey 钉住来源文档,
+  // 避免两份不同文档的配图因为同比例、同落盘时间被误并成一组。
+  for (const media of mediaFromDocs) {
+    const info = readImageInfo(media.localPath);
+    imageItems.push({
+      kind: 'image',
+      localPath: media.localPath,
+      width: info.width || null,
+      height: info.height || null,
+      mtimeMs: info.mtimeMs || null,
+      ocrText: '',
+      description: media.nearbyText || '',
+      label: `${media.sourceDoc || '文档'} 内嵌图`,
+      origin: 'document-media',
+      sourceKey: media.sourceDoc || '',
+      slide: media.slide,
+    });
+  }
+
+  // ---- 视频:抽帧 + 转写
+  const videoReports = [];
+  for (const file of buckets.video) {
+    if (options.skipVideo) {
+      warnings.push(`跳过视频(--skip-video):${path.basename(file)}`);
+      continue;
+    }
+    const report = await processVideo(file, {
+      workDir: path.join(mediaDir, 'video', path.basename(file, path.extname(file))),
+      frameCount: options.frameCount || 6,
+      diarize: Boolean(options.diarize),
+    });
+    videoReports.push(report);
+    for (const warning of report.warnings) warnings.push(`${path.basename(file)}:${warning}`);
+    if (report.transcription && report.transcription.text) {
+      textBundle.transcript = [textBundle.transcript, report.transcription.text].filter(Boolean).join('\n');
+    }
+    for (const frame of report.frames || []) {
+      imageItems.push({
+        kind: 'image',
+        localPath: frame.localPath,
+        width: null,
+        height: null,
+        mtimeMs: null,
+        ocrText: '',
+        description: `视频抽帧 @${frame.atMs != null ? `${Math.round(frame.atMs / 1000)}s` : '开头'}`,
+        label: `${path.basename(file)} 抽帧`,
+        origin: 'video-frame',
+        atMs: frame.atMs,
+        videoPath: path.resolve(file),
+      });
+    }
+  }
+
+  // ---- 分组(尺寸 / 时间 / 内容)
+  let groups = groupMaterials(imageItems, options.grouping || {});
+  groups = applyGroupOverrides(groups, options.groupOverrides);
+
+  const sourceType = options.sourceType || detectSourceType({
+    images: buckets.image.length,
+    documents: buckets.docx.length + buckets.text.length,
+    decks: buckets.pptx.length,
+    videos: buckets.video.length,
+  });
+
+  const plan = {
+    runId: options.runId || newRunId(),
+    generatedAt: new Date().toISOString(),
+    tenantId: options.tenantId || TENANT_DEFAULT,
+    sourceType,
+    sourceRef: options.sourceRef || '',
+    authorizationStatus: String(options.authorizationStatus || '').toLowerCase(),
+    inputs: {
+      images: buckets.image.map((p) => path.resolve(p)),
+      documents: buckets.docx.map((p) => path.resolve(p)),
+      decks: buckets.pptx.map((p) => path.resolve(p)),
+      videos: buckets.video.map((p) => path.resolve(p)),
+      texts: buckets.text.map((p) => path.resolve(p)),
+    },
+    counts: {
+      imageCount: imageItems.length,
+      documentCount: buckets.docx.length,
+      deckCount: buckets.pptx.length,
+      videoCount: buckets.video.length,
+      groupCount: groups.length,
+    },
+    textBundle,
+    groups,
+    videoReports,
+    warnings,
+    // 生成物绝对路径(后续 vision / submit 直接引用)
+    artifacts: {
+      extracted: path.join(outDir, 'extracted.json'),
+      groupPlan: path.join(outDir, 'group-plan.json'),
+    },
+  };
+  if (!groups.length) plan.warnings.push('没有识别到任何可用素材(没有任何图片/视频/文档内嵌图)');
+
+  writeJson(plan.artifacts.extracted, { textBundle, mediaFromDocs, videoReports });
+  writeJson(plan.artifacts.groupPlan, { sourceType, groups });
+  return plan;
+}
+
+// ---------------------------------------------------------------------------
+// ingest:plan + 图片理解 + 分类打标 → 案例包
+// ---------------------------------------------------------------------------
+
+export async function runIngest(options = {}) {
+  const plan = await runPlan(options);
+  const outDir = options.outDir;
+  const warnings = [...plan.warnings];
+
+  // 图片理解(OCR + 语义)——没有图片时跳过
+  const allImages = plan.groups.flatMap((g) => g.items.filter((i) => i.kind === 'image'));
+  let vision = { available: true, results: [], error: null, needsHostRead: false };
+
+  if (options.skipVision) {
+    warnings.push('跳过图片理解(--skip-vision):素材只有文件名,ocrText/description 为空');
+  } else if (allImages.length) {
+    const batchHint = [
+      plan.textBundle.titles.slice(0, 3).join(' / '),
+      plan.textBundle.paragraphs.slice(0, 3).join(' / '),
+    ].filter(Boolean).join(' | ');
+    vision = await analyzeImages(
+      allImages.map((item) => ({ localPath: item.localPath, url: item.url, label: item.label })),
+      { model: options.visionModel, batchHint },
+    );
+    if (!vision.available) {
+      warnings.push(vision.error);
+    } else if (vision.needsHostRead) {
+      warnings.push('宿主多模态模式:vision 返回了读图指令(instruction),请宿主 Agent 用 Read 工具读图后把 JSON 结果写入 vision-host.json,再执行 submit。');
+    } else if (vision.error) {
+      warnings.push(vision.error);
+    }
+  }
+
+  // 把 vision 结果并回素材
+  const visionByFile = new Map();
+  for (const result of vision.results || []) {
+    if (result.file) visionByFile.set(path.resolve(result.file), result);
+  }
+
+  const enrichedGroups = plan.groups.map((group) => ({
+    ...group,
+    items: group.items.map((item) => {
+      const enriched = { ...item };
+      if (item.kind === 'image' && item.localPath) {
+        const found = visionByFile.get(path.resolve(item.localPath));
+        if (found) {
+          enriched.ocrText = found.ocrText || enriched.ocrText || '';
+          enriched.description = found.description || enriched.description || '';
+          enriched.usageSuggestion = found.usageSuggestion || '';
+          enriched.label = found.label || enriched.label || '';
+          // 分析未产出结论(宿主读图待补 / 分析失败)时保持默认 role='material',
+          // 不能凭空把素材降级成说明配图。
+          enriched.role = found.role || 'material';
+          enriched.piiHints = found.piiHints || [];
+          enriched.visionPending = Boolean(found.needsHostRead);
+        } else {
+          enriched.role = enriched.role || 'material';
+        }
+      }
+      return enriched;
+    }),
+  }));
+
+  // 案例信息 vs 素材:分离后分类打标
+  const dict = loadDictionary(options.dictPath);
+  const visionResults = (vision.results || []).map((r) => ({ ...r }));
+
+  const classification = classifyCase({
+    dict,
+    groups: enrichedGroups,
+    visionResults,
+    textBundle: plan.textBundle,
+    authorizationStatus: plan.authorizationStatus,
+    hints: options.hints || {},
+  });
+
+  // ⑧ 字典自进化
+  const learn = learnTags(dict, classification.learned.tags);
+  const aliasLearn = learnAliases(dict, classification.learned.aliases);
+  const dictionaryChanges = { tags: learn.added, aliases: aliasLearn.added, skipped: learn.skipped };
+  const dictionaryFile = options.dictPath ? path.resolve(options.dictPath) : undefined;
+  if ((learn.added.length || aliasLearn.added.length) && !options.noLearn) {
+    saveDictionary(dict);
+  } else if (options.noLearn && (learn.added.length || aliasLearn.added.length)) {
+    warnings.push('字典自进化被 --no-learn 关闭:本次新增标签未落盘');
+    dictionaryChanges.tags = [];
+    dictionaryChanges.aliases = [];
+  }
+  if (dictionaryFile) writeJson(dictionaryFile, readJson(dictionaryFile, {}));
+
+  // ---- 组装 materialAssets[](逐字对齐 SHARED-CONTRACT.md)
+  const materialAssets = [];
+  const groupInfoList = [];
+  const assetProblems = [];
+  let missingUrlCount = 0;
+
+  for (const group of enrichedGroups) {
+    groupInfoList.push({
+      groupId: group.groupId,
+      layout: group.layout,
+      count: group.count,
+      orderingRule: group.orderingRule,
+    });
+    for (const item of sortAssets(group.items)) {
+      // 本地文件没有公网 URL 时,按 --asset-base-url 派生;没给基础地址则留空并警告
+      const derivedUrl = item.url || deriveAssetUrl(item.localPath, options.assetBaseUrl);
+      if (!derivedUrl && item.localPath) missingUrlCount += 1;
+      const { asset, problems } = normalizeMaterialAsset({
+        order: item.order,
+        groupId: group.groupId,
+        kind: item.kind,
+        url: derivedUrl,
+        localPath: item.localPath || '',
+        label: item.label || '',
+        ocrText: item.ocrText || '',
+        description: item.description || '',
+        usageSuggestion: item.usageSuggestion || '',
+        role: item.role || 'material',
+      }, materialAssets.length);
+      assetProblems.push(...problems);
+      materialAssets.push(asset);
+    }
+  }
+  for (const problem of assetProblems) warnings.push(problem);
+  if (missingUrlCount) {
+    warnings.push(`${missingUrlCount} 个素材只有本地路径、没有 URL:用 --asset-base-url <CDN前缀> 派生,或由归档方补齐 url 后再 submit(检索侧只看 url)。`);
+  }
+
+  // 全局重编号:materialAssets 已按「分组顺序 → 组内 order」构建,
+  // 这里只顺次改写 order。**不能**再按 order 全局排序——各组的 order 都从 1 起,
+  // 直接排会把四组的第一张图并到最前面。
+  const orderedAssets = materialAssets.map((asset, index) => ({ ...asset, order: index + 1 }));
+  const imageUrls = imageUrlsFrom(orderedAssets);
+
+  // ---- 组装案例包(幂等键:tenantId + 内容指纹)
+  const title = firstText(options.title, classification.caseFields.title);
+  const idempotencyKey = buildIdempotencyKey({
+    tenantId: plan.tenantId,
+    sourceType: plan.sourceType,
+    sourceRef: plan.sourceRef,
+    title,
+    materialAssets: orderedAssets,
+    explicit: options.idempotencyKey,
+  });
+
+  const primaryGroup = groupInfoList[0] || { groupId: 'g1', layout: '1x1', count: 0, orderingRule: DEFAULT_ORDERING_RULE };
+
+  const casePackage = {
+    schemaVersion: '1.0.0',
+    generatedAt: new Date().toISOString(),
+    runId: plan.runId,
+    contract: 'SHARED-CONTRACT.md',
+    // ---- 云函数 caseSubmit 入参(逐字对齐)
+    sourceType: plan.sourceType,
+    sourceRef: plan.sourceRef,
+    idempotencyKey,
+    title,
+    summary: firstText(options.summary, classification.caseFields.summary),
+    materialAssets: orderedAssets,
+    groupInfo: primaryGroup,
+    groupInfoList,
+    schoolCanonical: (classification.tags.schoolCanonical || [])[0] || '',
+    tags: uniq([
+      ...(classification.tags.country || []),
+      ...(classification.tags.schoolCanonical || []),
+      ...(classification.tags.major || []),
+      ...(classification.tags.stage || []),
+      ...(classification.tags.subject || []),
+      ...(classification.tags.highlightTypes || []),
+      ...(classification.tags.scenarioTags || []),
+      ...(classification.tags.objectionTags || []),
+      ...(classification.fit.fitStatus || []),
+      ...(classification.tags.productLine || []),
+    ]),
+    highlightTypes: classification.tags.highlightTypes || [],
+    scenarioTags: classification.tags.scenarioTags || [],
+    objectionTags: classification.tags.objectionTags || [],
+    riskFlags: classification.compliance.riskFlags || [],
+    authorizationStatus: plan.authorizationStatus || '',
+    // ---- 案例字段(案例说明,与素材分开)
+    targetCustomer: firstText(options.targetCustomer, classification.caseFields.targetCustomer),
+    usageSuggestion: firstText(options.usageSuggestion, classification.caseFields.usageSuggestion),
+    resultEvidence: classification.caseFields.resultEvidence,
+    outcome: classification.caseFields.outcome,
+    country: (classification.tags.country || [])[0] || '',
+    major: (classification.tags.major || [])[0] || '',
+    stage: (classification.tags.stage || [])[0] || '',
+    subject: (classification.tags.subject || [])[0] || '',
+    fitStatus: classification.fit.fitStatus || [],
+    schoolAliases: classification.tags.schoolAliases || [],
+    imageUrls,
+    hasImage: imageUrls.length > 0,
+    imageText: truncate(orderedAssets.map((a) => a.ocrText).filter(Boolean).join(' '), 400),
+    privacyFindings: classification.compliance.privacyFindings || [],
+    complianceBlockers: classification.compliance.complianceBlockers || [],
+    reviewHint: classification.compliance.reviewHint,
+    summaryLine: casePackageSummaryLine(classification, orderedAssets.length),
+    submittedBy: options.submittedBy || '',
+    // ---- 落库后由服务端决定的字段(此处只声明预期值,供 dry-run 展示)
+    expected: { reviewStatus: 'pending', readyForUse: false },
+  };
+
+  // ---- 产出物
+  const artifacts = {
+    extracted: plan.artifacts.extracted,
+    groupPlan: plan.artifacts.groupPlan,
+    vision: path.join(outDir, 'vision.json'),
+    materialAssets: path.join(outDir, 'material-assets.json'),
+    casePackage: path.join(outDir, 'case-package.json'),
+    report: path.join(outDir, 'report.md'),
+  };
+
+  writeJson(artifacts.vision, vision);
+  writeJson(artifacts.materialAssets, { groupInfoList, materialAssets: orderedAssets });
+  writeJson(artifacts.casePackage, casePackage);
+  fs.writeFileSync(artifacts.report, renderReport({ plan, casePackage, vision, dictionaryChanges, warnings }), 'utf8');
+
+  const blockers = casePackage.complianceBlockers || [];
+  const canSubmit = plan.authorizationStatus === 'authorized' && blockers.length === 0 && orderedAssets.length > 0;
+
+  return { plan, casePackage, vision, dictionaryChanges, warnings, artifacts, canSubmit };
+}
+
+function casePackageSummaryLine(classification, assetCount) {
+  const fits = (classification.fit.fitStatus || []).join('/') || '未命中';
+  const risks = (classification.compliance.riskFlags || []).join('/') || '无';
+  return `${assetCount} 个素材;适用跟进状态:${fits};风险:${risks}`;
+}
+
+// ---------------------------------------------------------------------------
+// 报告
+// ---------------------------------------------------------------------------
+
+export function renderReport({ plan, casePackage, vision, dictionaryChanges, warnings }) {
+  const lines = [];
+  lines.push(`# 案例采集报告 · ${casePackage.runId}`);
+  lines.push('');
+  lines.push(`- 生成时间:${casePackage.generatedAt}`);
+  lines.push(`- 来源类型:\`${casePackage.sourceType}\``);
+  lines.push(`- 授权状态:\`${casePackage.authorizationStatus || '缺失'}\``);
+  lines.push(`- 幂等键:\`${casePackage.idempotencyKey}\``);
+  lines.push(`- 素材数:${casePackage.materialAssets.length}(分组 ${casePackage.groupInfoList.length} 组)`);
+  lines.push(`- 图片理解:${vision.available ? `已启用(${vision.source}${vision.needsHostRead ? ',待宿主读图' : ''})` : '不可用'}`);
+  lines.push('');
+  lines.push('## 素材顺序(materialAssets)');
+  lines.push('');
+  lines.push('| order | groupId | kind | role | 标签 | 说明 |');
+  lines.push('| --- | --- | --- | --- | --- | --- |');
+  for (const asset of casePackage.materialAssets) {
+    lines.push(`| ${asset.order} | ${asset.groupId} | ${asset.kind} | ${asset.role} | ${truncate(asset.label, 20)} | ${truncate(asset.description || asset.ocrText, 60)} |`);
+  }
+  lines.push('');
+  lines.push('## 分组(九宫格阅读顺序:左上 → 右下)');
+  lines.push('');
+  for (const group of casePackage.groupInfoList) {
+    lines.push(`- \`${group.groupId}\` ${group.layout} · ${group.count} 张 · \`${group.orderingRule}\``);
+  }
+  lines.push('');
+  lines.push('## 标签');
+  lines.push('');
+  lines.push(`- 学校(归一):${casePackage.schoolCanonical || '未识别'}${casePackage.schoolAliases.length ? `(命中别名:${casePackage.schoolAliases.join('、')})` : ''}`);
+  lines.push(`- 国家:${casePackage.country || '未识别'} · 专业:${casePackage.major || '未识别'} · 阶段:${casePackage.stage || '未识别'} · 科目:${casePackage.subject || '未识别'}`);
+  lines.push(`- 亮点:${casePackage.highlightTypes.join('、') || '无'}`);
+  lines.push(`- 场景:${casePackage.scenarioTags.join('、') || '无'}`);
+  lines.push(`- 异议:${casePackage.objectionTags.join('、') || '无'}`);
+  lines.push(`- 适用跟进状态:${(casePackage.fitStatus || []).join('、') || '未命中'}`);
+  lines.push('');
+  lines.push('## 合规判断');
+  lines.push('');
+  lines.push(`- riskFlags:${casePackage.riskFlags.join('、') || '无'}`);
+  lines.push(`- privacyFindings:${casePackage.privacyFindings.length} 条(仅记录,本期不打码)`);
+  for (const finding of casePackage.privacyFindings) {
+    lines.push(`  - ${finding.field} / ${finding.rule}:\`${finding.snippet}\``);
+  }
+  lines.push(`- 结论:${casePackage.complianceBlockers.length ? '不入库' : '可入库待审'}`);
+  for (const blocker of casePackage.complianceBlockers) lines.push(`  - ⛔ ${blocker}`);
+  lines.push(`- 审核提示:${casePackage.reviewHint}`);
+  lines.push('');
+  lines.push('## 字典自进化');
+  lines.push('');
+  lines.push(`- 新增标签:${dictionaryChanges.tags.length ? dictionaryChanges.tags.map((t) => `${t.dimension}=${t.value}`).join('、') : '无'}`);
+  lines.push(`- 新增别名:${dictionaryChanges.aliases.length ? dictionaryChanges.aliases.map((a) => `${a.aliasText}→${a.canonicalName}`).join('、') : '无'}`);
+  lines.push('');
+  if (warnings.length) {
+    lines.push('## 警告');
+    lines.push('');
+    for (const warning of uniq(warnings)) lines.push(`- ⚠️ ${warning}`);
+    lines.push('');
+  }
+  lines.push('## 下一步');
+  lines.push('');
+  lines.push('1. 打开 `material-assets.json` 核对顺序(九宫格从左到右、从上到下),需要调整用 `reorder`;');
+  lines.push('2. 确认无误后执行 `submit --package case-package.json`,落到待审清单(reviewStatus=pending,不进公共素材库);');
+  lines.push('3. 由负责人用 `caseApprove` 审核放行后才会 `readyForUse=true`、进检索结果。');
+  return `${lines.join('\n')}\n`;
+}
+
+// ---------------------------------------------------------------------------
+// submit
+// ---------------------------------------------------------------------------
+
+export async function runSubmit(options = {}) {
+  const packagePath = path.resolve(options.package);
+  const casePackage = readJson(packagePath, null);
+  if (!casePackage) throw new Error(`读不到案例包:${packagePath}`);
+
+  const blockers = casePackage.complianceBlockers || [];
+  if (String(casePackage.authorizationStatus || '').toLowerCase() !== 'authorized') {
+    return {
+      submitted: false,
+      reason: `authorizationStatus=${casePackage.authorizationStatus || '缺失'},仅 authorized 才可入库(硬规则:无授权不入库)`,
+      blockers: ['UNAUTHORIZED'],
+    };
+  }
+  if (blockers.length) {
+    return { submitted: false, reason: 'complianceBlockers 非空,禁止入库', blockers };
+  }
+  if (!casePackage.materialAssets || !casePackage.materialAssets.length) {
+    return { submitted: false, reason: 'materialAssets 为空,无法构成案例', blockers: ['NO_MATERIAL'] };
+  }
+
+  try {
+    const response = await submitCase(casePackage, { submittedBy: options.submittedBy, token: options.token });
+    if (!response.ok) {
+      return { submitted: false, reason: `caseSubmit 返回错误:${response.error.code} ${response.error.message}`, error: response.error };
+    }
+    return {
+      submitted: true,
+      objectId: response.result.objectId,
+      reviewStatus: response.result.reviewStatus || 'pending',
+      readyForUse: response.result.readyForUse === true,
+    };
+  } catch (error) {
+    if (error instanceof CloudError) {
+      return {
+        submitted: false,
+        reason: error.message,
+        offline: true,
+        hint: '未提交成功(网络不可达/未登录)。案例包已保存在本地,可直接重跑 submit;幂等键保证重复提交只产生一条。',
+      };
+    }
+    throw error;
+  }
+}
+
+// ---------------------------------------------------------------------------
+// reorder
+// ---------------------------------------------------------------------------
+
+export function runReorder(options = {}) {
+  const assetsPath = path.resolve(options.materialAssets || options.package);
+  const bundle = readJson(assetsPath, null);
+  if (!bundle) throw new Error(`读不到素材文件:${assetsPath}`);
+  const list = Array.isArray(bundle) ? bundle : bundle.materialAssets;
+  if (!Array.isArray(list)) throw new Error('素材文件里没有 materialAssets 数组');
+
+  let next;
+  if (options.move) {
+    const [from, to] = String(options.move).split(/[:>,-]/).map((v) => Number(v.trim()));
+    if (!Number.isFinite(from) || !Number.isFinite(to)) throw new Error('--move 格式应为 from:to,如 3:1');
+    next = reorderAssets(list, from, to);
+  } else if (options.order) {
+    // --order "3,1,2" 表示「新顺序 = 原来第 3 项、第 1 项、第 2 项」(按位置,不是按 order 值)
+    const positions = String(options.order).split(/[,\s]+/).map((v) => Number(v.trim())).filter(Number.isFinite);
+    if (positions.length !== list.length) throw new Error(`--order 需要 ${list.length} 个位置序号,收到 ${positions.length} 个`);
+    const byPosition = sortAssets(list);
+    next = positions.map((position, index) => {
+      const picked = byPosition[position - 1];
+      if (!picked) throw new Error(`--order 位置越界:${position}(共 ${byPosition.length} 项)`);
+      return { ...picked, order: index + 1 };
+    });
+  } else {
+    throw new Error('需要 --move "3:1" 或 --order "3,1,2"');
+  }
+
+  const out = Array.isArray(bundle)
+    ? next
+    : { ...bundle, materialAssets: next, reorderedAt: new Date().toISOString() };
+  writeJson(assetsPath, out);
+  return { file: assetsPath, materialAssets: next };
+}
+
+// ---------------------------------------------------------------------------
+// 字典
+// ---------------------------------------------------------------------------
+
+export function runDictionary(options = {}) {
+  const dict = loadDictionary(options.dictPath);
+  if (options.add) {
+    const items = [];
+    for (const raw of options.add) {
+      const [dimension, value] = String(raw).split('=');
+      items.push({ dimension: (dimension || '').trim(), value: (value || '').trim() });
+    }
+    const result = learnTags(dict, items);
+    if (result.added.length) saveDictionary(dict);
+    return { added: result.added, skipped: result.skipped, file: dict.__file };
+  }
+  if (options.alias) {
+    const entries = [];
+    for (const raw of options.alias) {
+      const [aliasText, canonicalName, country] = String(raw).split('=');
+      entries.push({ aliasText: (aliasText || '').trim(), canonicalName: (canonicalName || '').trim(), country: (country || '').trim() });
+    }
+    const result = learnAliases(dict, entries);
+    if (result.added.length) saveDictionary(dict);
+    return { added: result.added, file: dict.__file };
+  }
+  return {
+    file: dict.__file,
+    dimensions: Object.entries(dict.dimensions || {}).map(([key, dim]) => ({
+      dimension: key,
+      label: dim.label,
+      count: Array.isArray(dim.values) ? dim.values.length : Array.isArray(dim.entries) ? dim.entries.length : 0,
+      values: dim.values || undefined,
+    })),
+    learnedCount: (dict.__learned || []).length,
+  };
+}
+
+// ---------------------------------------------------------------------------
+// selftest
+// ---------------------------------------------------------------------------
+
+export async function runSelftest(options = {}) {
+  const results = [];
+  const python = resolvePython();
+
+  // ① office 解包(python 标准库)
+  const officeScript = path.join(SKILL_ROOT, 'scripts', 'office_extract.py');
+  const officeRun = spawnSync(python, [officeScript, '--selftest'], { encoding: 'utf8' });
+  results.push({
+    name: 'office_extract(docx/pptx 标准库解包)',
+    ok: officeRun.status === 0,
+    detail: truncate(officeRun.stdout || officeRun.stderr || '', 400),
+  });
+
+  // ② 分组 / 打标 / 合规:离线合成样例
+  const tmpRoot = ensureDir(path.join(outputsDir(options.outDir), 'selftest', newRunId()));
+  const pngPath = path.join(tmpRoot, 'grid-01.png');
+  fs.writeFileSync(pngPath, makeTinyPng(400, 400));
+
+  const gridItems = [0, 1, 2, 3, 4, 5, 6, 7, 8].map((i) => ({
+    kind: 'image',
+    localPath: path.join(tmpRoot, `grid-${String(i + 1).padStart(2, '0')}.png`),
+    width: 400,
+    height: 400,
+    mtimeMs: 1_700_000_000_000 + i * 1000,
+    ocrText: '帝国理工 计算机科学 考前冲刺 提分',
+    description: '',
+    label: `朋友圈素材 ${i + 1}`,
+  }));
+  const groups = groupMaterials(gridItems, {});
+  const gridOk = groups.length === 1 && groups[0].count === 9 && groups[0].layout === '3x3'
+    && groups[0].items.every((item, index) => item.order === index + 1);
+  results.push({
+    name: '九宫格分组与顺序(3x3,order 1..9)',
+    ok: gridOk,
+    detail: `分组数=${groups.length},布局=${groups[0] && groups[0].layout},顺序=${groups[0] && groups[0].items.map((i) => i.order).join(',')}`,
+  });
+
+  const separated = groupMaterials([
+    { kind: 'image', localPath: path.join(tmpRoot, 'a1.png'), width: 400, height: 400, mtimeMs: 1, ocrText: '帝国理工 提分' },
+    { kind: 'image', localPath: path.join(tmpRoot, 'a2.png'), width: 400, height: 400, mtimeMs: 2000, ocrText: '帝国理工 提分' },
+    { kind: 'image', localPath: path.join(tmpRoot, 'b1.png'), width: 1600, height: 900, mtimeMs: 900_000_000, ocrText: '香港大学 申诉' },
+  ], {});
+  results.push({
+    name: '跨批次不误并组(时间 + 尺寸 + 内容三重信号)',
+    ok: separated.length === 2,
+    detail: `分组数=${separated.length},各组成员=${separated.map((g) => g.count).join('+')}`,
+  });
+
+  const dict = loadDictionary(options.dictPath);
+  const classified = classifyCase({
+    dict,
+    groups: [...groups, ...separated],
+    visionResults: [],
+    textBundle: { paragraphs: ['帝国理工 计算机科学 考前冲刺辅导案例,学生期末提分 18 分,家长很满意', '微信:student_demo 联系 13800000000'], notes: [], titles: [], transcript: '' },
+    authorizationStatus: 'authorized',
+    hints: {},
+  });
+  const classifyOk = classified.tags.schoolCanonical.includes('帝国理工')
+    && classified.tags.highlightTypes.includes('提分')
+    && classified.compliance.privacyFindings.length > 0
+    && classified.compliance.authorizationOk === true;
+  results.push({
+    name: '打标 + 别名归一(IC→帝国理工)+ 隐私发现',
+    ok: classifyOk,
+    detail: `学校=${classified.tags.schoolCanonical.join('/')},亮点=${classified.tags.highlightTypes.join('/')},隐私=${classified.compliance.privacyFindings.map((f) => f.field).join('/')}`,
+  });
+
+  const complaint = classifyCase({
+    dict,
+    groups: [],
+    visionResults: [],
+    textBundle: { paragraphs: ['学生投诉退费纠纷,要求赔偿'], notes: [], titles: [], transcript: '' },
+    authorizationStatus: 'pending',
+    hints: {},
+  });
+  const riskOk = complaint.compliance.riskFlags.includes('COMPLAINT')
+    && complaint.compliance.riskFlags.includes('UNAUTHORIZED')
+    && complaint.compliance.complianceBlockers.length >= 2
+    && complaint.compliance.authorizationOk === false;
+  results.push({
+    name: '合规判断(客诉/负面 → riskFlags;无授权 → 不入库)',
+    ok: riskOk,
+    detail: `riskFlags=${complaint.compliance.riskFlags.join('/')},blockers=${complaint.compliance.complianceBlockers.length}`,
+  });
+
+  const reorderOk = (() => {
+    const moved = reorderAssets(sortAssets(groups[0].items).map((i) => ({ ...i, kind: 'image' })), 9, 1);
+    return moved[0].order === 1 && moved.length === 9;
+  })();
+  results.push({ name: '素材顺序调整(reorder)', ok: reorderOk, detail: '把第 9 张移到第 1 位后重新编号 1..9' });
+
+  const failed = results.filter((r) => !r.ok);
+  return { ok: failed.length === 0, checked: results.length, failed: failed.length, results, workDir: tmpRoot };
+}
+
+/** 合成一个最小的合法 PNG(用于自检,不引用任何真实素材)。 */
+export function makeTinyPng(width = 1, height = 1) {
+  const crcTable = (() => {
+    const table = new Int32Array(256);
+    for (let n = 0; n < 256; n++) {
+      let c = n;
+      for (let k = 0; k < 8; k++) c = c & 1 ? 0xedb88320 ^ (c >>> 1) : c >>> 1;
+      table[n] = c;
+    }
+    return table;
+  })();
+  const crc32 = (buf) => {
+    let c = -1;
+    for (const byte of buf) c = crcTable[(c ^ byte) & 0xff] ^ (c >>> 8);
+    return (c ^ -1) >>> 0;
+  };
+  const chunk = (type, data) => {
+    const length = Buffer.alloc(4);
+    length.writeUInt32BE(data.length, 0);
+    const typeBuf = Buffer.from(type, 'ascii');
+    const crc = Buffer.alloc(4);
+    crc.writeUInt32BE(crc32(Buffer.concat([typeBuf, data])), 0);
+    return Buffer.concat([length, typeBuf, data, crc]);
+  };
+  const ihdr = Buffer.alloc(13);
+  ihdr.writeUInt32BE(width, 0);
+  ihdr.writeUInt32BE(height, 4);
+  ihdr[8] = 8;   // bit depth
+  ihdr[9] = 6;   // RGBA
+  const raw = Buffer.alloc((width * 4 + 1) * height);
+  return Buffer.concat([
+    Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]),
+    chunk('IHDR', ihdr),
+    chunk('IDAT', zlib.deflateSync(raw)),
+    chunk('IEND', Buffer.alloc(0)),
+  ]);
+}
+
+// ---------------------------------------------------------------------------
+// CLI
+// ---------------------------------------------------------------------------
+
+function parseCli(argv) {
+  const command = argv[0] && !argv[0].startsWith('-') ? argv[0] : 'ingest';
+  const rest = command === argv[0] ? argv.slice(1) : argv;
+  const { values, positionals } = parseArgs({
+    args: rest,
+    options: {
+      input: { type: 'string', multiple: true, default: [] },
+      image: { type: 'string', multiple: true, default: [] },
+      doc: { type: 'string', multiple: true, default: [] },
+      deck: { type: 'string', multiple: true, default: [] },
+      video: { type: 'string', multiple: true, default: [] },
+      'out-dir': { type: 'string' },
+      'source-type': { type: 'string' },
+      'source-ref': { type: 'string' },
+      authorization: { type: 'string', default: 'pending' },
+      title: { type: 'string' },
+      summary: { type: 'string' },
+      'target-customer': { type: 'string' },
+      'usage-suggestion': { type: 'string' },
+      'school-canonical': { type: 'string' },
+      'product-line': { type: 'string' },
+      major: { type: 'string' },
+      stage: { type: 'string' },
+      country: { type: 'string' },
+      highlight: { type: 'string', multiple: true, default: [] },
+      scenario: { type: 'string', multiple: true, default: [] },
+      objection: { type: 'string', multiple: true, default: [] },
+      'idempotency-key': { type: 'string' },
+      'asset-base-url': { type: 'string' },
+      'submitted-by': { type: 'string' },
+      'vision-model': { type: 'string' },
+      'skip-vision': { type: 'boolean', default: false },
+      'skip-video': { type: 'boolean', default: false },
+      'frame-count': { type: 'string' },
+      diarize: { type: 'boolean', default: false },
+      'no-learn': { type: 'boolean', default: false },
+      'dict-path': { type: 'string' },
+      package: { type: 'string' },
+      'material-assets': { type: 'string' },
+      move: { type: 'string' },
+      order: { type: 'string' },
+      add: { type: 'string', multiple: true, default: [] },
+      alias: { type: 'string', multiple: true, default: [] },
+      out: { type: 'string' },
+      json: { type: 'boolean', default: false },
+      help: { type: 'boolean', default: false, short: 'h' },
+    },
+    allowPositionals: true,
+  });
+  return { command, values, positionals };
+}
+
+function collectInputs(values, positionals) {
+  return [
+    ...values.input,
+    ...values.image,
+    ...values.doc,
+    ...values.deck,
+    ...values.video,
+    ...positionals,
+  ];
+}
+
+async function cliMain(argv) {
+  const { command, values, positionals } = parseCli(argv);
+
+  if (values.help) {
+    process.stdout.write(HELP_TEXT);
+    return 0;
+  }
+
+  if (command === 'selftest') {
+    const report = await runSelftest({ dictPath: values['dict-path'], outDir: values['out-dir'] });
+    process.stdout.write(`${JSON.stringify(report, null, 2)}\n`);
+    return report.ok ? 0 : 1;
+  }
+
+  if (command === 'dictionary') {
+    const report = runDictionary({ add: values.add, alias: values.alias, dictPath: values['dict-path'] });
+    process.stdout.write(`${JSON.stringify(report, null, 2)}\n`);
+    return 0;
+  }
+
+  if (command === 'reorder') {
+    const report = runReorder({
+      materialAssets: values['material-assets'] || values.package,
+      move: values.move,
+      order: values.order,
+    });
+    process.stdout.write(`${JSON.stringify({ ok: true, file: report.file, order: report.materialAssets.map((a) => a.order) }, null, 2)}\n`);
+    return 0;
+  }
+
+  if (command === 'submit') {
+    if (!values.package) {
+      process.stderr.write('需要 --package <case-package.json>\n');
+      return 2;
+    }
+    const report = await runSubmit({ package: values.package, submittedBy: values['submitted-by'] });
+    process.stdout.write(`${JSON.stringify(report, null, 2)}\n`);
+    return report.submitted ? 0 : 4;
+  }
+
+  // plan / ingest
+  const outDir = values['out-dir']
+    || path.join(outputsDir(), `${values['source-ref'] ? `${slug(values['source-ref'])}-` : ''}${newRunId()}`);
+
+  const options = {
+    inputs: collectInputs(values, positionals),
+    outDir,
+    sourceType: values['source-type'],
+    sourceRef: values['source-ref'],
+    authorizationStatus: values.authorization,
+    title: values.title,
+    summary: values.summary,
+    targetCustomer: values['target-customer'],
+    usageSuggestion: values['usage-suggestion'],
+    idempotencyKey: values['idempotency-key'],
+    assetBaseUrl: values['asset-base-url'] || process.env.CASE_ASSET_BASE_URL || '',
+    submittedBy: values['submitted-by'],
+    visionModel: values['vision-model'],
+    skipVision: values['skip-vision'],
+    skipVideo: values['skip-video'],
+    frameCount: values['frame-count'] ? Number(values['frame-count']) : undefined,
+    diarize: values.diarize,
+    noLearn: values['no-learn'],
+    dictPath: values['dict-path'],
+    hints: {
+      schoolCanonical: values['school-canonical'],
+      productLine: values['product-line'],
+      major: values.major,
+      stage: values.stage,
+      country: values.country,
+      highlightTypes: values.highlight,
+      scenarioTags: values.scenario,
+      objectionTags: values.objection,
+    },
+  };
+
+  if (!options.inputs.length) {
+    process.stderr.write('至少需要一个输入(位置参数或 --input / --image / --doc / --deck / --video)\n');
+    return 2;
+  }
+
+  if (command === 'plan') {
+    const plan = await runPlan(options);
+    process.stdout.write(`${JSON.stringify({
+      ok: true,
+      command: 'plan',
+      runId: plan.runId,
+      sourceType: plan.sourceType,
+      counts: plan.counts,
+      groups: plan.groups.map((g) => ({ groupId: g.groupId, layout: g.layout, count: g.count, orderingRule: g.orderingRule })),
+      warnings: plan.warnings,
+      outDir,
+    }, null, 2)}\n`);
+    return 0;
+  }
+
+  const report = await runIngest(options);
+  const summary = {
+    ok: true,
+    command: 'ingest',
+    runId: report.casePackage.runId,
+    sourceType: report.casePackage.sourceType,
+    idempotencyKey: report.casePackage.idempotencyKey,
+    materialAssets: report.casePackage.materialAssets.length,
+    groups: report.casePackage.groupInfoList,
+    title: report.casePackage.title,
+    tags: {
+      schoolCanonical: report.casePackage.schoolCanonical,
+      schoolAliases: report.casePackage.schoolAliases,
+      country: report.casePackage.country,
+      major: report.casePackage.major,
+      stage: report.casePackage.stage,
+      subject: report.casePackage.subject,
+      highlightTypes: report.casePackage.highlightTypes,
+      scenarioTags: report.casePackage.scenarioTags,
+      objectionTags: report.casePackage.objectionTags,
+      fitStatus: report.casePackage.fitStatus,
+    },
+    compliance: {
+      riskFlags: report.casePackage.riskFlags,
+      privacyFindings: report.casePackage.privacyFindings.length,
+      blockers: report.casePackage.complianceBlockers,
+      canSubmit: report.canSubmit,
+      expected: report.casePackage.expected,
+    },
+    dictionaryChanges: report.dictionaryChanges,
+    warnings: uniq(report.warnings),
+    artifacts: report.artifacts,
+  };
+  process.stdout.write(`${JSON.stringify(summary, null, 2)}\n`);
+  return 0;
+}
+
+function slug(value) {
+  return String(value).replace(/[^\w.-]+/g, '-').replace(/^-+|-+$/g, '').slice(0, 40);
+}
+
+const HELP_TEXT = `skill-case-get —— 案例采集 / 拆解 / 归档
+
+用法:
+  case-intake.mjs ingest [输入...] [选项]     拆解 + 打标,产出案例包(默认不提交)
+  case-intake.mjs plan   [输入...]            只识别 + 提取 + 分组(不调模型)
+  case-intake.mjs submit --package <file>     提交到云函数 caseSubmit(待审清单)
+  case-intake.mjs reorder --material-assets <file> --move "3:1"
+  case-intake.mjs dictionary [--add highlightTypes=退费保障] [--alias "IC=帝国理工=英国"]
+  case-intake.mjs selftest
+
+输入(可混用,支持目录):
+  --image a.jpg --image b.jpg ...       连续多图 / 裸图 / 九宫格
+  --doc case.docx                        Word
+  --deck deck.pptx                       PPT
+  --video clip.mp4                       视频(抽帧 + 转写)
+  --input <file|dir>                     自动按扩展名分流
+
+采集选项:
+  --authorization authorized|pending|denied   授权状态(硬门:仅 authorized 可入库,默认 pending)
+  --source-ref <ref>                          来源引用(不含 PII)
+  --source-type <t>                           image_group|document|presentation|mixed|single_image
+  --title / --summary / --target-customer / --usage-suggestion   覆盖自动抽取的案例字段
+  --school-canonical / --product-line / --major / --stage / --country
+  --highlight / --scenario / --objection      显式补标签(可重复)
+  --idempotency-key <key>                     幂等键(默认按内容指纹生成)
+  --asset-base-url <base>                     本地素材 URL 前缀(CDN / Parse Files),用于派生 materialAssets[].url
+                                               (也可用环境变量 CASE_ASSET_BASE_URL)
+  --out-dir <dir>                             产物目录(默认 <cwd>/case-get-outputs/<runId>)
+  --skip-vision / --skip-video / --no-learn   关闭图片理解 / 视频处理 / 字典自进化
+  --submitted-by <userId>                     提交人
+
+提交选项:
+  --package <case-package.json>   ingest 产出的案例包
+  --submitted-by <userId>
+
+产物(--out-dir):
+  extracted.json      文档文本 + 内嵌媒体
+  group-plan.json     分组与九宫格顺序
+  vision.json         每张图的 OCR / 描述 / 使用建议
+  material-assets.json materialAssets[](可编辑顺序)
+  case-package.json   云函数 caseSubmit 入参(逐字对齐共享契约)
+  report.md           人类可读报告
+`;
+
+const invokedDirectly = process.argv[1] && path.resolve(process.argv[1]).endsWith(path.join('scripts', 'case-intake.mjs'));
+if (invokedDirectly) {
+  cliMain(process.argv.slice(2))
+    .then((code) => process.exit(code))
+    .catch((error) => {
+      process.stderr.write(`skill-case-get 失败:${error.message}\n`);
+      process.exit(1);
+    });
+}
+
+export { cliMain };

+ 520 - 0
scripts/classify.mjs

@@ -0,0 +1,520 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 分类层:案例信息 vs 素材 分离、打标、合规判断
+ *
+ * 三类知识结构严格分开:
+ *   ① 案例说明(什么时候用 / 大概情况 / 背景)→ 案例字段
+ *      title / summary / targetCustomer / usageSuggestion / resultEvidence / outcome
+ *   ② 图片视频 → materialAssets[](本文件不动顺序,只做角色与标签)
+ *   ③ 合规判定 → riskFlags / privacyFindings
+ *
+ * 打标依据:docs/prd/dashboard/05-标签与别名管理.md + 10-案例推荐策略.md,
+ * 取值来源 references/tag-dictionary.json(技能会在采集过程中增量更新)。
+ */
+
+import { uniq, truncate } from './lib.mjs';
+
+// ---------------------------------------------------------------------------
+// 隐私发现(本期只记录,不打码;片段做掩码,绝不落完整 PII)
+// ---------------------------------------------------------------------------
+
+const PII_RULES = [
+  { field: '手机号', rule: 'phone', re: /(?:\+?86[-\s]?)?1[3-9]\d{9}\b/g, mask: (m) => `${m.slice(0, 3)}****${m.slice(-4)}` },
+  { field: '邮箱', rule: 'email', re: /[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}/gi, mask: (m) => `${m[0]}***@${m.split('@')[1]}` },
+  { field: '微信号', rule: 'wechat', re: /(?:微信|WeChat|wechat|VX|vx)\s*[::]?\s*[A-Za-z][-_A-Za-z0-9]{5,19}/g, mask: () => '微信:[已掩码]' },
+  { field: 'QQ', rule: 'qq', re: /(?:QQ|qq)\s*[::]?\s*\d{5,12}/g, mask: () => 'QQ:[已掩码]' },
+  { field: '身份证号', rule: 'id-card', re: /\b\d{17}[\dXx]\b/g, mask: (m) => `${m.slice(0, 3)}************${m.slice(-3)}` },
+  { field: '银行卡号', rule: 'bank-card', re: /\b\d{16,19}\b/g, mask: (m) => `${m.slice(0, 4)}********${m.slice(-4)}` },
+  { field: '称呼', rule: 'student-name', re: /(?:姓名|学员|学生|同学)\s*[::]\s*[一-龥]{2,4}/g, mask: (m) => `${m.split(/[::]/)[0]}:[已掩码]` },
+];
+
+/**
+ * 在文本中找隐私片段。**片段一律掩码后再返回**(真实 PII 不持久化,红线)。
+ * @returns {Array<{field:string, rule:string, snippet:string}>}
+ */
+export function findPrivacy(text) {
+  const findings = [];
+  const source = String(text || '');
+  if (!source) return findings;
+  for (const pii of PII_RULES) {
+    const matches = source.match(pii.re);
+    if (!matches) continue;
+    for (const match of uniq(matches).slice(0, 5)) {
+      findings.push({ field: pii.field, rule: pii.rule, snippet: truncate(pii.mask(match), 40) });
+    }
+  }
+  return findings;
+}
+
+// ---------------------------------------------------------------------------
+// 合规关键词
+// ---------------------------------------------------------------------------
+
+// 关键词只描述「我方服务出的负面问题」,不包含学生自身的前置困难。
+// 「挂科 / 补考 / 被拒 / 退款」本身是痛点与需求(甚至对应「挂科挽救」「退费保障」等亮点),
+// 命中它们会把最好的案例误判成负面事件——因此一律不在此列。
+const RISK_KEYWORDS = {
+  COMPLAINT: [
+    '投诉', '客诉', '差评', '维权', '举报', '曝光',
+    '退费纠纷', '退款纠纷', '要求退款', '要求赔偿', '协商赔偿', '黑猫',
+  ],
+  NEGATIVE_EVENT: [
+    '投诉老师', '要求换老师', '老师失联', '服务事故', '课时不符',
+    '没效果', '没有效果', '效果不好', '毫无效果',
+    '成绩造假', '伪造材料', '数据造假',
+  ],
+};
+
+const TOPIC_KEYWORDS = ['留学', '申请', '课程', '辅导', '补考', '提分', '录取', '申诉', '学校', '专业', '成绩', '成绩单', '老师', '答疑', 'GPA', '雅思', '托福', 'A-Level', 'AP', 'IB', 'case', 'offer', '期末', '考试', '论文'];
+
+// ---------------------------------------------------------------------------
+// 打标关键词(中文线索 → 字典取值)
+// ---------------------------------------------------------------------------
+
+const TAG_CLUES = {
+  highlightTypes: {
+    提分: ['提分', '提高了', '涨了', '分数上升', '从\d+分?到\d+分?'],
+    录取: ['录取', 'offer', '拿到了', '上岸', 'offer letter'],
+    申诉成功: ['申诉成功', '申诉通过', '撤销处分', '撤销学术不端', 'appeal 成功'],
+    补考通过: ['补考通过', '补考 pass', '补考过了', 'pass 了', '补考稳 pass'],
+    稳分: ['稳分', '稳住', '保分', '不再挂'],
+    高分冲刺: ['冲刺高分', '冲高分', 'distinction', '一等学位'],
+    挂科挽救: ['挂科挽救', '挽救', '差点挂', '补救'],
+    退费保障: ['退费', '保障', '不过退'],
+    家长认可: ['家长', '妈妈', '爸爸', '家长很满意', '家长认可'],
+    出分反馈: ['出分', '成绩出来了', '反馈成绩', '查分'],
+    首课体验: ['首课', '试听', '第一节课'],
+    导师匹配: ['匹配导师', '导师匹配', '换到合适的老师', '老师很对口'],
+    群内答疑: ['群内答疑', '答疑群', '随时问', '响应很快'],
+    课时反馈: ['课时反馈', '每节课反馈', '课后反馈', '课堂反馈'],
+    大课时囤课: ['大课时', '囤课', '续课', '加课'],
+  },
+  scenarioTags: {
+    考前冲刺: ['考前', '冲刺', '临近考试', '来不及'],
+    开学季: ['开学', '新学期'],
+    期中考试: ['期中'],
+    期末考试: ['期末', 'final'],
+    补考季: ['补考', 'resit', '补考季'],
+    论文季: ['论文', 'dissertation', '毕业论文'],
+    选课指导: ['选课', '选课指导'],
+    转专业: ['转专业', '换专业', '转系'],
+    申诉期: ['申诉', 'appeal'],
+    退费咨询: ['退费', '退款'],
+    家长陪同: ['家长', '陪读', '妈妈', '爸爸'],
+    多地时差: ['时差', '国外时间', '半夜'],
+    时间紧张: ['时间不够', '时间紧', '来不及', '任务重'],
+    基础薄弱: ['基础差', '基础薄弱', '零基础', '底子差'],
+    目标高分: ['高分', '冲分', 'distinction', '一等'],
+    临近毕业: ['毕业', '大四', '最后一年'],
+  },
+  objectionTags: {
+    贵: ['太贵', '贵了', '价格高', '预算不够', '便宜点'],
+    不放心老师: ['不放心老师', '老师行不行', '老师靠不靠谱', '换老师'],
+    犹豫: ['再想想', '犹豫', '考虑一下', '还没想好'],
+    怕没效果: ['没效果', '有用吗', '能提分吗', '真的能'],
+    怕时间不够: ['来不及', '时间不够', '太晚了'],
+    要和家里商量: ['和家里商量', '问下父母', '跟爸妈'],
+    对比其他机构: ['其他机构', '别家', '对比一下', '也在看'],
+    担心退费难: ['退费难', '能退吗', '退款麻烦'],
+  },
+};
+
+function matchesAny(text, needles) {
+  return needles.some((needle) => {
+    if (needle.startsWith('从') || needle.includes('\\d')) {
+      try {
+        return new RegExp(needle).test(text);
+      } catch {
+        return text.includes(needle);
+      }
+    }
+    return text.toLowerCase().includes(needle.toLowerCase());
+  });
+}
+
+/** 按字典 + 关键词线索打标,返回命中值与被用到但字典里没有的新值。 */
+export function classifyTags(text, dict) {
+  const corpus = String(text || '');
+  const dims = (dict && dict.dimensions) || {};
+  const out = {};
+  const learned = [];
+
+  const pickFromDictionary = (dimension) => {
+    const values = Array.isArray(dims[dimension] && dims[dimension].values) ? dims[dimension].values : [];
+    return values.filter((value) => corpus.includes(value));
+  };
+
+  // 硬性静态属性:字典值直接命中
+  out.country = pickFromDictionary('country').slice(0, 1);
+  out.schoolCanonical = pickFromDictionary('schoolCanonical');
+  out.major = pickFromDictionary('major');
+  out.subject = pickFromDictionary('subject');
+  out.stage = pickFromDictionary('stage').slice(0, 1);
+
+  // 产品线(硬规则:三选一,字典里没有则不猜)
+  out.productLine = pickFromDictionary('productLine').slice(0, 1);
+
+  // 关键词线索打标
+  for (const dimension of ['highlightTypes', 'scenarioTags', 'objectionTags']) {
+    const clues = TAG_CLUES[dimension] || {};
+    const hits = [];
+    for (const [value, needles] of Object.entries(clues)) {
+      if (matchesAny(corpus, needles)) hits.push(value);
+      else if (corpus.includes(value)) hits.push(value);
+    }
+    out[dimension] = uniq(hits);
+  }
+
+  // 字典里已有但关键词没覆盖到的补充命中(避免漏标字典取值)
+  for (const dimension of ['highlightTypes', 'scenarioTags', 'objectionTags']) {
+    const known = Array.isArray(dims[dimension] && dims[dimension].values) ? dims[dimension].values : [];
+    out[dimension] = uniq([...(out[dimension] || []), ...known.filter((v) => corpus.includes(v))]);
+  }
+
+  // 字典自进化候选:文本里出现的、形如标签的短语(保守:只在强线索命中后才收集)
+  return { tags: out, learned };
+}
+
+// ---------------------------------------------------------------------------
+// 学校别名归一
+// ---------------------------------------------------------------------------
+
+/**
+ * 把文本里的别名写法归一为标准校名。
+ * @returns {{schoolCanonical:string, schoolAliases:string[], country:string, hits:object[]}}
+ */
+export function normalizeSchool(text, dict) {
+  const corpus = String(text || '');
+  const entries = (dict && dict.dimensions && dict.dimensions.schoolAlias && dict.dimensions.schoolAlias.entries) || [];
+  const canonicalNames = (dict && dict.dimensions && dict.dimensions.schoolCanonical && dict.dimensions.schoolCanonical.values) || [];
+
+  const hits = [];
+  for (const entry of entries) {
+    const alias = String(entry.aliasText || '');
+    if (!alias) continue;
+    // 英文别名按词边界匹配,避免 "IC" 命中 "MAGIC";中文直接包含
+    const isAscii = /^[\x20-\x7e]+$/.test(alias);
+    const hit = isAscii
+      ? new RegExp(`(^|[^A-Za-z0-9])${alias.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}([^A-Za-z0-9]|$)`, 'i').test(corpus)
+      : corpus.includes(alias);
+    if (hit) hits.push({ aliasText: alias, canonicalName: entry.canonicalName, country: entry.country || '' });
+  }
+
+  // 标准名直接出现(此时国家从别名表里同名校名的条目推断)
+  const directHits = canonicalNames.filter((name) => corpus.includes(name));
+  const countryOf = (canonical) => {
+    const entry = entries.find((e) => String(e.canonicalName) === canonical);
+    return entry ? String(entry.country || '') : '';
+  };
+
+  const canonical = uniq([...hits.map((h) => h.canonicalName), ...directHits])[0] || '';
+  const country = (hits.find((h) => h.canonicalName === canonical) || {}).country || countryOf(canonical);
+  return {
+    schoolCanonical: canonical,
+    // 命中的别名写法(不含标准名本身),供写入 CaseAsset.schoolAliases
+    schoolAliases: uniq(hits.map((h) => h.aliasText)),
+    country,
+    hits,
+  };
+}
+
+// ---------------------------------------------------------------------------
+// 推荐适用跟进状态(docs/prd/dashboard/10-案例推荐策略.md)
+// ---------------------------------------------------------------------------
+
+const STRATEGY_MAP = [
+  { status: '异议处理中', when: (t) => t.objectionTags.length > 0 || ['出分反馈', '家长认可', '退费保障', '大课时囤课'].some((v) => t.highlightTypes.includes(v)) },
+  { status: '待决策', when: (t) => ['录取', '提分', '高分冲刺'].some((v) => t.highlightTypes.includes(v)) && t.scenarioTags.includes('考前冲刺') === false },
+  { status: '挖需中', when: (t) => ['补考通过', '挂科挽救', '稳分'].some((v) => t.highlightTypes.includes(v)) || t.scenarioTags.some((s) => ['时间紧张', '基础薄弱', '补考季'].includes(s)) },
+  { status: '方案推荐中', when: (t) => ['导师匹配', '群内答疑', '课时反馈'].some((v) => t.highlightTypes.includes(v)) },
+  { status: '沉默待跟进', when: (t) => t.scenarioTags.includes('考前冲刺') || t.highlightTypes.includes('出分反馈') },
+  { status: '已成交', when: (t) => ['首课体验', '群内答疑', '课时反馈'].some((v) => t.highlightTypes.includes(v)) },
+  { status: '新进线', when: (t) => t.subject.length > 0 },
+];
+
+export function inferFitStatus(tags) {
+  const t = {
+    highlightTypes: tags.highlightTypes || [],
+    scenarioTags: tags.scenarioTags || [],
+    objectionTags: tags.objectionTags || [],
+    subject: tags.subject || [],
+  };
+  const matched = STRATEGY_MAP.filter((rule) => rule.when(t)).map((rule) => rule.status);
+  const ordered = ['新进线', '挖需中', '方案推荐中', '异议处理中', '待决策', '沉默待跟进', '已成交'];
+  const fitStatus = ordered.filter((s) => matched.includes(s));
+  return {
+    fitStatus,
+    reason: fitStatus.length
+      ? `按「跟进状态 → 该发什么案例」策略命中:${fitStatus.join('、')}`
+      : '未命中策略规则,需人工指定适用跟进状态',
+  };
+}
+
+// ---------------------------------------------------------------------------
+// 案例 vs 素材 分离
+// ---------------------------------------------------------------------------
+
+/**
+ * 从各种来源抽出「案例说明」文本,与素材数组分开。
+ * @param {object} input
+ *   - textBundle: { paragraphs:string[], notes:string[], titles:string[], transcript:string }
+ *   - visionResults: [{ocrText, description, usageSuggestion, role, label, ok}]
+ */
+export function splitCaseAndMaterials(input = {}) {
+  const bundle = input.textBundle || {};
+  const descriptionTexts = [];
+  const materialTexts = [];
+
+  const push = (target, value) => {
+    const text = String(value || '').trim();
+    if (text) target.push(text);
+  };
+
+  for (const paragraph of bundle.paragraphs || []) push(descriptionTexts, paragraph);
+  for (const note of bundle.notes || []) push(descriptionTexts, typeof note === 'string' ? note : note.text);
+  push(descriptionTexts, bundle.transcript);
+
+  for (const result of input.visionResults || []) {
+    if (result.role === 'description') {
+      push(descriptionTexts, result.ocrText);
+      push(descriptionTexts, result.description);
+    } else {
+      push(materialTexts, result.ocrText);
+      push(materialTexts, result.description);
+    }
+  }
+
+  return {
+    // 案例说明语料(用于抽 title/summary/usageSuggestion/outcome)
+    descriptionCorpus: descriptionTexts.join('\n'),
+    descriptionTexts,
+    // 素材语料(用于打标、也用于素材 label)
+    materialCorpus: materialTexts.join('\n'),
+    materialTexts,
+  };
+}
+
+/** 从说明语料里抽一句话摘要、使用建议、目标客户、结果证据。 */
+export function deriveCaseFields(corpus, groups, tags, fit) {
+  const lines = String(corpus || '').split('\n').map((l) => l.trim()).filter(Boolean);
+
+  const title = truncate(
+    lines.find((l) => l.length >= 6 && l.length <= 40) || lines[0] || '未命名案例',
+    40,
+  );
+
+  const summaryLine = lines.find((l) => l.length >= 12 && /[,。!?]|提分|录取|通过|结果/.test(l)) || lines[1] || lines[0] || '';
+  const summary = truncate(summaryLine || title, 120);
+
+  const productLine = (tags.productLine || [])[0] || '';
+  const school = (tags.schoolCanonical || [])[0] || '';
+  const subject = (tags.subject || [])[0] || '';
+  const stage = (tags.stage || [])[0] || '';
+  const audience = [stage, school, subject, productLine].filter(Boolean).join(' · ');
+
+  const usageSuggestion = (() => {
+    const fits = fit.fitStatus || [];
+    if (fits.includes('异议处理中')) return `客户在顾虑(${(tags.objectionTags || []).join('/') || '犹豫、怕没效果'})时发,用结果打消顾虑`;
+    if (fits.includes('挖需中')) return `客户痛点对上(${(tags.scenarioTags || []).join('/') || '怕挂、时间紧'})时发,把痛点和方案对上`;
+    if (fits.includes('沉默待跟进')) return '客户不回复时做低成本唤醒,考前节点发';
+    if (fits.includes('待决策')) return '已报价、客户说考虑时发,给一条完整可参考的成单路径';
+    if (fits.includes('已成交')) return '成交后发,让客户安心(服务过程类)';
+    if (fits.includes('方案推荐中')) return '推产品阶段发,证明机制有效(导师匹配/群内答疑/课时反馈)';
+    if (fits.includes('新进线')) return `刚加上还没破冰时发,用${subject || '所学科目'}的案例建立专业感`;
+    return '按客户当前跟进状态选用';
+  })();
+
+  const resultEvidence = truncate(
+    lines.filter((l) => /提分|录取|通过|出分|反馈|成绩|offer|pass/i.test(l)).slice(0, 2).join(';'),
+    160,
+  );
+
+  const outcome = extractOutcome(corpus);
+
+  return {
+    title,
+    summary,
+    targetCustomer: audience || '有同类需求的在读学生',
+    usageSuggestion,
+    resultEvidence,
+    outcome,
+  };
+}
+
+/** 抽结果数据:提分幅度、周期、录取/出分结果、学校专业。 */
+export function extractOutcome(corpus) {
+  const text = String(corpus || '');
+  const outcome = {};
+
+  const scoreDelta = text.match(/(?:提分|提高|涨了|提升了?)\s*([\d.]+)\s*分/);
+  if (scoreDelta) outcome.scoreGain = `${scoreDelta[1]}分`;
+
+  const fromTo = text.match(/(\d{1,3})\s*分?\s*(?:到|→|->|至)\s*(\d{1,3})\s*分/);
+  if (fromTo) outcome.scoreRange = `${fromTo[1]} → ${fromTo[2]}`;
+
+  const duration = text.match(/(\d+)\s*(?:周|个月|月|天|课时)/);
+  if (duration) outcome.period = duration[0];
+
+  // 「被 / 收到 / 拿到 / 获得」后可选的「了 / 到」,再到校名 + 录取/offer
+  const admitted = text.match(/(?:被|收到|拿到|获得)\s*(?:了|到)?\s*([一-龥A-Za-z][一-龥A-Za-z\s]{1,19}?)(?:的)?(?:录取|offer)/i);
+  if (admitted) outcome.admittedTo = admitted[1].trim();
+
+  const grade = text.match(/(?:GPA|均分|成绩)\s*(?:从)?\s*([\d.]+)/i);
+  if (grade) outcome.grade = grade[1];
+
+  if (/pass|通过|及格/i.test(text)) outcome.result = 'pass';
+  else if (/distinction|一等/i.test(text)) outcome.result = 'distinction';
+
+  return outcome;
+}
+
+// ---------------------------------------------------------------------------
+// 合规判断
+// ---------------------------------------------------------------------------
+
+/**
+ * 能否进案例库?这是「客诉/负面事件」还是「可对外展示的好案例」?
+ *
+ * @param {object} input
+ *   - corpus: string 全部文本
+ *   - authorizationStatus: 'authorized' | 'pending' | 'denied' | 其它
+ *   - hasMaterials: boolean 是否有可用素材
+ *   - fields: 抽取出来的案例字段(用于判断信息完整度)
+ * @returns {{authorizationOk:boolean, riskFlags:string[], privacyFindings:object[], complianceBlockers:string[], reviewHint:string}}
+ */
+export function assessCompliance(input = {}) {
+  const corpus = String(input.corpus || '');
+  const riskFlags = [];
+  const complianceBlockers = [];
+
+  for (const [flag, needles] of Object.entries(RISK_KEYWORDS)) {
+    if (matchesAny(corpus, needles)) riskFlags.push(flag);
+  }
+
+  const privacyFindings = findPrivacy(corpus);
+  if (privacyFindings.length) riskFlags.push('PII_RISK');
+
+  // 主题相关性
+  const topicHit = matchesAny(corpus, TOPIC_KEYWORDS) || (input.hasMaterials && corpus.trim().length > 0);
+  if (!topicHit) riskFlags.push('OFF_TOPIC');
+
+  // 授权判定(硬门)
+  const authorizationStatus = String(input.authorizationStatus || '').toLowerCase();
+  const authorizationOk = authorizationStatus === 'authorized';
+  if (!authorizationOk) {
+    riskFlags.push('UNAUTHORIZED');
+    complianceBlockers.push(`无授权(authorizationStatus=${authorizationStatus || '缺失'})——一律不入库、不产生案例对象`);
+  }
+
+  if (!input.hasMaterials) {
+    complianceBlockers.push('没有任何可用素材(materialAssets 为空),无法构成案例');
+  }
+
+  if (['COMPLAINT', 'NEGATIVE_EVENT'].includes(riskFlags.find((f) => f === 'COMPLAINT' || f === 'NEGATIVE_EVENT'))) {
+    complianceBlockers.push('出现客诉/负面事件线索:这是内部复盘材料,不能作为对外展示的好案例进入公共素材库');
+  }
+
+  if (riskFlags.includes('OFF_TOPIC')) {
+    complianceBlockers.push('与留学/课程/辅导主题无关,不符合案例规范');
+  }
+
+  let reviewHint = '可入库待审(reviewStatus=pending,不进公共素材库,需人工审核后放行)';
+  if (riskFlags.includes('COMPLAINT') || riskFlags.includes('NEGATIVE_EVENT')) {
+    reviewHint = '建议驳回或转内部复盘,不得对外展示';
+  } else if (privacyFindings.length) {
+    reviewHint = '存在隐私片段,打码前必须人工抽检(未通过不得进入检索结果)';
+  }
+
+  return {
+    authorizationOk,
+    riskFlags: uniq(riskFlags),
+    privacyFindings,
+    complianceBlockers,
+    reviewHint,
+  };
+}
+
+// ---------------------------------------------------------------------------
+// 汇总:把以上拼成一个「案例包」
+// ---------------------------------------------------------------------------
+
+/**
+ * @param {object} input
+ *   - sourceType, sourceRef, authorizationStatus
+ *   - textBundle {paragraphs, notes, titles, transcript}
+ *   - visionResults [], groups []
+ *   - dict 已加载的标签字典
+ *   - hints { schoolCanonical, productLine, stage, tags:{...} } 人工/上游显式指定(优先级最高)
+ * @returns {{caseFields:object, tags:object, compliance:object, learned:{tags:[],aliases:[]}}}
+ */
+export function classifyCase(input = {}) {
+  const dict = input.dict;
+  const groups = input.groups || [];
+  const visionResults = input.visionResults || [];
+
+  const split = splitCaseAndMaterials({
+    textBundle: input.textBundle,
+    visionResults,
+  });
+
+  const groupsCorpus = groups.map((g) => g.preview || '').join('\n');
+  const fullCorpus = [split.descriptionCorpus, split.materialCorpus, groupsCorpus].filter(Boolean).join('\n');
+
+  // 打标:先关键词/字典,再用 hints 覆盖
+  const { tags: rawTags } = classifyTags(fullCorpus, dict);
+  const school = normalizeSchool(fullCorpus, dict);
+
+  const hints = input.hints || {};
+  const tags = {
+    country: uniq([hints.country, school.country, ...(rawTags.country || [])]).slice(0, 1),
+    productLine: uniq(hints.productLine ? [hints.productLine] : [], rawTags.productLine || []).slice(0, 1),
+    schoolCanonical: uniq(hints.schoolCanonical ? [hints.schoolCanonical] : [], school.schoolCanonical ? [school.schoolCanonical] : [], rawTags.schoolCanonical || []),
+    schoolAliases: uniq(school.schoolAliases),
+    major: uniq(hints.major ? [hints.major] : [], rawTags.major || []),
+    stage: uniq(hints.stage ? [hints.stage] : [], rawTags.stage || []).slice(0, 1),
+    subject: uniq(rawTags.subject || []),
+    highlightTypes: uniq([...(hints.highlightTypes || []), ...(rawTags.highlightTypes || [])]),
+    scenarioTags: uniq([...(hints.scenarioTags || []), ...(rawTags.scenarioTags || [])]),
+    objectionTags: uniq([...(hints.objectionTags || []), ...(rawTags.objectionTags || [])]),
+  };
+
+  const fit = inferFitStatus(tags);
+  const caseFields = deriveCaseFields(split.descriptionCorpus || split.materialCorpus, groups, tags, fit);
+
+  const compliance = assessCompliance({
+    corpus: fullCorpus,
+    authorizationStatus: input.authorizationStatus,
+    hasMaterials: groups.some((g) => g.count > 0),
+    fields: caseFields,
+  });
+
+  // 字典自进化候选:文本里出现、字典里没有、但强线索命中的高价值短语
+  const learnedTags = [];
+  for (const dimension of ['highlightTypes', 'scenarioTags', 'objectionTags']) {
+    for (const value of tags[dimension] || []) {
+      learnedTags.push({ dimension, value });
+    }
+  }
+
+  return {
+    caseFields,
+    tags,
+    fit,
+    school,
+    compliance,
+    learned: { tags: learnedTags, aliases: school.hits.map((h) => ({ aliasText: h.aliasText, canonicalName: h.canonicalName, country: h.country })) },
+    texts: split,
+  };
+}
+
+export { TAG_CLUES, RISK_KEYWORDS, STRATEGY_MAP, PII_RULES };

+ 131 - 0
scripts/cloud-client.mjs

@@ -0,0 +1,131 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 云函数客户端:把案例包提交到 caseSubmit,落入待审清单。
+ *
+ * 契约(SHARED-CONTRACT.md 第 3 节):
+ *   入参 {sourceType, idempotencyKey, title, summary, materialAssets[], groupInfo?,
+ *         schoolCanonical?, tags[], highlightTypes[], scenarioTags[], objectionTags[],
+ *         riskFlags[], authorizationStatus, submittedBy?}
+ *   返回 {ok:true, result:{objectId, reviewStatus, readyForUse}}
+ *   落库 reviewStatus='pending'、readyForUse=false —— 新案例先进待审清单,不进公共素材库。
+ *
+ * 安全:
+ *   - 云函数基址来自项目 CLAUDE.md:https://server.lumistedu.com/api/functions
+ *   - session token 从环境变量 / 本地配置读取,**仅内存持有**,不写文件、不进日志、不进报告。
+ *   - tenantId 由服务端从会话上下文解析,客户端不传(传了也会被服务端拒绝不匹配)。
+ */
+
+const FUNCTIONS_URL = process.env.CASE_FUNCTIONS_URL
+  || 'https://server.lumistedu.com/api/functions';
+
+const TOKEN_SOURCES = [
+  { env: 'CASE_PARSE_SESSION_TOKEN', label: 'env:CASE_PARSE_SESSION_TOKEN' },
+  { env: 'PARSE_SESSION_TOKEN', label: 'env:PARSE_SESSION_TOKEN' },
+];
+
+function resolveToken(explicit) {
+  if (explicit) return { token: explicit, source: 'explicit' };
+  for (const item of TOKEN_SOURCES) {
+    const value = process.env[item.env];
+    if (value && value.trim()) return { token: value.trim(), source: item.label };
+  }
+  return { token: '', source: 'missing' };
+}
+
+export class CloudError extends Error {
+  constructor(message, code, requestId) {
+    super(message);
+    this.name = 'CloudError';
+    this.code = code || 'INTERNAL_ERROR';
+    this.requestId = requestId || '';
+  }
+}
+
+/**
+ * 调用云函数。
+ * @returns {Promise<{ok:true, result:object}|{ok:false, error:{code,message,requestId}}>}
+ *          网络不可达时抛 CloudError(调用方转成 dry-run 提示,不伪造成功)。
+ */
+export async function callFunction(name, params, options = {}) {
+  const { token, source } = resolveToken(options.token);
+  const headers = { 'Content-Type': 'application/json' };
+  if (token) headers['X-Parse-Session-Token'] = token;
+
+  const controller = new AbortController();
+  const timer = setTimeout(() => controller.abort(), options.timeoutMs || 20000);
+
+  let response;
+  try {
+    response = await fetch(FUNCTIONS_URL, {
+      method: 'POST',
+      headers,
+      body: JSON.stringify({ path: `/${name}`, params: params || {}, token }),
+      signal: controller.signal,
+    });
+  } catch (error) {
+    clearTimeout(timer);
+    const reason = error.name === 'AbortError' ? '请求超时' : error.message;
+    throw new CloudError(`云函数 ${name} 不可达:${reason}(token 来源:${source})`, 'NETWORK_UNAVAILABLE');
+  }
+  clearTimeout(timer);
+
+  let payload;
+  try {
+    payload = await response.json();
+  } catch {
+    throw new CloudError(`云函数 ${name} 返回不是 JSON(HTTP ${response.status})`, 'INTERNAL_ERROR');
+  }
+
+  if (!response.ok || payload.error) {
+    const error = payload.error || {};
+    return {
+      ok: false,
+      error: {
+        code: error.code || 'INTERNAL_ERROR',
+        message: error.message || `云函数 ${name} 调用失败(HTTP ${response.status})`,
+        requestId: error.requestId || '',
+      },
+    };
+  }
+  return { ok: true, result: payload.result || payload.data || {} };
+}
+
+/**
+ * 提交案例到待审清单。
+ * @param {object} casePackage 组装好的案例包(见 case-intake.mjs buildPayload)
+ */
+export async function submitCase(casePackage, options = {}) {
+  const payload = {
+    sourceType: casePackage.sourceType,
+    idempotencyKey: casePackage.idempotencyKey,
+    title: casePackage.title,
+    summary: casePackage.summary,
+    materialAssets: casePackage.materialAssets,
+    groupInfo: casePackage.groupInfo,
+    schoolCanonical: casePackage.schoolCanonical,
+    tags: casePackage.tags,
+    highlightTypes: casePackage.highlightTypes,
+    scenarioTags: casePackage.scenarioTags,
+    objectionTags: casePackage.objectionTags,
+    riskFlags: casePackage.riskFlags,
+    authorizationStatus: casePackage.authorizationStatus,
+  };
+  if (options.submittedBy) payload.submittedBy = options.submittedBy;
+
+  return callFunction('caseSubmit', payload, options);
+}
+
+/** 待审清单(只读,便于采集完立刻确认落库)。 */
+export async function pendingList(params = {}, options = {}) {
+  return callFunction('casePendingList', params, options);
+}

+ 285 - 0
scripts/grouping.mjs

@@ -0,0 +1,285 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 素材分组与阅读顺序
+ *
+ * 同事连发 6 张 / 9 张图,**不是多个素材,而是同一个素材**(朋友圈九宫格)。
+ * 这里用三种信号判定「是否属于同一组」,并给出显式 order:
+ *   ① 尺寸(宽高比)——九宫格同一批图通常同源同比例;
+ *   ② 时间(mtime / 收图时间)——同一批连发的时间戳接近;
+ *   ③ 内容相似度——OCR 文本 + 描述的词/字面重叠度。
+ * 任两种信号一致即归为一组;只有一种信号可用时按该信号判定,并在 reasons 里说明。
+ *
+ * 顺序:同组内按「九宫格阅读顺序」= 左上 → 右下 = 从左到右、从上到下。
+ * 由于原始排版位置通常已经丢失,用「收图时间升序 + 文件名自然序」近似还原发送顺序,
+ * 并显式写出 orderingRule,供人工在 materialAssets[] 上再调整(reorderAssets)。
+ */
+
+import { uniq, truncate } from './lib.mjs';
+
+export const DEFAULT_ORDERING_RULE = 'left-to-right,top-to-bottom';
+
+// ---------------------------------------------------------------------------
+// 相似度
+// ---------------------------------------------------------------------------
+
+function bigrams(text) {
+  const clean = String(text || '').replace(/\s+/g, '').toLowerCase();
+  if (clean.length < 2) return clean ? new Set([clean]) : new Set();
+  const out = new Set();
+  for (let i = 0; i < clean.length - 1; i++) out.add(clean.slice(i, i + 2));
+  return out;
+}
+
+/** 字符二元组 Jaccard 相似度(0–1)。中文、英文都适用,无需分词。 */
+export function textSimilarity(a, b) {
+  const setA = bigrams(a);
+  const setB = bigrams(b);
+  if (!setA.size || !setB.size) return 0;
+  let inter = 0;
+  for (const g of setA) if (setB.has(g)) inter += 1;
+  return inter / (setA.size + setB.size - inter);
+}
+
+/** 宽高比相似度(0–1):同源同排版的图接近 1。 */
+export function aspectSimilarity(a, b) {
+  const ratioA = Number(a && a.width) > 0 && Number(a && a.height) > 0 ? a.width / a.height : null;
+  const ratioB = Number(b && b.width) > 0 && Number(b && b.height) > 0 ? b.width / b.height : null;
+  if (ratioA === null || ratioB === null) return null;
+  const diff = Math.abs(ratioA - ratioB) / Math.max(ratioA, ratioB);
+  return Math.max(0, 1 - diff);
+}
+
+/** 时间邻近度(0–1):窗口内线性衰减。 */
+export function timeProximity(a, b, windowMs) {
+  const ta = Number(a && (a.mtimeMs ?? a.capturedAtMs));
+  const tb = Number(b && (b.mtimeMs ?? b.capturedAtMs));
+  if (!Number.isFinite(ta) || !Number.isFinite(tb)) return null;
+  const gap = Math.abs(ta - tb);
+  if (gap >= windowMs) return 0;
+  return 1 - gap / windowMs;
+}
+
+// ---------------------------------------------------------------------------
+// 排序:把「发送顺序」近似还原成九宫格阅读顺序
+// ---------------------------------------------------------------------------
+
+function naturalCompare(a, b) {
+  return String(a).localeCompare(String(b), 'zh-Hans-CN', { numeric: true, sensitivity: 'base' });
+}
+
+function sourceName(item) {
+  const p = item.localPath || item.url || '';
+  return p.split(/[\\/]/).pop() || p;
+}
+
+/**
+ * 同组内排序:时间优先(同一批连发的时间戳能还原发送顺序),
+ * 时间相同/缺失时回落文件名自然序。
+ */
+export function orderWithinGroup(items) {
+  return items
+    .map((item, index) => ({ item, index }))
+    .sort((a, b) => {
+      const ta = Number(a.item.mtimeMs ?? a.item.capturedAtMs);
+      const tb = Number(b.item.mtimeMs ?? b.item.capturedAtMs);
+      const hasTa = Number.isFinite(ta);
+      const hasTb = Number.isFinite(tb);
+      if (hasTa && hasTb && ta !== tb) return ta - tb;
+      const byName = naturalCompare(sourceName(a.item), sourceName(b.item));
+      if (byName !== 0) return byName;
+      return a.index - b.index;
+    })
+    .map((entry) => entry.item);
+}
+
+/**
+ * 九宫格排版推断。
+ * 9 → 3x3;4 → 2x2;6 → 3x2;2 → 2x1;其它按最接近方阵的因数对。
+ */
+export function inferLayout(count) {
+  const n = Math.max(1, Number(count) || 1);
+  const presets = { 1: [1, 1], 2: [2, 1], 3: [3, 1], 4: [2, 2], 6: [3, 2], 9: [3, 3], 12: [4, 3] };
+  if (presets[n]) return `${presets[n][0]}x${presets[n][1]}`;
+  for (let cols = Math.ceil(Math.sqrt(n)); cols <= n; cols++) {
+    if (n % cols === 0) return `${cols}x${n / cols}`;
+  }
+  return `${n}x1`;
+}
+
+// ---------------------------------------------------------------------------
+// 分组
+// ---------------------------------------------------------------------------
+
+const DEFAULTS = {
+  timeWindowMs: 15 * 60 * 1000, // 连发 15 分钟内算同一批
+  aspectThreshold: 0.88,        // 宽高比相似度阈值
+  textThreshold: 0.25,          // 文本相似度阈值(同批截图常有共同话术)
+};
+
+/**
+ * 把散图切成「素材组」。
+ *
+ * @param {Array<object>} items 每项:{localPath|url, kind, width, height, mtimeMs, ocrText, description}
+ * @param {object} [options] timeWindowMs / aspectThreshold / textThreshold / groupIdPrefix
+ * @returns {Array<{groupId, layout, count, orderingRule, kind, signals, reasons, items}>}
+ */
+export function groupMaterials(items, options = {}) {
+  const opts = { ...DEFAULTS, ...options };
+  const list = (items || []).filter(Boolean).map((item, index) => ({
+    kind: 'image',
+    ...item,
+    _index: index,
+  }));
+  if (!list.length) return [];
+
+  const groups = [];
+
+  for (const item of list) {
+    let best = null;
+    let bestScore = -1;
+    let bestReasons = [];
+
+    for (const group of groups) {
+      const { score, reasons } = groupAffinity(group, item, opts);
+      if (score > bestScore) {
+        best = group;
+        bestScore = score;
+        bestReasons = reasons;
+      }
+    }
+
+    if (best && bestScore >= 0.5) {
+      best.items.push(item);
+      best.reasons = uniq([...best.reasons, ...bestReasons]);
+      continue;
+    }
+
+    groups.push({
+      groupId: `${opts.groupIdPrefix || 'g'}${groups.length + 1}`,
+      items: [item],
+      reasons: ['seed'],
+      signals: [],
+    });
+  }
+
+  // 组内排序 + 排版 + 显式 order
+  return groups.map((group) => {
+    const ordered = orderWithinGroup(group.items);
+    const withOrder = ordered.map((item, index) => ({
+      ...item,
+      groupId: group.groupId,
+      order: index + 1,
+    }));
+    const kind = withOrder.every((i) => i.kind === 'video') ? 'video' : 'image';
+    const signals = uniq(withOrder.flatMap((i) => i._signals || []));
+    return {
+      groupId: group.groupId,
+      layout: inferLayout(withOrder.length),
+      count: withOrder.length,
+      orderingRule: DEFAULT_ORDERING_RULE,
+      kind,
+      signals,
+      reasons: group.reasons,
+      items: withOrder,
+      preview: withOrder.map((i) => truncate(i.ocrText || i.description || i.label || '', 40)).filter(Boolean).join(' / '),
+    };
+  });
+}
+
+/** 单张图与一个已有组的「亲缘度」:≥2 种信号一致 → 强;仅 1 种 → 弱。 */
+function groupAffinity(group, item, opts) {
+  const reasons = [];
+  let aspectHits = 0;
+  let aspectTotal = 0;
+  let textBest = 0;
+  let timeBest = 0;
+
+  for (const member of group.items) {
+    // 来自不同文档的内嵌图永不并组:它们是各自文档的配图,不是同一批发图。
+    // (同一文档的两张图 sourceKey 相同,仍可正常成组。)
+    if (item.sourceKey && member.sourceKey && item.sourceKey !== member.sourceKey) continue;
+    const aspect = aspectSimilarity(member, item);
+    if (aspect !== null) {
+      aspectTotal += 1;
+      if (aspect >= opts.aspectThreshold) aspectHits += 1;
+    }
+    const text = textSimilarity(member.ocrText || member.description || '', item.ocrText || item.description || '');
+    if (text > textBest) textBest = text;
+    const time = timeProximity(member, item, opts.timeWindowMs);
+    if (time !== null && time > timeBest) timeBest = time;
+  }
+
+  const aspectOk = aspectTotal > 0 && aspectHits / aspectTotal >= 0.6;
+  const textOk = textBest >= opts.textThreshold;
+  const timeOk = timeBest > 0;
+
+  const signals = [];
+  if (aspectOk) signals.push('aspect');
+  if (textOk) signals.push('text');
+  if (timeOk) signals.push('time');
+
+  if (aspectOk) reasons.push('aspect:同一排版比例');
+  if (textOk) reasons.push(`text:内容相似度 ${textBest.toFixed(2)}`);
+  if (timeOk) reasons.push(`time:${Math.round((1 - timeBest) * opts.timeWindowMs / 1000)}s 内连发`);
+
+  // 亲缘分:3 信号 → 0.9;2 信号 → 0.7;1 信号 → 0.4(不足以成组)
+  const score = signals.length >= 3 ? 0.9 : signals.length === 2 ? 0.7 : signals.length === 1 ? 0.4 : 0;
+  item._signals = uniq([...(item._signals || []), ...signals]);
+  return { score, reasons, signals };
+}
+
+/**
+ * 应用显式分组/顺序覆盖(人工或上游 manifest 指定)。
+ * override 形如:{ groups: [ { groupId, layout?, orderingRule?, items: [{ localPath|url, order }] } ] }
+ */
+export function applyGroupOverrides(groups, override) {
+  if (!override || !Array.isArray(override.groups) || !override.groups.length) return groups;
+  const result = [];
+  const consumed = new Set();
+
+  for (const spec of override.groups) {
+    const wanted = new Set((spec.items || []).map((i) => i.localPath || i.url).filter(Boolean));
+    const collected = [];
+    for (const group of groups) {
+      for (const item of group.items) {
+        const key = item.localPath || item.url;
+        if (wanted.has(key) && !consumed.has(key)) {
+          consumed.add(key);
+          const explicit = (spec.items || []).find((i) => (i.localPath || i.url) === key);
+          collected.push({ ...item, order: Number(explicit && explicit.order) || collected.length + 1 });
+        }
+      }
+    }
+    if (!collected.length) continue;
+    const ordered = collected.slice().sort((a, b) => a.order - b.order).map((item, index) => ({ ...item, order: index + 1 }));
+    result.push({
+      groupId: spec.groupId || `g${result.length + 1}`,
+      layout: spec.layout || inferLayout(ordered.length),
+      count: ordered.length,
+      orderingRule: spec.orderingRule || DEFAULT_ORDERING_RULE,
+      kind: ordered.every((i) => i.kind === 'video') ? 'video' : 'image',
+      signals: uniq(ordered.flatMap((i) => i._signals || [])),
+      reasons: ['manual-override'],
+      items: ordered,
+      preview: ordered.map((i) => truncate(i.ocrText || i.description || '', 40)).filter(Boolean).join(' / '),
+    });
+  }
+
+  // 没被覆盖到的组,原样保留
+  for (const group of groups) {
+    const rest = group.items.filter((item) => !consumed.has(item.localPath || item.url));
+    if (!rest.length) continue;
+    result.push({ ...group, items: rest.map((item, index) => ({ ...item, order: index + 1 })), count: rest.length });
+  }
+  return result;
+}

+ 131 - 0
scripts/image-size.mjs

@@ -0,0 +1,131 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 零依赖图片尺寸/类型探测(分组要用「尺寸」信号,不值得为此引第三方库)。
+ * 支持 PNG / JPEG / GIF / WebP(VP8/VP8L/VP8X)/ BMP。读不出返回 null。
+ */
+
+import fs from 'node:fs';
+
+const SIGS = [
+  { ext: 'png', mime: 'image/png', test: (b) => b.length > 8 && b.readUInt32BE(0) === 0x89504e47 },
+  { ext: 'jpg', mime: 'image/jpeg', test: (b) => b.length > 3 && b[0] === 0xff && b[1] === 0xd8 },
+  { ext: 'gif', mime: 'image/gif', test: (b) => b.length > 6 && b.slice(0, 3).toString('ascii') === 'GIF' },
+  { ext: 'webp', mime: 'image/webp', test: (b) => b.length > 12 && b.slice(0, 4).toString('ascii') === 'RIFF' && b.slice(8, 12).toString('ascii') === 'WEBP' },
+  { ext: 'bmp', mime: 'image/bmp', test: (b) => b.length > 2 && b[0] === 0x42 && b[1] === 0x4d },
+];
+
+function detectFormat(buffer) {
+  for (const sig of SIGS) {
+    try {
+      if (sig.test(buffer)) return sig;
+    } catch { /* 太短,继续 */ }
+  }
+  return null;
+}
+
+function pngSize(buffer) {
+  if (buffer.length < 24) return null;
+  return { width: buffer.readUInt32BE(16), height: buffer.readUInt32BE(20) };
+}
+
+function gifSize(buffer) {
+  if (buffer.length < 10) return null;
+  return { width: buffer.readUInt16LE(6), height: buffer.readUInt16LE(8) };
+}
+
+function bmpSize(buffer) {
+  if (buffer.length < 26) return null;
+  return { width: Math.abs(buffer.readInt32LE(18)), height: Math.abs(buffer.readInt32LE(22)) };
+}
+
+function jpegSize(buffer) {
+  let offset = 2;
+  while (offset + 9 < buffer.length) {
+    if (buffer[offset] !== 0xff) { offset += 1; continue; }
+    const marker = buffer[offset + 1];
+    // SOF0..SOF15,排除 DHT(c4) / JPG(c8) / DAC(cc)
+    if (marker >= 0xc0 && marker <= 0xcf && ![0xc4, 0xc8, 0xcc].includes(marker)) {
+      return { height: buffer.readUInt16BE(offset + 5), width: buffer.readUInt16BE(offset + 7) };
+    }
+    const segmentLength = buffer.readUInt16BE(offset + 2);
+    if (segmentLength <= 0) return null;
+    offset += 2 + segmentLength;
+  }
+  return null;
+}
+
+function webpSize(buffer) {
+  const chunk = buffer.slice(12, 16).toString('ascii');
+  if (chunk === 'VP8 ' && buffer.length > 30) {
+    return { width: buffer.readUInt16LE(26) & 0x3fff, height: buffer.readUInt16LE(28) & 0x3fff };
+  }
+  if (chunk === 'VP8L' && buffer.length > 25) {
+    const bits = buffer.readUInt32LE(21);
+    return { width: (bits & 0x3fff) + 1, height: ((bits >> 14) & 0x3fff) + 1 };
+  }
+  if (chunk === 'VP8X' && buffer.length > 30) {
+    const width = 1 + (buffer[24] | (buffer[25] << 8) | (buffer[26] << 16));
+    const height = 1 + (buffer[27] | (buffer[28] << 8) | (buffer[29] << 16));
+    return { width, height };
+  }
+  return null;
+}
+
+/**
+ * 读取图片信息。
+ * @returns {{ok:boolean, ext:string, mime:string, width:number|null, height:number|null, byteSize:number, error:string|null}}
+ */
+export function readImageInfo(filePath) {
+  const base = { ok: false, ext: '', mime: '', width: null, height: null, byteSize: 0, error: null };
+  let fd;
+  try {
+    const stat = fs.statSync(filePath);
+    base.byteSize = stat.size;
+    const head = Buffer.alloc(Math.min(64 * 1024, Math.max(16, stat.size)));
+    fd = fs.openSync(filePath, 'r');
+    fs.readSync(fd, head, 0, head.length, 0);
+
+    const format = detectFormat(head);
+    if (!format) return { ...base, error: 'UNKNOWN_IMAGE_FORMAT' };
+
+    let size = null;
+    if (format.ext === 'png') size = pngSize(head);
+    else if (format.ext === 'gif') size = gifSize(head);
+    else if (format.ext === 'bmp') size = bmpSize(head);
+    else if (format.ext === 'jpg') size = jpegSize(head);
+    else if (format.ext === 'webp') size = webpSize(head);
+
+    return {
+      ...base,
+      ok: true,
+      ext: format.ext,
+      mime: format.mime,
+      width: size ? size.width : null,
+      height: size ? size.height : null,
+      mtimeMs: stat.mtimeMs,
+      error: size ? null : 'SIZE_NOT_FOUND',
+    };
+  } catch (error) {
+    return { ...base, error: `READ_FAILED: ${error.message}` };
+  } finally {
+    if (fd !== undefined) {
+      try { fs.closeSync(fd); } catch { /* ignore */ }
+    }
+  }
+}
+
+export const IMAGE_EXTENSIONS = new Set(['.png', '.jpg', '.jpeg', '.gif', '.webp', '.bmp', '.tif', '.tiff', '.heic']);
+export const VIDEO_EXTENSIONS = new Set(['.mp4', '.mov', '.m4v', '.avi', '.webm', '.mkv', '.wmv']);
+export const DOC_EXTENSIONS = new Set(['.docx', '.docm']);
+export const DECK_EXTENSIONS = new Set(['.pptx', '.pptm']);
+export const TEXT_EXTENSIONS = new Set(['.md', '.txt', '.markdown', '.csv']);

+ 352 - 0
scripts/lib.mjs

@@ -0,0 +1,352 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * skill-case-get 共享工具层
+ *
+ * 零依赖。集中处理:
+ *   - 仓库内路径解析(技能可能被拷到 ~/.claude/skills/,也可能在仓库内直接跑)
+ *   - 标签字典读写(自进化)
+ *   - 幂等键生成(tenantId + idempotencyKey)
+ *   - 素材/案例字段规范化(与 SHARED-CONTRACT.md 逐字对齐)
+ *   - 绝对不做的事:不写任何真实 PII、不落 token、不碰 .env
+ *
+ * ⚠️ 本仓库禁止出现任何密钥/token/私网地址。token 只从用户环境读取,仅内存持有。
+ */
+
+import fs from 'node:fs';
+import path from 'node:path';
+import os from 'node:os';
+import crypto from 'node:crypto';
+import { fileURLToPath } from 'node:url';
+
+export const SCRIPT_DIR = path.dirname(fileURLToPath(import.meta.url));
+export const SKILL_ROOT = path.resolve(SCRIPT_DIR, '..');
+
+export const TENANT_DEFAULT = process.env.CASE_TENANT_ID || 'lumi-demo';
+
+// ---------------------------------------------------------------------------
+// 文件与 JSON
+// ---------------------------------------------------------------------------
+
+export function readJson(filePath, fallback = null) {
+  try {
+    if (!filePath || !fs.existsSync(filePath)) return fallback;
+    // 手写的 json 可能带 UTF-8 BOM
+    const raw = fs.readFileSync(filePath, 'utf8').replace(/^/, '');
+    return JSON.parse(raw);
+  } catch {
+    return fallback;
+  }
+}
+
+export function writeJson(filePath, value) {
+  fs.mkdirSync(path.dirname(filePath), { recursive: true });
+  fs.writeFileSync(filePath, `${JSON.stringify(value, null, 2)}\n`, 'utf8');
+}
+
+export function ensureDir(dir) {
+  fs.mkdirSync(dir, { recursive: true });
+  return dir;
+}
+
+export function sha1(value) {
+  return crypto.createHash('sha1').update(String(value ?? '')).digest('hex');
+}
+
+export function shortHash(value, length = 10) {
+  return sha1(value).slice(0, length);
+}
+
+// ---------------------------------------------------------------------------
+// 兄弟技能脚本解析(skill-vision / skill-listen)
+// ---------------------------------------------------------------------------
+
+function candidateRoots(envVar) {
+  const roots = [];
+  if (envVar && process.env[envVar]) roots.push(process.env[envVar]);
+  roots.push(path.join(os.homedir(), '.claude', 'skills'));
+  roots.push(path.join(process.cwd(), '.claude', 'skills'));
+  roots.push(path.join(SKILL_ROOT, '..'));
+  return roots;
+}
+
+/**
+ * 定位兄弟技能里的脚本。找不到返回 null(调用方给可执行的降级提示,不抛栈)。
+ * @param {string} envVar 允许用环境变量覆盖的变量名
+ * @param {string} relativePath 相对 skills 根目录的路径,如 skill-vision/scripts/vision-client.mjs
+ */
+export function resolveSiblingScript(envVar, relativePath) {
+  for (const root of candidateRoots(envVar)) {
+    const full = path.join(root, relativePath);
+    if (fs.existsSync(full)) return full;
+  }
+  return null;
+}
+
+export function resolvePython() {
+  return process.env.CASE_PYTHON || 'python3';
+}
+
+// ---------------------------------------------------------------------------
+// 标签字典
+// ---------------------------------------------------------------------------
+
+export function dictionaryPath(custom) {
+  if (custom) return path.resolve(custom);
+  return path.join(SKILL_ROOT, 'references', 'tag-dictionary.json');
+}
+
+export function loadDictionary(custom) {
+  const file = dictionaryPath(custom);
+  const dict = readJson(file, null);
+  if (!dict || !dict.dimensions) {
+    throw new Error(`标签字典不可用或结构不对:${file}`);
+  }
+  if (!Array.isArray(dict.__learned)) dict.__learned = [];
+  dict.__file = file;
+  return dict;
+}
+
+/** 取某维度的候选取值(数组;schoolAlias 这类结构特殊,单独处理)。 */
+export function dimensionValues(dict, dimension) {
+  const dim = dict && dict.dimensions ? dict.dimensions[dimension] : null;
+  if (!dim) return [];
+  if (Array.isArray(dim.values)) return dim.values;
+  return [];
+}
+
+/**
+ * 字典自进化:把采集过程中遇到的新取值追加进字典。
+ * 只追加、不改写既有条目;带 source/learnedAt 便于追溯。
+ *
+ * @returns {{added: Array<{dimension:string, value:string}>, skipped: Array<{dimension:string,value:string,reason:string}>}}
+ */
+export function learnTags(dict, incoming) {
+  const added = [];
+  const skipped = [];
+  const now = new Date().toISOString();
+
+  for (const item of incoming || []) {
+    const dimension = String(item && item.dimension || '').trim();
+    const value = String(item && item.value || '').trim();
+    if (!dimension || !value) {
+      skipped.push({ dimension, value, reason: 'empty' });
+      continue;
+    }
+    const dim = dict.dimensions ? dict.dimensions[dimension] : null;
+    if (!dim) {
+      skipped.push({ dimension, value, reason: 'unknown-dimension' });
+      continue;
+    }
+    if (!Array.isArray(dim.values)) {
+      // 结构型维度(schoolAlias)走 addAlias
+      skipped.push({ dimension, value, reason: 'structured-dimension' });
+      continue;
+    }
+    if (dim.values.includes(value)) {
+      skipped.push({ dimension, value, reason: 'exists' });
+      continue;
+    }
+    dim.values.push(value);
+    added.push({ dimension, value });
+  }
+
+  if (added.length) {
+    dict.updatedAt = now;
+    dict.__learned = Array.isArray(dict.__learned) ? dict.__learned : [];
+    for (const entry of added) {
+      dict.__learned.push({ ...entry, source: 'learned', learnedAt: now });
+    }
+  }
+  return { added, skipped };
+}
+
+/** 学校别名自进化:追加到 schoolAlias.entries。 */
+export function learnAliases(dict, entries) {
+  const added = [];
+  const dim = dict.dimensions && dict.dimensions.schoolAlias;
+  if (!dim || !Array.isArray(dim.entries)) return { added };
+  const existing = new Set(dim.entries.map((e) => String(e.aliasText || '').toLowerCase()));
+  const now = new Date().toISOString();
+  for (const item of entries || []) {
+    const aliasText = String(item && item.aliasText || '').trim();
+    const canonicalName = String(item && item.canonicalName || '').trim();
+    if (!aliasText || !canonicalName) continue;
+    if (existing.has(aliasText.toLowerCase())) continue;
+    const entry = { aliasText, canonicalName, country: item.country || '', source: 'learned', learnedAt: now };
+    dim.entries.push(entry);
+    existing.add(aliasText.toLowerCase());
+    added.push(entry);
+  }
+  if (added.length) {
+    dict.updatedAt = now;
+    dict.__learned = Array.isArray(dict.__learned) ? dict.__learned : [];
+    for (const entry of added) dict.__learned.push({ dimension: 'schoolAlias', value: entry.aliasText, source: 'learned', learnedAt: now });
+  }
+  return { added };
+}
+
+export function saveDictionary(dict) {
+  const { __file, __learned, ...clean } = dict;
+  writeJson(dict.__file || dictionaryPath(), clean);
+  return dict.__file || dictionaryPath();
+}
+
+// ---------------------------------------------------------------------------
+// 幂等
+// ---------------------------------------------------------------------------
+
+/**
+ * 生成幂等键。同一批素材重复采集(同 sourceRef + 同标题 + 同素材序列)必须得到同一个键,
+ * 保证 `tenantId + idempotencyKey` 去重、重复提交只产生一条。
+ */
+export function buildIdempotencyKey({ tenantId, sourceType, sourceRef, title, materialAssets, explicit }) {
+  if (explicit) return String(explicit);
+  const assetFingerprint = (materialAssets || [])
+    .map((a) => `${a.order || ''}:${a.kind || ''}:${a.url || a.localPath || ''}`)
+    .join('|');
+  const seed = [tenantId || TENANT_DEFAULT, sourceType || '', sourceRef || '', title || '', assetFingerprint].join('::');
+  return `caseget-${shortHash(seed, 16)}`;
+}
+
+/**
+ * 为本地素材推导可访问 URL。
+ *
+ * 技能本机不存储/上传二进制(不新建业务服务器),因此:
+ *   - 调用了 `--asset-base-url <base>`(已有 CDN / Parse Files 前缀)→ 用 base + 文件名派生 URL;
+ *   - 没提供 → 返回空串,并在调用处给出警告,由归档方补齐 URL 后再入库。
+ * 绝不伪造 URL,也不把本地绝对路径当 URL 写进案例包。
+ */
+export function deriveAssetUrl(localPath, baseUrl) {
+  if (!baseUrl) return '';
+  const name = String(localPath || '').split(/[\\/]/).pop() || '';
+  if (!name) return '';
+  // 路径片段保留、文件名做 URL 安全化,避免空格 / 中文导致的坏链接
+  const safe = encodeURIComponent(name).replace(/%2F/g, '_');
+  return `${String(baseUrl).replace(/\/+$/, '')}/${safe}`;
+}
+
+// ---------------------------------------------------------------------------
+// 素材/案例字段规范化(逐字对齐 SHARED-CONTRACT.md)
+// ---------------------------------------------------------------------------
+
+export const MATERIAL_KINDS = new Set(['image', 'video']);
+export const MATERIAL_ROLES = new Set(['material', 'description']);
+
+/**
+ * 规范化单个素材项。order 1 起、同组内按九宫格阅读顺序。
+ * @returns {{asset: object|null, problems: string[]}}
+ */
+export function normalizeMaterialAsset(raw, index) {
+  const problems = [];
+  const item = raw || {};
+  const kind = String(item.kind || 'image').toLowerCase();
+  if (!MATERIAL_KINDS.has(kind)) problems.push(`素材 #${index + 1} 的 kind 非法:${item.kind}`);
+
+  const order = Number.isFinite(Number(item.order)) && Number(item.order) > 0
+    ? Math.trunc(Number(item.order))
+    : index + 1;
+
+  const url = item.url || item.imageUrl || item.videoUrl || '';
+  const localPath = item.localPath || item.path || '';
+  if (!url && !localPath) problems.push(`素材 #${index + 1} 既没有 url 也没有 localPath`);
+
+  const role = String(item.role || 'material').toLowerCase();
+  if (!MATERIAL_ROLES.has(role)) problems.push(`素材 #${index + 1} 的 role 非法:${item.role}`);
+
+  const asset = {
+    order,
+    groupId: item.groupId || '',
+    kind,
+    url: String(url || ''),
+    localPath: String(localPath || ''),
+    label: String(item.label || ''),
+    ocrText: String(item.ocrText || ''),
+    description: String(item.description || ''),
+    usageSuggestion: String(item.usageSuggestion || ''),
+    role,
+  };
+  return { asset, problems };
+}
+
+/** 按 order 升序排序(稳定;同 order 保持原相对顺序)。 */
+export function sortAssets(assets) {
+  return (assets || [])
+    .map((asset, index) => ({ asset, index }))
+    .sort((a, b) => (a.asset.order - b.asset.order) || (a.index - b.index))
+    .map((entry) => entry.asset);
+}
+
+/**
+ * 手动调整顺序:把指定素材移动到目标位置,其余顺次重排(order 从 1 连续)。
+ * 支持「数组 + 顺序调整」的显式要求。
+ */
+export function reorderAssets(assets, from, to) {
+  const list = sortAssets(assets);
+  const start = Math.trunc(from) - 1;
+  const end = Math.trunc(to) - 1;
+  if (start < 0 || start >= list.length || end < 0 || end >= list.length) {
+    throw new Error(`reorder 越界:from=${from} to=${to},共 ${list.length} 项`);
+  }
+  const [moved] = list.splice(start, 1);
+  list.splice(end, 0, moved);
+  return list.map((asset, index) => ({ ...asset, order: index + 1 }));
+}
+
+/** 汇总兼容字段:imageUrls = kind='image' 的 url,按 order 排序。 */
+export function imageUrlsFrom(assets) {
+  return sortAssets(assets)
+    .filter((a) => a.kind === 'image' && a.url)
+    .map((a) => a.url);
+}
+
+// ---------------------------------------------------------------------------
+// 输出目录
+// ---------------------------------------------------------------------------
+
+export function outputsDir(custom) {
+  return path.resolve(custom || process.env.CASE_OUTPUT_DIR || path.join(SKILL_ROOT, 'outputs'));
+}
+
+export function newRunId() {
+  const now = new Date();
+  const stamp = now.toISOString().replace(/[-:T]/g, '').slice(0, 14);
+  return `${stamp}-${shortHash(`${now.getTime()}-${Math.random()}`, 6)}`;
+}
+
+// ---------------------------------------------------------------------------
+// 通用小工具
+// ---------------------------------------------------------------------------
+
+/**
+ * 去重并丢弃空值。同时接受多个数组(或嵌套数组):
+ *   uniq(['a']), uniq(['a'], ['b']) 都成立,避免调用处到处 spread。
+ */
+export function uniq(...args) {
+  return [...new Set(
+    args
+      .flat(Infinity)
+      .filter((item) => item !== undefined && item !== null && item !== ''),
+  )];
+}
+
+export function firstText(...values) {
+  for (const value of values) {
+    if (typeof value === 'string' && value.trim()) return value.trim();
+  }
+  return '';
+}
+
+export function truncate(value, max = 160) {
+  const text = String(value ?? '').replace(/\s+/g, ' ').trim();
+  return text.length > max ? `${text.slice(0, max)}…` : text;
+}

+ 251 - 0
scripts/listen-bridge.mjs

@@ -0,0 +1,251 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 视频桥:抽帧(ffmpeg)+ 音频转写(复用兄弟技能 skill-listen)
+ *
+ * 产出:
+ *   - frames[]      按时间轴均匀抽出的关键帧(落地到 --frames-out,交给 vision 做理解)
+ *   - transcript    音轨转写文本(skill-listen / listen-runner.mjs,走 Fmode 网关,凭据仅服务端)
+ *   - durationMs    时长(ffprobe;探测不到则 null)
+ *
+ * 缺 ffmpeg / ffprobe 或 skill-listen 时,返回 available:false + 明确安装提示,绝不伪造结果。
+ * 绝不写任何 token / 凭据。
+ */
+
+import fs from 'node:fs';
+import path from 'node:path';
+import { spawnSync } from 'node:child_process';
+import { parseArgs } from 'node:util';
+import { ensureDir, resolveSiblingScript, truncate } from './lib.mjs';
+
+const LISTEN_REL = path.join('skill-listen', 'scripts', 'listen-runner.mjs');
+
+function hasBinary(name) {
+  const probe = spawnSync(name, ['-version'], { encoding: 'utf8' });
+  return probe.status === 0 || Boolean(probe.stdout);
+}
+
+export function resolveFfmpeg() {
+  return process.env.CASE_FFMPEG || 'ffmpeg';
+}
+
+export function probeDurationMs(videoPath) {
+  const probe = spawnSync(process.env.CASE_FFPROBE || 'ffprobe', [
+    '-v', 'error',
+    '-show_entries', 'format=duration',
+    '-of', 'default=noprint_wrappers=1:nokey=1',
+    videoPath,
+  ], { encoding: 'utf8' });
+  const seconds = Number.parseFloat(String(probe.stdout || '').trim());
+  return Number.isFinite(seconds) && seconds > 0 ? Math.round(seconds * 1000) : null;
+}
+
+/**
+ * 均匀抽帧。
+ * @returns {{ok:boolean, frames:object[], error:string|null}}
+ */
+export function extractFrames(videoPath, framesOut, options = {}) {
+  const count = Math.max(1, Math.min(24, Number(options.frameCount) || 6));
+  ensureDir(framesOut);
+  const durationMs = options.durationMs || probeDurationMs(videoPath);
+  const ffmpeg = options.ffmpeg || resolveFfmpeg();
+
+  const frames = [];
+  if (durationMs && durationMs > 0) {
+    // 均匀取 count 个时间点
+    for (let i = 0; i < count; i++) {
+      const tMs = Math.round((durationMs * (i + 0.5)) / count);
+      const target = path.join(framesOut, `frame-${String(i + 1).padStart(3, '0')}.jpg`);
+      const run = spawnSync(ffmpeg, [
+        '-y', '-loglevel', 'error',
+        '-ss', (tMs / 1000).toFixed(3),
+        '-i', videoPath,
+        '-frames:v', '1', '-q:v', '3',
+        target,
+      ], { encoding: 'utf8' });
+      if (run.status === 0 && fs.existsSync(target)) {
+        frames.push({ index: i + 1, atMs: tMs, localPath: target });
+      }
+    }
+  } else {
+    // 拿不到时长:按 fps 抽前 count 帧
+    const run = spawnSync(ffmpeg, [
+      '-y', '-loglevel', 'error',
+      '-i', videoPath,
+      '-frames:v', String(count), '-vf', 'fps=1', '-q:v', '3',
+      path.join(framesOut, 'frame-%03d.jpg'),
+    ], { encoding: 'utf8' });
+    if (run.status === 0) {
+      for (const name of fs.readdirSync(framesOut).sort().slice(0, count)) {
+        frames.push({ index: frames.length + 1, atMs: null, localPath: path.join(framesOut, name) });
+      }
+    }
+  }
+  if (!frames.length) {
+    return { ok: false, frames: [], error: 'FFMPEG_FRAME_EXTRACT_FAILED(检查 ffmpeg 是否可用及视频是否可解码)' };
+  }
+  return { ok: true, frames, error: null };
+}
+
+/** 抽出 16kHz 单声道 wav 音轨(转写前压体积,更稳)。 */
+export function extractAudio(videoPath, audioOut, options = {}) {
+  ensureDir(path.dirname(audioOut));
+  const ffmpeg = options.ffmpeg || resolveFfmpeg();
+  const run = spawnSync(ffmpeg, [
+    '-y', '-loglevel', 'error',
+    '-i', videoPath,
+    '-vn', '-ar', '16000', '-ac', '1',
+    '-f', 'wav', audioOut,
+  ], { encoding: 'utf8' });
+  if (run.status === 0 && fs.existsSync(audioOut) && fs.statSync(audioOut).size > 1024) {
+    return { ok: true, audioPath: audioOut, error: null };
+  }
+  return { ok: false, audioPath: '', error: 'FFMPEG_AUDIO_EXTRACT_FAILED(视频可能没有音轨)' };
+}
+
+/** 调 skill-listen 转写音轨。缺技能时返回 available:false,不抛栈。 */
+export async function transcribeAudio(audioPath, options = {}) {
+  const listenPath = resolveSiblingScript('CASE_LISTEN_SCRIPT', LISTEN_REL);
+  if (!listenPath) {
+    return {
+      available: false,
+      text: '',
+      segments: [],
+      error: `未找到 skill-listen(期望 ${LISTEN_REL})。请先安装:npx skill-listen@latest install`,
+    };
+  }
+  try {
+    const runner = await import(new URL(`file://${listenPath.split(path.sep).join('/')}`).href);
+    const durationMs = probeDurationMs(audioPath);
+    const out = await runner.transcribeFile({
+      filePath: audioPath,
+      language: options.language || 'autodialect',
+      diarize: Boolean(options.diarize),
+      speakers: options.speakers,
+      durationMs: durationMs || undefined,
+    });
+    return {
+      available: true,
+      text: String(out.text || ''),
+      segments: out.segments || [],
+      error: null,
+    };
+  } catch (error) {
+    return { available: true, text: '', segments: [], error: `TRANSCRIBE_FAILED: ${error.message}` };
+  }
+}
+
+/**
+ * 视频处理总入口:抽帧 → 抽音轨 → 转写。
+ */
+export async function processVideo(videoPath, options = {}) {
+  const workDir = ensureDir(options.workDir || path.join(process.cwd(), 'case-get-outputs', 'video'));
+  const durationMs = probeDurationMs(videoPath);
+  const ffmpegOk = hasBinary(options.ffmpeg || resolveFfmpeg());
+
+  const report = {
+    video: path.resolve(videoPath),
+    durationMs,
+    frameCount: 0,
+    frames: [],
+    transcription: { available: false, text: '', segments: [], error: null },
+    warnings: [],
+  };
+
+  if (!ffmpegOk) {
+    report.warnings.push('未找到 ffmpeg:无法抽帧/抽音轨。请安装 ffmpeg 后重试。');
+    return report;
+  }
+
+  const frameResult = extractFrames(videoPath, path.join(workDir, 'frames'), { frameCount: options.frameCount, durationMs });
+  report.frames = frameResult.frames;
+  report.frameCount = frameResult.frames.length;
+  if (frameResult.error) report.warnings.push(frameResult.error);
+
+  const audio = extractAudio(videoPath, path.join(workDir, 'audio.wav'));
+  if (audio.ok) {
+    report.transcription = await transcribeAudio(audio.audioPath, options);
+    if (report.transcription.error) report.warnings.push(report.transcription.error);
+  } else {
+    report.transcription = { available: false, text: '', segments: [], error: audio.error };
+    report.warnings.push(audio.error);
+  }
+
+  report.summary = truncate(report.transcription.text || '', 200);
+  return report;
+}
+
+// ---------------------------------------------------------------------------
+// CLI
+// ---------------------------------------------------------------------------
+
+async function main() {
+  const { values, positionals } = parseArgs({
+    options: {
+      video: { type: 'string', multiple: true, default: [] },
+      out: { type: 'string' },
+      'work-dir': { type: 'string' },
+      'frame-count': { type: 'string' },
+      diarize: { type: 'boolean', default: false },
+      speakers: { type: 'string' },
+      language: { type: 'string' },
+      help: { type: 'boolean', default: false },
+    },
+    allowPositionals: true,
+  });
+
+  if (values.help) {
+    process.stdout.write([
+      'skill-case-get / listen-bridge — 视频抽帧 + 音轨转写(复用 skill-listen)',
+      '',
+      '  node listen-bridge.mjs --video a.mp4 [--frame-count 6] [--diarize] [--out video.json]',
+      '',
+      '依赖:ffmpeg / ffprobe(抽帧、抽音轨)、skill-listen(转写)。缺任一都会在 warnings 里说明,不伪造结果。',
+    ].join('\n'));
+    return 0;
+  }
+
+  const inputs = [...values.video, ...positionals];
+  if (!inputs.length) {
+    process.stderr.write('至少需要一个 --video <file>\n');
+    return 2;
+  }
+
+  const reports = [];
+  for (const video of inputs) {
+    reports.push(await processVideo(video, {
+      workDir: values['work-dir'] || path.join(path.dirname(path.resolve(video)), 'case-get-video'),
+      frameCount: Number(values['frame-count']) || 6,
+      diarize: values.diarize,
+      speakers: values.speakers ? Number(values.speakers) : undefined,
+      language: values.language,
+    }));
+  }
+
+  const payload = { ok: reports.every((r) => !r.warnings.length), count: reports.length, reports };
+  const text = JSON.stringify(payload, null, 2);
+  if (values.out) {
+    fs.mkdirSync(path.dirname(path.resolve(values.out)), { recursive: true });
+    fs.writeFileSync(path.resolve(values.out), `${text}\n`, 'utf8');
+  } else {
+    process.stdout.write(`${text}\n`);
+  }
+  return payload.ok ? 0 : 1;
+}
+
+const invokedDirectly = process.argv[1] && path.resolve(process.argv[1]).endsWith(path.join('scripts', 'listen-bridge.mjs'));
+if (invokedDirectly) {
+  main().then((code) => process.exit(code)).catch((error) => {
+    process.stderr.write(`listen-bridge 失败:${error.message}\n`);
+    process.exit(1);
+  });
+}

+ 665 - 0
scripts/office_extract.py

@@ -0,0 +1,665 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""skill-case-get —— .docx / .pptx 解包抽取器(纯标准库)
+
+本机没有 python-docx / python-pptx,Node 也没有对应库;
+.docx / .pptx 本质是 zip,因此用标准库 `zipfile` + `xml.etree.ElementTree` 解包解析。
+
+抽取内容:
+  - 文字段落(按文档/幻灯片顺序,保持阅读序)
+  - 表格文本(.docx 的 w:tbl,逐行拼接)
+  - 内嵌媒体(word/media/*、ppt/media/*、xl/media/*)落地到 --media-out
+  - 超链接(document.xml 的 w:hyperlink r:id → rels 目标;slide 的 a:hlinkClick)
+  - .pptx 的幻灯片切换顺序(presentation.xml 的 sldIdLst + rels)
+  - .pptx 图解/备注:p:sp 的 a:t 文本(标题、正文、备注页)
+  - .docx 页眉页脚(可选,--no-headers 关闭)
+
+用法:
+  python3 office_extract.py --input <file.docx|file.pptx> [--media-out <dir>] [--out <result.json>]
+  python3 office_extract.py --input a.docx --input b.pptx --out bundle.json
+  python3 office_extract.py --selftest
+
+输出:一份 JSON(stdout 或 --out),结构见 `result` 组装处。
+不写日志、不落任何 PII 到 stdout 之外的位置;媒体文件按原字节落地到 media-out。
+"""
+
+import argparse
+import json
+import os
+import re
+import sys
+import zipfile
+import xml.etree.ElementTree as ET
+
+NS = {
+    # WordprocessingML
+    "w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main",
+    "r": "http://schemas.openxmlformats.org/officeDocument/2006/relationships",
+    "wp": "http://schemas.openxmlformats.org/drawingml/2006/wordprocessingDrawing",
+    "a": "http://schemas.openxmlformats.org/drawingml/2006/main",
+    # PresentationML
+    "p": "http://schemas.openxmlformats.org/presentationml/2006/main",
+    # package relationships
+    "pr": "http://schemas.openxmlformats.org/package/2006/relationships",
+    "ct": "http://schemas.openxmlformats.org/package/2006/content-types",
+}
+
+DOCX_EXT = {".docx", ".docm"}
+PPTX_EXT = {".pptx", ".pptm"}
+MEDIA_PREFIXES = ("word/media/", "ppt/media/", "xl/media/", "media/")
+IMAGE_EXTS = {".png", ".jpg", ".jpeg", ".gif", ".bmp", ".webp", ".tiff", ".emf", ".wmf"}
+VIDEO_EXTS = {".mp4", ".mov", ".m4v", ".avi", ".webm", ".mkv", ".wmv"}
+AUDIO_EXTS = {".mp3", ".m4a", ".wav", ".aac", ".flac", ".ogg", ".wma"}
+
+REL_NS = "{http://schemas.openxmlformats.org/officeDocument/2006/relationships}"
+PR_NS = "{http://schemas.openxmlformats.org/package/2006/relationships}"
+
+
+def qname(prefix, local):
+    return "{%s}%s" % (NS[prefix], local)
+
+
+def safe_relpath(name):
+    """把 zip 内条目名规整成安全的相对路径(防 zip slip)。
+
+    正确处理 `..`:向上弹一层(而不是直接丢弃),这样
+    `ppt/slides/../media/a.png` → `ppt/media/a.png`;
+    越出根的 `..` 一律忽略,路径不可能逃出包外。
+    """
+    clean = name.replace("\\", "/").lstrip("/")
+    parts = []
+    for segment in clean.split("/"):
+        if not segment or segment == ".":
+            continue
+        if segment == "..":
+            if parts:
+                parts.pop()
+            continue
+        parts.append(segment)
+    return "/".join(parts)
+
+
+def local_name(tag):
+    return tag.split("}", 1)[1] if "}" in tag else tag
+
+
+# ---------------------------------------------------------------------------
+# relationships
+# ---------------------------------------------------------------------------
+
+def parse_rels(zf, rels_path):
+    """解析 _rels 部件 → {rId: {target, type, mode}}。"""
+    out = {}
+    try:
+        raw = zf.read(rels_path)
+    except (KeyError, OSError):
+        return out
+    try:
+        root = ET.fromstring(raw)
+    except ET.ParseError:
+        return out
+    for rel in root:
+        rid = rel.get("Id") or ""
+        target = rel.get("Target") or ""
+        if not rid:
+            continue
+        out[rid] = {
+            "id": rid,
+            "target": target,
+            "type": (rel.get("Type") or "").rsplit("/", 1)[-1],
+            "mode": rel.get("TargetMode") or "Internal",
+        }
+    return out
+
+
+def rel_target_to_part(rels_path, target):
+    """把相对 Target 规整成 zip 内条目路径。"""
+    if target.startswith("/"):
+        return safe_relpath(target)
+    base = posix_dirname(posix_dirname(rels_path))  # _rels 的上一级
+    return safe_relpath(posix_join(base, target))
+
+
+def posix_dirname(p):
+    return p.rsplit("/", 1)[0] if "/" in p else ""
+
+
+def posix_join(a, b):
+    if not a:
+        return safe_relpath(b)
+    return safe_relpath(a + "/" + b)
+
+
+# ---------------------------------------------------------------------------
+# text helpers
+# ---------------------------------------------------------------------------
+
+def collect_text_runs(element, run_tag, text_tag):
+    """按阅读序收集某个节点下的全部文本片段。"""
+    parts = []
+    for node in element.iter():
+        if node.tag == run_tag or local_name(node.tag) == local_name(text_tag):
+            if node.tag == text_tag and node.text:
+                parts.append(node.text)
+    return "".join(parts)
+
+
+def paragraph_text(p_node):
+    """一个 w:p / a:p 的文本。"""
+    parts = []
+    for node in p_node.iter():
+        ln = local_name(node.tag)
+        if ln == "t" and node.text:
+            parts.append(node.text)
+        elif ln == "tab":
+            parts.append("\t")
+        elif ln in ("br", "cr"):
+            parts.append("\n")
+    return "".join(parts).strip()
+
+
+def is_heading(p_node):
+    """w:p 是否带 Heading 样式。"""
+    for node in p_node.iter(qname("w", "pStyle")):
+        val = node.get(qname("w", "val")) or ""
+        if val.lower().startswith("heading") or val in ("Title", "标题"):
+            return True
+    return False
+
+
+# ---------------------------------------------------------------------------
+# docx
+# ---------------------------------------------------------------------------
+
+def extract_docx(zf, media_out, include_headers=True):
+    result = {
+        "kind": "document",
+        "format": "docx",
+        "paragraphs": [],
+        "outline": [],
+        "tables": [],
+        "hyperlinks": [],
+        "headings": [],
+        "media": [],
+        "embeddedObjects": [],
+        "notes": [],
+    }
+
+    doc_part = "word/document.xml"
+    if doc_part not in zf.namelist():
+        for name in zf.namelist():
+            if name.startswith("word/") and name.endswith("document.xml"):
+                doc_part = name
+                break
+    if doc_part not in zf.namelist():
+        result["error"] = "DOCX_MISSING_DOCUMENT_XML"
+        return result
+
+    rels = parse_rels(zf, "word/_rels/document.xml.rels")
+    root = ET.fromstring(zf.read(doc_part))
+    body = root.find(qname("w", "body"))
+    if body is None:
+        body = root
+
+    def walk_blocks(container, scope):
+        for child in container:
+            ln = local_name(child.tag)
+            if ln == "p":
+                text = paragraph_text(child)
+                if text:
+                    entry = {"index": len(result["paragraphs"]), "scope": scope, "text": text}
+                    if is_heading(child):
+                        entry["heading"] = True
+                        result["headings"].append(text)
+                    result["paragraphs"].append(entry)
+                    result["outline"].append({"level": 1 if entry.get("heading") else 0, "text": text})
+                # 段落里的图片(drawing / pict)
+                for blip in child.iter():
+                    if local_name(blip.tag) == "blip":
+                        rid = blip.get(REL_NS + "embed") or blip.get(REL_NS + "link")
+                        if rid and rid in rels:
+                            result["media"].append({
+                                "rId": rid,
+                                "part": rel_target_to_part("word/_rels/document.xml.rels", rels[rid]["target"]),
+                                "anchor": "inline",
+                                "nearbyText": text[:80],
+                            })
+            elif ln == "tbl":
+                rows = []
+                for tr in child.findall(qname("w", "tr")):
+                    cells = []
+                    for tc in tr.findall(qname("w", "tc")):
+                        cells.append(" ".join(paragraph_text(p) for p in tc.findall(qname("w", "p"))).strip())
+                    rows.append(cells)
+                if rows:
+                    result["tables"].append({"scope": scope, "rows": rows, "rowCount": len(rows)})
+                walk_blocks(child, scope)
+            elif ln in ("sdt", "sdtContent", "tc"):
+                walk_blocks(child, scope)
+
+    walk_blocks(body, "body")
+
+    # 超链接:w:hyperlink r:id → rels
+    for link in root.iter(qname("w", "hyperlink")):
+        rid = link.get(REL_NS + "id") or link.get(qname("r", "id")) or ""
+        target = rels.get(rid, {}).get("target", "")
+        text = paragraph_text(link)
+        if target or text:
+            result["hyperlinks"].append({"text": text, "target": target, "scope": "body"})
+
+    # 内嵌 OLE / 附件对象
+    for obj in root.iter(qname("w", "object")):
+        for node in obj.iter():
+            if local_name(node.tag) == "OLEObject":
+                rid = node.get(REL_NS + "id") or node.get(qname("r", "id")) or ""
+                if rid in rels:
+                    result["embeddedObjects"].append({"rId": rid, "part": rel_target_to_part("word/_rels/document.xml.rels", rels[rid]["target"]), "progId": node.get("ProgID") or ""})
+
+    if include_headers:
+        for name in sorted(zf.namelist()):
+            if re.match(r"^word/(header|footer)\d*\.xml$", name):
+                try:
+                    hroot = ET.fromstring(zf.read(name))
+                except ET.ParseError:
+                    continue
+                texts = [paragraph_text(p) for p in hroot.iter(qname("w", "p"))]
+                texts = [t for t in texts if t]
+                if texts:
+                    result["notes"].append({"scope": name, "text": "\n".join(texts)})
+
+    result["media"] = attach_media(zf, result["media"], media_out)
+    return result
+
+
+# ---------------------------------------------------------------------------
+# pptx
+# ---------------------------------------------------------------------------
+
+def slide_sort_key(name):
+    m = re.search(r"slide(\d+)\.xml$", name)
+    return int(m.group(1)) if m else 10 ** 6
+
+
+def extract_pptx(zf, media_out):
+    result = {
+        "kind": "presentation",
+        "format": "pptx",
+        "slides": [],
+        "paragraphs": [],
+        "outline": [],
+        "tables": [],
+        "hyperlinks": [],
+        "headings": [],
+        "media": [],
+        "embeddedObjects": [],
+        "notes": [],
+        "slideCount": 0,
+    }
+
+    pres_rels = parse_rels(zf, "ppt/_rels/presentation.xml.rels")
+    ordered_parts = []
+
+    # 首选:presentation.xml 的 sldIdLst(真实放映顺序)
+    if "ppt/presentation.xml" in zf.namelist():
+        try:
+            proot = ET.fromstring(zf.read("ppt/presentation.xml"))
+            for sld in proot.iter(qname("p", "sldId")):
+                rid = sld.get(REL_NS + "id") or sld.get(qname("r", "id")) or ""
+                if rid in pres_rels:
+                    ordered_parts.append(rel_target_to_part("ppt/_rels/presentation.xml.rels", pres_rels[rid]["target"]))
+        except ET.ParseError:
+            pass
+
+    if not ordered_parts:
+        ordered_parts = sorted(
+            [n for n in zf.namelist() if re.match(r"^ppt/slides/slide\d+\.xml$", n)],
+            key=slide_sort_key,
+        )
+
+    for idx, part in enumerate(ordered_parts, start=1):
+        if part not in zf.namelist():
+            continue
+        try:
+            sroot = ET.fromstring(zf.read(part))
+        except ET.ParseError:
+            continue
+        rels = parse_rels(zf, posix_join(posix_dirname(part), "_rels/" + part.rsplit("/", 1)[-1] + ".rels"))
+
+        title = ""
+        body_lines = []
+        slide_texts = []
+        pic_count = 0
+
+        for sp in sroot.iter(qname("p", "sp")):
+            ph = None
+            for node in sp.iter(qname("p", "ph")):
+                ph = node.get("type") or "body"
+                break
+            texts = []
+            for p_node in sp.iter(qname("a", "p")):
+                t = paragraph_text(p_node)
+                if t:
+                    texts.append(t)
+            if not texts:
+                continue
+            if ph == "title" or ph == "ctrTitle" or (not title and ph is None and len(texts) == 1 and len(texts[0]) < 60):
+                if not title:
+                    title = texts[0]
+                    result["headings"].append(texts[0])
+                    continue
+            body_lines.extend(texts)
+
+        # 表格
+        for tbl in sroot.iter(qname("a", "tbl")):
+            rows = []
+            for tr in tbl.iter(qname("a", "tr")):
+                cells = []
+                for tc in tr.iter(qname("a", "tc")):
+                    cell_text = " ".join(paragraph_text(p) for p in tc.iter(qname("a", "p"))).strip()
+                    cells.append(cell_text)
+                if cells:
+                    rows.append(cells)
+            if rows:
+                result["tables"].append({"scope": part, "slide": idx, "rows": rows, "rowCount": len(rows)})
+
+        # 图片
+        for blip in sroot.iter(qname("a", "blip")):
+            rid = blip.get(REL_NS + "embed") or blip.get(REL_NS + "link")
+            if rid and rid in rels:
+                pic_count += 1
+                result["media"].append({
+                    "rId": rid,
+                    "part": rel_target_to_part(posix_join(posix_dirname(part), "_rels/" + part.rsplit("/", 1)[-1] + ".rels"), rels[rid]["target"]),
+                    "anchor": "slide",
+                    "slide": idx,
+                    "nearbyText": title or (body_lines[0] if body_lines else ""),
+                })
+
+        # 超链接
+        for link in sroot.iter(qname("a", "hlinkClick")):
+            rid = link.get(REL_NS + "id")
+            if rid and rid in rels:
+                result["hyperlinks"].append({"text": title, "target": rels[rid]["target"], "scope": part, "slide": idx})
+
+        # 备注页
+        notes_part = posix_join(posix_dirname(part), "notesSlides/notesSlide%d.xml" % idx)
+        notes_text = ""
+        if notes_part in zf.namelist():
+            try:
+                nroot = ET.fromstring(zf.read(notes_part))
+                notes_text = "\n".join(t for t in (paragraph_text(p) for p in nroot.iter(qname("a", "p"))) if t)
+            except ET.ParseError:
+                notes_text = ""
+        if notes_text:
+            result["notes"].append({"scope": notes_part, "slide": idx, "text": notes_text})
+
+        slide_entry = {
+            "index": idx,
+            "part": part,
+            "title": title,
+            "texts": body_lines,
+            "text": "\n".join([title] + body_lines).strip(),
+            "notes": notes_text,
+            "imageCount": pic_count,
+        }
+        result["slides"].append(slide_entry)
+        for line in slide_entry["text"].split("\n"):
+            if line.strip():
+                result["paragraphs"].append({"index": len(result["paragraphs"]), "scope": part, "slide": idx, "text": line.strip()})
+                result["outline"].append({"level": 1 if line.strip() == title else 0, "text": line.strip()})
+
+    # 内嵌文件(ppt/embeddings/*)
+    for name in zf.namelist():
+        if name.startswith("ppt/embeddings/") and not name.endswith("/"):
+            result["embeddedObjects"].append({"part": safe_relpath(name), "progId": ""})
+
+    result["slideCount"] = len(result["slides"])
+    result["media"] = attach_media(zf, result["media"], media_out)
+    return result
+
+
+# ---------------------------------------------------------------------------
+# media
+# ---------------------------------------------------------------------------
+
+def classify_media(part):
+    ext = os.path.splitext(part)[1].lower()
+    if ext in IMAGE_EXTS or ext == "":
+        return "image"
+    if ext in VIDEO_EXTS:
+        return "video"
+    if ext in AUDIO_EXTS:
+        return "audio"
+    return "other"
+
+
+def attach_media(zf, media_entries, media_out):
+    """把媒体条目去重、补上 kind/size,并按需落地到 media_out。"""
+    by_part = {}
+    order = []
+    for entry in media_entries:
+        part = safe_relpath(entry.get("part") or "")
+        if not part:
+            continue
+        if part not in by_part:
+            by_part[part] = {**entry, "part": part, "kind": classify_media(part), "occurrences": 0}
+            order.append(part)
+        by_part[part]["occurrences"] += 1
+
+    out = []
+    for part in order:
+        entry = by_part[part]
+        entry["size"] = zf.getinfo(part).file_size if part in zf.namelist() else 0
+        entry["localPath"] = ""
+        if media_out and part in zf.namelist():
+            safe_name = part.replace("/", "__")
+            target = os.path.join(media_out, safe_name)
+            os.makedirs(media_out, exist_ok=True)
+            with zf.open(part) as src, open(target, "wb") as dst:
+                dst.write(src.read())
+            entry["localPath"] = os.path.abspath(target)
+        out.append(entry)
+    return out
+
+
+# ---------------------------------------------------------------------------
+# entry
+# ---------------------------------------------------------------------------
+
+def package_kind(path):
+    ext = os.path.splitext(path)[1].lower()
+    if ext in DOCX_EXT:
+        return "docx"
+    if ext in PPTX_EXT:
+        return "pptx"
+    # 扩展名不认识时,看 package 内容
+    try:
+        with zipfile.ZipFile(path) as zf:
+            names = zf.namelist()
+            if any(n.startswith("word/") for n in names):
+                return "docx"
+            if any(n.startswith("ppt/") for n in names):
+                return "pptx"
+    except (zipfile.BadZipFile, OSError):
+        pass
+    return ""
+
+
+def extract_one(path, media_out=None, include_headers=True):
+    kind = package_kind(path)
+    base = {
+        "input": os.path.abspath(path),
+        "fileName": os.path.basename(path),
+        "byteSize": os.path.getsize(path) if os.path.exists(path) else 0,
+    }
+    if not kind:
+        return {**base, "ok": False, "error": "UNSUPPORTED_PACKAGE", "message": "不是 .docx/.pptx 包(zip 内无 word/ 或 ppt/)"}
+    try:
+        with zipfile.ZipFile(path) as zf:
+            if kind == "docx":
+                payload = extract_docx(zf, media_out, include_headers=include_headers)
+            else:
+                payload = extract_pptx(zf, media_out)
+    except zipfile.BadZipFile:
+        return {**base, "ok": False, "error": "BAD_ZIP", "message": "文件损坏或不是有效 zip 包"}
+    except ET.ParseError as exc:
+        return {**base, "ok": False, "error": "BAD_XML", "message": "XML 解析失败:%s" % exc}
+    return {**base, "ok": True, **payload}
+
+
+def selftest():
+    """自检:现场构造一个最小 docx + pptx,解包并断言关键结果。"""
+    import io
+    import tempfile
+
+    doc_xml = (
+        '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
+        '<w:document xmlns:w="%s" xmlns:r="%s" xmlns:wp="%s" xmlns:a="%s">'
+        "<w:body>"
+        '<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr><w:r><w:t>帝国理工 计算机科学 辅导案例</w:t></w:r></w:p>'
+        "<w:p><w:r><w:t>学生考前冲刺,最终</w:t></w:r><w:r><w:t>提分 18 分</w:t></w:r></w:p>"
+        "<w:p><w:r><w:drawing><wp:inline><a:graphic><a:blip r:embed=\"rId5\"/></a:graphic></wp:inline></w:drawing></w:r></w:p>"
+        '<w:hyperlink r:id="rId9"><w:r><w:t>查看反馈</w:t></w:r></w:hyperlink>'
+        "<w:tbl><w:tr><w:tc><w:p><w:r><w:t>A</w:t></w:r></w:p></w:tc><w:tc><w:p><w:r><w:t>B</w:t></w:r></w:p></w:tc></w:tr></w:tbl>"
+        "</w:body></w:document>"
+    ) % (NS["w"], NS["r"], NS["wp"], NS["a"])
+
+    doc_rels = (
+        '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
+        '<Relationships xmlns="%s">'
+        '<Relationship Id="rId5" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image" Target="media/image1.png"/>'
+        '<Relationship Id="rId9" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/hyperlink" Target="https://example.invalid/case" TargetMode="External"/>'
+        "</Relationships>"
+    ) % NS["pr"]
+
+    slide_xml = (
+        '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
+        '<p:sld xmlns:p="%s" xmlns:a="%s" xmlns:r="%s"><p:cSld><p:spTree>'
+        '<p:sp><p:nvSpPr><p:nvPr><p:ph type="title"/></p:nvPr></p:nvSpPr><p:txBody><a:p><a:r><a:t>案例合集</a:t></a:r></a:p></p:txBody></p:sp>'
+        "<p:sp><p:nvSpPr/><p:txBody><a:p><a:r><a:t>正文一</a:t></a:r></a:p></p:txBody></p:sp>"
+        '<p:pic><p:blipFill><a:blip r:embed="rId2"/></p:blipFill></p:pic>'
+        "</p:spTree></p:cSld></p:sld>"
+    ) % (NS["p"], NS["a"], NS["r"])
+
+    slide_rels = (
+        '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
+        '<Relationships xmlns="%s">'
+        '<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image" Target="../media/image1.png"/>'
+        "</Relationships>"
+    ) % NS["pr"]
+
+    pres_xml = (
+        '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
+        '<p:presentation xmlns:p="%s" xmlns:r="%s">'
+        '<p:sldIdLst><p:sldId id="256" r:id="rId1"/></p:sldIdLst></p:presentation>'
+    ) % (NS["p"], NS["r"])
+
+    pres_rels = (
+        '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
+        '<Relationships xmlns="%s">'
+        '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/slide" Target="slides/slide1.xml"/>'
+        "</Relationships>"
+    ) % NS["pr"]
+
+    png_bytes = bytes.fromhex(
+        "89504e470d0a1a0a0000000d49484452000000010000000108060000001f15c489"
+        "0000000a49444154789c6300010000050001"
+        "0d0a2db40000000049454e44ae426082"
+    )
+
+    failures = []
+    with tempfile.TemporaryDirectory() as tmp:
+        docx_path = os.path.join(tmp, "case.docx")
+        with zipfile.ZipFile(docx_path, "w") as zf:
+            zf.writestr("[Content_Types].xml", "<Types/>")
+            zf.writestr("word/document.xml", doc_xml)
+            zf.writestr("word/_rels/document.xml.rels", doc_rels)
+            zf.writestr("word/media/image1.png", png_bytes)
+
+        res = extract_one(docx_path, media_out=os.path.join(tmp, "media"))
+        if not res.get("ok"):
+            failures.append("docx 解包失败:%s" % res.get("error"))
+        else:
+            texts = [p["text"] for p in res["paragraphs"]]
+            if "帝国理工 计算机科学 辅导案例" not in texts:
+                failures.append("docx 标题段落缺失")
+            if not any("提分 18 分" in t for t in texts):
+                failures.append("docx 跨 run 文本拼接失败")
+            if not res["headings"]:
+                failures.append("docx 标题识别失败")
+            if not res["tables"] or res["tables"][0]["rows"][0] != ["A", "B"]:
+                failures.append("docx 表格抽取失败")
+            if not res["hyperlinks"] or res["hyperlinks"][0]["target"] != "https://example.invalid/case":
+                failures.append("docx 超链接抽取失败")
+            media = [m for m in res["media"] if m["kind"] == "image"]
+            if not media:
+                failures.append("docx 内嵌图片未识别")
+            elif not os.path.exists(media[0]["localPath"]):
+                failures.append("docx 内嵌图片未落地")
+
+        pptx_path = os.path.join(tmp, "deck.pptx")
+        with zipfile.ZipFile(pptx_path, "w") as zf:
+            zf.writestr("[Content_Types].xml", "<Types/>")
+            zf.writestr("ppt/presentation.xml", pres_xml)
+            zf.writestr("ppt/_rels/presentation.xml.rels", pres_rels)
+            zf.writestr("ppt/slides/slide1.xml", slide_xml)
+            zf.writestr("ppt/slides/_rels/slide1.xml.rels", slide_rels)
+            zf.writestr("ppt/media/image1.png", png_bytes)
+
+        res = extract_one(pptx_path, media_out=os.path.join(tmp, "media2"))
+        if not res.get("ok"):
+            failures.append("pptx 解包失败:%s" % res.get("error"))
+        else:
+            if res.get("slideCount") != 1:
+                failures.append("pptx 页数不对:%s" % res.get("slideCount"))
+            if not res["slides"] or res["slides"][0]["title"] != "案例合集":
+                failures.append("pptx 标题占位符识别失败")
+            if not any("正文一" in t for t in res["slides"][0]["texts"]):
+                failures.append("pptx 正文抽取失败")
+            if not [m for m in res["media"] if m["kind"] == "image"]:
+                failures.append("pptx 图片未识别")
+            elif res["media"][0]["part"] != "ppt/media/image1.png":
+                failures.append("pptx 图片路径未正确解析 ..:%s" % res["media"][0]["part"])
+            elif not os.path.exists(res["media"][0]["localPath"]):
+                failures.append("pptx 内嵌图片未落地")
+
+    if failures:
+        print(json.dumps({"ok": False, "failures": failures}, ensure_ascii=False, indent=2))
+        return 1
+    print(json.dumps({"ok": True, "checked": ["docx", "pptx"], "message": "office 解包自检通过"}, ensure_ascii=False, indent=2))
+    return 0
+
+
+def main(argv=None):
+    parser = argparse.ArgumentParser(description="skill-case-get .docx/.pptx 解包抽取器(标准库实现)")
+    parser.add_argument("--input", action="append", default=[], help="待解包的 .docx/.pptx(可重复)")
+    parser.add_argument("--media-out", default="", help="内嵌媒体落地目录")
+    parser.add_argument("--out", default="", help="结果 JSON 输出路径(默认 stdout)")
+    parser.add_argument("--no-headers", action="store_true", help=".docx 不解页眉页脚")
+    parser.add_argument("--selftest", action="store_true", help="运行内置自检")
+    args = parser.parse_args(argv)
+
+    if args.selftest:
+        return selftest()
+
+    if not args.input:
+        parser.error("至少需要一个 --input")
+
+    results = [extract_one(p, args.media_out or None, include_headers=not args.no_headers) for p in args.input]
+    payload = {
+        "ok": all(r.get("ok") for r in results),
+        "generator": "skill-case-get/office_extract.py",
+        "count": len(results),
+        "results": results,
+    }
+    text = json.dumps(payload, ensure_ascii=False, indent=2)
+    if args.out:
+        with open(args.out, "w", encoding="utf-8") as fh:
+            fh.write(text + "\n")
+    else:
+        sys.stdout.write(text + "\n")
+    return 0 if payload["ok"] else 2
+
+
+if __name__ == "__main__":
+    sys.exit(main())

+ 121 - 0
scripts/smoke.mjs

@@ -0,0 +1,121 @@
+#!/usr/bin/env node
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 冒烟检查:不联网、不调模型、不写仓库。
+ *
+ * 检查项:
+ *   1. 必需文件齐全(含标签字典可解析)
+ *   2. 运行器自检(office 解包 + 分组 + 打标 + 合规 + reorder)
+ *   3. 单元测试(node --test)
+ *   4. 未把 token / 私钥 / 内部地址写进包内任何文件
+ */
+
+import fs from 'node:fs';
+import path from 'node:path';
+import { spawnSync } from 'node:child_process';
+import { fileURLToPath } from 'node:url';
+
+const ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
+const failures = [];
+const checks = [];
+
+function record(name, ok, detail) {
+  checks.push({ name, ok, detail: String(detail ?? '') });
+  if (!ok) failures.push(name);
+}
+
+// 1. 必需文件
+const REQUIRED = [
+  'SKILL.md', 'README.md', 'package.json', 'skill-package-manifest.json',
+  'references/tag-dictionary.json',
+  'scripts/case-intake.mjs', 'scripts/office_extract.py', 'scripts/vision-bridge.mjs',
+  'scripts/listen-bridge.mjs', 'scripts/cloud-client.mjs', 'scripts/grouping.mjs',
+  'scripts/classify.mjs', 'scripts/image-size.mjs', 'scripts/lib.mjs',
+  'bin/skill-case-get.js',
+];
+{
+  const missing = REQUIRED.filter((rel) => !fs.existsSync(path.join(ROOT, rel)));
+  record('必需文件齐全', missing.length === 0, missing.length ? `缺少:${missing.join(', ')}` : `${REQUIRED.length} 个文件`);
+}
+
+// 2. 标签字典可解析
+{
+  try {
+    const dict = JSON.parse(fs.readFileSync(path.join(ROOT, 'references', 'tag-dictionary.json'), 'utf8'));
+    const dims = Object.keys(dict.dimensions || {});
+    record('标签字典可解析', dims.length >= 8, `维度:${dims.length} 个(${dims.join(', ')})`);
+  } catch (error) {
+    record('标签字典可解析', false, error.message);
+  }
+}
+
+// 3. 运行器自检
+{
+  const run = spawnSync(process.execPath, [path.join(ROOT, 'scripts', 'case-intake.mjs'), 'selftest'], { encoding: 'utf8' });
+  record('case-intake selftest', run.status === 0, run.stdout ? run.stdout.slice(0, 200) : run.stderr);
+}
+
+// 4. 单元测试
+{
+  const run = spawnSync(process.execPath, ['--test', path.join(ROOT, 'scripts', 'tests')], { encoding: 'utf8' });
+  // node --test 传目录在部分版本上会当模块解析;退回 glob
+  const globRun = run.status === 0 ? run : spawnSync(
+    process.execPath,
+    ['--test', path.join(ROOT, 'scripts', 'tests', 'grouping.test.mjs'), path.join(ROOT, 'scripts', 'tests', 'classify.test.mjs')],
+    { encoding: 'utf8' },
+  );
+  const output = `${globRun.stdout || ''}${globRun.stderr || ''}`;
+  const passMatch = output.match(/# pass (\d+)/) || output.match(/ℹ pass (\d+)/);
+  const failMatch = output.match(/# fail (\d+)/) || output.match(/ℹ fail (\d+)/);
+  record('单元测试', globRun.status === 0, `pass=${passMatch ? passMatch[1] : '?'} fail=${failMatch ? failMatch[1] : '?'}`);
+}
+
+// 5. 包内不得出现凭据 / 内部地址
+{
+  const FORBIDDEN = [
+    { label: 'internal-address', re: new RegExp(`\\b${['local', 'host'].join('')}\\b`, 'i') },
+    { label: 'loopback', re: /\b(127\.0\.0\.1|0\.0\.0\.0)\b/ },
+    { label: 'parse-session-token', re: /\br:[a-f0-9]{20,}\b/i },
+    { label: 'api-key', re: /\bsk-[a-zA-Z0-9-]{20,}\b/ },
+    { label: 'mongo-uri', re: /mongodb(\+srv)?:\/\//i },
+  ];
+  const BINARY = new Set(['.png', '.jpg', '.jpeg', '.gif', '.zip', '.pdf']);
+  // 本文件自身就写着这些正则(和仓库 self-check.mjs 一样自跳过),否则会自己命中自己
+  const SELF = new Set([path.resolve(fileURLToPath(import.meta.url))]);
+  const offenders = [];
+  const walk = (dir) => {
+    for (const entry of fs.readdirSync(dir)) {
+      if (['node_modules', 'outputs', '.git', '.claude', 'case-get-outputs'].includes(entry)) continue;
+      const full = path.join(dir, entry);
+      const stat = fs.statSync(full);
+      if (stat.isDirectory()) { walk(full); continue; }
+      if (SELF.has(path.resolve(full))) continue;
+      if (BINARY.has(path.extname(full).toLowerCase())) continue;
+      let content = '';
+      try { content = fs.readFileSync(full, 'utf8'); } catch { continue; }
+      for (const item of FORBIDDEN) {
+        if (item.re.test(content)) offenders.push(`${path.relative(ROOT, full)}:${item.label}`);
+      }
+    }
+  };
+  walk(ROOT);
+  record('包内无凭据 / 无内网地址', offenders.length === 0, offenders.length ? offenders.join(', ') : '干净');
+}
+
+// 输出
+const passed = checks.filter((c) => c.ok).length;
+for (const check of checks) {
+  process.stdout.write(`${check.ok ? '✔' : '✖'} ${check.name} — ${check.detail}\n`);
+}
+process.stdout.write(`\n冒烟检查:${passed}/${checks.length} 通过\n`);
+process.exit(failures.length ? 1 : 0);

+ 336 - 0
scripts/tests/classify.test.mjs

@@ -0,0 +1,336 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 拆解 / 打标 / 合规 / 入库契约 测试:`node --test`
+ */
+
+import test from 'node:test';
+import assert from 'node:assert/strict';
+import path from 'node:path';
+import fs from 'node:fs';
+import os from 'node:os';
+
+import { loadDictionary, buildIdempotencyKey, normalizeMaterialAsset, imageUrlsFrom, sortAssets, uniq } from '../lib.mjs';
+import {
+  assessCompliance, classifyCase, deriveCaseFields, extractOutcome, findPrivacy,
+  inferFitStatus, normalizeSchool, splitCaseAndMaterials,
+} from '../classify.mjs';
+import { runIngest, makeTinyPng } from '../case-intake.mjs';
+import { makeFixtures } from './fixtures.mjs';
+
+const dict = loadDictionary();
+
+// ---------------------------------------------------------------------------
+// 1. 输入识别 / 拆解
+// ---------------------------------------------------------------------------
+
+test('splitCaseAndMaterials:案例说明 与 素材 严格分开', () => {
+  const split = splitCaseAndMaterials({
+    textBundle: { paragraphs: ['案例说明段落'], notes: ['备注'], titles: [], transcript: '视频转写内容' },
+    visionResults: [
+      { role: 'material', ocrText: '素材图里的字', description: '成绩单截图' },
+      { role: 'description', ocrText: '说明配图的字', description: '流程图' },
+    ],
+  });
+  assert.ok(split.descriptionCorpus.includes('案例说明段落'));
+  assert.ok(split.descriptionCorpus.includes('视频转写内容'));
+  assert.ok(split.descriptionCorpus.includes('流程图'), 'description 角色归案例说明');
+  assert.ok(split.materialCorpus.includes('成绩单截图'), 'material 角色归素材');
+  assert.ok(!split.materialCorpus.includes('案例说明段落'));
+});
+
+// ---------------------------------------------------------------------------
+// 2. 打标
+// ---------------------------------------------------------------------------
+
+test('学校别名归一:IC → 帝国理工,UCL → 伦敦大学学院;不误伤 MAGIC', () => {
+  const a = normalizeSchool('IC 的计算机课,考前冲刺', dict);
+  assert.equal(a.schoolCanonical, '帝国理工');
+  assert.deepEqual(a.schoolAliases, ['IC']);
+  assert.equal(a.country, '英国');
+
+  const b = normalizeSchool('UCL 的学生找我辅导', dict);
+  assert.equal(b.schoolCanonical, '伦敦大学学院');
+
+  assert.equal(normalizeSchool('MAGIC 系统很好用', dict).schoolCanonical, '');
+});
+
+test('打标:亮点 / 场景 / 异议 / 静态属性 按字典与策略命中', () => {
+  const { tags } = classifyCase({
+    dict,
+    groups: [],
+    visionResults: [],
+    textBundle: {
+      paragraphs: ['帝国理工 计算机科学 本科二年级,期末考前冲刺,提分 18 分,家长很满意', '客户说太贵了,不放心老师'],
+      notes: [], titles: [], transcript: '',
+    },
+    authorizationStatus: 'authorized',
+  });
+
+  assert.deepEqual(tags.schoolCanonical, ['帝国理工']);
+  assert.deepEqual(tags.country, ['英国']);
+  assert.ok(tags.major.includes('计算机科学'));
+  assert.ok(tags.highlightTypes.includes('提分'));
+  assert.ok(tags.highlightTypes.includes('家长认可'));
+  assert.ok(tags.scenarioTags.includes('考前冲刺'));
+  assert.ok(tags.scenarioTags.includes('期末考试'));
+  assert.ok(tags.objectionTags.includes('贵'));
+  assert.ok(tags.objectionTags.includes('不放心老师'));
+});
+
+test('适用跟进状态:按 docs/prd/dashboard/10 策略命中(异议处理中)', () => {
+  const fit = inferFitStatus({
+    highlightTypes: ['提分', '家长认可'],
+    scenarioTags: ['考前冲刺'],
+    objectionTags: ['贵'],
+    subject: [],
+  });
+  assert.ok(fit.fitStatus.includes('异议处理中'));
+  assert.match(fit.reason, /跟进状态/);
+});
+
+test('适用跟进状态:命中七个枚举之一(全部在契约枚举内)', () => {
+  const allowed = new Set(['新进线', '挖需中', '方案推荐中', '异议处理中', '待决策', '沉默待跟进', '已成交']);
+  for (const scenario of [
+    { highlightTypes: ['补考通过'], scenarioTags: [], objectionTags: [], subject: [] },
+    { highlightTypes: ['导师匹配'], scenarioTags: [], objectionTags: [], subject: [] },
+    { highlightTypes: [], scenarioTags: ['考前冲刺'], objectionTags: [], subject: [] },
+    { highlightTypes: ['首课体验'], scenarioTags: [], objectionTags: [], subject: ['高等数学'] },
+  ]) {
+    for (const status of inferFitStatus(scenario).fitStatus) {
+      assert.ok(allowed.has(status), `非法跟进状态:${status}`);
+    }
+  }
+});
+
+// ---------------------------------------------------------------------------
+// 3. 隐私 / 合规
+// ---------------------------------------------------------------------------
+
+test('findPrivacy:片段一律掩码,绝不返回完整 PII', () => {
+  const findings = findPrivacy('联系 13812345678 邮箱 a.b@example.com 微信:student_abc');
+  const fields = findings.map((f) => f.field);
+  assert.ok(fields.includes('手机号'));
+  assert.ok(fields.includes('邮箱'));
+  assert.ok(fields.includes('微信号'));
+  const phone = findings.find((f) => f.field === '手机号');
+  assert.ok(!phone.snippet.includes('13812345678'), '手机号必须已掩码');
+  assert.match(phone.snippet, /^138\*{4}5678$/);
+});
+
+test('合规:客诉 / 负面事件 → riskFlags,且禁止入库', () => {
+  const result = assessCompliance({
+    corpus: '客户投诉退费纠纷,要求赔偿,还发了黑猫',
+    authorizationStatus: 'authorized',
+    hasMaterials: true,
+  });
+  assert.ok(result.riskFlags.includes('COMPLAINT'));
+  assert.equal(result.authorizationOk, true);
+  assert.ok(result.complianceBlockers.some((b) => b.includes('客诉')));
+  assert.match(result.reviewHint, /驳回|复盘/);
+});
+
+test('合规:无授权一律不入库(硬规则)', () => {
+  for (const status of ['pending', 'denied', '']) {
+    const result = assessCompliance({
+      corpus: '帝国理工 提分案例,很好',
+      authorizationStatus: status,
+      hasMaterials: true,
+    });
+    assert.equal(result.authorizationOk, false, `status=${status} 必须判为无授权`);
+    assert.ok(result.riskFlags.includes('UNAUTHORIZED'));
+    assert.ok(result.complianceBlockers.some((b) => b.includes('无授权')));
+  }
+});
+
+test('合规:无素材 / 跑题 也拦下来', () => {
+  assert.ok(assessCompliance({ corpus: '帝国理工 提分', authorizationStatus: 'authorized', hasMaterials: false })
+    .complianceBlockers.some((b) => b.includes('可用素材')));
+  assert.ok(assessCompliance({ corpus: '今天天气不错,出门散步', authorizationStatus: 'authorized', hasMaterials: false })
+    .riskFlags.includes('OFF_TOPIC'));
+});
+
+test('合规:正常好案例 → 可入库待审,无 blocker', () => {
+  const result = assessCompliance({
+    corpus: '帝国理工 计算机 考前冲刺,期末提分 18 分,家长很满意',
+    authorizationStatus: 'authorized',
+    hasMaterials: true,
+  });
+  assert.equal(result.authorizationOk, true);
+  assert.equal(result.complianceBlockers.length, 0);
+  assert.match(result.reviewHint, /待审/);
+});
+
+// ---------------------------------------------------------------------------
+// 4. 案例字段抽取
+// ---------------------------------------------------------------------------
+
+test('extractOutcome:提分幅度 / 周期 / 结果 / 录取', () => {
+  assert.equal(extractOutcome('期末提分 18 分').scoreGain, '18分');
+  assert.equal(extractOutcome('从 52 分到 78 分').scoreRange, '52 → 78');
+  assert.equal(extractOutcome('辅导 6 周').period, '6 周');
+  assert.equal(extractOutcome('拿到了帝国理工的录取').admittedTo, '帝国理工');
+  assert.equal(extractOutcome('最终 pass 了').result, 'pass');
+});
+
+test('deriveCaseFields:标题 / 摘要 / 使用建议 / 结果证据 都产出', () => {
+  const fields = deriveCaseFields(
+    'IC 计算机科学考前冲刺案例\n客户在读本科二年级,期末前两周找到我们,担心挂科。\n结果:期末提分 18 分,家长很满意。',
+    [{ count: 3, preview: '' }],
+    { productLine: ['大学课程'], schoolCanonical: ['帝国理工'], subject: ['高等数学'], stage: ['本科二年级'], scenarioTags: ['考前冲刺'], objectionTags: [] },
+    { fitStatus: ['挖需中'] },
+  );
+  assert.ok(fields.title.length > 0);
+  assert.ok(fields.summary.length > 0);
+  assert.match(fields.usageSuggestion, /痛点|挂|时间/);
+  assert.match(fields.targetCustomer, /帝国理工/);
+  assert.match(fields.resultEvidence, /提分/);
+});
+
+// ---------------------------------------------------------------------------
+// 5. materialAssets 契约(逐字对齐 SHARED-CONTRACT.md)
+// ---------------------------------------------------------------------------
+
+test('normalizeMaterialAsset:字段齐全、order 1 起、非法值报问题', () => {
+  const { asset, problems } = normalizeMaterialAsset({ kind: 'image', url: 'https://x/a.jpg', order: 3 }, 0);
+  assert.deepEqual(Object.keys(asset).sort(), [
+    'description', 'groupId', 'kind', 'label', 'localPath', 'ocrText', 'order', 'role', 'url', 'usageSuggestion',
+  ]);
+  assert.equal(asset.order, 3);
+  assert.equal(asset.role, 'material');
+  assert.equal(problems.length, 0);
+
+  const bad = normalizeMaterialAsset({ kind: 'pdf' }, 0);
+  assert.ok(bad.problems.some((p) => p.includes('kind')));
+  assert.ok(bad.problems.some((p) => p.includes('url')));
+});
+
+test('imageUrlsFrom:= kind=image 的 url,按 order 排序(兼容字段)', () => {
+  const urls = imageUrlsFrom([
+    { order: 2, kind: 'image', url: 'https://x/b.jpg' },
+    { order: 1, kind: 'image', url: 'https://x/a.jpg' },
+    { order: 3, kind: 'video', url: 'https://x/c.mp4' },
+  ]);
+  assert.deepEqual(urls, ['https://x/a.jpg', 'https://x/b.jpg']);
+});
+
+test('buildIdempotencyKey:同内容同键,内容变则变(tenantId + idempotencyKey 幂等)', () => {
+  const base = { tenantId: 'lumi-demo', sourceType: 'image_group', sourceRef: 'wx-2026-10', title: 'IC 提分', materialAssets: [{ order: 1, kind: 'image', url: 'u1' }] };
+  const k1 = buildIdempotencyKey(base);
+  const k2 = buildIdempotencyKey({ ...base });
+  const k3 = buildIdempotencyKey({ ...base, title: '别的标题' });
+  assert.equal(k1, k2);
+  assert.notEqual(k1, k3);
+  assert.match(k1, /^caseget-[0-9a-f]{16}$/);
+  assert.equal(buildIdempotencyKey({ ...base, explicit: 'manual-key' }), 'manual-key');
+});
+
+test('uniq:多参数/嵌套都支持,并丢空值', () => {
+  assert.deepEqual(uniq(['a'], ['b', 'a'], [], ['c']), ['a', 'b', 'c']);
+  assert.deepEqual(uniq([[1, 2], [2, 3]]), [1, 2, 3]);
+  assert.deepEqual(uniq([null, undefined, '', 'x']), ['x']);
+});
+
+test('sortAssets:按 order 升序且稳定', () => {
+  const sorted = sortAssets([{ order: 2, label: 'b' }, { order: 1, label: 'a' }, { order: 2, label: 'c' }]);
+  assert.deepEqual(sorted.map((a) => a.label), ['a', 'b', 'c']);
+});
+
+// ---------------------------------------------------------------------------
+// 6. 端到端:夹具 → ingest → 案例包
+// ---------------------------------------------------------------------------
+
+test('端到端 ingest:多图(9)+docx+pptx → 案例包,字段对齐共享契约,不自动提交', async (t) => {
+  const workRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'caseget-e2e-'));
+  t.after(() => fs.rmSync(workRoot, { recursive: true, force: true }));
+
+  const fixtures = makeFixtures(path.join(workRoot, 'fixtures'));
+  const outDir = path.join(workRoot, 'out');
+
+  const report = await runIngest({
+    inputs: [...fixtures.images, fixtures.docx, fixtures.pptx],
+    outDir,
+    sourceRef: 'e2e-fixture',
+    authorizationStatus: 'authorized',
+    assetBaseUrl: 'https://cdn.example.invalid/cases/e2e',
+    skipVision: true, // 不调模型:OCR 留空,其余链路全跑
+    noLearn: true,
+  });
+
+  const pkg = report.casePackage;
+
+  // 输入识别:多图 + 文档 + PPT → mixed
+  assert.equal(pkg.sourceType, 'mixed');
+
+  // 素材:9 张九宫格 + 1 张异批 + docx 内嵌 1 + pptx 内嵌 1
+  assert.equal(pkg.materialAssets.length, 12);
+  assert.deepEqual(pkg.materialAssets.map((a) => a.order), Array.from({ length: 12 }, (_, i) => i + 1));
+
+  // 四个分组:九宫格 / 异批单图 / docx 内嵌图 / pptx 内嵌图
+  // (两份文档的内嵌图不会互相并组——sourceKey 不同)
+  assert.equal(pkg.groupInfoList.length, 4);
+  assert.equal(pkg.groupInfoList.filter((g) => g.count === 1).length, 3);
+  assert.ok(pkg.materialAssets.some((a) => a.description.includes('文档') || a.label.includes('内嵌')), '应包含文档内嵌图素材');
+
+  // 九宫格分组
+  const gridGroup = pkg.groupInfoList.find((g) => g.count === 9);
+  assert.ok(gridGroup, '应识别出 9 张的九宫格组');
+  assert.equal(gridGroup.layout, '3x3');
+  assert.equal(gridGroup.orderingRule, 'left-to-right,top-to-bottom');
+  assert.equal(pkg.groupInfo.layout, '3x3');
+
+  // 契约必填字段
+  for (const key of ['sourceType', 'idempotencyKey', 'title', 'summary', 'materialAssets', 'groupInfo', 'tags', 'highlightTypes', 'scenarioTags', 'objectionTags', 'riskFlags', 'authorizationStatus']) {
+    assert.ok(key in pkg, `案例包缺少契约字段 ${key}`);
+  }
+
+  // 落库预期:进待审清单,不进公共素材库
+  assert.deepEqual(pkg.expected, { reviewStatus: 'pending', readyForUse: false });
+
+  // 文档抽取的文本进入打标语料
+  assert.ok(pkg.schoolCanonical === '帝国理工', `学校归一失败:${pkg.schoolCanonical}`);
+  assert.ok(pkg.highlightTypes.includes('提分'));
+  assert.ok(pkg.imageUrls.length >= 10);
+  assert.ok(pkg.imageUrls.every((u) => u.startsWith('https://cdn.example.invalid/cases/e2e/')), 'url 应按 --asset-base-url 派生');
+  assert.equal(pkg.hasImage, true);
+  assert.equal(pkg.materialAssets.filter((a) => !a.url).length, 0, '提供了 base url 后不应再有缺 url 的素材');
+
+  // 产物落盘
+  for (const file of Object.values(report.artifacts)) {
+    assert.ok(fs.existsSync(file), `产物缺失:${file}`);
+  }
+  const pkgOnDisk = JSON.parse(fs.readFileSync(report.artifacts.casePackage, 'utf8'));
+  assert.equal(pkgOnDisk.idempotencyKey, pkg.idempotencyKey);
+
+  // 报告提到九宫格顺序与待审
+  const md = fs.readFileSync(report.artifacts.report, 'utf8');
+  assert.match(md, /materialAssets/);
+  assert.match(md, /3x3/);
+});
+
+test('端到端 ingest:未授权时不产出可提交案例包(canSubmit=false)', async (t) => {
+  const workRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'caseget-unauth-'));
+  t.after(() => fs.rmSync(workRoot, { recursive: true, force: true }));
+
+  const fixtures = makeFixtures(path.join(workRoot, 'fixtures'));
+  const report = await runIngest({
+    inputs: [fixtures.images[0]],
+    outDir: path.join(workRoot, 'out'),
+    authorizationStatus: 'pending',
+    skipVision: true,
+    noLearn: true,
+  });
+
+  assert.equal(report.canSubmit, false);
+  assert.ok(report.casePackage.riskFlags.includes('UNAUTHORIZED'));
+  assert.ok(report.casePackage.complianceBlockers.length > 0);
+});

+ 127 - 0
scripts/tests/fixtures.mjs

@@ -0,0 +1,127 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 测试夹具生成器:现场合成 .docx / .pptx / .png,落盘到指定目录。
+ *
+ * 刻意不在仓库里存任何二进制素材(.gitignore 也排除了图片)——夹具全部即时生成,
+ * 内容为合成的假数据,不含任何真实 PII。
+ */
+
+import fs from 'node:fs';
+import path from 'node:path';
+import { spawnSync } from 'node:child_process';
+import { ensureDir, resolvePython } from '../lib.mjs';
+import { makeTinyPng } from '../case-intake.mjs';
+
+const PY_BUILD = `
+import sys, zipfile, os
+target = sys.argv[1]
+
+doc_xml = '''<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
+<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"
+            xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships"
+            xmlns:wp="http://schemas.openxmlformats.org/drawingml/2006/wordprocessingDrawing"
+            xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main"><w:body>
+<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr><w:r><w:t>IC 计算机科学 考前冲刺案例</w:t></w:r></w:p>
+<w:p><w:r><w:t>客户在读帝国理工本科二年级,期末前两周找到我们,担心挂科。</w:t></w:r></w:p>
+<w:p><w:r><w:t>跟进状态:异议处理中。客户觉得贵,家长不放心老师。</w:t></w:r></w:p>
+<w:p><w:r><w:t>结果:期末提分 18 分,家长很满意,已续课。</w:t></w:r></w:p>
+<w:p><w:r><w:drawing><wp:inline><a:graphic><a:blip r:embed="rId5"/></a:graphic></wp:inline></w:drawing></w:r></w:p>
+<w:p><w:hyperlink r:id="rId9"><w:r><w:t>查看成绩单</w:t></w:r></w:hyperlink></w:p>
+</w:body></w:document>'''
+
+doc_rels = '''<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
+<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">
+<Relationship Id="rId5" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image" Target="media/image1.png"/>
+<Relationship Id="rId9" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/hyperlink" Target="https://example.invalid/score" TargetMode="External"/>
+</Relationships>'''
+
+pres_xml = '''<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
+<p:presentation xmlns:p="http://schemas.openxmlformats.org/presentationml/2006/main"
+                xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">
+<p:sldIdLst><p:sldId id="256" r:id="rId1"/><p:sldId id="257" r:id="rId2"/></p:sldIdLst></p:presentation>'''
+
+pres_rels = '''<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
+<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">
+<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/slide" Target="slides/slide1.xml"/>
+<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/slide" Target="slides/slide2.xml"/>
+</Relationships>'''
+
+def slide(title, body, with_pic, rid):
+    pic = ('<p:pic><p:blipFill><a:blip r:embed="%s"/></p:blipFill></p:pic>' % rid) if with_pic else ''
+    return '''<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
+<p:sld xmlns:p="http://schemas.openxmlformats.org/presentationml/2006/main"
+       xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main"
+       xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships"><p:cSld><p:spTree>
+<p:sp><p:nvSpPr><p:nvPr><p:ph type="title"/></p:nvPr></p:nvSpPr><p:txBody><a:p><a:r><a:t>%s</a:t></a:r></a:p></p:txBody></p:sp>
+<p:sp><p:nvSpPr/><p:txBody><a:p><a:r><a:t>%s</a:t></a:r></a:p></p:txBody></p:sp>
+%s
+</p:spTree></p:cSld></p:sld>''' % (title, body, pic)
+
+slide_rels = '''<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
+<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">
+<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image" Target="../media/image1.png"/>
+</Relationships>'''
+
+png = bytes.fromhex(
+    "89504e470d0a1a0a0000000d49484452000000010000000108060000001f15c489"
+    "0000000a49444154789c6300010000050001"
+    "0d0a2db40000000049454e44ae426082")
+
+with zipfile.ZipFile(os.path.join(target, 'case.docx'), 'w') as zf:
+    zf.writestr('[Content_Types].xml', '<Types/>')
+    zf.writestr('word/document.xml', doc_xml)
+    zf.writestr('word/_rels/document.xml.rels', doc_rels)
+    zf.writestr('word/media/image1.png', png)
+
+with zipfile.ZipFile(os.path.join(target, 'deck.pptx'), 'w') as zf:
+    zf.writestr('[Content_Types].xml', '<Types/>')
+    zf.writestr('ppt/presentation.xml', pres_xml)
+    zf.writestr('ppt/_rels/presentation.xml.rels', pres_rels)
+    zf.writestr('ppt/slides/slide1.xml', slide('案例合集', '英国本科 方案推荐中 导师匹配', True, 'rId2'))
+    zf.writestr('ppt/slides/slide2.xml', slide('结果与反馈', '期末提分 出分反馈 家长认可', False, 'rId2'))
+    zf.writestr('ppt/slides/_rels/slide1.xml.rels', slide_rels)
+    zf.writestr('ppt/media/image1.png', png)
+
+print('ok')
+`;
+
+/**
+ * 生成夹具。
+ * @param {string} dir 目标目录
+ * @returns {{docx:string, pptx:string, images:string[]}}
+ */
+export function makeFixtures(dir) {
+  const target = ensureDir(dir);
+  const python = resolvePython();
+  const run = spawnSync(python, ['-c', PY_BUILD, target], { encoding: 'utf8' });
+  if (run.status !== 0) {
+    throw new Error(`夹具生成失败:${run.stderr || run.stdout}`);
+  }
+
+  const images = [];
+  // 九宫格:9 张同尺寸、同一批连发(时间戳相邻)→ 应判为同一组 3x3
+  const base = 1_700_000_000_000;
+  for (let i = 0; i < 9; i++) {
+    const file = path.join(target, `grid-${String(i + 1).padStart(2, '0')}.png`);
+    fs.writeFileSync(file, makeTinyPng(400, 400));
+    fs.utimesSync(file, new Date(base + i * 1000), new Date(base + i * 1000));
+    images.push(file);
+  }
+  // 另一批:不同尺寸 + 不同时间 → 不应并进九宫格
+  const other = path.join(target, 'other-01.png');
+  fs.writeFileSync(other, makeTinyPng(1600, 900));
+  fs.utimesSync(other, new Date(base - 3 * 86_400_000), new Date(base - 3 * 86_400_000));
+  images.push(other);
+
+  return { docx: path.join(target, 'case.docx'), pptx: path.join(target, 'deck.pptx'), images };
+}

+ 176 - 0
scripts/tests/grouping.test.mjs

@@ -0,0 +1,176 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 分组 / 顺序 单元测试:`node --test`
+ */
+
+import test from 'node:test';
+import assert from 'node:assert/strict';
+import {
+  aspectSimilarity, groupMaterials, inferLayout, orderWithinGroup,
+  textSimilarity, timeProximity, applyGroupOverrides, DEFAULT_ORDERING_RULE,
+} from '../grouping.mjs';
+import { reorderAssets } from '../lib.mjs';
+import { makeTinyPng } from '../case-intake.mjs';
+
+const BASE = 1_700_000_000_000;
+
+function grid(count, overrides = {}) {
+  return Array.from({ length: count }, (_, i) => ({
+    kind: 'image',
+    localPath: `/tmp/g${String(i + 1).padStart(2, '0')}.png`,
+    width: 400,
+    height: 400,
+    mtimeMs: BASE + i * 1000,
+    ocrText: '帝国理工 计算机科学 考前冲刺 提分',
+    ...overrides,
+  }));
+}
+
+test('textSimilarity:完全相同为 1,无关文本接近 0', () => {
+  assert.equal(textSimilarity('帝国理工提分', '帝国理工提分'), 1);
+  assert.ok(textSimilarity('帝国理工提分案例', '帝国理工提分反馈') > 0.3);
+  assert.ok(textSimilarity('帝国理工提分', '完全无关的另一段内容') < 0.2);
+});
+
+test('aspectSimilarity:同比例为 1,缺失尺寸返回 null', () => {
+  assert.equal(aspectSimilarity({ width: 400, height: 400 }, { width: 800, height: 800 }), 1);
+  assert.equal(aspectSimilarity({ width: 400, height: 400 }, {}), null);
+  assert.ok(aspectSimilarity({ width: 400, height: 400 }, { width: 1600, height: 900 }) < 0.8);
+});
+
+test('timeProximity:窗口内衰减,窗口外为 0', () => {
+  const a = { mtimeMs: BASE };
+  assert.equal(timeProximity(a, { mtimeMs: BASE }, 60_000), 1);
+  assert.ok(timeProximity(a, { mtimeMs: BASE + 30_000 }, 60_000) === 0.5);
+  assert.equal(timeProximity(a, { mtimeMs: BASE + 120_000 }, 60_000), 0);
+});
+
+test('inferLayout:9→3x3 / 4→2x2 / 6→3x2 / 1→1x1', () => {
+  assert.equal(inferLayout(9), '3x3');
+  assert.equal(inferLayout(4), '2x2');
+  assert.equal(inferLayout(6), '3x2');
+  assert.equal(inferLayout(1), '1x1');
+});
+
+test('九宫格:9 张同批图 → 1 组,layout 3x3,order 1..9(左上→右下)', () => {
+  const groups = groupMaterials(grid(9));
+  assert.equal(groups.length, 1);
+  assert.equal(groups[0].layout, '3x3');
+  assert.equal(groups[0].count, 9);
+  assert.equal(groups[0].orderingRule, DEFAULT_ORDERING_RULE);
+  assert.deepEqual(groups[0].items.map((i) => i.order), [1, 2, 3, 4, 5, 6, 7, 8, 9]);
+  assert.deepEqual(groups[0].items.map((i) => i.groupId), Array(9).fill(groups[0].groupId));
+});
+
+test('6 张连发图 → 1 组 3x2(同事连发 6 张是同一个素材)', () => {
+  const groups = groupMaterials(grid(6));
+  assert.equal(groups.length, 1);
+  assert.equal(groups[0].layout, '3x2');
+});
+
+test('不同批次不误并组:尺寸/时间/内容都不同 → 2 组', () => {
+  const items = [
+    ...grid(3),
+    {
+      kind: 'image', localPath: '/tmp/z1.png', width: 1600, height: 900,
+      mtimeMs: BASE - 5 * 86_400_000, ocrText: '香港大学 申诉',
+    },
+  ];
+  const groups = groupMaterials(items);
+  assert.equal(groups.length, 2);
+  assert.deepEqual(groups.map((g) => g.count).sort(), [1, 3]);
+});
+
+test('尺寸 + 内容两个信号一致 → 并组(时间是弱信号,不足即被 2/3 规则覆盖)', () => {
+  const items = grid(2);
+  items[1] = { ...items[1], mtimeMs: BASE + 5 * 86_400_000 };
+  const groups = groupMaterials(items);
+  assert.equal(groups.length, 1, 'aspect + text 两个信号一致即同组');
+});
+
+test('只有 1 个信号一致 → 不并组(避免把不同素材误并成一组)', () => {
+  const items = [
+    ...grid(1),
+    {
+      kind: 'image', localPath: '/tmp/other.png', width: 1600, height: 900,
+      mtimeMs: BASE + 5 * 86_400_000, ocrText: '帝国理工 计算机科学 考前冲刺 提分',
+    },
+  ];
+  assert.equal(groupMaterials(items).length, 2, '仅内容相似、尺寸与时间都不同 → 判为两组');
+});
+
+test('裸图无 OCR:靠尺寸 + 时间仍能成组', () => {
+  const items = grid(4, { ocrText: '' });
+  const groups = groupMaterials(items);
+  assert.equal(groups.length, 1);
+  assert.equal(groups[0].count, 4);
+});
+
+test('orderWithinGroup:按时间升序还原发送顺序;无时间时文件名自然序', () => {
+  const withTime = [
+    { localPath: '/tmp/b.png', mtimeMs: BASE + 2000 },
+    { localPath: '/tmp/a.png', mtimeMs: BASE + 1000 },
+  ];
+  assert.deepEqual(orderWithinGroup(withTime).map((i) => i.localPath), ['/tmp/a.png', '/tmp/b.png']);
+
+  const noTime = [
+    { localPath: '/tmp/img10.png' },
+    { localPath: '/tmp/img2.png' },
+  ];
+  assert.deepEqual(orderWithinGroup(noTime).map((i) => i.localPath), ['/tmp/img2.png', '/tmp/img10.png']);
+});
+
+test('reorderAssets:把第 9 张移到第 1 位后重新编号 1..9,元素完整', () => {
+  const groups = groupMaterials(grid(9));
+  const before = groups[0].items;
+  const after = reorderAssets(before, 9, 1);
+  assert.equal(after[0].localPath, '/tmp/g09.png');
+  assert.deepEqual(after.map((a) => a.order), [1, 2, 3, 4, 5, 6, 7, 8, 9]);
+  assert.equal(after.length, 9);
+  assert.deepEqual(new Set(after.map((a) => a.localPath)).size, 9);
+});
+
+test('reorderAssets:越界报错', () => {
+  assert.throws(() => reorderAssets(grid(3), 5, 1), /越界/);
+});
+
+test('applyGroupOverrides:人工指定分组与顺序优先', () => {
+  const groups = groupMaterials(grid(4));
+  const overridden = applyGroupOverrides(groups, {
+    groups: [{
+      groupId: 'manual',
+      items: [
+        { localPath: '/tmp/g04.png', order: 1 },
+        { localPath: '/tmp/g01.png', order: 2 },
+      ],
+    }],
+  });
+  const manual = overridden.find((g) => g.groupId === 'manual');
+  assert.ok(manual);
+  assert.deepEqual(manual.items.map((i) => i.localPath), ['/tmp/g04.png', '/tmp/g01.png']);
+  assert.deepEqual(manual.reasons, ['manual-override']);
+});
+
+test('makeTinyPng 产出可被 image-size 识别的合法 PNG', async () => {
+  const fs = await import('node:fs');
+  const os = await import('node:os');
+  const path = await import('node:path');
+  const { readImageInfo } = await import('../image-size.mjs');
+  const file = path.join(fs.mkdtempSync(path.join(os.tmpdir(), 'caseget-')), 'p.png');
+  fs.writeFileSync(file, makeTinyPng(320, 240));
+  const info = readImageInfo(file);
+  assert.equal(info.ok, true);
+  assert.equal(info.ext, 'png');
+  assert.equal(info.width, 320);
+  assert.equal(info.height, 240);
+});

+ 220 - 0
scripts/vision-bridge.mjs

@@ -0,0 +1,220 @@
+// Copyright (c) 未来飞马
+//
+// This Source Code Form is subject to the terms of the Mozilla Public
+// License, v. 2.0. If a copy of the MPL was not distributed with this
+// file, You can obtain one at https://mozilla.org/MPL/2.0/.
+//
+// Trademark Notice:
+// The MPL-2.0 license grants copyright permissions for source code only.
+// It does NOT grant any rights to use trademarks including "未来飞马",
+// "Harness Loop", "RSI", and associated slogan "让AI进化提前发生,让AI落地快人一步".
+// Any use of these trademarks requires separate written permission.
+/**
+ * 图片理解桥(复用兄弟技能 skill-vision)
+ *
+ * 职责:把一张图交给 skill-vision 的 analyze(),产出结构化结果:
+ *   - ocrText        图里出现的文字(尽量逐字)
+ *   - description    画面/内容语义描述
+ *   - usageSuggestion 能不能作为素材、适合发给谁
+ *   - label          给素材起一个短标签
+ *   - canBeMaterial  是否可作为可发送素材(false → role=description)
+ *   - piiHints       疑似 PII 片段(姓名/手机/微信/邮箱/头像),仅记录不外发
+ *
+ * analyze() 有两种返回:
+ *   1) 宿主多模态:{ provider:'host', instruction, imagePath } —— 由本脚本打印指令,
+ *      由宿主 Agent 用自己的 Read 工具读图后按提示词产出 JSON(不调 Fmode API);
+ *   2) 网关:{ provider:'fmode', raw, parsed } —— 直接拿到结构化 JSON。
+ * 两条路径的输出契约一致,调用方(SKILL.md 工作流)无感知。
+ *
+ * 绝不打印 token、绝不把图片内容写进仓库。
+ */
+
+import fs from 'node:fs';
+import path from 'node:path';
+import { parseArgs } from 'node:util';
+import {
+  readJson, writeJson, ensureDir, resolveSiblingScript, truncate, shortHash,
+} from './lib.mjs';
+
+const VISION_REL = path.join('skill-vision', 'scripts', 'vision-client.mjs');
+
+const SYSTEM_PROMPT = [
+  '你是案例库采集助手。你看到的图片来自留学咨询/课程辅导机构的真实沟通素材(聊天截图、成绩单、反馈截图、海报、笔记等)。',
+  '你要做两件事:① 如实 OCR 出图里的文字;② 判断这张图能不能作为「可发给客户的案例素材」,并给出使用建议。',
+  '严格输出 JSON,不要输出多余文字。字段:',
+  '{"ocrText":"图内文字,逐字,保留换行","description":"画面与内容说明,2-4 句",',
+  ' "label":"不超过 16 字的短标签","usageSuggestion":"什么时候发给什么客户,一句话",',
+  ' "canBeMaterial":true, "role":"material|description",',
+  ' "piiHints":[{"field":"姓名|手机号|邮箱|微信号|头像|其它","snippet":"疑似片段"}]}',
+  '判据:canBeMaterial=false 的典型情况——纯说明性配图、logo、无信息量的装饰图、与课程辅导无关。',
+  'piiHints 只记录疑似片段,不要脑补补全。',
+].join('\n');
+
+export function buildUserPrompt(context = {}) {
+  const lines = [
+    '请分析这张图片,按系统提示词输出 JSON。',
+    context.index ? `这是同一批素材中的第 ${context.index} 张${context.total ? `(共 ${context.total} 张)` : ''}。` : '',
+    context.batchHint ? `同批其它图的初步内容:${truncate(context.batchHint, 300)}。请据此判断这张图在整组里的作用。` : '',
+    context.extra ? String(context.extra) : '',
+  ].filter(Boolean);
+  return lines.join('\n');
+}
+
+/** 从 vision 返回里取出结构化对象(宿主路径拿不到时回落为一个空壳)。 */
+export function normalizeVisionResult(result, fallbackLabel = '') {
+  const parsed = (result && result.parsed) || null;
+  if (!parsed) {
+    const needsHostRead = Boolean(result && result.instruction);
+    return {
+      ok: false,
+      provider: result && result.provider,
+      model: result && result.model,
+      ocrText: '',
+      description: '',
+      label: fallbackLabel,
+      usageSuggestion: '',
+      // 还没真正分析过:不下结论,role 留空交给调用方保持默认(material)
+      canBeMaterial: null,
+      role: '',
+      piiHints: [],
+      needsHostRead,
+      instruction: result && result.instruction,
+      raw: (result && result.raw) || '',
+      error: (result && result.error) || (needsHostRead ? 'HOST_READ_PENDING' : 'NO_PARSED_OUTPUT'),
+    };
+  }
+  const role = String(parsed.role || (parsed.canBeMaterial === false ? 'description' : 'material')).toLowerCase();
+  return {
+    ok: true,
+    provider: result && result.provider,
+    model: result && result.model,
+    ocrText: String(parsed.ocrText || ''),
+    description: String(parsed.description || ''),
+    label: String(parsed.label || fallbackLabel || ''),
+    usageSuggestion: String(parsed.usageSuggestion || ''),
+    canBeMaterial: parsed.canBeMaterial !== false && role === 'material',
+    role: role === 'description' ? 'description' : 'material',
+    piiHints: Array.isArray(parsed.piiHints) ? parsed.piiHints.filter((h) => h && h.field) : [],
+    needsHostRead: false,
+    instruction: '',
+    raw: result && result.raw ? String(result.raw) : '',
+    error: null,
+  };
+}
+
+/**
+ * 分析一张图(或一组图)。
+ * @returns {Promise<{available:boolean, source:string, skillPath:string|null, results:object[], error:string|null}>}
+ */
+export async function analyzeImages(items, options = {}) {
+  const visionPath = resolveSiblingScript('CASE_VISION_SCRIPT', VISION_REL);
+  if (!visionPath) {
+    return {
+      available: false,
+      source: 'missing',
+      skillPath: null,
+      results: [],
+      error: `未找到 skill-vision(期望 ${VISION_REL})。请先安装:npx skill-vision@latest install`,
+    };
+  }
+
+  const analyze = (await import(pathToFileUrl(visionPath))).analyze;
+  const results = [];
+  for (let i = 0; i < items.length; i++) {
+    const item = items[i];
+    const context = { index: i + 1, total: items.length, batchHint: options.batchHint || '' };
+    let result;
+    try {
+      result = await analyze({
+        imagePath: item.localPath || undefined,
+        imageUrl: !item.localPath ? item.url || undefined : undefined,
+        systemPrompt: SYSTEM_PROMPT,
+        userPrompt: buildUserPrompt(context),
+        model: options.model || undefined,
+        maxTokens: options.maxTokens || 1800,
+      });
+    } catch (error) {
+      result = { provider: 'error', parsed: null, error: error.message };
+    }
+    const normalized = normalizeVisionResult(result, item.label || `素材${i + 1}`);
+    results.push({ file: item.localPath || item.url || '', index: i + 1, ...normalized });
+  }
+
+  const provider = results.find((r) => r.ok)?.provider || results[0]?.provider || 'unknown';
+  const needsHost = results.some((r) => r.needsHostRead);
+  return {
+    available: true,
+    source: provider,
+    skillPath: visionPath,
+    needsHostRead: needsHost,
+    results,
+    error: results.every((r) => !r.ok) ? '全部图片分析未产出结构化结果' : null,
+  };
+}
+
+function pathToFileUrl(p) {
+  return new URL(`file://${p.split(path.sep).join('/')}`).href;
+}
+
+// ---------------------------------------------------------------------------
+// CLI
+// ---------------------------------------------------------------------------
+
+async function main() {
+  const { values } = parseArgs({
+    options: {
+      image: { type: 'string', multiple: true, default: [] },
+      'image-list': { type: 'string' },
+      out: { type: 'string' },
+      model: { type: 'string' },
+      'batch-hint': { type: 'string' },
+      help: { type: 'boolean', default: false },
+    },
+    allowPositionals: true,
+  });
+
+  if (values.help) {
+    process.stdout.write([
+      'skill-case-get / vision-bridge — 逐图 OCR + 语义理解(复用 skill-vision)',
+      '',
+      '  node vision-bridge.mjs --image a.jpg --image b.jpg [--batch-hint "同批都是考前冲刺截图"] [--out vision.json]',
+      '  node vision-bridge.mjs --image-list batch.json [--out vision.json]   # batch.json: {items:[{localPath|url,label}], batchHint}',
+      '',
+      '说明:若宿主(FmodeCode / Claude Code)自带多模态模型,analyze() 会返回读图指令,',
+      '      此时输出里的 needsHostRead=true,由宿主 Agent 用自己的 Read 工具读图后补全。',
+    ].join('\n'));
+    return 0;
+  }
+
+  const items = values.image.map((p) => ({ localPath: p }));
+  if (values['image-list']) {
+    const bundle = readJson(values['image-list'], {});
+    for (const item of bundle.items || []) items.push(item);
+    if (bundle.batchHint && !values['batch-hint']) values['batch-hint'] = bundle.batchHint;
+  }
+  if (!items.length) {
+    process.stderr.write('至少需要一个 --image 或 --image-list\n');
+    return 2;
+  }
+
+  const payload = await analyzeImages(items, { model: values.model, batchHint: values['batch-hint'] });
+  const text = JSON.stringify(payload, null, 2);
+  if (values.out) {
+    ensureDir(path.dirname(path.resolve(values.out)));
+    writeJson(path.resolve(values.out), payload);
+  } else {
+    process.stdout.write(`${text}\n`);
+  }
+  if (!payload.available) return 3;
+  return 0;
+}
+
+const invokedDirectly = process.argv[1] && path.resolve(process.argv[1]).endsWith(path.join('scripts', 'vision-bridge.mjs'));
+if (invokedDirectly) {
+  main().then((code) => process.exit(code)).catch((error) => {
+    process.stderr.write(`vision-bridge 失败:${error.message}\n`);
+    process.exit(1);
+  });
+}
+
+export { SYSTEM_PROMPT };

+ 34 - 0
skill-package-manifest.json

@@ -0,0 +1,34 @@
+{
+  "name": "skill-case-get",
+  "version": "1.0.0",
+  "description": "FmodeCode / Claude Code 独立技能包:案例采集 / 拆解 / 归档。把连续多图(九宫格)、Word(.docx)、PPT(.pptx)、图文混合、裸图、视频等异构素材拆解为「案例信息」与「素材数组 materialAssets[]」,做合规判断(riskFlags / privacyFindings)、打标(标签字典 + 学校别名归一),最后经云函数 caseSubmit 入库到待审清单(reviewStatus=pending, readyForUse=false)。.docx/.pptx 用 python3 标准库 zipfile + ElementTree 解包,无第三方依赖;图片复用 skill-vision,音视频复用 skill-listen。",
+  "plugin": "skill-case-get",
+  "skills": [
+    "skill-case-get"
+  ],
+  "entrySkill": "skill-case-get",
+  "npmPackage": "skill-case-get",
+  "smokeCommand": "npm run smoke",
+  "installCommand": "npx skill-case-get@latest install",
+  "workspaceInstallCommand": "npx skill-case-get@latest workspace",
+  "workspaceSkillPath": ".claude/skills/skill-case-get/SKILL.md",
+  "globalSkillPath": "%USERPROFILE%/.claude/skills/skill-case-get/SKILL.md",
+  "installHint": "工作区安装:npx skill-case-get@latest workspace,会写入 ./.claude/skills/skill-case-get/。用户级安装:npx skill-case-get@latest install,会写入 ~/.claude/skills/skill-case-get/。图片理解复用 skill-vision(analyze()),音视频转写复用 skill-listen(listen-runner.mjs),两者与云函数 token 都从本地配置自动解析,凭据仅服务端,客户端不落盘任何密钥。",
+  "requires": {
+    "python3": ">=3.8(.docx/.pptx 解包,仅标准库)",
+    "node": ">=18",
+    "optional": {
+      "ffmpeg": "视频抽帧/抽音轨",
+      "ffprobe": "音视频时长探测"
+    }
+  },
+  "cloudFunctions": {
+    "submit": "caseSubmit",
+    "pendingList": "casePendingList",
+    "approve": "caseApprove",
+    "search": "caseSearch"
+  },
+  "contract": "SHARED-CONTRACT.md(字段名逐字对齐;枚举小写)",
+  "license": "MPL-2.0",
+  "copyright": "Copyright (c) 2026 未来飞马 Fmode"
+}