5.1 KiB
5.1 KiB
name, description, metadata
| name | description | metadata | ||||
|---|---|---|---|---|---|---|
| course-transcript-notes | Use when 课程转录文件批量提炼为结构化 markdown 笔记。 |
|
Course Transcript Notes — 批量课程转录提炼
触发器:"帮我整理课程笔记" / "提炼转录文件" / 用户对 knowledgebase/course/ 下某课程目录要求结构化总结。
工作流
-
扫描目录结构:
find <course_dir> -name "*.txt" | sort识别章节分组(通常按子目录或文件名前缀编号)。
-
并行委派子任务(每章节一个 subagent):
- 用
delegate_task并行派发,每个 subagent 读取该章节所有 txt 文件 - 要求输出格式统一:
## Chapter N — [number]. [title]+ Core Insights / Detailed Notes / Action Items / Key Terms - 设定
output_schema: {"properties": {"markdown": {"type": "string"}}, "required": ["markdown"]}
- 用
-
接收结果后立即落盘(关键步骤):
- delegate_task 完成后,结果直接返回在对话中 — 已经是解析好的 markdown 字符串
- 立即将每个 subagent 返回的 markdown 写入独立临时文件:
/tmp/ch<N>_notes.md - 不要尝试从 subagent-summary 或 log 文件重新提取 — 这些文件可能有 JSON 编码问题或会被清理
-
组装最终文件并双写:
# 简单字符串拼接,不做任何 JSON 处理 chapters = [] for i in range(num_chapters): with open(f'/tmp/ch{i}_notes.md') as f: chapters.append(f.read()) full = header + '\n\n---\n\n'.join(chapters) # 写入两个位置 # 1. 原始 course 目录 with open(output_path, 'w') as f: f.write(full) # 2. nexus knowledgebase 目录 import os, shutil nexus_dir = os.path.expanduser('~/Workspace/nexus/knowledgebase/course-notes') os.makedirs(nexus_dir, exist_ok=True) nexus_path = os.path.join(nexus_dir, os.path.basename(output_path)) shutil.copy2(output_path, nexus_path) print(f"Written to: {output_path}") print(f"Copied to: {nexus_path}") -
立即验证完整性(在告知用户完成之前):
# 统计每个章节的视频数 for ch in 1 2 3 4 5; do count=$(grep -c "^## Chapter $ch —" <output_file>) echo "Chapter $ch: $count videos" done- 确认每个章节的视频数 == 该章节的 txt 文件数
- 不匹配则定位缺失章节,直接读取原始转录文件重新生成(不要重试 subagent)
-
回退策略(如果 subagent 结果丢失或格式异常):
- 直接从原始 txt 文件读取内容
- 在本地生成结构化摘要
- 写入最终文件
坑
- 不要从 log/transcript 文件重新提取 subagent 输出:
delegate_task返回的结果在对话中已经是可用的 markdown。如果当时没落盘,不要试图从subagent-summary-*.txt或live/*.log文件中用正则/JSON 解析恢复 — 这些文件中的 JSON 可能包含控制字符导致解析失败,且文件会被清理。正确做法:直接重新读取原始 txt 转录文件,本地生成摘要。 - 验证必须按章节逐个计数:
grep -c "^## Chapter"只给总数,掩盖单章缺失。必须按章节分别grep -c "^## Chapter N —"并与该章 txt 文件数对比。 - subagent-summary 文件会被清理:delegation cache 中的
subagent-summary-*.txt不持久。收到结果后必须立即写入/tmp/或目标目录,不要依赖后续重新读取。 - 结果过大落入 spillover 缓存:
delegate_task返回超 ~100KB 时结果落在~/.hermes/cache/spillover/call_*.txt(JSON 数组,每个 result 的 summary 可能是双重编码 JSON 或原始字符串)。此时可安全恢复:外层 json.loads 后在每个 summary 中找{"markdown"位置用json.JSONDecoder().raw_decode提取并 unescape,得到的就是完整 markdown(别对整个文件做 replace 修复)。比重新跑 subagent 快得多。 - 转录文件可能是 .srt 而非 .txt:每段带序号+时间戳行,subagent 的 context 里要写明剥离序号/时间戳、按对话文本处理。
- 组装时不要做 JSON unescape 修复:如果某章节内容看起来格式异常(全在一行),正确做法是重新读取原始转录文件生成该章节,不要用
json.loads或.replace('\\n', '\n')修复整个文件 — 这会丢失其他章节。 - 大批量章节超出 subagent 并发限制:
delegate_task最多 10 个并行。超过时分批派发,每批完成后立即落盘再派下一批。 - 输出文件体积:39 个视频提炼后约 70-80KB / 1500+ 行。如果用户通过
read_file查看,超过 100K 字符会被截断 — 保持简洁,每个视频控制在 1.5-2KB。
输出格式
每个视频一个 section:
## Chapter N — [video_number]. [video_title]
**Core Insights**
- bullet 1
- bullet 2
- bullet 3
**Detailed Notes**
*Topic grouping*
- detail point
**Action Items**
- actionable step
**Key Terms/Concepts**
- **Term**: definition
语言:英语(除非用户要求中文)。保持简洁,跳过开场白和重复内容。