diff --git a/blogwatcher-daily/blogwatcher.db b/blogwatcher-daily/blogwatcher.db index 045dc77..f53fc17 100644 Binary files a/blogwatcher-daily/blogwatcher.db and b/blogwatcher-daily/blogwatcher.db differ diff --git a/course-transcript-notes/SKILL.md b/course-transcript-notes/SKILL.md new file mode 100644 index 0000000..1318cb9 --- /dev/null +++ b/course-transcript-notes/SKILL.md @@ -0,0 +1,103 @@ +--- +name: course-transcript-notes +description: Use when 课程转录文件批量提炼为结构化 markdown 笔记。 +metadata: + author: Hermes Agent + last_updated: "2026-09-19" +--- + +# Course Transcript Notes — 批量课程转录提炼 + +触发器:**"帮我整理课程笔记"** / **"提炼转录文件"** / 用户对 `knowledgebase/course/` 下某课程目录要求结构化总结。 + +## 工作流 + +1. **扫描目录结构**: + ```bash + find -name "*.txt" | sort + ``` + 识别章节分组(通常按子目录或文件名前缀编号)。 + +2. **并行委派子任务**(每章节一个 subagent): + - 用 `delegate_task` 并行派发,每个 subagent 读取该章节所有 txt 文件 + - 要求输出格式统一:`## Chapter N — [number]. [title]` + Core Insights / Detailed Notes / Action Items / Key Terms + - 设定 `output_schema: {"properties": {"markdown": {"type": "string"}}, "required": ["markdown"]}` + +3. **接收结果后立即落盘**(关键步骤): + - delegate_task 完成后,结果直接返回在对话中 — 已经是解析好的 markdown 字符串 + - 立即将每个 subagent 返回的 markdown 写入独立临时文件:`/tmp/ch_notes.md` + - **不要尝试从 subagent-summary 或 log 文件重新提取** — 这些文件可能有 JSON 编码问题或会被清理 + +4. **组装最终文件并双写**: + ```python + # 简单字符串拼接,不做任何 JSON 处理 + chapters = [] + for i in range(num_chapters): + with open(f'/tmp/ch{i}_notes.md') as f: + chapters.append(f.read()) + full = header + '\n\n---\n\n'.join(chapters) + + # 写入两个位置 + # 1. 原始 course 目录 + with open(output_path, 'w') as f: + f.write(full) + + # 2. nexus knowledgebase 目录 + import os, shutil + nexus_dir = os.path.expanduser('~/Workspace/nexus/knowledgebase/course-note') + os.makedirs(nexus_dir, exist_ok=True) + nexus_path = os.path.join(nexus_dir, os.path.basename(output_path)) + shutil.copy2(output_path, nexus_path) + print(f"Written to: {output_path}") + print(f"Copied to: {nexus_path}") + ``` + +5. **立即验证完整性**(在告知用户完成之前): + ```bash + # 统计每个章节的视频数 + for ch in 1 2 3 4 5; do + count=$(grep -c "^## Chapter $ch —" ) + echo "Chapter $ch: $count videos" + done + ``` + - 确认每个章节的视频数 == 该章节的 txt 文件数 + - 不匹配则定位缺失章节,**直接读取原始转录文件重新生成**(不要重试 subagent) + +6. **回退策略**(如果 subagent 结果丢失或格式异常): + - 直接从原始 txt 文件读取内容 + - 在本地生成结构化摘要 + - 写入最终文件 + +## 坑 + +- **不要从 log/transcript 文件重新提取 subagent 输出**:`delegate_task` 返回的结果在对话中已经是可用的 markdown。如果当时没落盘,不要试图从 `subagent-summary-*.txt` 或 `live/*.log` 文件中用正则/JSON 解析恢复 — 这些文件中的 JSON 可能包含控制字符导致解析失败,且文件会被清理。正确做法:直接重新读取原始 txt 转录文件,本地生成摘要。 +- **验证必须按章节逐个计数**:`grep -c "^## Chapter"` 只给总数,掩盖单章缺失。必须按章节分别 `grep -c "^## Chapter N —"` 并与该章 txt 文件数对比。 +- **subagent-summary 文件会被清理**:delegation cache 中的 `subagent-summary-*.txt` 不持久。收到结果后必须立即写入 `/tmp/` 或目标目录,不要依赖后续重新读取。 +- **组装时不要做 JSON unescape 修复**:如果某章节内容看起来格式异常(全在一行),正确做法是重新读取原始转录文件生成该章节,不要用 `json.loads` 或 `.replace('\\n', '\n')` 修复整个文件 — 这会丢失其他章节。 +- **大批量章节超出 subagent 并发限制**:`delegate_task` 最多 10 个并行。超过时分批派发,每批完成后立即落盘再派下一批。 +- **输出文件体积**:39 个视频提炼后约 70-80KB / 1500+ 行。如果用户通过 `read_file` 查看,超过 100K 字符会被截断 — 保持简洁,每个视频控制在 1.5-2KB。 + +## 输出格式 + +每个视频一个 section: + +```markdown +## Chapter N — [video_number]. [video_title] + +**Core Insights** +- bullet 1 +- bullet 2 +- bullet 3 + +**Detailed Notes** +*Topic grouping* +- detail point + +**Action Items** +- actionable step + +**Key Terms/Concepts** +- **Term**: definition +``` + +语言:英语(除非用户要求中文)。保持简洁,跳过开场白和重复内容。