Files
atlas/course-transcript-notes/SKILL.md

106 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: course-transcript-notes
description: Use when 课程转录文件批量提炼为结构化 markdown 笔记。
metadata:
author: Hermes Agent
last_updated: "2026-09-19"
---
# Course Transcript Notes — 批量课程转录提炼
触发器:**"帮我整理课程笔记"** / **"提炼转录文件"** / 用户对 `knowledgebase/course/` 下某课程目录要求结构化总结。
## 工作流
1. **扫描目录结构**:
```bash
find <course_dir> -name "*.txt" | sort
```
识别章节分组(通常按子目录或文件名前缀编号)。
2. **并行委派子任务**(每章节一个 subagent):
- 用 `delegate_task` 并行派发,每个 subagent 读取该章节所有 txt 文件
- 要求输出格式统一:`## Chapter N — [number]. [title]` + Core Insights / Detailed Notes / Action Items / Key Terms
- 设定 `output_schema: {"properties": {"markdown": {"type": "string"}}, "required": ["markdown"]}`
3. **接收结果后立即落盘**(关键步骤):
- delegate_task 完成后,结果直接返回在对话中 — 已经是解析好的 markdown 字符串
- 立即将每个 subagent 返回的 markdown 写入独立临时文件:`/tmp/ch<N>_notes.md`
- **不要尝试从 subagent-summary 或 log 文件重新提取** — 这些文件可能有 JSON 编码问题或会被清理
4. **组装最终文件并双写**:
```python
# 简单字符串拼接,不做任何 JSON 处理
chapters = []
for i in range(num_chapters):
with open(f'/tmp/ch{i}_notes.md') as f:
chapters.append(f.read())
full = header + '\n\n---\n\n'.join(chapters)
# 写入两个位置
# 1. 原始 course 目录
with open(output_path, 'w') as f:
f.write(full)
# 2. nexus knowledgebase 目录
import os, shutil
nexus_dir = os.path.expanduser('~/Workspace/nexus/knowledgebase/course-notes')
os.makedirs(nexus_dir, exist_ok=True)
nexus_path = os.path.join(nexus_dir, os.path.basename(output_path))
shutil.copy2(output_path, nexus_path)
print(f"Written to: {output_path}")
print(f"Copied to: {nexus_path}")
```
5. **立即验证完整性**(在告知用户完成之前):
```bash
# 统计每个章节的视频数
for ch in 1 2 3 4 5; do
count=$(grep -c "^## Chapter $ch —" <output_file>)
echo "Chapter $ch: $count videos"
done
```
- 确认每个章节的视频数 == 该章节的 txt 文件数
- 不匹配则定位缺失章节,**直接读取原始转录文件重新生成**(不要重试 subagent)
6. **回退策略**(如果 subagent 结果丢失或格式异常):
- 直接从原始 txt 文件读取内容
- 在本地生成结构化摘要
- 写入最终文件
## 坑
- **不要从 log/transcript 文件重新提取 subagent 输出**:`delegate_task` 返回的结果在对话中已经是可用的 markdown。如果当时没落盘,不要试图从 `subagent-summary-*.txt` 或 `live/*.log` 文件中用正则/JSON 解析恢复 — 这些文件中的 JSON 可能包含控制字符导致解析失败,且文件会被清理。正确做法:直接重新读取原始 txt 转录文件,本地生成摘要。
- **验证必须按章节逐个计数**:`grep -c "^## Chapter"` 只给总数,掩盖单章缺失。必须按章节分别 `grep -c "^## Chapter N —"` 并与该章 txt 文件数对比。
- **subagent-summary 文件会被清理**:delegation cache 中的 `subagent-summary-*.txt` 不持久。收到结果后必须立即写入 `/tmp/` 或目标目录,不要依赖后续重新读取。
- **结果过大落入 spillover 缓存**:`delegate_task` 返回超 ~100KB 时结果落在 `~/.hermes/cache/spillover/call_*.txt`(JSON 数组,每个 result 的 summary 可能是双重编码 JSON 或原始字符串)。此时可安全恢复:外层 json.loads 后在每个 summary 中找 `{"markdown"` 位置用 `json.JSONDecoder().raw_decode` 提取并 unescape,得到的就是完整 markdown(别对整个文件做 replace 修复)。比重新跑 subagent 快得多。
- **转录文件可能是 .srt 而非 .txt**:每段带序号+时间戳行,subagent 的 context 里要写明剥离序号/时间戳、按对话文本处理。
- **组装时不要做 JSON unescape 修复**:如果某章节内容看起来格式异常(全在一行),正确做法是重新读取原始转录文件生成该章节,不要用 `json.loads` 或 `.replace('\\n', '\n')` 修复整个文件 — 这会丢失其他章节。
- **大批量章节超出 subagent 并发限制**:`delegate_task` 最多 10 个并行。超过时分批派发,每批完成后立即落盘再派下一批。
- **输出文件体积**:39 个视频提炼后约 70-80KB / 1500+ 行。如果用户通过 `read_file` 查看,超过 100K 字符会被截断 — 保持简洁,每个视频控制在 1.5-2KB。
## 输出格式
每个视频一个 section:
```markdown
## Chapter N — [video_number]. [video_title]
**Core Insights**
- bullet 1
- bullet 2
- bullet 3
**Detailed Notes**
*Topic grouping*
- detail point
**Action Items**
- actionable step
**Key Terms/Concepts**
- **Term**: definition
```
语言:英语(除非用户要求中文)。保持简洁,跳过开场白和重复内容。