Files
atlas/course-transcript-notes/SKILL.md

5.1 KiB
Raw Blame History

name, description, metadata
name description metadata
course-transcript-notes Use when 课程转录文件批量提炼为结构化 markdown 笔记。
author last_updated
Hermes Agent 2026-09-19

Course Transcript Notes — 批量课程转录提炼

触发器:"帮我整理课程笔记" / "提炼转录文件" / 用户对 knowledgebase/course/ 下某课程目录要求结构化总结。

工作流

  1. 扫描目录结构:

    find <course_dir> -name "*.txt" | sort
    

    识别章节分组(通常按子目录或文件名前缀编号)。

  2. 并行委派子任务(每章节一个 subagent):

    • 用 delegate_task 并行派发,每个 subagent 读取该章节所有 txt 文件
    • 要求输出格式统一:## Chapter N — [number]. [title] + Core Insights / Detailed Notes / Action Items / Key Terms
    • 设定 output_schema: {"properties": {"markdown": {"type": "string"}}, "required": ["markdown"]}
  3. 接收结果后立即落盘(关键步骤):

    • delegate_task 完成后,结果直接返回在对话中 — 已经是解析好的 markdown 字符串
    • 立即将每个 subagent 返回的 markdown 写入独立临时文件:/tmp/ch<N>_notes.md
    • 不要尝试从 subagent-summary 或 log 文件重新提取 — 这些文件可能有 JSON 编码问题或会被清理
  4. 组装最终文件并双写:

    # 简单字符串拼接,不做任何 JSON 处理
    chapters = []
    for i in range(num_chapters):
        with open(f'/tmp/ch{i}_notes.md') as f:
            chapters.append(f.read())
    full = header + '\n\n---\n\n'.join(chapters)
    
    # 写入两个位置
    # 1. 原始 course 目录
    with open(output_path, 'w') as f:
        f.write(full)
    
    # 2. nexus knowledgebase 目录
    import os, shutil
    nexus_dir = os.path.expanduser('~/Workspace/nexus/knowledgebase/course-notes')
    os.makedirs(nexus_dir, exist_ok=True)
    nexus_path = os.path.join(nexus_dir, os.path.basename(output_path))
    shutil.copy2(output_path, nexus_path)
    print(f"Written to: {output_path}")
    print(f"Copied to: {nexus_path}")
    
  5. 立即验证完整性(在告知用户完成之前):

    # 统计每个章节的视频数
    for ch in 1 2 3 4 5; do
      count=$(grep -c "^## Chapter $ch —" <output_file>)
      echo "Chapter $ch: $count videos"
    done
    
    • 确认每个章节的视频数 == 该章节的 txt 文件数
    • 不匹配则定位缺失章节,直接读取原始转录文件重新生成(不要重试 subagent)
  6. 回退策略(如果 subagent 结果丢失或格式异常):

    • 直接从原始 txt 文件读取内容
    • 在本地生成结构化摘要
    • 写入最终文件

坑

  • 不要从 log/transcript 文件重新提取 subagent 输出:delegate_task 返回的结果在对话中已经是可用的 markdown。如果当时没落盘,不要试图从 subagent-summary-*.txt 或 live/*.log 文件中用正则/JSON 解析恢复 — 这些文件中的 JSON 可能包含控制字符导致解析失败,且文件会被清理。正确做法:直接重新读取原始 txt 转录文件,本地生成摘要。
  • 验证必须按章节逐个计数:grep -c "^## Chapter" 只给总数,掩盖单章缺失。必须按章节分别 grep -c "^## Chapter N —" 并与该章 txt 文件数对比。
  • subagent-summary 文件会被清理:delegation cache 中的 subagent-summary-*.txt 不持久。收到结果后必须立即写入 /tmp/ 或目标目录,不要依赖后续重新读取。
  • 结果过大落入 spillover 缓存:delegate_task 返回超 ~100KB 时结果落在 ~/.hermes/cache/spillover/call_*.txt(JSON 数组,每个 result 的 summary 可能是双重编码 JSON 或原始字符串)。此时可安全恢复:外层 json.loads 后在每个 summary 中找 {"markdown" 位置用 json.JSONDecoder().raw_decode 提取并 unescape,得到的就是完整 markdown(别对整个文件做 replace 修复)。比重新跑 subagent 快得多。
  • 转录文件可能是 .srt 而非 .txt:每段带序号+时间戳行,subagent 的 context 里要写明剥离序号/时间戳、按对话文本处理。
  • 组装时不要做 JSON unescape 修复:如果某章节内容看起来格式异常(全在一行),正确做法是重新读取原始转录文件生成该章节,不要用 json.loads 或 .replace('\\n', '\n') 修复整个文件 — 这会丢失其他章节。
  • 大批量章节超出 subagent 并发限制:delegate_task 最多 10 个并行。超过时分批派发,每批完成后立即落盘再派下一批。
  • 输出文件体积:39 个视频提炼后约 70-80KB / 1500+ 行。如果用户通过 read_file 查看,超过 100K 字符会被截断 — 保持简洁,每个视频控制在 1.5-2KB。

输出格式

每个视频一个 section:

## Chapter N — [video_number]. [video_title]

**Core Insights**
- bullet 1
- bullet 2
- bullet 3

**Detailed Notes**
*Topic grouping*
- detail point

**Action Items**
- actionable step

**Key Terms/Concepts**
- **Term**: definition

语言:英语(除非用户要求中文)。保持简洁,跳过开场白和重复内容。