feat(skills): 新增 course-transcript-notes 技能并更新 blogwatcher 数据库

This commit is contained in:
2026-09-19 10:58:19 +08:00
parent 72e6b9efdd
commit 467bc620a9
2 changed files with 103 additions and 0 deletions

View File

@@ -0,0 +1,103 @@
---
name: course-transcript-notes
description: Use when 课程转录文件批量提炼为结构化 markdown 笔记。
metadata:
author: Hermes Agent
last_updated: "2026-09-19"
---
# Course Transcript Notes — 批量课程转录提炼
触发器:**"帮我整理课程笔记"** / **"提炼转录文件"** / 用户对 `knowledgebase/course/` 下某课程目录要求结构化总结。
## 工作流
1. **扫描目录结构**:
```bash
find <course_dir> -name "*.txt" | sort
```
识别章节分组(通常按子目录或文件名前缀编号)。
2. **并行委派子任务**(每章节一个 subagent):
- 用 `delegate_task` 并行派发,每个 subagent 读取该章节所有 txt 文件
- 要求输出格式统一:`## Chapter N — [number]. [title]` + Core Insights / Detailed Notes / Action Items / Key Terms
- 设定 `output_schema: {"properties": {"markdown": {"type": "string"}}, "required": ["markdown"]}`
3. **接收结果后立即落盘**(关键步骤):
- delegate_task 完成后,结果直接返回在对话中 — 已经是解析好的 markdown 字符串
- 立即将每个 subagent 返回的 markdown 写入独立临时文件:`/tmp/ch<N>_notes.md`
- **不要尝试从 subagent-summary 或 log 文件重新提取** — 这些文件可能有 JSON 编码问题或会被清理
4. **组装最终文件并双写**:
```python
# 简单字符串拼接,不做任何 JSON 处理
chapters = []
for i in range(num_chapters):
with open(f'/tmp/ch{i}_notes.md') as f:
chapters.append(f.read())
full = header + '\n\n---\n\n'.join(chapters)
# 写入两个位置
# 1. 原始 course 目录
with open(output_path, 'w') as f:
f.write(full)
# 2. nexus knowledgebase 目录
import os, shutil
nexus_dir = os.path.expanduser('~/Workspace/nexus/knowledgebase/course-note')
os.makedirs(nexus_dir, exist_ok=True)
nexus_path = os.path.join(nexus_dir, os.path.basename(output_path))
shutil.copy2(output_path, nexus_path)
print(f"Written to: {output_path}")
print(f"Copied to: {nexus_path}")
```
5. **立即验证完整性**(在告知用户完成之前):
```bash
# 统计每个章节的视频数
for ch in 1 2 3 4 5; do
count=$(grep -c "^## Chapter $ch —" <output_file>)
echo "Chapter $ch: $count videos"
done
```
- 确认每个章节的视频数 == 该章节的 txt 文件数
- 不匹配则定位缺失章节,**直接读取原始转录文件重新生成**(不要重试 subagent)
6. **回退策略**(如果 subagent 结果丢失或格式异常):
- 直接从原始 txt 文件读取内容
- 在本地生成结构化摘要
- 写入最终文件
## 坑
- **不要从 log/transcript 文件重新提取 subagent 输出**:`delegate_task` 返回的结果在对话中已经是可用的 markdown。如果当时没落盘,不要试图从 `subagent-summary-*.txt` 或 `live/*.log` 文件中用正则/JSON 解析恢复 — 这些文件中的 JSON 可能包含控制字符导致解析失败,且文件会被清理。正确做法:直接重新读取原始 txt 转录文件,本地生成摘要。
- **验证必须按章节逐个计数**:`grep -c "^## Chapter"` 只给总数,掩盖单章缺失。必须按章节分别 `grep -c "^## Chapter N —"` 并与该章 txt 文件数对比。
- **subagent-summary 文件会被清理**:delegation cache 中的 `subagent-summary-*.txt` 不持久。收到结果后必须立即写入 `/tmp/` 或目标目录,不要依赖后续重新读取。
- **组装时不要做 JSON unescape 修复**:如果某章节内容看起来格式异常(全在一行),正确做法是重新读取原始转录文件生成该章节,不要用 `json.loads` 或 `.replace('\\n', '\n')` 修复整个文件 — 这会丢失其他章节。
- **大批量章节超出 subagent 并发限制**:`delegate_task` 最多 10 个并行。超过时分批派发,每批完成后立即落盘再派下一批。
- **输出文件体积**:39 个视频提炼后约 70-80KB / 1500+ 行。如果用户通过 `read_file` 查看,超过 100K 字符会被截断 — 保持简洁,每个视频控制在 1.5-2KB。
## 输出格式
每个视频一个 section:
```markdown
## Chapter N — [video_number]. [video_title]
**Core Insights**
- bullet 1
- bullet 2
- bullet 3
**Detailed Notes**
*Topic grouping*
- detail point
**Action Items**
- actionable step
**Key Terms/Concepts**
- **Term**: definition
```
语言:英语(除非用户要求中文)。保持简洁,跳过开场白和重复内容。