feat(skills): 新增 course-transcript-notes 技能并更新 blogwatcher 数据库
This commit is contained in:
Binary file not shown.
103
course-transcript-notes/SKILL.md
Normal file
103
course-transcript-notes/SKILL.md
Normal file
@@ -0,0 +1,103 @@
|
||||
---
|
||||
name: course-transcript-notes
|
||||
description: Use when 课程转录文件批量提炼为结构化 markdown 笔记。
|
||||
metadata:
|
||||
author: Hermes Agent
|
||||
last_updated: "2026-09-19"
|
||||
---
|
||||
|
||||
# Course Transcript Notes — 批量课程转录提炼
|
||||
|
||||
触发器:**"帮我整理课程笔记"** / **"提炼转录文件"** / 用户对 `knowledgebase/course/` 下某课程目录要求结构化总结。
|
||||
|
||||
## 工作流
|
||||
|
||||
1. **扫描目录结构**:
|
||||
```bash
|
||||
find <course_dir> -name "*.txt" | sort
|
||||
```
|
||||
识别章节分组(通常按子目录或文件名前缀编号)。
|
||||
|
||||
2. **并行委派子任务**(每章节一个 subagent):
|
||||
- 用 `delegate_task` 并行派发,每个 subagent 读取该章节所有 txt 文件
|
||||
- 要求输出格式统一:`## Chapter N — [number]. [title]` + Core Insights / Detailed Notes / Action Items / Key Terms
|
||||
- 设定 `output_schema: {"properties": {"markdown": {"type": "string"}}, "required": ["markdown"]}`
|
||||
|
||||
3. **接收结果后立即落盘**(关键步骤):
|
||||
- delegate_task 完成后,结果直接返回在对话中 — 已经是解析好的 markdown 字符串
|
||||
- 立即将每个 subagent 返回的 markdown 写入独立临时文件:`/tmp/ch<N>_notes.md`
|
||||
- **不要尝试从 subagent-summary 或 log 文件重新提取** — 这些文件可能有 JSON 编码问题或会被清理
|
||||
|
||||
4. **组装最终文件并双写**:
|
||||
```python
|
||||
# 简单字符串拼接,不做任何 JSON 处理
|
||||
chapters = []
|
||||
for i in range(num_chapters):
|
||||
with open(f'/tmp/ch{i}_notes.md') as f:
|
||||
chapters.append(f.read())
|
||||
full = header + '\n\n---\n\n'.join(chapters)
|
||||
|
||||
# 写入两个位置
|
||||
# 1. 原始 course 目录
|
||||
with open(output_path, 'w') as f:
|
||||
f.write(full)
|
||||
|
||||
# 2. nexus knowledgebase 目录
|
||||
import os, shutil
|
||||
nexus_dir = os.path.expanduser('~/Workspace/nexus/knowledgebase/course-note')
|
||||
os.makedirs(nexus_dir, exist_ok=True)
|
||||
nexus_path = os.path.join(nexus_dir, os.path.basename(output_path))
|
||||
shutil.copy2(output_path, nexus_path)
|
||||
print(f"Written to: {output_path}")
|
||||
print(f"Copied to: {nexus_path}")
|
||||
```
|
||||
|
||||
5. **立即验证完整性**(在告知用户完成之前):
|
||||
```bash
|
||||
# 统计每个章节的视频数
|
||||
for ch in 1 2 3 4 5; do
|
||||
count=$(grep -c "^## Chapter $ch —" <output_file>)
|
||||
echo "Chapter $ch: $count videos"
|
||||
done
|
||||
```
|
||||
- 确认每个章节的视频数 == 该章节的 txt 文件数
|
||||
- 不匹配则定位缺失章节,**直接读取原始转录文件重新生成**(不要重试 subagent)
|
||||
|
||||
6. **回退策略**(如果 subagent 结果丢失或格式异常):
|
||||
- 直接从原始 txt 文件读取内容
|
||||
- 在本地生成结构化摘要
|
||||
- 写入最终文件
|
||||
|
||||
## 坑
|
||||
|
||||
- **不要从 log/transcript 文件重新提取 subagent 输出**:`delegate_task` 返回的结果在对话中已经是可用的 markdown。如果当时没落盘,不要试图从 `subagent-summary-*.txt` 或 `live/*.log` 文件中用正则/JSON 解析恢复 — 这些文件中的 JSON 可能包含控制字符导致解析失败,且文件会被清理。正确做法:直接重新读取原始 txt 转录文件,本地生成摘要。
|
||||
- **验证必须按章节逐个计数**:`grep -c "^## Chapter"` 只给总数,掩盖单章缺失。必须按章节分别 `grep -c "^## Chapter N —"` 并与该章 txt 文件数对比。
|
||||
- **subagent-summary 文件会被清理**:delegation cache 中的 `subagent-summary-*.txt` 不持久。收到结果后必须立即写入 `/tmp/` 或目标目录,不要依赖后续重新读取。
|
||||
- **组装时不要做 JSON unescape 修复**:如果某章节内容看起来格式异常(全在一行),正确做法是重新读取原始转录文件生成该章节,不要用 `json.loads` 或 `.replace('\\n', '\n')` 修复整个文件 — 这会丢失其他章节。
|
||||
- **大批量章节超出 subagent 并发限制**:`delegate_task` 最多 10 个并行。超过时分批派发,每批完成后立即落盘再派下一批。
|
||||
- **输出文件体积**:39 个视频提炼后约 70-80KB / 1500+ 行。如果用户通过 `read_file` 查看,超过 100K 字符会被截断 — 保持简洁,每个视频控制在 1.5-2KB。
|
||||
|
||||
## 输出格式
|
||||
|
||||
每个视频一个 section:
|
||||
|
||||
```markdown
|
||||
## Chapter N — [video_number]. [video_title]
|
||||
|
||||
**Core Insights**
|
||||
- bullet 1
|
||||
- bullet 2
|
||||
- bullet 3
|
||||
|
||||
**Detailed Notes**
|
||||
*Topic grouping*
|
||||
- detail point
|
||||
|
||||
**Action Items**
|
||||
- actionable step
|
||||
|
||||
**Key Terms/Concepts**
|
||||
- **Term**: definition
|
||||
```
|
||||
|
||||
语言:英语(除非用户要求中文)。保持简洁,跳过开场白和重复内容。
|
||||
Reference in New Issue
Block a user