extract_audio and transcribe_audio scripts
This commit is contained in:
631
synthmind/README.md
Normal file
631
synthmind/README.md
Normal file
@@ -0,0 +1,631 @@
|
||||
# SynthMind - 视频音频提取与转录工具
|
||||
|
||||
将视频文件转换为文字稿的自动化工具链,为后续 AI 内容分析与总结提供数据基础。
|
||||
|
||||
**流水线**:视频文件 → MP3 音频 → TXT 文字稿 → *(后续 AI 分析)*
|
||||
|
||||
---
|
||||
|
||||
## 目录
|
||||
|
||||
- [功能特性](#功能特性)
|
||||
- [环境要求](#环境要求)
|
||||
- [安装](#安装)
|
||||
- [快速开始](#快速开始)
|
||||
- [脚本一:extract_audio.py](#脚本一extract_audiopy)
|
||||
- [脚本二:transcribe_audio.py](#脚本二transcribe_audiopy)
|
||||
- [多模型精度对比](#多模型精度对比)
|
||||
- [Manifest 文件说明](#manifest-文件说明)
|
||||
- [高级用法](#高级用法)
|
||||
- [故障排查](#故障排查)
|
||||
- [测试报告](#测试报告)
|
||||
|
||||
---
|
||||
|
||||
## 功能特性
|
||||
|
||||
### 通用特性
|
||||
- ✅ 支持单文件或递归扫描目录
|
||||
- ✅ 完善的**断点续传**:中断后重新运行自动跳过已完成项
|
||||
- ✅ **JSON manifest** 进度追踪,两个脚本共享同一文件
|
||||
- ✅ **失败自动记录 + 重试机制**(`--retry-failed`)
|
||||
- ✅ **状态查询**(`--status`)实时查看进度
|
||||
- ✅ 完全支持**含空格、中文的文件名**
|
||||
- ✅ 详细的时间戳日志
|
||||
|
||||
### extract_audio.py 特性
|
||||
- 使用 FFmpeg 提取音频:**MP3 / 64kbps / 22050Hz / 单声道**(语音场景优化,体积小)
|
||||
- 支持视频**切分模式**:超大视频按时长切分为多段分别处理
|
||||
- 支持多种输入格式:`.mp4/.mkv/.avi/.mov/.flv/.wmv/.webm/.m4v`
|
||||
|
||||
### transcribe_audio.py 特性
|
||||
- 基于 **openai-whisper** 命令行工具
|
||||
- 支持 4 种模型:**`tiny / base / small / medium`**(详见 [Whisper 模型选择](#whisper-模型选择))
|
||||
- **多模型批量转录(`--all-models`)**:一次运行对每个音频用 4 个模型分别转录,输出 `<stem>.<model>.txt` 便于精度对比
|
||||
- **模型预下载(`--preload-models`)**:一次性下载全部支持的模型到本地缓存
|
||||
- **自定义输出后缀(`--output-suffix`)**:便于同一模型多次转录(不同参数)不覆盖
|
||||
- 支持指定语言(`--language zh/en/...`)或自动检测
|
||||
- 支持音频**切分转录**:超长音频切分后逐段转录再合并
|
||||
- 支持多种音频格式:`.mp3/.m4a/.wav/.flac/.ogg/.aac`
|
||||
- **Manifest key 含模型维度**:多模型多参数结果共存,互不覆盖
|
||||
|
||||
---
|
||||
|
||||
## 环境要求
|
||||
|
||||
| 依赖 | 版本 | 用途 |
|
||||
|------|------|------|
|
||||
| Python | 3.8+ | 脚本运行 |
|
||||
| FFmpeg | 4.0+ | 音视频处理(含 `ffprobe`)|
|
||||
| openai-whisper | 最新 | 音频转录 |
|
||||
|
||||
**当前实测环境**:Ubuntu 22.04 (WSL2) + Python 3.13 + FFmpeg 6.0 + Whisper 20250625
|
||||
|
||||
---
|
||||
|
||||
## 安装
|
||||
|
||||
### 1. 安装 FFmpeg
|
||||
|
||||
```bash
|
||||
sudo apt update
|
||||
sudo apt install -y ffmpeg
|
||||
ffmpeg -version # 验证
|
||||
```
|
||||
|
||||
### 2. 安装 Whisper
|
||||
|
||||
```bash
|
||||
pip install openai-whisper
|
||||
whisper --help # 验证
|
||||
```
|
||||
|
||||
### 3. 预下载所有 Whisper 模型(推荐)
|
||||
|
||||
首次运行前,建议一次性下载所需的 4 个模型(共约 2.1GB):
|
||||
|
||||
```bash
|
||||
python3 transcribe_audio.py --preload-models
|
||||
```
|
||||
|
||||
**下载耗时参考**(本机测试):
|
||||
```
|
||||
tiny (~75MB) → 数秒
|
||||
base (~140MB) → 数秒
|
||||
small (~460MB) → 数十秒
|
||||
medium (~1.5GB) → 75s
|
||||
```
|
||||
|
||||
模型会缓存到 `~/.cache/whisper/`。
|
||||
|
||||
### 4. 下载脚本
|
||||
|
||||
将 `extract_audio.py` 和 `transcribe_audio.py` 放到同一目录即可,无需额外配置。
|
||||
|
||||
---
|
||||
|
||||
## 快速开始
|
||||
|
||||
### 场景 1:处理某目录下所有视频(简单流水线)
|
||||
|
||||
```bash
|
||||
# Step 1: 提取音频(生成 .mp3)
|
||||
python3 extract_audio.py /path/to/videos/
|
||||
|
||||
# Step 2: 转录文字(生成 .txt,默认 base 模型)
|
||||
python3 transcribe_audio.py /path/to/videos/ --language zh
|
||||
```
|
||||
|
||||
### 场景 2:对比 4 个模型的精度
|
||||
|
||||
```bash
|
||||
# 一次运行,每个音频用 4 个模型都跑一遍
|
||||
python3 transcribe_audio.py /path/to/audio/ --all-models --language zh
|
||||
```
|
||||
|
||||
输出(每个 mp3 生成 4 个 txt):
|
||||
```
|
||||
audio1.mp3
|
||||
audio1.tiny.txt
|
||||
audio1.base.txt
|
||||
audio1.small.txt
|
||||
audio1.medium.txt
|
||||
```
|
||||
|
||||
### 场景 3:高精度转录(推荐配置)
|
||||
|
||||
```bash
|
||||
python3 transcribe_audio.py /path/to/audio/ --model medium --language zh
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 脚本一:extract_audio.py
|
||||
|
||||
### 用法
|
||||
|
||||
```bash
|
||||
python3 extract_audio.py <source> [选项]
|
||||
```
|
||||
|
||||
### 参数说明
|
||||
|
||||
| 参数 | 简写 | 默认 | 说明 |
|
||||
|------|------|------|------|
|
||||
| `source` | - | 必填 | 视频文件路径 或 包含视频的目录 |
|
||||
| `--output` | `-o` | 与源同目录 | MP3 输出目录 |
|
||||
| `--manifest` | `-m` | source 目录下 | manifest.json 路径 |
|
||||
| `--split-size` | - | 不切分 | 视频时长超过此秒数则切分(单位:秒) |
|
||||
| `--split-segment` | - | 600 | 每个切分片段时长(单位:秒) |
|
||||
| `--retry-failed` | - | false | 重新处理上次失败的文件 |
|
||||
| `--status` | - | false | 仅查看进度状态,不执行提取 |
|
||||
|
||||
### 使用示例
|
||||
|
||||
**1. 处理目录下所有视频**
|
||||
```bash
|
||||
python3 extract_audio.py /home/user/videos/
|
||||
```
|
||||
|
||||
**2. 处理单个视频**
|
||||
```bash
|
||||
python3 extract_audio.py /home/user/videos/lesson1.mp4
|
||||
```
|
||||
|
||||
**3. 指定 MP3 输出目录**
|
||||
```bash
|
||||
python3 extract_audio.py /home/user/videos/ -o /home/user/audio/
|
||||
```
|
||||
|
||||
**4. 超过 1 小时的视频自动切分为 10 分钟片段**
|
||||
```bash
|
||||
python3 extract_audio.py /home/user/videos/ --split-size 3600 --split-segment 600
|
||||
```
|
||||
|
||||
**5. 查看进度**
|
||||
```bash
|
||||
python3 extract_audio.py /home/user/videos/ --status
|
||||
```
|
||||
|
||||
**6. 重试失败的文件**
|
||||
```bash
|
||||
python3 extract_audio.py /home/user/videos/ --retry-failed
|
||||
```
|
||||
|
||||
### 输出规格(FFmpeg 参数)
|
||||
|
||||
```
|
||||
-vn 去除视频轨道
|
||||
-acodec libmp3lame MP3 编码
|
||||
-ab 64k 比特率 64kbps
|
||||
-ar 22050 采样率 22050Hz
|
||||
-ac 1 单声道
|
||||
```
|
||||
|
||||
> **场景优化**:这套参数针对语音课程/讲座优化,体积约为原视频的 **20%**,同时保留清晰的人声。
|
||||
|
||||
---
|
||||
|
||||
## 脚本二:transcribe_audio.py
|
||||
|
||||
### 用法
|
||||
|
||||
```bash
|
||||
python3 transcribe_audio.py <source> [选项]
|
||||
python3 transcribe_audio.py --preload-models # 仅下载模型(source 可省略)
|
||||
```
|
||||
|
||||
### 参数说明
|
||||
|
||||
| 参数 | 简写 | 默认 | 说明 |
|
||||
|------|------|------|------|
|
||||
| `source` | - | 见备注 | MP3 文件路径 或 包含音频的目录 |
|
||||
| `--output` | `-o` | 与源同目录 | TXT 输出目录 |
|
||||
| `--manifest` | `-m` | source 目录下 | manifest.json 路径 |
|
||||
| `--model` | - | `base` | Whisper 模型:`tiny/base/small/medium`(4 选一) |
|
||||
| `--language` | - | 自动检测 | 语言代码。推荐 **`zh`**(中文)或 **`en`**(英文) |
|
||||
| `--all-models` | - | false | 每个音频用**全部 4 个模型**转录,输出 `<stem>.<model>.txt` |
|
||||
| `--output-suffix` | - | 空 | 输出 txt 附加后缀:`<stem><suffix>.txt` |
|
||||
| `--split-size` | - | 不切分 | 音频超过此秒数则切分(需要 ffmpeg) |
|
||||
| `--split-segment` | - | 1800 | 每个切分片段时长(单位:秒) |
|
||||
| `--retry-failed` | - | false | 重新处理上次失败的文件 |
|
||||
| `--status` | - | false | 仅查看进度状态 |
|
||||
| `--preload-models` | - | false | 预下载全部支持模型,无需 source |
|
||||
|
||||
> **备注**:`--preload-models` 或 `--status` 模式下可省略 `source`。
|
||||
|
||||
### Whisper 模型选择
|
||||
|
||||
| 模型 | 参数量 | 磁盘 | 显存 | CPU 转录 890s 音频 | 中文精度 | 语言输出 |
|
||||
|------|--------|------|------|------|------|------|
|
||||
| `tiny` | 39M | ~75MB | ~1GB | **47s** | ⭐⭐ | 繁体 |
|
||||
| `base` | 74M | ~140MB | ~1GB | **58s** | ⭐⭐⭐ | 简体 |
|
||||
| `small` | 244M | ~460MB | ~2GB | **93s** | ⭐⭐⭐⭐ | 简体 |
|
||||
| `medium` | 769M | ~1.5GB | ~5GB | **170s** | ⭐⭐⭐⭐⭐ | 简体 + 标点 |
|
||||
|
||||
> **实测数据来源**:本项目 [测试报告](#测试报告)(14 分 50 秒中文课程视频)
|
||||
> **推荐**:日常速稿用 `base`,重要内容用 `medium`。
|
||||
|
||||
### 使用示例
|
||||
|
||||
**1. 转录目录下所有 MP3(默认 base 模型)**
|
||||
```bash
|
||||
python3 transcribe_audio.py /home/user/audio/
|
||||
```
|
||||
|
||||
**2. 中文课程使用 medium 模型(高精度)**
|
||||
```bash
|
||||
python3 transcribe_audio.py /home/user/audio/ --model medium --language zh
|
||||
```
|
||||
|
||||
**3. 一次跑 4 个模型对比精度**
|
||||
```bash
|
||||
python3 transcribe_audio.py /home/user/audio/ --all-models --language zh
|
||||
```
|
||||
|
||||
**4. 预下载模型(首次安装后一次性下载全部)**
|
||||
```bash
|
||||
python3 transcribe_audio.py --preload-models
|
||||
```
|
||||
|
||||
**5. 自定义输出后缀(避免覆盖)**
|
||||
```bash
|
||||
# 用 base 模型转录,输出为 <stem>_v1.txt
|
||||
python3 transcribe_audio.py /home/user/audio/ --model base --output-suffix _v1
|
||||
|
||||
# 稍后再用 base + 不同参数转录,输出为 <stem>_v2.txt(不覆盖 v1)
|
||||
python3 transcribe_audio.py /home/user/audio/ --model base --output-suffix _v2
|
||||
```
|
||||
|
||||
**6. 超长音频(>30 分钟)切分转录**
|
||||
```bash
|
||||
python3 transcribe_audio.py /home/user/audio/ \
|
||||
--model medium --language zh \
|
||||
--split-size 1800 --split-segment 600
|
||||
```
|
||||
|
||||
**7. 查看进度(含每个模型的记录)**
|
||||
```bash
|
||||
python3 transcribe_audio.py /home/user/audio/ --status
|
||||
```
|
||||
|
||||
**8. 重试失败的文件**
|
||||
```bash
|
||||
python3 transcribe_audio.py /home/user/audio/ --retry-failed
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 多模型精度对比
|
||||
|
||||
### 实测样本
|
||||
|
||||
- **音频**:SketchUp 软件优化设置课程(14 分 50 秒,中文普通话)
|
||||
- **命令**:`python3 transcribe_audio.py <mp3> --all-models --language zh`
|
||||
|
||||
### 精度对比(关键词识别)
|
||||
|
||||
| 关键词 | 原意 | tiny | base | small | medium |
|
||||
|--------|------|------|------|-------|--------|
|
||||
| SketchUp 中文名 | 草图大师 | 草毒大師 ❌ | 吵毒大师 ❌ | 草读大师 ❌ | 草毒大师 ❌ |
|
||||
| 出厂设置 | 常规的设置 | 殘酷 ❌ | 残酷 ❌ | 常规 ✅ | 常规 ✅ |
|
||||
| 行业 | 行业 | 含業 ❌ | 行业 ✅ | 行业 ✅ | 行业 ✅ |
|
||||
| 轮廓线 | 轮廓线 | 能夠像 ❌ | 能扩限 ❌ | 轮廓线 ✅ | 轮廓线 ✅ |
|
||||
| 场景 | 场景 | 殘景 ❌ | 场景 ✅ | 场景 ✅ | 场景 ✅ |
|
||||
| 输出语言 | 简体中文 | **繁体** ⚠️ | 简体 ✅ | 简体 ✅ | 简体 ✅ |
|
||||
| 标点符号 | 有 | 无 ❌ | 无 ❌ | 少量 ⚠️ | 有 ✅ |
|
||||
|
||||
### 各模型样例输出(首行)
|
||||
|
||||
**原音频**:"大家好,这一节我们对**草图大师**来进行一个软件的一些优化设置方面的一些操作"
|
||||
|
||||
| 模型 | 输出 | 备注 |
|
||||
|------|------|------|
|
||||
| tiny | 大家好,這一節我們對**草毒大師**來進行一個軟件的一些優化設置方面的一些操作 | 繁体,专有名词错 |
|
||||
| base | 大家好,这一节我们对**吵毒大师**来进行一个软件的一些优化设置方面的一些操作 | 简体,专有名词错 |
|
||||
| small | 大家好,这一节我们对**草读大师**来进行一个软件的一些优化设置方面的一些操作 | 简体,专有名词接近 |
|
||||
| medium | 大家好,这一节我们对**草毒大师**来进行一个软件的一些优化设置方面的一些操作 | 简体,专有名词接近,有逗号 |
|
||||
|
||||
### 结论
|
||||
|
||||
| 场景 | 推荐模型 |
|
||||
|------|----------|
|
||||
| 快速内容概览、草稿 | **tiny** 或 **base**(快,可接受) |
|
||||
| 日常转录、笔记 | **base** 或 **small**(平衡) |
|
||||
| 正式转录、发布 | **medium**(最好精度) |
|
||||
| 专有名词很多的领域 | 结合 `--initial_prompt` 或人工校对 |
|
||||
|
||||
> **提示**:所有 whisper 模型对**专有名词**(品牌名、人名、技术术语)识别都容易出错。生产环境建议:
|
||||
> 1. 使用 medium 模型
|
||||
> 2. 通过 `whisper --initial_prompt "本视频讨论 SketchUp / 草图大师"` 提示 whisper(本脚本暂未透传此参数,可直接调 whisper CLI)
|
||||
> 3. 人工二次校对
|
||||
|
||||
---
|
||||
|
||||
## Manifest 文件说明
|
||||
|
||||
两个脚本**共享同一个 `manifest.json`**,通过 key 前缀区分:
|
||||
|
||||
- **视频提取记录**:key = 视频绝对路径
|
||||
- **音频转录记录**:key = `transcribe:<model>[:<suffix>]:<音频绝对路径>`
|
||||
|
||||
### key 结构详解
|
||||
|
||||
| Key 格式 | 场景 |
|
||||
|---------|------|
|
||||
| `/path/to/x.mp4` | 视频提取记录 |
|
||||
| `transcribe:tiny:/path/to/x.mp3` | tiny 模型转录(无 suffix) |
|
||||
| `transcribe:medium:/path/to/x.mp3` | medium 模型转录(无 suffix) |
|
||||
| `transcribe:base:_v1:/path/to/x.mp3` | base 模型 + `_v1` 后缀 |
|
||||
| `transcribe:base:_v2:/path/to/x.mp3` | base 模型 + `_v2` 后缀 |
|
||||
|
||||
多个 key 可并存于同一 manifest,互不覆盖。
|
||||
|
||||
### manifest.json 完整示例(含多模型对比)
|
||||
|
||||
```json
|
||||
{
|
||||
"version": 1,
|
||||
"files": {
|
||||
"/path/to/lesson1.mp4": {
|
||||
"audio_extracted": true,
|
||||
"audio_path": "/path/to/lesson1.mp3",
|
||||
"split_segments": [],
|
||||
"error": null,
|
||||
"updated_at": "2026-08-23T10:30:00"
|
||||
},
|
||||
"transcribe:tiny:/path/to/lesson1.mp3": {
|
||||
"transcribed": true,
|
||||
"txt_path": "/path/to/lesson1.tiny.txt",
|
||||
"split_segments": [],
|
||||
"model": "tiny",
|
||||
"language": "zh",
|
||||
"output_suffix": null,
|
||||
"error": null,
|
||||
"updated_at": "2026-08-23T10:35:00"
|
||||
},
|
||||
"transcribe:medium:/path/to/lesson1.mp3": {
|
||||
"transcribed": true,
|
||||
"txt_path": "/path/to/lesson1.medium.txt",
|
||||
"split_segments": [],
|
||||
"model": "medium",
|
||||
"language": "zh",
|
||||
"output_suffix": null,
|
||||
"error": null,
|
||||
"updated_at": "2026-08-23T10:50:00"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 字段说明
|
||||
|
||||
**视频提取记录字段**
|
||||
- `audio_extracted` - 是否已提取音频(布尔)
|
||||
- `audio_path` - 生成的 MP3 路径(切分模式下为第一个片段)
|
||||
- `split_segments` - 切分片段列表(非切分模式为空)
|
||||
- `error` - 失败时的错误信息,成功为 `null`
|
||||
- `updated_at` - 最后更新时间(ISO 8601)
|
||||
|
||||
**音频转录记录字段**
|
||||
- `transcribed` - 是否已转录(布尔)
|
||||
- `txt_path` - 生成的 TXT 路径
|
||||
- `split_segments` - 切分片段列表
|
||||
- `model` - 使用的 whisper 模型(`tiny/base/small/medium`)
|
||||
- `language` - 使用的语言代码(`zh/en/...`)或 `null`(自动检测)
|
||||
- `output_suffix` - 输出后缀(默认 `null`)
|
||||
- `error` - 失败错误信息
|
||||
- `updated_at` - 最后更新时间
|
||||
|
||||
### 断点续传逻辑
|
||||
|
||||
脚本判断是否需要处理某文件时,**同时校验**:
|
||||
1. manifest 中标记为已完成
|
||||
2. 目标文件实际存在且非空
|
||||
|
||||
**任一条件不满足则重新处理**(例如手动删除了 mp3 后重跑,会重新提取)。
|
||||
|
||||
**多模型模式**:每个 (音频 × 模型 × 后缀) 组合独立跟踪。已完成的组合被跳过,未完成的继续。
|
||||
|
||||
---
|
||||
|
||||
## 高级用法
|
||||
|
||||
### 1. 完整流水线脚本
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
VIDEO_DIR="/path/to/videos"
|
||||
|
||||
echo "=== Step 0: 预下载模型(仅首次)==="
|
||||
python3 transcribe_audio.py --preload-models
|
||||
|
||||
echo "=== Step 1: 提取音频 ==="
|
||||
python3 extract_audio.py "$VIDEO_DIR" --split-size 3600 --split-segment 600
|
||||
|
||||
echo "=== Step 2: 转录文字(medium 模型高精度)==="
|
||||
python3 transcribe_audio.py "$VIDEO_DIR" \
|
||||
--model medium --language zh \
|
||||
--split-size 1800 --split-segment 600
|
||||
|
||||
echo "=== Step 3: 查看最终状态 ==="
|
||||
python3 extract_audio.py "$VIDEO_DIR" --status
|
||||
python3 transcribe_audio.py "$VIDEO_DIR" --status
|
||||
```
|
||||
|
||||
### 2. 多模型对比工作流
|
||||
|
||||
```bash
|
||||
# 1) 提取音频
|
||||
python3 extract_audio.py /videos/
|
||||
|
||||
# 2) 一次跑 4 个模型(此步骤会跑 4x 时长,请预留时间)
|
||||
python3 transcribe_audio.py /videos/ --all-models --language zh
|
||||
|
||||
# 3) 查看所有结果
|
||||
python3 transcribe_audio.py /videos/ --status
|
||||
|
||||
# 4) 对比不同模型输出(例如用 diff)
|
||||
diff /videos/lesson1.base.txt /videos/lesson1.medium.txt
|
||||
```
|
||||
|
||||
### 3. 独立分离 MP3 与 TXT 输出
|
||||
|
||||
```bash
|
||||
python3 extract_audio.py /videos/ -o /audios/
|
||||
python3 transcribe_audio.py /audios/ -o /transcripts/ --model medium
|
||||
```
|
||||
|
||||
### 4. 自定义 manifest 位置(多套并行)
|
||||
|
||||
```bash
|
||||
python3 extract_audio.py /videos/ -m /var/run/task1_manifest.json
|
||||
python3 extract_audio.py /videos2/ -m /var/run/task2_manifest.json
|
||||
```
|
||||
|
||||
### 5. 同一模型不同参数的多轮转录
|
||||
|
||||
```bash
|
||||
# 第一版:自动检测语言
|
||||
python3 transcribe_audio.py audio.mp3 --model base --output-suffix _auto
|
||||
|
||||
# 第二版:强制中文
|
||||
python3 transcribe_audio.py audio.mp3 --model base --language zh --output-suffix _zh
|
||||
|
||||
# 两个 txt 独立保存,manifest 各自记录
|
||||
ls audio_*.txt
|
||||
# audio_auto.txt audio_zh.txt
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 故障排查
|
||||
|
||||
### FFmpeg 相关
|
||||
|
||||
| 症状 | 原因 | 解决 |
|
||||
|------|------|------|
|
||||
| 未找到 ffmpeg 命令 | 未安装 | `sudo apt install ffmpeg` |
|
||||
| 视频时长为 0 | ffprobe 无法解析 | 检查视频是否损坏 `ffprobe <file>` |
|
||||
| 切分片段为空 | 编码不兼容 codec copy | 手动尝试 `ffmpeg -i in.mp4 -c copy out.mp4` |
|
||||
|
||||
### Whisper 相关
|
||||
|
||||
| 症状 | 原因 | 解决 |
|
||||
|------|------|------|
|
||||
| 未找到 whisper 命令 | 未安装 | `pip install openai-whisper` |
|
||||
| 首次转录卡住 | 正在下载模型 | 提前用 `--preload-models` 一次下完 |
|
||||
| CUDA out of memory | 显存不足 | 使用更小模型;whisper 自动 fallback 到 CPU |
|
||||
| 转录内容为繁体 | tiny 模型对中文倾向输出繁体 | 使用 `base` 及以上模型 |
|
||||
| 专有名词识别错 | 模型未训练过 | 使用 `medium`;未来可支持 `--initial-prompt` |
|
||||
|
||||
### Manifest 相关
|
||||
|
||||
| 症状 | 解决 |
|
||||
|------|------|
|
||||
| 想强制重跑所有文件 | 删除 `manifest.json` |
|
||||
| 想强制重跑某个文件 | 编辑 manifest 或删除对应 `.mp3/.txt` 文件 |
|
||||
| 想重跑失败的文件 | 加 `--retry-failed` 参数 |
|
||||
| 想同一音频用同模型换参数再跑一次 | 用 `--output-suffix` 区分 |
|
||||
|
||||
### 文件名相关
|
||||
|
||||
- 中文文件名:完全支持
|
||||
- 空格文件名:完全支持(内部用 subprocess 参数列表,非 shell 拼接)
|
||||
- 特殊字符(如 `&`、`|`):应可正常处理,遇到问题请提 issue
|
||||
|
||||
---
|
||||
|
||||
## 测试报告
|
||||
|
||||
### 测试环境
|
||||
- OS: Ubuntu 22.04 (WSL2 on Windows)
|
||||
- Python: 3.13
|
||||
- FFmpeg: 6.0
|
||||
- Whisper: 20250625 (全部 4 个模型)
|
||||
|
||||
### 测试用例
|
||||
|
||||
**测试视频**:`03-SU软件的优化设置.mp4`
|
||||
- 时长:890 秒(14 分 50 秒)
|
||||
- 大小:37 MB
|
||||
- 视频:H.264 1280x720
|
||||
- 音频:AAC 44.1kHz 立体声
|
||||
- 语言:中文(普通话)
|
||||
- 内容:SketchUp 软件优化设置教学
|
||||
|
||||
### 提取音频 & 通用功能测试
|
||||
|
||||
| # | 测试项 | 结果 | 耗时 | 备注 |
|
||||
|---|--------|------|------|------|
|
||||
| 1 | 基本音频提取 | ✅ | 4s | 输出 6.9MB MP3 (22050Hz mono 64kbps) |
|
||||
| 2 | 断点续传(重跑跳过) | ✅ | <1s | manifest 校验通过 |
|
||||
| 3 | `--status` 状态查看 | ✅ | <1s | 正确显示 1/1 完成 |
|
||||
| 4 | 视频切分模式(500s 阈值) | ✅ | 4s | 3 片段完美衔接(修复 pipe:0 bug) |
|
||||
| 5 | 失败重试 + 含空格文件名 | ✅ | - | 失败记录、跳过、`--retry-failed` 均正常 |
|
||||
| 6 | 端到端流水线(视频→MP3→TXT) | ✅ | 54s | 从零开始生成完整产物 |
|
||||
|
||||
### 转录 & 多模型功能测试
|
||||
|
||||
| # | 测试项 | 结果 | 耗时 | 备注 |
|
||||
|---|--------|------|------|------|
|
||||
| 7 | `--preload-models` 预下载 4 模型 | ✅ | ~120s | 共 2.1GB,缓存到 `~/.cache/whisper/` |
|
||||
| 8 | 单模型转录(tiny + zh) | ✅ | 47s | 输出 9KB TXT |
|
||||
| 9 | 单模型转录(base + zh) | ✅ | 58s | 输出 9KB TXT |
|
||||
| 10 | 单模型转录(small + zh) | ✅ | 93s | 输出 9KB TXT |
|
||||
| 11 | 单模型转录(medium + zh) | ✅ | 170s | 输出 9KB TXT,有标点符号 |
|
||||
| 12 | `--all-models` 4 模型一次跑 | ✅ | 368s | 生成 4 个独立 txt |
|
||||
| 13 | 多模型 manifest 记录 | ✅ | - | 4 条独立 key,互不覆盖 |
|
||||
| 14 | 多模型 `--status` 展示 | ✅ | <1s | 每个模型独立行显示 |
|
||||
| 15 | 断点续传(重跑跳过) | ✅ | <1s | 4/4 全部跳过 |
|
||||
| 16 | `--output-suffix` 自定义后缀 | ✅ | 46s | `_custom` 生成独立 key + 独立 txt |
|
||||
| 17 | 音频切分转录 | ✅ | 59s | 3 片段并转录后合并成完整 TXT |
|
||||
| 18 | manifest 共享验证 | ✅ | - | 提取+转录记录并存互不干扰 |
|
||||
|
||||
### 已修复 Bug
|
||||
|
||||
**BUG-01**:切分模式下 MP3 为空文件(148 字节)
|
||||
- **原因**:使用 `cat <file> | ffmpeg -i pipe:0` 时,MP4 的 `moov` atom 位于文件末尾需要 seek,pipe 不支持 seek 导致 ffmpeg 静默失败
|
||||
- **修复**:改用 `ffmpeg -i <file>` 直接读取本地文件,同时增加输出文件非空校验(> 1KB)
|
||||
- **验证**:切分后 3 个片段合计 890.4s ≈ 原视频 890.2s
|
||||
|
||||
**BUG-02**:`--output-suffix` 被 manifest 误跳过
|
||||
- **原因**:初版 manifest key 只含 `model + path`,未含 `output_suffix`,导致同一模型换 suffix 后被判定为已完成
|
||||
- **修复**:key 结构改为 `transcribe:<model>[:<suffix>]:<path>`,suffix 参与 key 判定
|
||||
- **验证**:同一 tiny 模型,无 suffix 和 `_custom` suffix 生成独立 manifest 条目和独立 txt
|
||||
|
||||
### 性能参考(本次测试环境,CPU 转录)
|
||||
|
||||
| 阶段 | 输入 | 输出 | 耗时 | 吞吐 |
|
||||
|------|------|------|------|------|
|
||||
| 音频提取 | 37MB / 890s 视频 | 6.9MB MP3 | 4s | 223x realtime |
|
||||
| 转录(tiny) | 6.9MB / 890s 音频 | 9KB TXT | 47s | 19x realtime |
|
||||
| 转录(base) | 6.9MB / 890s 音频 | 9KB TXT | 58s | 15x realtime |
|
||||
| 转录(small) | 6.9MB / 890s 音频 | 9KB TXT | 93s | 9.6x realtime |
|
||||
| 转录(medium) | 6.9MB / 890s 音频 | 9KB TXT | 170s | 5.2x realtime |
|
||||
| **多模型一次跑(4 model)** | 6.9MB / 890s 音频 | 4 个 9KB TXT | 368s | - |
|
||||
|
||||
> 上述数据为 CPU 转录性能。GPU 转录速度可提升 5-10x。
|
||||
|
||||
---
|
||||
|
||||
## 项目路线图
|
||||
|
||||
- [x] Phase 1: 音频提取(extract_audio.py)
|
||||
- [x] Phase 2: 音频转录(transcribe_audio.py)
|
||||
- [x] 单模型转录
|
||||
- [x] 多模型批量对比(`--all-models`)
|
||||
- [x] 模型预下载(`--preload-models`)
|
||||
- [x] 输出后缀自定义(`--output-suffix`)
|
||||
- [x] 语言指定(`--language zh/en`)
|
||||
- [ ] Phase 3: 内容归纳整理(summarize.py)- 待定
|
||||
- AI 分析文字内容 + 视频内容
|
||||
- 生成脑图、结构化文档
|
||||
|
||||
---
|
||||
|
||||
## 相关脚本
|
||||
|
||||
- `nas_audio_extract_v3.py` - 针对 NAS 场景的原始脚本(通过 SSH 直接处理群晖 NAS 上的视频)
|
||||
- `extract_audio.py` - 本地通用视频音频提取(本项目 Phase 1)
|
||||
- `transcribe_audio.py` - 本地音频转录(本项目 Phase 2)
|
||||
384
synthmind/extract_audio.py
Normal file
384
synthmind/extract_audio.py
Normal file
@@ -0,0 +1,384 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
extract_audio.py - 视频音频批量提取脚本
|
||||
功能:
|
||||
- 扫描单个视频文件或遍历指定目录中的所有视频文件
|
||||
- 维护 manifest.json 进度文件,支持断点续传
|
||||
- 使用 FFmpeg 提取音频(MP3, 64k, 22050Hz, 单声道)
|
||||
- 支持视频切分(处理超大文件)
|
||||
- 支持含空格和中文的文件名
|
||||
|
||||
用法:
|
||||
python extract_audio.py <视频文件或目录> [选项]
|
||||
|
||||
示例:
|
||||
python extract_audio.py /path/to/videos/
|
||||
python extract_audio.py /path/to/video.mp4
|
||||
python extract_audio.py /path/to/videos/ --output /path/to/output/
|
||||
python extract_audio.py /path/to/videos/ --split-size 600 # 切分为600秒片段
|
||||
python extract_audio.py /path/to/videos/ --retry-failed # 重新处理失败项
|
||||
python extract_audio.py /path/to/videos/ --status # 查看进度状态
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
# ─── 常量 ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
MANIFEST_NAME = "manifest.json"
|
||||
VIDEO_EXTS = {".mp4", ".mkv", ".avi", ".mov", ".flv", ".wmv", ".webm", ".m4v"}
|
||||
|
||||
# FFmpeg 音频提取参数(语音课程优化:体积小)
|
||||
FFMPEG_AUDIO_OPTS = [
|
||||
"-vn", # 去除视频轨道
|
||||
"-acodec", "libmp3lame",
|
||||
"-ab", "64k",
|
||||
"-ar", "22050",
|
||||
"-ac", "1", # 单声道
|
||||
"-f", "mp3",
|
||||
]
|
||||
|
||||
# 视频切分:默认超过此时长(秒)才切分
|
||||
DEFAULT_SPLIT_THRESHOLD_SECS = 3600 # 1小时
|
||||
DEFAULT_SPLIT_SIZE_SECS = 600 # 每片10分钟
|
||||
|
||||
|
||||
# ─── 日志 ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
def log(msg: str, level: str = "INFO"):
|
||||
ts = time.strftime("%H:%M:%S")
|
||||
prefix = {"INFO": " ", "OK": "✅", "ERR": "❌", "WARN": "⚠️ ", "STEP": "▶ "}.get(level, " ")
|
||||
print(f"[{ts}] {prefix} {msg}", flush=True)
|
||||
|
||||
|
||||
# ─── Manifest 操作 ─────────────────────────────────────────────────────────────
|
||||
|
||||
def load_manifest(manifest_path: Path) -> dict:
|
||||
if manifest_path.exists():
|
||||
with open(manifest_path, encoding="utf-8") as f:
|
||||
return json.load(f)
|
||||
return {"version": 1, "files": {}}
|
||||
|
||||
|
||||
def save_manifest(manifest_path: Path, manifest: dict):
|
||||
manifest_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
with open(manifest_path, "w", encoding="utf-8") as f:
|
||||
json.dump(manifest, f, ensure_ascii=False, indent=2)
|
||||
|
||||
|
||||
def get_file_entry(manifest: dict, video_path: str) -> dict:
|
||||
"""获取或初始化某个视频文件的 manifest 条目"""
|
||||
if video_path not in manifest["files"]:
|
||||
manifest["files"][video_path] = {
|
||||
"audio_extracted": False,
|
||||
"audio_path": None,
|
||||
"split_segments": [], # 切分后的片段列表(若有)
|
||||
"error": None,
|
||||
"updated_at": None,
|
||||
}
|
||||
return manifest["files"][video_path]
|
||||
|
||||
|
||||
def mark_done(manifest: dict, manifest_path: Path, video_path: str,
|
||||
audio_path: str, segments: list = None):
|
||||
entry = get_file_entry(manifest, video_path)
|
||||
entry["audio_extracted"] = True
|
||||
entry["audio_path"] = audio_path
|
||||
entry["split_segments"] = segments or []
|
||||
entry["error"] = None
|
||||
entry["updated_at"] = time.strftime("%Y-%m-%dT%H:%M:%S")
|
||||
save_manifest(manifest_path, manifest)
|
||||
|
||||
|
||||
def mark_failed(manifest: dict, manifest_path: Path, video_path: str, error: str):
|
||||
entry = get_file_entry(manifest, video_path)
|
||||
entry["audio_extracted"] = False
|
||||
entry["error"] = error
|
||||
entry["updated_at"] = time.strftime("%Y-%m-%dT%H:%M:%S")
|
||||
save_manifest(manifest_path, manifest)
|
||||
|
||||
|
||||
# ─── FFmpeg 工具函数 ────────────────────────────────────────────────────────────
|
||||
|
||||
def check_ffmpeg():
|
||||
try:
|
||||
subprocess.run(["ffmpeg", "-version"], capture_output=True, check=True)
|
||||
return True
|
||||
except (subprocess.CalledProcessError, FileNotFoundError):
|
||||
log("未找到 ffmpeg 命令,请确认已安装:sudo apt install ffmpeg", "ERR")
|
||||
return False
|
||||
|
||||
|
||||
def get_video_duration(video_path: str) -> float:
|
||||
"""用 ffprobe 获取视频时长(秒),失败返回 0"""
|
||||
try:
|
||||
r = subprocess.run(
|
||||
["ffprobe", "-v", "error", "-show_entries", "format=duration",
|
||||
"-of", "default=noprint_wrappers=1:nokey=1", video_path],
|
||||
capture_output=True, text=True
|
||||
)
|
||||
return float(r.stdout.strip())
|
||||
except Exception:
|
||||
return 0.0
|
||||
|
||||
|
||||
def extract_audio_simple(video_path: str, mp3_path: str) -> bool:
|
||||
"""
|
||||
从视频提取音频到 mp3 文件。
|
||||
注意:MP4 的 moov atom 可能在文件尾部,需要 seek,因此不能用 pipe:0 输入。
|
||||
直接用 -i <file> 让 ffmpeg 自行 seek,同时用 -f mp3 明确输出格式。
|
||||
"""
|
||||
cmd = ["ffmpeg", "-y", "-i", video_path] + FFMPEG_AUDIO_OPTS + [mp3_path]
|
||||
r = subprocess.run(cmd, capture_output=True)
|
||||
if r.returncode != 0:
|
||||
return False
|
||||
return os.path.exists(mp3_path) and os.path.getsize(mp3_path) > 1024
|
||||
|
||||
|
||||
def split_video(video_path: str, output_dir: str, segment_secs: int) -> list:
|
||||
"""
|
||||
将视频切分为多个片段。
|
||||
返回切分后的 mp4 片段路径列表;若失败返回空列表。
|
||||
"""
|
||||
stem = Path(video_path).stem
|
||||
ext = Path(video_path).suffix
|
||||
out_pattern = os.path.join(output_dir, f"{stem}_seg%03d{ext}")
|
||||
|
||||
cmd = [
|
||||
"ffmpeg", "-y", "-i", video_path,
|
||||
"-c", "copy",
|
||||
"-segment_time", str(segment_secs),
|
||||
"-f", "segment",
|
||||
"-reset_timestamps", "1",
|
||||
out_pattern
|
||||
]
|
||||
r = subprocess.run(cmd, capture_output=True)
|
||||
if r.returncode != 0:
|
||||
return []
|
||||
|
||||
# 收集生成的片段文件(按名称排序)
|
||||
parent = Path(output_dir)
|
||||
segments = sorted(
|
||||
str(p) for p in parent.glob(f"{stem}_seg*{ext}")
|
||||
)
|
||||
return segments
|
||||
|
||||
|
||||
def extract_audio_from_segments(segments: list, mp3_dir: str, stem: str) -> tuple:
|
||||
"""
|
||||
对切分后的每个片段提取音频,返回 (segment_mp3s, all_ok)
|
||||
"""
|
||||
segment_mp3s = []
|
||||
all_ok = True
|
||||
for seg in segments:
|
||||
seg_stem = Path(seg).stem
|
||||
seg_mp3 = os.path.join(mp3_dir, f"{seg_stem}.mp3")
|
||||
ok = extract_audio_simple(seg, seg_mp3)
|
||||
if ok:
|
||||
segment_mp3s.append(seg_mp3)
|
||||
log(f" 片段 {Path(seg).name} → {Path(seg_mp3).name}", "OK")
|
||||
else:
|
||||
log(f" 片段 {Path(seg).name} 提取失败", "ERR")
|
||||
all_ok = False
|
||||
return segment_mp3s, all_ok
|
||||
|
||||
|
||||
# ─── 文件扫描 ──────────────────────────────────────────────────────────────────
|
||||
|
||||
def scan_videos(source: str) -> list:
|
||||
"""扫描单文件或目录,返回视频文件路径列表"""
|
||||
p = Path(source)
|
||||
if p.is_file():
|
||||
if p.suffix.lower() in VIDEO_EXTS:
|
||||
return [str(p.resolve())]
|
||||
else:
|
||||
log(f"不支持的文件格式: {p.suffix}", "ERR")
|
||||
return []
|
||||
elif p.is_dir():
|
||||
files = []
|
||||
for f in sorted(p.rglob("*")):
|
||||
if f.suffix.lower() in VIDEO_EXTS:
|
||||
files.append(str(f.resolve()))
|
||||
return files
|
||||
else:
|
||||
log(f"路径不存在: {source}", "ERR")
|
||||
return []
|
||||
|
||||
|
||||
def mp3_exists(mp3_path: str) -> bool:
|
||||
return os.path.exists(mp3_path) and os.path.getsize(mp3_path) > 0
|
||||
|
||||
|
||||
# ─── 状态展示 ──────────────────────────────────────────────────────────────────
|
||||
|
||||
def show_status(manifest: dict, all_videos: list):
|
||||
print("\n" + "=" * 60)
|
||||
print("📊 提取音频进度状态")
|
||||
print("=" * 60)
|
||||
done = sum(1 for v in all_videos
|
||||
if manifest["files"].get(v, {}).get("audio_extracted"))
|
||||
failed = sum(1 for v in all_videos
|
||||
if manifest["files"].get(v, {}).get("error"))
|
||||
pending = len(all_videos) - done
|
||||
print(f" 总文件数 : {len(all_videos)}")
|
||||
print(f" 已完成 : {done}")
|
||||
print(f" 失败 : {failed}")
|
||||
print(f" 待处理 : {pending}")
|
||||
print("=" * 60)
|
||||
for v in all_videos:
|
||||
entry = manifest["files"].get(v, {})
|
||||
if entry.get("audio_extracted"):
|
||||
status = "✅ 完成"
|
||||
elif entry.get("error"):
|
||||
status = f"❌ 失败: {entry['error']}"
|
||||
else:
|
||||
status = "⏳ 待处理"
|
||||
print(f" {status} {Path(v).name}")
|
||||
print()
|
||||
|
||||
|
||||
# ─── 主流程 ────────────────────────────────────────────────────────────────────
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(
|
||||
description="视频音频批量提取工具",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog=__doc__
|
||||
)
|
||||
parser.add_argument("source", help="视频文件路径或包含视频的目录")
|
||||
parser.add_argument("--output", "-o", default=None,
|
||||
help="MP3 输出目录(默认与视频文件同目录)")
|
||||
parser.add_argument("--manifest", "-m", default=None,
|
||||
help="manifest.json 路径(默认在 source 目录下)")
|
||||
parser.add_argument("--split-size", type=int, default=None,
|
||||
help=f"切分阈值(秒):超过此时长的视频会被切分(默认不切分)")
|
||||
parser.add_argument("--split-segment", type=int, default=DEFAULT_SPLIT_SIZE_SECS,
|
||||
help=f"每个切分片段的时长(秒,默认 {DEFAULT_SPLIT_SIZE_SECS})")
|
||||
parser.add_argument("--retry-failed", action="store_true",
|
||||
help="重新处理上次失败的文件")
|
||||
parser.add_argument("--status", action="store_true",
|
||||
help="仅查看进度状态,不执行提取")
|
||||
args = parser.parse_args()
|
||||
|
||||
if not check_ffmpeg():
|
||||
sys.exit(1)
|
||||
|
||||
# 确定 manifest 路径
|
||||
source_path = Path(args.source).resolve()
|
||||
base_dir = source_path if source_path.is_dir() else source_path.parent
|
||||
manifest_path = Path(args.manifest) if args.manifest else base_dir / MANIFEST_NAME
|
||||
|
||||
manifest = load_manifest(manifest_path)
|
||||
|
||||
# 扫描视频文件
|
||||
all_videos = scan_videos(args.source)
|
||||
if not all_videos:
|
||||
log("未找到任何视频文件", "WARN")
|
||||
sys.exit(0)
|
||||
|
||||
log(f"扫描到 {len(all_videos)} 个视频文件", "INFO")
|
||||
|
||||
# 仅查看状态
|
||||
if args.status:
|
||||
show_status(manifest, all_videos)
|
||||
return
|
||||
|
||||
# 确定待处理列表
|
||||
pending = []
|
||||
for v in all_videos:
|
||||
entry = get_file_entry(manifest, v)
|
||||
already_done = entry.get("audio_extracted") and entry.get("audio_path") and \
|
||||
mp3_exists(entry["audio_path"])
|
||||
is_failed = bool(entry.get("error"))
|
||||
|
||||
if already_done:
|
||||
continue
|
||||
if is_failed and not args.retry_failed:
|
||||
log(f"跳过(上次失败,用 --retry-failed 重试): {Path(v).name}", "WARN")
|
||||
continue
|
||||
pending.append(v)
|
||||
|
||||
done_count = len(all_videos) - len(pending)
|
||||
log(f"📊 总体: {done_count}/{len(all_videos)} 已完成,{len(pending)} 待处理")
|
||||
|
||||
if not pending:
|
||||
log("全部已完成 ✅", "OK")
|
||||
return
|
||||
|
||||
success, failed_list = 0, []
|
||||
|
||||
for i, video in enumerate(pending, 1):
|
||||
video_name = Path(video).name
|
||||
stem = Path(video).stem
|
||||
video_dir = Path(video).parent
|
||||
|
||||
# 确定输出目录
|
||||
out_dir = Path(args.output).resolve() if args.output else video_dir
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
mp3_path = str(out_dir / f"{stem}.mp3")
|
||||
|
||||
log(f"\n[{i}/{len(pending)}] {video_name}", "STEP")
|
||||
|
||||
try:
|
||||
# 判断是否需要切分
|
||||
duration = get_video_duration(video)
|
||||
need_split = args.split_size and duration > 0 and duration > args.split_size
|
||||
|
||||
if need_split:
|
||||
log(f" 视频时长 {duration:.0f}s > {args.split_size}s,启动切分模式")
|
||||
seg_dir = out_dir / f"{stem}_segments"
|
||||
seg_dir.mkdir(exist_ok=True)
|
||||
|
||||
segments = split_video(video, str(seg_dir), args.split_segment)
|
||||
if not segments:
|
||||
raise RuntimeError("视频切分失败")
|
||||
log(f" 切分为 {len(segments)} 个片段")
|
||||
|
||||
seg_mp3s, all_ok = extract_audio_from_segments(
|
||||
segments, str(out_dir), stem
|
||||
)
|
||||
if not all_ok:
|
||||
raise RuntimeError("部分片段音频提取失败")
|
||||
|
||||
# 记录切分片段到 manifest(主 mp3_path 设为第一个片段)
|
||||
mark_done(manifest, manifest_path, video, seg_mp3s[0] if seg_mp3s else mp3_path,
|
||||
segments=seg_mp3s)
|
||||
else:
|
||||
# 普通提取
|
||||
log(f" 📥 提取音频中...")
|
||||
if not extract_audio_simple(video, mp3_path):
|
||||
raise RuntimeError("FFmpeg 音频提取失败")
|
||||
|
||||
size = os.path.getsize(mp3_path)
|
||||
log(f" 输出: {Path(mp3_path).name} ({size // 1024} KB)", "OK")
|
||||
mark_done(manifest, manifest_path, video, mp3_path)
|
||||
|
||||
success += 1
|
||||
log(f" 进度: {done_count + success}/{len(all_videos)}", "OK")
|
||||
|
||||
except Exception as e:
|
||||
err_msg = str(e)
|
||||
log(f" {err_msg}", "ERR")
|
||||
mark_failed(manifest, manifest_path, video, err_msg)
|
||||
failed_list.append(video_name)
|
||||
# 清理可能残留的不完整 mp3
|
||||
if os.path.exists(mp3_path) and os.path.getsize(mp3_path) == 0:
|
||||
os.remove(mp3_path)
|
||||
|
||||
print(f"\n{'='*60}")
|
||||
log(f"🏁 完成: {success}/{len(pending)} 失败: {len(failed_list)}", "INFO")
|
||||
if failed_list:
|
||||
log(f"失败文件列表:", "ERR")
|
||||
for f in failed_list:
|
||||
print(f" - {f}")
|
||||
print(" 运行时加 --retry-failed 可重新处理失败项")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,171 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
NAS 视频音频批量提取 v3
|
||||
流程: ssh cat .mp4 → 本地文件 → FFmpeg → shell重定向上传 → 删除本地
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import os
|
||||
import time
|
||||
|
||||
NAS_BASE = "/volume2/work/Public Cloud Learning Sessions"
|
||||
MAC_MP4 = os.path.expanduser("~/.openclaw/temp/xingshu/nas_mp4")
|
||||
MAC_MP3 = os.path.expanduser("~/.openclaw/temp/xingshu/nas_mp3_out")
|
||||
MAC_LOG = os.path.expanduser("~/.openclaw/temp/xingshu/logs/nas_audio_v3.log")
|
||||
PROGRESS = os.path.expanduser("~/.openclaw/temp/xingshu/nas_audio_v3.done")
|
||||
SSH_USER = "shenwei"
|
||||
NAS_HOST = "192.168.3.17"
|
||||
|
||||
os.makedirs(MAC_MP4, exist_ok=True)
|
||||
os.makedirs(MAC_MP3, exist_ok=True)
|
||||
|
||||
FFMPEG_BIN = "/opt/homebrew/bin/ffmpeg"
|
||||
FFMPEG_CMD = [FFMPEG_BIN, "-y", "-i", "pipe:0",
|
||||
"-vn", "-acodec", "libmp3lame", "-ab", "64k", "-ar", "22050", "-ac", "1",
|
||||
"-f", "mp3", "pipe:1"]
|
||||
|
||||
def log(msg):
|
||||
ts = time.strftime("%H:%M:%S")
|
||||
line = f"[{ts}] {msg}"
|
||||
print(line)
|
||||
with open(MAC_LOG, "a") as f:
|
||||
f.write(line + "\n")
|
||||
|
||||
def ssh_download(nas_path, local_path):
|
||||
"""NAS 文件通过 ssh cat 落地到本地(1MB 分块)"""
|
||||
cmd = ["ssh", f"{SSH_USER}@{NAS_HOST}", f"cat '{nas_path}'"]
|
||||
with open(local_path, "wb") as fout:
|
||||
proc = subprocess.Popen(cmd, stdout=subprocess.PIPE)
|
||||
while True:
|
||||
chunk = proc.stdout.read(1024 * 1024)
|
||||
if not chunk:
|
||||
break
|
||||
fout.write(chunk)
|
||||
proc.stdout.close()
|
||||
proc.wait()
|
||||
return proc.returncode == 0
|
||||
|
||||
def ssh_upload(local_path, nas_path):
|
||||
"""本地文件通过 shell 重定向上传到 NAS(本地路径需加引号防空格)"""
|
||||
cmd = f'ssh {SSH_USER}@{NAS_HOST} "cat > \'{nas_path}\'" < "{local_path}"'
|
||||
return subprocess.run(cmd, shell=True, stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL).returncode == 0
|
||||
|
||||
def get_nas_mp3s():
|
||||
r = subprocess.run(["ssh", f"{SSH_USER}@{NAS_HOST}",
|
||||
f"ls '{NAS_BASE}'/*.mp3 2>/dev/null"],
|
||||
capture_output=True, text=True)
|
||||
if r.returncode != 0:
|
||||
return set()
|
||||
return set(l.strip() for l in r.stdout.split("\n") if l.strip())
|
||||
|
||||
def get_nas_mp4s():
|
||||
r = subprocess.run(["ssh", f"{SSH_USER}@{NAS_HOST}",
|
||||
f"ls '{NAS_BASE}'/*.mp4 2>/dev/null"],
|
||||
capture_output=True, text=True)
|
||||
if r.returncode != 0:
|
||||
return []
|
||||
return [l.strip() for l in r.stdout.split("\n") if l.strip()]
|
||||
|
||||
def load_done():
|
||||
if os.path.exists(PROGRESS):
|
||||
with open(PROGRESS) as f:
|
||||
return set(f.read().splitlines())
|
||||
return set()
|
||||
|
||||
def save_done(done_set):
|
||||
with open(PROGRESS, "w") as f:
|
||||
f.write("\n".join(done_set))
|
||||
|
||||
def main():
|
||||
log("=" * 60)
|
||||
log("🚀 v3 开始批量提取")
|
||||
|
||||
done = load_done()
|
||||
nas_mp3s = get_nas_mp3s()
|
||||
all_mp4s = get_nas_mp4s()
|
||||
|
||||
pending = []
|
||||
for v in all_mp4s:
|
||||
name = os.path.basename(v)
|
||||
mp3_name = name.replace(".mp4", ".mp3")
|
||||
nas_mp3_path = f"{NAS_BASE}/{mp3_name}"
|
||||
if name not in done and nas_mp3_path not in nas_mp3s:
|
||||
pending.append(name)
|
||||
|
||||
total = len(all_mp4s)
|
||||
done_count = total - len(pending)
|
||||
log(f"📊 总体: {done_count}/{total} 已完成,{len(pending)} 待处理")
|
||||
|
||||
if not pending:
|
||||
log("✅ 全部已完成")
|
||||
return
|
||||
|
||||
success, failed = 0, []
|
||||
|
||||
for i, video in enumerate(pending, 1):
|
||||
mp3_name = video.replace(".mp4", ".mp3")
|
||||
nas_mp4 = f"{NAS_BASE}/{video}"
|
||||
nas_mp3 = f"{NAS_BASE}/{mp3_name}"
|
||||
local_mp4 = f"{MAC_MP4}/{video}"
|
||||
local_mp3 = f"{MAC_MP3}/{mp3_name}"
|
||||
|
||||
log(f"\n[{i}/{len(pending)}] ▶ {video}")
|
||||
|
||||
try:
|
||||
# Step 1: 下载 .mp4 → 本地
|
||||
log(f" 📥 downloading...")
|
||||
if not ssh_download(nas_mp4, local_mp4):
|
||||
raise RuntimeError("download failed")
|
||||
size_mp4 = os.path.getsize(local_mp4)
|
||||
log(f" ✅ downloaded {size_mp4//1024//1024}MB")
|
||||
|
||||
# Step 2: FFmpeg 提取 .mp3
|
||||
log(f" 🔊 extracting...")
|
||||
with open(local_mp3, "wb") as fout:
|
||||
proc = subprocess.Popen(["/bin/cat", local_mp4],
|
||||
stdout=subprocess.PIPE, stderr=subprocess.DEVNULL)
|
||||
ffmpeg = subprocess.Popen(FFMPEG_CMD,
|
||||
stdin=proc.stdout, stdout=fout, stderr=subprocess.DEVNULL)
|
||||
proc.stdout.close()
|
||||
ret = ffmpeg.wait()
|
||||
proc.wait()
|
||||
|
||||
if ret != 0:
|
||||
raise RuntimeError("ffmpeg failed")
|
||||
|
||||
# 删除 .mp4 释放空间
|
||||
os.remove(local_mp4)
|
||||
|
||||
size_mp3 = os.path.getsize(local_mp3)
|
||||
log(f" ✅ extracted {size_mp3//1024//1024}MB ({size_mp3//1024}KB)")
|
||||
|
||||
# Step 3: 上传 .mp3 → NAS
|
||||
log(f" 📤 uploading...")
|
||||
if not ssh_upload(local_mp3, nas_mp3):
|
||||
raise RuntimeError("upload failed")
|
||||
log(f" ✅ uploaded to NAS")
|
||||
|
||||
# 删除 .mp3
|
||||
os.remove(local_mp3)
|
||||
|
||||
done.add(video)
|
||||
save_done(done)
|
||||
success += 1
|
||||
|
||||
log(f" 📊 进度: {done_count + success}/{total}")
|
||||
|
||||
except Exception as e:
|
||||
log(f" ❌ {e}")
|
||||
failed.append(video)
|
||||
for p in [local_mp4, local_mp3]:
|
||||
if os.path.exists(p):
|
||||
os.remove(p)
|
||||
|
||||
log(f"\n{'='*60}")
|
||||
log(f"🏁 完成: {success}/{len(pending)},失败: {len(failed)}")
|
||||
if failed:
|
||||
log(f"失败列表: {failed}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
553
synthmind/transcribe_audio.py
Normal file
553
synthmind/transcribe_audio.py
Normal file
@@ -0,0 +1,553 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
transcribe_audio.py - MP3 音频批量转录文字脚本
|
||||
功能:
|
||||
- 扫描单个 MP3 文件或遍历指定目录中的所有 MP3 文件
|
||||
- 维护 manifest.json 进度文件,支持断点续传(key 含模型维度,多模型互不覆盖)
|
||||
- 使用系统已安装的 whisper 命令行工具转录音频
|
||||
- 支持 MP3 切分(处理可能生成超大 txt 的音频文件)
|
||||
- 支持含空格和中文的文件名
|
||||
- 支持预下载模型 (--preload-models)
|
||||
- 支持多模型对比 (--all-models)
|
||||
- 支持自定义输出后缀 (--output-suffix)
|
||||
|
||||
支持的模型: tiny / base / small / medium
|
||||
推荐语言: zh (中文) / en (英文),也接受其他 whisper 支持的语言代码
|
||||
|
||||
用法:
|
||||
python transcribe_audio.py <MP3文件或目录> [选项]
|
||||
|
||||
示例:
|
||||
# 预下载所有支持的模型
|
||||
python transcribe_audio.py --preload-models
|
||||
|
||||
# 使用 medium 模型转录中文
|
||||
python transcribe_audio.py /path/to/audio/ --model medium --language zh
|
||||
|
||||
# 一次跑 4 个模型对比精度(生成 file.tiny.txt / file.base.txt / ...)
|
||||
python transcribe_audio.py /path/to/audio/ --all-models --language zh
|
||||
|
||||
# 自定义输出后缀
|
||||
python transcribe_audio.py /path/to/audio.mp3 --model base --output-suffix _v1
|
||||
|
||||
# 超1800秒切分
|
||||
python transcribe_audio.py /path/to/audio/ --split-size 1800
|
||||
|
||||
# 重新处理失败项
|
||||
python transcribe_audio.py /path/to/audio/ --retry-failed
|
||||
|
||||
# 查看进度状态
|
||||
python transcribe_audio.py /path/to/audio/ --status
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
# ─── 常量 ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
MANIFEST_NAME = "manifest.json"
|
||||
AUDIO_EXTS = {".mp3", ".m4a", ".wav", ".flac", ".ogg", ".aac"}
|
||||
|
||||
SUPPORTED_MODELS = ["tiny", "base", "small", "medium"]
|
||||
DEFAULT_WHISPER_MODEL = "base"
|
||||
DEFAULT_SPLIT_SEGMENT_SECS = 1800 # 每片30分钟
|
||||
|
||||
|
||||
# ─── 日志 ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
def log(msg: str, level: str = "INFO"):
|
||||
ts = time.strftime("%H:%M:%S")
|
||||
prefix = {"INFO": " ", "OK": "✅", "ERR": "❌", "WARN": "⚠️ ", "STEP": "▶ "}.get(level, " ")
|
||||
print(f"[{ts}] {prefix} {msg}", flush=True)
|
||||
|
||||
|
||||
# ─── Manifest 操作 (key 包含 model,支持同一音频多模型共存) ────────────────────
|
||||
|
||||
def load_manifest(manifest_path: Path) -> dict:
|
||||
if manifest_path.exists():
|
||||
with open(manifest_path, encoding="utf-8") as f:
|
||||
return json.load(f)
|
||||
return {"version": 1, "files": {}}
|
||||
|
||||
|
||||
def save_manifest(manifest_path: Path, manifest: dict):
|
||||
manifest_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
with open(manifest_path, "w", encoding="utf-8") as f:
|
||||
json.dump(manifest, f, ensure_ascii=False, indent=2)
|
||||
|
||||
|
||||
def transcribe_key(model: str, audio_path: str, output_suffix: str = "") -> str:
|
||||
"""
|
||||
manifest key 由 model + output_suffix + 音频路径共同决定,
|
||||
因此改变 --output-suffix 会形成新的 key(不会误跳过)。
|
||||
"""
|
||||
if output_suffix:
|
||||
return f"transcribe:{model}:{output_suffix}:{audio_path}"
|
||||
return f"transcribe:{model}:{audio_path}"
|
||||
|
||||
|
||||
def get_transcribe_entry(manifest: dict, audio_path: str, model: str,
|
||||
output_suffix: str = "") -> dict:
|
||||
key = transcribe_key(model, audio_path, output_suffix)
|
||||
if key not in manifest["files"]:
|
||||
manifest["files"][key] = {
|
||||
"transcribed": False,
|
||||
"txt_path": None,
|
||||
"split_segments": [],
|
||||
"error": None,
|
||||
"model": model,
|
||||
"language": None,
|
||||
"output_suffix": output_suffix or None,
|
||||
"updated_at": None,
|
||||
}
|
||||
return manifest["files"][key]
|
||||
|
||||
|
||||
def mark_transcribed(manifest: dict, manifest_path: Path, audio_path: str,
|
||||
txt_path: str, model: str, language: str = None,
|
||||
output_suffix: str = "", segments: list = None):
|
||||
entry = get_transcribe_entry(manifest, audio_path, model, output_suffix)
|
||||
entry["transcribed"] = True
|
||||
entry["txt_path"] = txt_path
|
||||
entry["split_segments"] = segments or []
|
||||
entry["model"] = model
|
||||
entry["language"] = language
|
||||
entry["output_suffix"] = output_suffix or None
|
||||
entry["error"] = None
|
||||
entry["updated_at"] = time.strftime("%Y-%m-%dT%H:%M:%S")
|
||||
save_manifest(manifest_path, manifest)
|
||||
|
||||
|
||||
def mark_transcribe_failed(manifest: dict, manifest_path: Path,
|
||||
audio_path: str, model: str, error: str,
|
||||
output_suffix: str = ""):
|
||||
entry = get_transcribe_entry(manifest, audio_path, model, output_suffix)
|
||||
entry["transcribed"] = False
|
||||
entry["error"] = error
|
||||
entry["updated_at"] = time.strftime("%Y-%m-%dT%H:%M:%S")
|
||||
save_manifest(manifest_path, manifest)
|
||||
|
||||
|
||||
# ─── 模型预下载 ────────────────────────────────────────────────────────────────
|
||||
|
||||
def preload_models(models: list) -> bool:
|
||||
"""
|
||||
调用 whisper Python API 预下载模型到 ~/.cache/whisper/。
|
||||
比调用 whisper CLI 更快(无需额外音频输入)。
|
||||
"""
|
||||
try:
|
||||
import whisper as whisper_lib
|
||||
except ImportError:
|
||||
log("openai-whisper Python 库未安装,无法预下载", "ERR")
|
||||
log("请安装:pip install openai-whisper", "ERR")
|
||||
return False
|
||||
|
||||
log(f"开始预下载 {len(models)} 个模型:{', '.join(models)}", "STEP")
|
||||
for m in models:
|
||||
log(f" 下载 {m}...")
|
||||
t0 = time.time()
|
||||
try:
|
||||
whisper_lib.load_model(m)
|
||||
log(f" {m} 就绪 ({time.time() - t0:.1f}s)", "OK")
|
||||
except Exception as e:
|
||||
log(f" {m} 下载失败: {e}", "ERR")
|
||||
return False
|
||||
log(f"所有模型下载完成 ✅", "OK")
|
||||
return True
|
||||
|
||||
|
||||
# ─── Whisper / FFmpeg 工具函数 ─────────────────────────────────────────────────
|
||||
|
||||
def check_whisper() -> bool:
|
||||
try:
|
||||
subprocess.run(["whisper", "--help"], capture_output=True, check=False)
|
||||
return True
|
||||
except FileNotFoundError:
|
||||
log("未找到 whisper 命令,请确认已安装:pip install openai-whisper", "ERR")
|
||||
return False
|
||||
|
||||
|
||||
def check_ffmpeg() -> bool:
|
||||
try:
|
||||
subprocess.run(["ffmpeg", "-version"], capture_output=True, check=True)
|
||||
return True
|
||||
except (subprocess.CalledProcessError, FileNotFoundError):
|
||||
return False
|
||||
|
||||
|
||||
def get_audio_duration(audio_path: str) -> float:
|
||||
try:
|
||||
r = subprocess.run(
|
||||
["ffprobe", "-v", "error", "-show_entries", "format=duration",
|
||||
"-of", "default=noprint_wrappers=1:nokey=1", audio_path],
|
||||
capture_output=True, text=True
|
||||
)
|
||||
return float(r.stdout.strip())
|
||||
except Exception:
|
||||
return 0.0
|
||||
|
||||
|
||||
def split_audio(audio_path: str, output_dir: str, segment_secs: int) -> list:
|
||||
"""
|
||||
将 MP3 切分为多个片段,返回片段路径列表。
|
||||
使用 ffmpeg segment 模式,避免重编码(codec copy)。
|
||||
"""
|
||||
stem = Path(audio_path).stem
|
||||
ext = Path(audio_path).suffix
|
||||
out_pattern = os.path.join(output_dir, f"{stem}_seg%03d{ext}")
|
||||
|
||||
cmd = [
|
||||
"ffmpeg", "-y", "-i", audio_path,
|
||||
"-f", "segment",
|
||||
"-segment_time", str(segment_secs),
|
||||
"-c", "copy",
|
||||
"-reset_timestamps", "1",
|
||||
out_pattern
|
||||
]
|
||||
r = subprocess.run(cmd, capture_output=True)
|
||||
if r.returncode != 0:
|
||||
return []
|
||||
|
||||
return sorted(str(p) for p in Path(output_dir).glob(f"{stem}_seg*{ext}"))
|
||||
|
||||
|
||||
def run_whisper(audio_path: str, output_dir: str, model: str,
|
||||
language: str = None) -> tuple:
|
||||
"""
|
||||
调用 whisper 命令转录音频,返回 (success: bool, txt_path_or_err: str)。
|
||||
whisper 会自动在 output_dir 生成 <stem>.txt 等文件。
|
||||
注意: 多次调用同 output_dir 且相同 stem 时会覆盖,需由调用方即时重命名。
|
||||
"""
|
||||
cmd = [
|
||||
"whisper", audio_path,
|
||||
"--model", model,
|
||||
"--output_dir", output_dir,
|
||||
"--output_format", "txt",
|
||||
"--verbose", "False",
|
||||
]
|
||||
if language:
|
||||
cmd += ["--language", language]
|
||||
|
||||
r = subprocess.run(cmd, capture_output=True, text=True)
|
||||
|
||||
stem = Path(audio_path).stem
|
||||
txt_path = os.path.join(output_dir, f"{stem}.txt")
|
||||
|
||||
if r.returncode == 0 and os.path.exists(txt_path):
|
||||
return True, txt_path
|
||||
|
||||
err = r.stderr[-500:] if r.stderr else "unknown error"
|
||||
return False, err
|
||||
|
||||
|
||||
def merge_txt_files(txt_files: list, merged_path: str) -> bool:
|
||||
"""将多个转录片段 txt 按顺序合并为一个完整文件"""
|
||||
try:
|
||||
with open(merged_path, "w", encoding="utf-8") as fout:
|
||||
for i, tf in enumerate(txt_files):
|
||||
if i > 0:
|
||||
fout.write("\n\n")
|
||||
with open(tf, encoding="utf-8") as fin:
|
||||
fout.write(fin.read().strip())
|
||||
return True
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
|
||||
# ─── 文件扫描 ──────────────────────────────────────────────────────────────────
|
||||
|
||||
def scan_audios(source: str) -> list:
|
||||
p = Path(source)
|
||||
if p.is_file():
|
||||
if p.suffix.lower() in AUDIO_EXTS:
|
||||
return [str(p.resolve())]
|
||||
else:
|
||||
log(f"不支持的音频格式: {p.suffix}", "ERR")
|
||||
return []
|
||||
elif p.is_dir():
|
||||
return sorted(str(f.resolve()) for f in p.rglob("*")
|
||||
if f.suffix.lower() in AUDIO_EXTS)
|
||||
else:
|
||||
log(f"路径不存在: {source}", "ERR")
|
||||
return []
|
||||
|
||||
|
||||
def txt_exists(txt_path: str) -> bool:
|
||||
return bool(txt_path) and os.path.exists(txt_path) and os.path.getsize(txt_path) > 0
|
||||
|
||||
|
||||
# ─── 输出文件命名 ──────────────────────────────────────────────────────────────
|
||||
|
||||
def compute_target_txt(out_dir: Path, stem: str, model: str,
|
||||
output_suffix: str, tag_model: bool) -> str:
|
||||
"""
|
||||
生成最终 txt 路径。命名规则:
|
||||
- tag_model=True (多模型模式): <stem>.<model>.txt
|
||||
- output_suffix 非空: <stem><output_suffix>.txt
|
||||
- 都无: <stem>.txt (默认)
|
||||
"""
|
||||
if tag_model:
|
||||
return str(out_dir / f"{stem}.{model}.txt")
|
||||
if output_suffix:
|
||||
return str(out_dir / f"{stem}{output_suffix}.txt")
|
||||
return str(out_dir / f"{stem}.txt")
|
||||
|
||||
|
||||
# ─── 状态展示 ──────────────────────────────────────────────────────────────────
|
||||
|
||||
def show_status(manifest: dict, all_audios: list):
|
||||
print("\n" + "=" * 70)
|
||||
print("📊 音频转录进度状态(按 模型 x 音频 展示)")
|
||||
print("=" * 70)
|
||||
|
||||
entries = {k: v for k, v in manifest["files"].items() if k.startswith("transcribe:")}
|
||||
done = sum(1 for v in entries.values() if v.get("transcribed"))
|
||||
failed = sum(1 for v in entries.values() if v.get("error"))
|
||||
print(f" Manifest 记录: {len(entries)} 已完成: {done} 失败: {failed}")
|
||||
print("=" * 70)
|
||||
|
||||
for a in all_audios:
|
||||
print(f"\n 📄 {Path(a).name}")
|
||||
matching = [(k, v) for k, v in entries.items() if k.endswith(f":{a}")]
|
||||
if not matching:
|
||||
print(f" ⏳ 无记录(未转录)")
|
||||
continue
|
||||
for key, entry in matching:
|
||||
m = entry.get("model", "?")
|
||||
suffix = entry.get("output_suffix") or ""
|
||||
suffix_tag = f" suffix={suffix}" if suffix else ""
|
||||
if entry.get("transcribed"):
|
||||
size = ""
|
||||
if entry.get("txt_path") and os.path.exists(entry["txt_path"]):
|
||||
size = f" {os.path.getsize(entry['txt_path']) // 1024}KB"
|
||||
lang = entry.get("language") or "auto"
|
||||
txt_name = Path(entry['txt_path']).name if entry.get('txt_path') else '?'
|
||||
print(f" ✅ [{m:6s}] lang={lang}{suffix_tag}{size} → {txt_name}")
|
||||
elif entry.get("error"):
|
||||
print(f" ❌ [{m:6s}]{suffix_tag} {entry['error'][:60]}")
|
||||
print()
|
||||
|
||||
|
||||
# ─── 核心:单个音频 + 单个模型 转录 ─────────────────────────────────────────────
|
||||
|
||||
def transcribe_one(audio: str, model: str, out_dir: Path, args,
|
||||
manifest: dict, manifest_path: Path,
|
||||
has_ffmpeg: bool, tag_model: bool) -> str:
|
||||
"""
|
||||
返回状态字符串:'ok' / 'skip' / 'skip-failed' / 'error:<msg>'
|
||||
"""
|
||||
stem = Path(audio).stem
|
||||
final_txt = compute_target_txt(out_dir, stem, model, args.output_suffix, tag_model)
|
||||
|
||||
# 多模型模式用 .<model>.txt 命名,output_suffix 在这种情况下强制为空
|
||||
effective_suffix = "" if tag_model else args.output_suffix
|
||||
entry = get_transcribe_entry(manifest, audio, model, effective_suffix)
|
||||
already_done = entry.get("transcribed") and txt_exists(entry.get("txt_path"))
|
||||
is_failed = bool(entry.get("error"))
|
||||
|
||||
if already_done:
|
||||
return "skip"
|
||||
if is_failed and not args.retry_failed:
|
||||
return "skip-failed"
|
||||
|
||||
try:
|
||||
duration = get_audio_duration(audio) if has_ffmpeg else 0
|
||||
need_split = args.split_size and duration > 0 and duration > args.split_size
|
||||
|
||||
if need_split:
|
||||
log(f" 音频时长 {duration:.0f}s > {args.split_size}s,启动切分转录")
|
||||
seg_dir = out_dir / f"{stem}_segments"
|
||||
seg_dir.mkdir(exist_ok=True)
|
||||
|
||||
segments = split_audio(audio, str(seg_dir), args.split_segment)
|
||||
if not segments:
|
||||
raise RuntimeError("音频切分失败")
|
||||
|
||||
# 每个模型独立子目录,避免片段 txt 互相覆盖
|
||||
seg_txt_dir = seg_dir / model
|
||||
seg_txt_dir.mkdir(exist_ok=True)
|
||||
log(f" 切分为 {len(segments)} 个片段,逐段转录中...")
|
||||
|
||||
seg_txts = []
|
||||
for j, seg in enumerate(segments, 1):
|
||||
ok, result = run_whisper(seg, str(seg_txt_dir), model, args.language)
|
||||
if ok:
|
||||
seg_txts.append(result)
|
||||
log(f" [{j}/{len(segments)}] {Path(seg).name} → done", "OK")
|
||||
else:
|
||||
raise RuntimeError(f"片段 {Path(seg).name} 转录失败: {result}")
|
||||
|
||||
if not merge_txt_files(seg_txts, final_txt):
|
||||
raise RuntimeError("合并转录片段失败")
|
||||
|
||||
size = os.path.getsize(final_txt)
|
||||
log(f" 合并 → {Path(final_txt).name} ({size // 1024} KB)", "OK")
|
||||
mark_transcribed(manifest, manifest_path, audio, final_txt,
|
||||
model, language=args.language,
|
||||
output_suffix=effective_suffix, segments=seg_txts)
|
||||
else:
|
||||
log(f" 🎙 转录中(模型: {model}, 语言: {args.language or 'auto'})...")
|
||||
ok, whisper_out = run_whisper(audio, str(out_dir), model, args.language)
|
||||
if not ok:
|
||||
raise RuntimeError(f"转录失败: {whisper_out}")
|
||||
|
||||
if whisper_out != final_txt:
|
||||
os.replace(whisper_out, final_txt)
|
||||
|
||||
size = os.path.getsize(final_txt)
|
||||
log(f" 输出: {Path(final_txt).name} ({size // 1024} KB)", "OK")
|
||||
mark_transcribed(manifest, manifest_path, audio, final_txt,
|
||||
model, language=args.language,
|
||||
output_suffix=effective_suffix)
|
||||
|
||||
return "ok"
|
||||
|
||||
except Exception as e:
|
||||
err_msg = str(e)
|
||||
mark_transcribe_failed(manifest, manifest_path, audio, model, err_msg,
|
||||
output_suffix=effective_suffix)
|
||||
# 清理不完整的输出
|
||||
if os.path.exists(final_txt) and os.path.getsize(final_txt) == 0:
|
||||
os.remove(final_txt)
|
||||
return f"error: {err_msg}"
|
||||
|
||||
|
||||
# ─── 主流程 ────────────────────────────────────────────────────────────────────
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(
|
||||
description="MP3 音频批量转录工具(基于 openai-whisper)",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog=__doc__
|
||||
)
|
||||
parser.add_argument("source", nargs="?", default=None,
|
||||
help="MP3 文件路径或包含音频的目录(--preload-models 时可省略)")
|
||||
parser.add_argument("--output", "-o", default=None,
|
||||
help="TXT 输出目录(默认与音频文件同目录)")
|
||||
parser.add_argument("--manifest", "-m", default=None,
|
||||
help="manifest.json 路径(默认在 source 目录下)")
|
||||
parser.add_argument("--model", default=DEFAULT_WHISPER_MODEL,
|
||||
choices=SUPPORTED_MODELS,
|
||||
help=f"Whisper 模型(默认 {DEFAULT_WHISPER_MODEL},可选:{'/'.join(SUPPORTED_MODELS)})")
|
||||
parser.add_argument("--language", default=None,
|
||||
help="音频语言代码,推荐 zh(中文) / en(英文)(默认自动检测)")
|
||||
parser.add_argument("--all-models", action="store_true",
|
||||
help=f"对每个音频依次使用所有支持的模型转录 ({'/'.join(SUPPORTED_MODELS)}),"
|
||||
f"输出为 <stem>.<model>.txt 便于对比精度")
|
||||
parser.add_argument("--output-suffix", default="",
|
||||
help="输出文件名后缀,如 _v1 → <stem>_v1.txt(与 --all-models 互斥时被忽略)")
|
||||
parser.add_argument("--split-size", type=int, default=None,
|
||||
help="切分阈值(秒):超过此时长的音频会被切分后分别转录")
|
||||
parser.add_argument("--split-segment", type=int, default=DEFAULT_SPLIT_SEGMENT_SECS,
|
||||
help=f"每个切分片段时长(秒,默认 {DEFAULT_SPLIT_SEGMENT_SECS})")
|
||||
parser.add_argument("--retry-failed", action="store_true",
|
||||
help="重新处理上次失败的文件")
|
||||
parser.add_argument("--status", action="store_true",
|
||||
help="仅查看进度状态,不执行转录")
|
||||
parser.add_argument("--preload-models", action="store_true",
|
||||
help=f"预下载所有支持的模型({'/'.join(SUPPORTED_MODELS)})到本地缓存后退出")
|
||||
args = parser.parse_args()
|
||||
|
||||
# ── 模式1: 仅预下载模型 ──
|
||||
if args.preload_models:
|
||||
if not preload_models(SUPPORTED_MODELS):
|
||||
sys.exit(1)
|
||||
return
|
||||
|
||||
# 从这里开始 source 是必需的
|
||||
if not args.source:
|
||||
parser.error("需要指定 source 参数(音频文件或目录),除非使用 --preload-models")
|
||||
|
||||
if not check_whisper():
|
||||
sys.exit(1)
|
||||
|
||||
has_ffmpeg = check_ffmpeg()
|
||||
if args.split_size and not has_ffmpeg:
|
||||
log("--split-size 需要 ffmpeg,但未找到 ffmpeg 命令", "ERR")
|
||||
sys.exit(1)
|
||||
|
||||
source_path = Path(args.source).resolve()
|
||||
base_dir = source_path if source_path.is_dir() else source_path.parent
|
||||
manifest_path = Path(args.manifest) if args.manifest else base_dir / MANIFEST_NAME
|
||||
|
||||
manifest = load_manifest(manifest_path)
|
||||
|
||||
all_audios = scan_audios(args.source)
|
||||
if not all_audios:
|
||||
log("未找到任何音频文件", "WARN")
|
||||
sys.exit(0)
|
||||
|
||||
# 决定使用哪些模型
|
||||
if args.all_models:
|
||||
models_to_run = SUPPORTED_MODELS
|
||||
if args.output_suffix:
|
||||
log(f"--all-models 已启用,--output-suffix 将被忽略(自动使用 .<model>.txt 命名)", "WARN")
|
||||
else:
|
||||
models_to_run = [args.model]
|
||||
|
||||
tag_model = args.all_models
|
||||
log(f"扫描到 {len(all_audios)} 个音频文件,将使用 {len(models_to_run)} 个模型:{', '.join(models_to_run)}")
|
||||
|
||||
# ── 模式2: 仅查看状态 ──
|
||||
if args.status:
|
||||
show_status(manifest, all_audios)
|
||||
return
|
||||
|
||||
# ── 模式3: 执行转录 ──
|
||||
total_jobs = len(all_audios) * len(models_to_run)
|
||||
stats = {"ok": 0, "skip": 0, "skip-failed": 0, "error": 0}
|
||||
failed_details = []
|
||||
|
||||
log(f"共 {total_jobs} 个 (音频 × 模型) 任务待处理")
|
||||
|
||||
job_idx = 0
|
||||
for audio in all_audios:
|
||||
audio_name = Path(audio).name
|
||||
audio_dir = Path(audio).parent
|
||||
out_dir = Path(args.output).resolve() if args.output else audio_dir
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
log(f"\n📄 {audio_name}", "STEP")
|
||||
|
||||
for model in models_to_run:
|
||||
job_idx += 1
|
||||
log(f" [{job_idx}/{total_jobs}] 模型: {model}")
|
||||
|
||||
t0 = time.time()
|
||||
result = transcribe_one(audio, model, out_dir, args, manifest,
|
||||
manifest_path, has_ffmpeg, tag_model)
|
||||
elapsed = time.time() - t0
|
||||
|
||||
if result == "ok":
|
||||
stats["ok"] += 1
|
||||
log(f" ⏱ 耗时 {elapsed:.1f}s", "OK")
|
||||
elif result == "skip":
|
||||
stats["skip"] += 1
|
||||
log(f" ⏭ 已完成,跳过", "INFO")
|
||||
elif result == "skip-failed":
|
||||
stats["skip-failed"] += 1
|
||||
log(f" ⚠ 上次失败,加 --retry-failed 可重试", "WARN")
|
||||
else:
|
||||
stats["error"] += 1
|
||||
failed_details.append((audio_name, model, result))
|
||||
log(f" ❌ {result}", "ERR")
|
||||
|
||||
# ── 汇总 ──
|
||||
print(f"\n{'='*70}")
|
||||
log(f"🏁 完成: 成功 {stats['ok']} 已跳过 {stats['skip']} "
|
||||
f"失败跳过 {stats['skip-failed']} 失败 {stats['error']}", "INFO")
|
||||
|
||||
if failed_details:
|
||||
log("失败详情:", "ERR")
|
||||
for audio_name, model, err in failed_details:
|
||||
print(f" - {audio_name} [模型:{model}] : {err[:80]}")
|
||||
print(" 运行时加 --retry-failed 可重新处理失败项")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user