添加 blogwatcher-daily 目录内容
This commit is contained in:
239
blogwatcher-daily/SKILL.md
Normal file
239
blogwatcher-daily/SKILL.md
Normal file
@@ -0,0 +1,239 @@
|
|||||||
|
---
|
||||||
|
name: blogwatcher-daily
|
||||||
|
description: RSS 订阅监控 + 每日笔记生成。使用 RSSHub + feedparser 抓取 31 个订阅,自动去重并存入 SQLite,新文章追加写入 Markdown 笔记。
|
||||||
|
version: 1.0
|
||||||
|
category: custom
|
||||||
|
tags: [rss, blog, monitoring, automation]
|
||||||
|
metadata:
|
||||||
|
author: Hermes Agent
|
||||||
|
last_updated: 2025-04-19
|
||||||
|
platform: macos, ubuntu
|
||||||
|
custom_skill_path: /Users/weishen/.hermes/skills/custom/
|
||||||
|
installation_note: "blogwatcher-daily 脚本在 Mac mini 本地运行(依赖 feedparser)。需要 RSSHub 服务(http://192.168.3.45:1200)访问 YouTube/Bilibili 等被墙源。"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Blogwatcher Daily
|
||||||
|
|
||||||
|
RSS 订阅监控 + 每日笔记生成自动化。
|
||||||
|
|
||||||
|
## 依赖
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip3 install feedparser
|
||||||
|
```
|
||||||
|
|
||||||
|
> feedparser 是 Python 最成熟的 RSS 解析库,支持 RSS 1.0/2.0/Atom、任意编码、畸形 XML。
|
||||||
|
|
||||||
|
## 实际抓取架构(重要)
|
||||||
|
|
||||||
|
```
|
||||||
|
YouTube 频道 URL
|
||||||
|
↓ 脚本自动识别并转为 RSSHub 格式
|
||||||
|
http://192.168.3.45:1200/youtube/channel/{id}
|
||||||
|
↓
|
||||||
|
curl → RSSHub → YouTube
|
||||||
|
|
||||||
|
普通 RSS(Engadget, Slashdot 等)
|
||||||
|
↓ 脚本直接访问(绕过 RSSHub /rss/ 路由)
|
||||||
|
原始 RSS URL
|
||||||
|
↓
|
||||||
|
curl → 目标网站
|
||||||
|
```
|
||||||
|
|
||||||
|
**关键发现**:RSSHub 的 `/rss/{url}` 路由不稳定(返回 RSSHub 欢迎页),
|
||||||
|
因此普通 RSS 源直接访问,不走 RSSHub 代理。
|
||||||
|
|
||||||
|
脚本内部 `build_fetch_url()` 根据 URL 类型自动选择路由:
|
||||||
|
- YouTube → RSSHub
|
||||||
|
- 已有的 RSSHub URL → 直接使用
|
||||||
|
- 其他 → 直接访问原始 URL
|
||||||
|
|
||||||
|
```
|
||||||
|
YouTube 频道 URL
|
||||||
|
↓ 脚本自动识别并转为 RSSHub 格式
|
||||||
|
http://192.168.3.45:1200/youtube/channel/{id}
|
||||||
|
↓
|
||||||
|
curl → RSSHub → YouTube
|
||||||
|
|
||||||
|
普通 RSS(Engadget, Slashdot 等)
|
||||||
|
↓ 脚本直接访问(绕过 RSSHub /rss/ 路由)
|
||||||
|
原始 RSS URL
|
||||||
|
↓
|
||||||
|
curl → 目标网站
|
||||||
|
```
|
||||||
|
|
||||||
|
**关键发现**:RSSHub 的 `/rss/{url}` 路由不稳定(返回 RSSHub 欢迎页),
|
||||||
|
因此普通 RSS 源直接访问,不走 RSSHub 代理。
|
||||||
|
|
||||||
|
脚本内部 `build_fetch_url()` 根据 URL 类型自动选择路由:
|
||||||
|
- YouTube → RSSHub
|
||||||
|
- 已有的 RSSHub URL → 直接使用
|
||||||
|
- 其他 → 直接访问原始 URL
|
||||||
|
|
||||||
|
## 目录结构
|
||||||
|
|
||||||
|
```
|
||||||
|
~/.hermes/skills/custom/blogwatcher-daily/
|
||||||
|
├── SKILL.md # 本文件
|
||||||
|
├── scripts/
|
||||||
|
│ └── blogwatcher-daily.py # 主脚本
|
||||||
|
├── subscriptions.txt # 订阅列表(name|URL)
|
||||||
|
└── blogwatcher.db # SQLite 数据库(自动创建)
|
||||||
|
```
|
||||||
|
|
||||||
|
## 使用方法
|
||||||
|
|
||||||
|
### 扫描订阅(默认)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py
|
||||||
|
```
|
||||||
|
|
||||||
|
- 扫描所有订阅,新增文章存入数据库
|
||||||
|
- 生成今日笔记:`~/Workspace/nexus/ishenwei/blogwatcher/YYYY-MM-DD.md`
|
||||||
|
|
||||||
|
### 添加订阅
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py \
|
||||||
|
--add "频道名" "URL"
|
||||||
|
```
|
||||||
|
|
||||||
|
**YouTube 频道自动转换**:直接贴 YouTube 频道 URL 或 feed URL,
|
||||||
|
脚本自动识别并转为 RSSHub 格式,无需手动拼接。
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 以下三种方式效果相同:
|
||||||
|
--add "Tech With Tim" "https://www.youtube.com/channel/UC4JX40jDee_tINbkjycV4Sg"
|
||||||
|
--add "Tech With Tim" "https://www.youtube.com/feeds/videos.xml?channel_id=UC4JX40jDee_tINbkjycV4Sg"
|
||||||
|
--add "Tech With Tim" "http://192.168.3.45:1200/youtube/channel/UC4JX40jDee_tINbkjycV4Sg"
|
||||||
|
```
|
||||||
|
|
||||||
|
### 列出订阅
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py --list
|
||||||
|
```
|
||||||
|
|
||||||
|
### 仅扫描(不写文件)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py --scan-only
|
||||||
|
```
|
||||||
|
|
||||||
|
### 强制回扫(`--all`)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 强制抓取每频道10篇(忽略已读状态),写入独立文件
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py --all
|
||||||
|
|
||||||
|
# 测试模式:只打印,不写文件
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py --scan-only --all
|
||||||
|
```
|
||||||
|
|
||||||
|
- 输出文件:`~/Workspace/nexus/ishenwei/blogwatcher/all-YYYY-MM-DD.md`(**覆盖模式**,每次运行覆盖)
|
||||||
|
- **不混入**每日报告 `YYYY-MM-DD.md`
|
||||||
|
- 每频道最多10篇
|
||||||
|
- 用途:历史内容回扫、测试订阅状态
|
||||||
|
|
||||||
|
### 标记已读
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py --mark-read
|
||||||
|
```
|
||||||
|
|
||||||
|
## 添加订阅示例
|
||||||
|
|
||||||
|
### YouTube 频道
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py \
|
||||||
|
--add "Tech With Tim" "http://192.168.3.45:1200/youtube/channel/UC4JX40jDee_tINbkjycV4Sg"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Bilibili
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py \
|
||||||
|
--add "B站频道" "http://192.168.3.45:1200/bilibili/user/{uid}"
|
||||||
|
```
|
||||||
|
|
||||||
|
### 普通 RSS
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py \
|
||||||
|
--add "博客名" "http://192.168.3.45:1200/rss/https://example.com/feed.xml"
|
||||||
|
```
|
||||||
|
|
||||||
|
## RSSHub vs 直接 RSS
|
||||||
|
|
||||||
|
脚本自动判断路由,不需要手动选择:
|
||||||
|
|
||||||
|
| 源类型 | 路由方式 | 示例 |
|
||||||
|
|--------|----------|------|
|
||||||
|
| YouTube 频道/用户/feed | RSSHub | 自动识别并代理 |
|
||||||
|
| RSSHub URL(已存储) | 直接使用 | `http://192.168.3.45:1200/youtube/channel/...` |
|
||||||
|
| 普通 RSS(Engadget, Slashdot 等) | 直接访问 | `https://www.engadget.com/rss.xml` |
|
||||||
|
|
||||||
|
> ⚠️ RSSHub 的 `/rss/{url}` 代理路由**不稳定**(实测试返回欢迎页),普通 RSS 不要走 RSSHub。
|
||||||
|
|
||||||
|
## 配置
|
||||||
|
|
||||||
|
| 配置项 | 默认值 |
|
||||||
|
|--------|--------|
|
||||||
|
| RSSHub 地址 | `http://192.168.3.45:1200`(可设置 `RSSHUB_URL` 环境变量覆盖) |
|
||||||
|
| 笔记输出目录 | `~/Workspace/nexus/ishenwei/blogwatcher/` |
|
||||||
|
| 数据库 | `~/.hermes/skills/custom/blogwatcher-daily/blogwatcher.db` |
|
||||||
|
| 订阅列表 | `~/.hermes/skills/custom/blogwatcher-daily/subscriptions.txt` |
|
||||||
|
|
||||||
|
## Cron Job 设置
|
||||||
|
|
||||||
|
推荐每天早上 6:00 自动执行:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cronjob --create \
|
||||||
|
--name "Blogwatcher Daily" \
|
||||||
|
--schedule "0 6 * * *" \
|
||||||
|
--repeat 999 \
|
||||||
|
--deliver "telegram:5038825565" \
|
||||||
|
--prompt "使用 blogwatcher-daily 技能执行每日 RSS 扫描。
|
||||||
|
|
||||||
|
执行步骤:
|
||||||
|
1. 加载 blogwatcher-daily 技能
|
||||||
|
2. 运行:python3 ~/.hermes/skills/custom/blogwatcher-daily/scripts/blogwatcher-daily.py
|
||||||
|
3. 检查输出,确认写入的笔记文件路径
|
||||||
|
4. 如果有新文章,简短汇总发给我"
|
||||||
|
```
|
||||||
|
|
||||||
|
## 数据流程
|
||||||
|
|
||||||
|
1. **加载订阅**:从 `subscriptions.txt` 读取所有 name|URL 对
|
||||||
|
2. **URL 路由**:`build_fetch_url()` 自动判断:
|
||||||
|
- YouTube → 转为 `http://192.168.3.45:1200/youtube/channel/{id}`
|
||||||
|
- RSSHub URL → 直接使用
|
||||||
|
- 其他 → 直接访问原始 URL(绕过不稳定的 `/rss/` 路由)
|
||||||
|
3. **抓取**:curl 获取 XML,SSL 跳过验证
|
||||||
|
4. **解析**:feedparser 提取 title、link、description、pub_date
|
||||||
|
5. **去重**:按 link 去重,已存在则跳过
|
||||||
|
6. **存储**:新文章存入 SQLite
|
||||||
|
7. **输出**:**追加写入** Markdown 笔记(同一文件多次扫描不会覆盖,仅追加新链接)
|
||||||
|
- 读取已有文件,用正则提取现有链接 → 去重
|
||||||
|
- 新文章追加 `## 📦 新增 N 篇` 分隔块
|
||||||
|
|
||||||
|
## 已知问题
|
||||||
|
|
||||||
|
| 问题 | 说明 | 状态 |
|
||||||
|
|------|------|------|
|
||||||
|
| How to of the Day | wikiHow 完全封了爬虫 | ❌ 不可用 |
|
||||||
|
| How-To Geek | 远程服务器关闭连接 | ❌ 不可用 |
|
||||||
|
|
||||||
|
> feedparser 已修复:電腦玩物 ✅、阿榮福利味 ✅、异次元软件世界 ✅、Slashdot ✅
|
||||||
|
|
||||||
|
## 注意事项
|
||||||
|
|
||||||
|
- 首次使用需 `pip3 install feedparser`
|
||||||
|
- YouTube 路由需要 RSSHub 容器内配置 `HTTP_PROXY`/`HTTPS_PROXY` 环境变量
|
||||||
|
- 每次最多取每源 10 篇(TED Talks 等除外,取 20 篇)
|
||||||
|
- 数据库自动创建,无需手动初始化
|
||||||
|
- OPML 文件可用 Python 解析后批量导入(参考 /wiki-ingest 流程中的 OPML 解析代码)
|
||||||
|
- **Markdown 追加模式**:同一天多次扫描不会覆盖旧内容,仅追加新链接(去重基于 URL)
|
||||||
BIN
blogwatcher-daily/blogwatcher.db
Normal file
BIN
blogwatcher-daily/blogwatcher.db
Normal file
Binary file not shown.
Binary file not shown.
436
blogwatcher-daily/scripts/blogwatcher-daily.py
Normal file
436
blogwatcher-daily/scripts/blogwatcher-daily.py
Normal file
@@ -0,0 +1,436 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""
|
||||||
|
Blogwatcher Daily - RSS Feed 监控脚本
|
||||||
|
放在 ~/.hermes/skills/research/blogwatcher-daily/scripts/blogwatcher-daily.py
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
python3 blogwatcher-daily.py # 扫描所有订阅,写入今日笔记
|
||||||
|
python3 blogwatcher-daily.py --list # 列出所有订阅
|
||||||
|
python3 blogwatcher-daily.py --add "频道名" "RSS URL" # 添加订阅
|
||||||
|
python3 blogwatcher-daily.py --scan-only # 只扫描不写入文件
|
||||||
|
python3 blogwatcher-daily.py --mark-read # 标记所有文章为已读
|
||||||
|
"""
|
||||||
|
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import sqlite3
|
||||||
|
import argparse
|
||||||
|
import feedparser
|
||||||
|
from datetime import datetime
|
||||||
|
from html import unescape
|
||||||
|
from urllib.request import Request, urlopen
|
||||||
|
from urllib.error import URLError, HTTPError
|
||||||
|
import ssl
|
||||||
|
|
||||||
|
# ========== 配置 ==========
|
||||||
|
SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
SKILL_DIR = os.path.dirname(SCRIPT_DIR) # .../research/blogwatcher-daily
|
||||||
|
DB_PATH = os.path.join(SKILL_DIR, "blogwatcher.db")
|
||||||
|
SUBSCRIPTIONS_FILE = os.path.join(SKILL_DIR, "subscriptions.txt")
|
||||||
|
OUTPUT_DIR = os.path.expanduser("~/Workspace/nexus/ishenwei/blogwatcher")
|
||||||
|
|
||||||
|
# RSSHub 基础地址(可配置)
|
||||||
|
RSSHUB_BASE = os.environ.get("RSSHUB_URL", "http://192.168.3.45:1200")
|
||||||
|
|
||||||
|
# YouTube channel ID 提取正则
|
||||||
|
YOUTUBE_CHANNEL_RE = re.compile(r'youtube\.com/channel/([\w-]+)', re.IGNORECASE)
|
||||||
|
YOUTUBE_FEED_RE = re.compile(r'youtube\.com/feeds/videos\.xml\?channel_id=([\w-]+)', re.IGNORECASE)
|
||||||
|
YOUTUBE_USER_RE = re.compile(r'youtube\.com/user/([\w-]+)', re.IGNORECASE)
|
||||||
|
|
||||||
|
# ========== 数据库 ==========
|
||||||
|
def get_db():
|
||||||
|
"""获取数据库连接"""
|
||||||
|
os.makedirs(os.path.dirname(DB_PATH), exist_ok=True)
|
||||||
|
conn = sqlite3.connect(DB_PATH)
|
||||||
|
conn.execute("""
|
||||||
|
CREATE TABLE IF NOT EXISTS articles (
|
||||||
|
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||||
|
feed_url TEXT NOT NULL,
|
||||||
|
channel_title TEXT,
|
||||||
|
title TEXT NOT NULL,
|
||||||
|
link TEXT UNIQUE NOT NULL,
|
||||||
|
description TEXT,
|
||||||
|
pub_date TEXT,
|
||||||
|
is_read INTEGER DEFAULT 0,
|
||||||
|
fetched_at TEXT DEFAULT CURRENT_TIMESTAMP
|
||||||
|
)
|
||||||
|
""")
|
||||||
|
conn.execute("""
|
||||||
|
CREATE TABLE IF NOT EXISTS subscriptions (
|
||||||
|
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||||
|
name TEXT NOT NULL,
|
||||||
|
feed_url TEXT UNIQUE NOT NULL,
|
||||||
|
enabled INTEGER DEFAULT 1,
|
||||||
|
created_at TEXT DEFAULT CURRENT_TIMESTAMP
|
||||||
|
)
|
||||||
|
""")
|
||||||
|
return conn
|
||||||
|
|
||||||
|
def is_article_new(conn, link):
|
||||||
|
"""检查文章是否已存在"""
|
||||||
|
cursor = conn.execute("SELECT is_read FROM articles WHERE link = ?", (link,))
|
||||||
|
row = cursor.fetchone()
|
||||||
|
return row is None
|
||||||
|
|
||||||
|
def save_article(conn, feed_url, channel_title, title, link, description, pub_date):
|
||||||
|
"""保存文章到数据库"""
|
||||||
|
try:
|
||||||
|
conn.execute("""
|
||||||
|
INSERT OR IGNORE INTO articles
|
||||||
|
(feed_url, channel_title, title, link, description, pub_date)
|
||||||
|
VALUES (?, ?, ?, ?, ?, ?)
|
||||||
|
""", (feed_url, channel_title, title, link, description, pub_date))
|
||||||
|
conn.commit()
|
||||||
|
except Exception as e:
|
||||||
|
print(f" ⚠️ 保存失败: {e}")
|
||||||
|
|
||||||
|
def mark_all_read(conn, feed_url=None):
|
||||||
|
"""标记所有文章为已读"""
|
||||||
|
if feed_url:
|
||||||
|
conn.execute("UPDATE articles SET is_read = 1 WHERE feed_url = ?", (feed_url,))
|
||||||
|
else:
|
||||||
|
conn.execute("UPDATE articles SET is_read = 1")
|
||||||
|
conn.commit()
|
||||||
|
|
||||||
|
# ========== RSS 获取 ==========
|
||||||
|
def build_fetch_url(url):
|
||||||
|
"""根据 URL 类型决定获取方式:
|
||||||
|
- YouTube channel/user/feed URL → 路由到 RSSHub
|
||||||
|
- RSSHub URL → 直接使用
|
||||||
|
- 其他 → 直接访问(绕过 RSSHub /rss/ 路由不稳定)
|
||||||
|
"""
|
||||||
|
if url.startswith(RSSHUB_BASE):
|
||||||
|
return url
|
||||||
|
# YouTube channel
|
||||||
|
m = YOUTUBE_CHANNEL_RE.search(url)
|
||||||
|
if m:
|
||||||
|
return f"{RSSHUB_BASE}/youtube/channel/{m.group(1)}"
|
||||||
|
# YouTube feed
|
||||||
|
m = YOUTUBE_FEED_RE.search(url)
|
||||||
|
if m:
|
||||||
|
return f"{RSSHUB_BASE}/youtube/channel/{m.group(1)}"
|
||||||
|
# YouTube user
|
||||||
|
m = YOUTUBE_USER_RE.search(url)
|
||||||
|
if m:
|
||||||
|
return f"{RSSHUB_BASE}/youtube/user/{m.group(1)}"
|
||||||
|
return url
|
||||||
|
|
||||||
|
def convert_to_stored_url(url):
|
||||||
|
"""将用户输入的 YouTube URL 转换为 RSSHub 存储格式"""
|
||||||
|
if url.startswith(RSSHUB_BASE):
|
||||||
|
return url
|
||||||
|
m = YOUTUBE_CHANNEL_RE.search(url)
|
||||||
|
if m:
|
||||||
|
return f"{RSSHUB_BASE}/youtube/channel/{m.group(1)}"
|
||||||
|
m = YOUTUBE_FEED_RE.search(url)
|
||||||
|
if m:
|
||||||
|
return f"{RSSHUB_BASE}/youtube/channel/{m.group(1)}"
|
||||||
|
m = YOUTUBE_USER_RE.search(url)
|
||||||
|
if m:
|
||||||
|
return f"{RSSHUB_BASE}/youtube/user/{m.group(1)}"
|
||||||
|
return url
|
||||||
|
|
||||||
|
def fetch_rss(url, timeout=15):
|
||||||
|
"""获取并解析 RSS feed"""
|
||||||
|
ctx = ssl.create_default_context()
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
|
||||||
|
fetch_url = build_fetch_url(url)
|
||||||
|
req = Request(fetch_url, headers={'User-Agent': 'Mozilla/5.0 (compatible; Blogwatcher/1.0)'})
|
||||||
|
|
||||||
|
try:
|
||||||
|
with urlopen(req, timeout=timeout, context=ctx) as response:
|
||||||
|
content = response.read().decode('utf-8', errors='replace')
|
||||||
|
return parse_rss(content, fetch_url)
|
||||||
|
except (URLError, HTTPError, Exception) as e:
|
||||||
|
print(f" ❌ 获取失败: {e}")
|
||||||
|
return None, []
|
||||||
|
|
||||||
|
def parse_rss(xml_content, fetch_url):
|
||||||
|
"""解析 RSS/Atom XML(支持 RSS 1.0/2.0/Atom)"""
|
||||||
|
parsed = feedparser.parse(xml_content, response_headers={'fetch_url': fetch_url})
|
||||||
|
channel_title = parsed.feed.get('title', 'Unknown') if parsed.feed else 'Unknown'
|
||||||
|
|
||||||
|
items = []
|
||||||
|
for entry in parsed.entries[:20]: # 最多取20篇
|
||||||
|
# 提取 link
|
||||||
|
link_text = ''
|
||||||
|
if hasattr(entry, 'link') and entry.link:
|
||||||
|
link_text = entry.link
|
||||||
|
elif hasattr(entry, 'id') and entry.id:
|
||||||
|
link_text = entry.id
|
||||||
|
|
||||||
|
# 提取 title
|
||||||
|
title_text = entry.get('title', 'No Title') or 'No Title'
|
||||||
|
|
||||||
|
# 提取 description/summary
|
||||||
|
desc_text = ''
|
||||||
|
if hasattr(entry, 'summary') and entry.summary:
|
||||||
|
desc_text = re.sub(r'<[^>]+>', '', entry.summary)
|
||||||
|
desc_text = unescape(desc_text)
|
||||||
|
desc_text = re.sub(r'\s+', ' ', desc_text).strip()[:200]
|
||||||
|
elif hasattr(entry, 'description') and entry.description:
|
||||||
|
desc_text = re.sub(r'<[^>]+>', '', entry.description)
|
||||||
|
desc_text = unescape(desc_text)
|
||||||
|
desc_text = re.sub(r'\s+', ' ', desc_text).strip()[:200]
|
||||||
|
|
||||||
|
# 提取 pubDate
|
||||||
|
pub_text = ''
|
||||||
|
if hasattr(entry, 'published') and entry.published:
|
||||||
|
pub_text = entry.published
|
||||||
|
elif hasattr(entry, 'updated') and entry.updated:
|
||||||
|
pub_text = entry.updated
|
||||||
|
|
||||||
|
items.append({
|
||||||
|
'title': title_text,
|
||||||
|
'link': link_text,
|
||||||
|
'description': desc_text,
|
||||||
|
'pub_date': pub_text
|
||||||
|
})
|
||||||
|
|
||||||
|
return channel_title, items
|
||||||
|
|
||||||
|
# ========== 订阅管理 ==========
|
||||||
|
def load_subscriptions():
|
||||||
|
"""从文件加载订阅列表"""
|
||||||
|
subs = []
|
||||||
|
if os.path.exists(SUBSCRIPTIONS_FILE):
|
||||||
|
with open(SUBSCRIPTIONS_FILE, 'r') as f:
|
||||||
|
for line in f:
|
||||||
|
line = line.strip()
|
||||||
|
if line and not line.startswith('#'):
|
||||||
|
parts = line.split('|')
|
||||||
|
if len(parts) >= 2:
|
||||||
|
name = parts[0].strip()
|
||||||
|
url = parts[1].strip()
|
||||||
|
enabled = parts[2].strip() != '0' if len(parts) > 2 else True
|
||||||
|
if enabled:
|
||||||
|
subs.append({'name': name, 'url': url})
|
||||||
|
return subs
|
||||||
|
|
||||||
|
def save_subscription(name, url):
|
||||||
|
"""添加订阅到文件"""
|
||||||
|
# 检查是否已存在
|
||||||
|
subs = load_subscriptions()
|
||||||
|
for sub in subs:
|
||||||
|
if sub['url'] == url:
|
||||||
|
print(f"⚠️ 订阅已存在: {name}")
|
||||||
|
return False
|
||||||
|
|
||||||
|
with open(SUBSCRIPTIONS_FILE, 'a') as f:
|
||||||
|
f.write(f"{name}|{url}\n")
|
||||||
|
print(f"✅ 已添加订阅: {name}")
|
||||||
|
return True
|
||||||
|
|
||||||
|
def list_subscriptions():
|
||||||
|
"""列出所有订阅"""
|
||||||
|
subs = load_subscriptions()
|
||||||
|
if not subs:
|
||||||
|
print("📭 暂无订阅")
|
||||||
|
return
|
||||||
|
|
||||||
|
print(f"\n📡 当前订阅 ({len(subs)} 个):\n")
|
||||||
|
for i, sub in enumerate(subs, 1):
|
||||||
|
print(f" [{i}] {sub['name']}")
|
||||||
|
print(f" {sub['url']}\n")
|
||||||
|
|
||||||
|
# ========== Markdown 输出 ==========
|
||||||
|
def generate_markdown(results, date_str=None):
|
||||||
|
"""生成 Markdown 格式"""
|
||||||
|
if date_str is None:
|
||||||
|
date_str = datetime.now().strftime("%Y-%m-%d")
|
||||||
|
|
||||||
|
md = f"# Blogwatcher Daily - {date_str}\n\n"
|
||||||
|
md += f"_生成时间: {datetime.now().strftime('%H:%M:%S')}_\n\n"
|
||||||
|
md += "---\n\n"
|
||||||
|
|
||||||
|
if not results:
|
||||||
|
md += "_今日无新文章_\n"
|
||||||
|
return md
|
||||||
|
|
||||||
|
# 按频道分组
|
||||||
|
channels = {}
|
||||||
|
for item in results:
|
||||||
|
ch = item['channel']
|
||||||
|
if ch not in channels:
|
||||||
|
channels[ch] = []
|
||||||
|
channels[ch].append(item)
|
||||||
|
|
||||||
|
for channel, items in channels.items():
|
||||||
|
md += f"## 【{channel}】\n\n"
|
||||||
|
for item in items:
|
||||||
|
md += f"- [{item['title']}]({item['link']})\n"
|
||||||
|
if item['description']:
|
||||||
|
md += f" {item['description'][:150]}...\n"
|
||||||
|
md += "\n"
|
||||||
|
md += "---\n\n"
|
||||||
|
|
||||||
|
return md
|
||||||
|
|
||||||
|
def save_daily_report(new_results, date_str=None):
|
||||||
|
"""保存每日报告(追加模式,避免覆盖)"""
|
||||||
|
if date_str is None:
|
||||||
|
date_str = datetime.now().strftime("%Y-%m-%d")
|
||||||
|
|
||||||
|
os.makedirs(OUTPUT_DIR, exist_ok=True)
|
||||||
|
filepath = os.path.join(OUTPUT_DIR, f"{date_str}.md")
|
||||||
|
|
||||||
|
# 读取已有内容,提取已有链接避免重复
|
||||||
|
existing_links = set()
|
||||||
|
if os.path.exists(filepath):
|
||||||
|
with open(filepath, 'r') as f:
|
||||||
|
existing_content = f.read()
|
||||||
|
# 简单提取已有链接(用于去重),去掉末尾的 ) 等标点
|
||||||
|
import re as _re
|
||||||
|
raw_links = _re.findall(r'https?://\S+', existing_content)
|
||||||
|
existing_links = set()
|
||||||
|
for l in raw_links:
|
||||||
|
# 去掉末尾的 ) 等非 URL 字符
|
||||||
|
while l and l[-1] in '),;:':
|
||||||
|
l = l[:-1]
|
||||||
|
existing_links.add(l)
|
||||||
|
else:
|
||||||
|
existing_content = ""
|
||||||
|
|
||||||
|
# 如果没有新结果,且文件已存在,直接返回
|
||||||
|
if not new_results and existing_content:
|
||||||
|
return filepath
|
||||||
|
|
||||||
|
# 生成新内容的 header
|
||||||
|
new_md = ""
|
||||||
|
new_articles = [a for a in new_results if a['link'] not in existing_links]
|
||||||
|
|
||||||
|
if new_articles:
|
||||||
|
new_md += f"\n## 📦 新增 {len(new_articles)} 篇 ({datetime.now().strftime('%H:%M:%S')})\n\n"
|
||||||
|
channels = {}
|
||||||
|
for item in new_articles:
|
||||||
|
ch = item['channel']
|
||||||
|
if ch not in channels:
|
||||||
|
channels[ch] = []
|
||||||
|
channels[ch].append(item)
|
||||||
|
for channel, items in channels.items():
|
||||||
|
new_md += f"### 【{channel}】\n\n"
|
||||||
|
for item in items:
|
||||||
|
new_md += f"- [{item['title']}]({item['link']})\n"
|
||||||
|
if item['description']:
|
||||||
|
new_md += f" {item['description'][:150]}...\n"
|
||||||
|
new_md += "\n"
|
||||||
|
|
||||||
|
# 追加写入
|
||||||
|
with open(filepath, 'a') as f:
|
||||||
|
f.write(new_md)
|
||||||
|
|
||||||
|
return filepath
|
||||||
|
|
||||||
|
# ========== 主流程 ==========
|
||||||
|
def scan_all(force_all=False, write_file=True):
|
||||||
|
"""扫描所有订阅
|
||||||
|
force_all: True 则忽略已读状态,每个频道强制抓10篇
|
||||||
|
write_file: True 则写入文件
|
||||||
|
"""
|
||||||
|
print("=" * 50)
|
||||||
|
print("Blogwatcher Daily Scan")
|
||||||
|
print("=" * 50)
|
||||||
|
|
||||||
|
subs = load_subscriptions()
|
||||||
|
if not subs:
|
||||||
|
print("\n📭 暂无订阅,请先添加:")
|
||||||
|
print(f" python3 {__file__} --add \"频道名\" \"RSS URL\"")
|
||||||
|
return
|
||||||
|
|
||||||
|
conn = get_db()
|
||||||
|
all_new_articles = []
|
||||||
|
new_count = 0
|
||||||
|
|
||||||
|
print(f"\n📡 开始扫描 {len(subs)} 个订阅...\n")
|
||||||
|
|
||||||
|
for sub in subs:
|
||||||
|
print(f"🔍 扫描: {sub['name']}")
|
||||||
|
|
||||||
|
channel_title, items = fetch_rss(sub['url'])
|
||||||
|
|
||||||
|
if items is None:
|
||||||
|
print(f" ⏭️ 跳过\n")
|
||||||
|
continue
|
||||||
|
|
||||||
|
new_in_feed = 0
|
||||||
|
items_to_save = items[:10] if force_all else [item for item in items[:10] if is_article_new(conn, item['link'])]
|
||||||
|
|
||||||
|
for item in items_to_save:
|
||||||
|
save_article(conn, sub['url'], channel_title, item['title'], item['link'],
|
||||||
|
item['description'], item['pub_date'])
|
||||||
|
all_new_articles.append({
|
||||||
|
'channel': channel_title,
|
||||||
|
**item
|
||||||
|
})
|
||||||
|
new_in_feed += 1
|
||||||
|
new_count += 1
|
||||||
|
|
||||||
|
print(f" ✅ {channel_title}: {len(items)} 篇, 新增 {new_in_feed} 篇\n")
|
||||||
|
|
||||||
|
conn.close()
|
||||||
|
|
||||||
|
# 生成报告
|
||||||
|
print("-" * 50)
|
||||||
|
print(f"📊 扫描完成: 共发现 {new_count} 篇新文章\n")
|
||||||
|
|
||||||
|
if all_new_articles and write_file:
|
||||||
|
if force_all:
|
||||||
|
# --all 模式:写入独立文件(不追加日常报告)
|
||||||
|
md = generate_markdown(all_new_articles)
|
||||||
|
filepath = os.path.join(OUTPUT_DIR, f"all-{datetime.now().strftime('%Y-%m-%d')}.md")
|
||||||
|
with open(filepath, 'w') as f:
|
||||||
|
f.write(md)
|
||||||
|
print(f"📝 已写入(force-all 模式): {filepath}")
|
||||||
|
else:
|
||||||
|
filepath = save_daily_report(all_new_articles)
|
||||||
|
print(f"📝 已写入: {filepath}")
|
||||||
|
elif not all_new_articles:
|
||||||
|
print("📭 今日无新文章")
|
||||||
|
|
||||||
|
return all_new_articles
|
||||||
|
|
||||||
|
# ========== CLI ==========
|
||||||
|
def main():
|
||||||
|
parser = argparse.ArgumentParser(description='Blogwatcher Daily RSS 监控脚本')
|
||||||
|
parser.add_argument('--list', '-l', action='store_true', help='列出所有订阅')
|
||||||
|
parser.add_argument('--add', nargs=2, metavar=('NAME', 'URL'), help='添加订阅')
|
||||||
|
parser.add_argument('--scan-only', action='store_true', help='仅扫描不写入文件')
|
||||||
|
parser.add_argument('--mark-read', action='store_true', help='标记所有为已读')
|
||||||
|
parser.add_argument('--date', '-d', help='指定日期 (YYYY-MM-DD)')
|
||||||
|
parser.add_argument('--rsshub', help='设置 RSSHub 地址')
|
||||||
|
parser.add_argument('--all', action='store_true', help='忽略已读状态,强制抓取每频道最新10篇(写入 all-YYYY-MM-DD.md,不追加日常报告)')
|
||||||
|
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
global RSSHUB_BASE
|
||||||
|
if args.rsshub:
|
||||||
|
RSSHUB_BASE = args.rsshub
|
||||||
|
|
||||||
|
if args.list:
|
||||||
|
list_subscriptions()
|
||||||
|
|
||||||
|
elif args.add:
|
||||||
|
name, url = args.add
|
||||||
|
# 如果 URL 不是完整地址,尝试添加 RSSHub 前缀
|
||||||
|
if not url.startswith('http'):
|
||||||
|
url = f"{RSSHUB_BASE}/{url}"
|
||||||
|
# YouTube URL 自动转为 RSSHub 格式
|
||||||
|
stored_url = convert_to_stored_url(url)
|
||||||
|
save_subscription(name, stored_url)
|
||||||
|
|
||||||
|
elif args.mark_read:
|
||||||
|
conn = get_db()
|
||||||
|
mark_all_read(conn)
|
||||||
|
conn.close()
|
||||||
|
print("✅ 已标记所有文章为已读")
|
||||||
|
|
||||||
|
elif args.scan_only:
|
||||||
|
scan_all(force_all=args.all, write_file=False)
|
||||||
|
|
||||||
|
else:
|
||||||
|
scan_all(force_all=args.all, write_file=True)
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
33
blogwatcher-daily/subscriptions.txt
Normal file
33
blogwatcher-daily/subscriptions.txt
Normal file
@@ -0,0 +1,33 @@
|
|||||||
|
Tech With Tim|http://192.168.3.45:1200/youtube/channel/UC4JX40jDee_tINbkjycV4Sg
|
||||||
|
Jon Law|http://192.168.3.45:1200/youtube/channel/UCQM_HxoKmza1simMUchYfVA
|
||||||
|
小白AI笔记|http://192.168.3.45:1200/youtube/channel/UCEhgrCbJ0eD1PQIoSPWxzMA
|
||||||
|
大有牧森 Austin Chou|http://192.168.3.45:1200/youtube/channel/UC3hsgc8SHJs1RDMEZBCAccA
|
||||||
|
huangyihe|http://192.168.3.45:1200/youtube/channel/UCPpdGTNbIKdiWgxCrbka4Zw
|
||||||
|
陶淵小明|http://192.168.3.45:1200/youtube/channel/UCqccJHWokUkv2ZaQjgRkB3A
|
||||||
|
惫懒の欧阳川|http://192.168.3.45:1200/youtube/channel/UCqh6eVqfg9l4C0c-q9ZBUog
|
||||||
|
Brayden Chen|http://192.168.3.45:1200/youtube/channel/UCFX6lKn8w8bIbZCdLc2PBpQ
|
||||||
|
TEDx Talks|http://192.168.3.45:1200/youtube/channel/UCsT0YIqwnpJCM-mx7-gSA4Q
|
||||||
|
灵姐说AI|http://192.168.3.45:1200/youtube/channel/UCMenHpvUet8myDntrTh7Ckw
|
||||||
|
Greyson Zhang|http://192.168.3.45:1200/youtube/channel/UCCNfpRADcNqqPK8PJZgHRNQ
|
||||||
|
李哈利Harry|http://192.168.3.45:1200/youtube/channel/UCEA4ZfPzWDHp72mlq7IvUcw
|
||||||
|
coursera|http://192.168.3.45:1200/youtube/channel/UCZ50rYSkYQG31YDEJm9Di_g
|
||||||
|
無遠弗屆教學教室|http://192.168.3.45:1200/youtube/channel/UCXDP8XCQyoldEiaIRhHz-Vw
|
||||||
|
零度解说|http://192.168.3.45:1200/youtube/channel/UCvijahEyGtvMpmMHBu4FS2w
|
||||||
|
Bloomberg Business|http://192.168.3.45:1200/youtube/channel/UCUMZ7gohGI9HcU9VNsr2FJQ
|
||||||
|
Reuters|http://192.168.3.45:1200/youtube/channel/UChqUTb7kYRX8-EiaN3XFrSQ
|
||||||
|
BBC中文网|http://192.168.3.45:1200/youtube/channel/UCb3TZ4SD_Ys3j4z0-8o6auA
|
||||||
|
理想生活实验室|https://www.toodaylab.com/feed
|
||||||
|
阿榮福利味|https://www.azofreeware.com/feeds/posts/default
|
||||||
|
重灌狂人|https://feeds.feedburner.com/briian
|
||||||
|
Engadget|https://www.engadget.com/rss.xml
|
||||||
|
How to of the Day|https://www.wikihow.com/feed.rss
|
||||||
|
异次元软件世界|https://feed.iplaysoft.com/
|
||||||
|
小众软件|https://feeds.appinn.com/appinns/
|
||||||
|
SaltTiger|https://salttiger.com/feed/
|
||||||
|
TED Talks Daily|https://feeds.acast.com/public/shows/67587e77c705e441797aff96
|
||||||
|
Slashdot|https://rss.slashdot.org/Slashdot/slashdot
|
||||||
|
AWS DevOps Blog|https://aws.amazon.com/blogs/devops/feed/
|
||||||
|
SRE WEEKLY|https://sreweekly.com/feed/
|
||||||
|
電腦玩物|https://feeds.feedburner.com/playpc
|
||||||
|
The Guardian AI|https://www.theguardian.com/technology/artificialintelligenceai/rss
|
||||||
|
DowJones World News|https://feeds.content.dowjones.io/public/rss/RSSWorldNews
|
||||||
Reference in New Issue
Block a user