Add baoyu-skills package

This commit is contained in:
2026-09-13 07:37:15 +08:00
commit 832004bd6f
782 changed files with 251339 additions and 0 deletions

View File

@@ -0,0 +1,107 @@
# EXTEND.md Schema for baoyu-translate
## Format
EXTEND.md uses YAML format:
```yaml
# Default target language (ISO code or common name)
target_language: zh-CN
# Default translation mode
default_mode: normal # quick | normal | refined
# Target audience (affects annotation depth and register)
audience: general # general | technical | academic | business | or custom string
# Translation style preference
style: storytelling # storytelling | formal | technical | literal | academic | business | humorous | conversational | elegant | or custom string
# Word count threshold to trigger chunked translation
chunk_threshold: 4000
# Max words per chunk
chunk_max_words: 5000
# Custom glossary (merged with built-in glossary)
# CLI --glossary flag overrides these
# Supports inline entries and/or file paths
glossary:
- from: "Reinforcement Learning"
to: "强化学习"
- from: "Transformer"
to: "Transformer"
note: "Keep English"
# Load glossary from external file(s)
# Supports absolute path or relative to EXTEND.md location
# File format: markdown table with | from | to | note | columns,
# or YAML list of {from, to, note} entries
glossary_files:
- ./my-glossary.md
- /path/to/shared-glossary.yaml
# Language-pair specific glossaries
glossaries:
en-zh:
- from: "AI Agent"
to: "AI 智能体"
ja-zh:
- from: "人工知能"
to: "人工智能"
```
## Fields
| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `target_language` | string | `zh-CN` | Default target language code |
| `default_mode` | string | `normal` | Default translation mode (`quick` / `normal` / `refined`) |
| `audience` | string | `general` | Target reader profile (`general` / `technical` / `academic` / `business` / custom) |
| `style` | string | `storytelling` | Translation style (`storytelling` / `formal` / `technical` / `literal` / `academic` / `business` / `humorous` / `conversational` / `elegant` / custom) |
| `chunk_threshold` | number | `4000` | Word count threshold to trigger chunked translation |
| `chunk_max_words` | number | `5000` | Max words per chunk |
| `glossary` | array | `[]` | Universal glossary entries (inline) |
| `glossary_files` | array | `[]` | External glossary file paths (absolute or relative to EXTEND.md) |
| `glossaries` | object | `{}` | Language-pair specific glossary entries |
## Glossary Entry
| Field | Required | Description |
|-------|----------|-------------|
| `from` | yes | Source term |
| `to` | yes | Target translation |
| `note` | no | Usage note (e.g., "Keep English", "Only in tech context") |
## Glossary File Format
External glossary files (`glossary_files`) support two formats:
**Markdown table** (`.md`):
```markdown
| from | to | note |
|------|----|------|
| Reinforcement Learning | 强化学习 | |
| Transformer | Transformer | Keep English |
```
**YAML list** (`.yaml` / `.yml`):
```yaml
- from: "Reinforcement Learning"
to: "强化学习"
- from: "Transformer"
to: "Transformer"
note: "Keep English"
```
Paths can be absolute or relative to the EXTEND.md file location.
## Priority
1. CLI `--glossary` file entries
2. EXTEND.md `glossaries[pair]` entries
3. EXTEND.md `glossary` entries (inline)
4. EXTEND.md `glossary_files` entries (in listed order, later files override earlier)
5. Built-in glossary (e.g., `references/glossary-en-zh.md`)
Later entries override earlier ones for the same source term.

View File

@@ -0,0 +1,169 @@
---
name: first-time-setup
description: First-time setup flow for baoyu-translate preferences
---
# First-Time Setup
## Overview
When no EXTEND.md is found, guide user through preference setup.
**BLOCKING OPERATION**: This setup MUST complete before ANY translation. Do NOT:
- Start translating content
- Ask about files or output paths
- Proceed to any workflow steps
ONLY ask the questions in this setup flow, save EXTEND.md, then continue.
## Setup Flow
```
No EXTEND.md found
|
v
+---------------------+
| AskUserQuestion |
| (all questions) |
+---------------------+
|
v
+---------------------+
| Create EXTEND.md |
+---------------------+
|
v
Continue translation
```
## Questions
**Language**: Use user's input language or saved language preference.
Use AskUserQuestion with ALL questions in ONE call:
### Question 1: Target Language
```yaml
header: "Target Language"
question: "Default target language?"
options:
- label: "简体中文 zh-CN (Recommended)"
description: "Translate to Simplified Chinese"
- label: "繁體中文 zh-TW"
description: "Translate to Traditional Chinese"
- label: "English en"
description: "Translate to English"
- label: "日本語 ja"
description: "Translate to Japanese"
```
Note: User may type a custom language code.
### Question 2: Translation Mode
```yaml
header: "Mode"
question: "Default translation mode?"
options:
- label: "Normal (Recommended)"
description: "Analyze content first, then translate"
- label: "Quick"
description: "Direct translation, no analysis"
- label: "Refined"
description: "Full workflow: analyze → translate → review → polish"
```
### Question 3: Target Audience
```yaml
header: "Audience"
question: "Default target audience?"
options:
- label: "General readers (Recommended)"
description: "Plain language, more translator's notes for jargon"
- label: "Technical"
description: "Developers/engineers, less annotation on tech terms"
- label: "Academic"
description: "Formal register, precise terminology"
- label: "Business"
description: "Business-friendly tone, explain tech concepts"
```
Note: User may type a custom audience description.
### Question 4: Translation Style
```yaml
header: "Style"
question: "Translation style?"
options:
- label: "Storytelling (Recommended)"
description: "Engaging narrative flow, smooth transitions"
- label: "Formal"
description: "Professional, structured, neutral tone"
- label: "Technical"
description: "Precise, documentation-style, concise"
- label: "Literal"
description: "Close to original structure"
- label: "Academic"
description: "Scholarly, rigorous, formal register"
- label: "Business"
description: "Concise, results-focused, action-oriented"
- label: "Humorous"
description: "Preserves humor, witty, playful"
- label: "Conversational"
description: "Casual, friendly, spoken-like"
- label: "Elegant"
description: "Literary, polished, aesthetically refined"
```
Note: User may type a custom style description.
### Question 5: Save Location
```yaml
header: "Save"
question: "Where to save preferences?"
options:
- label: "User (Recommended)"
description: "$HOME/.baoyu-skills/ (all projects)"
- label: "Project"
description: ".baoyu-skills/ (this project only)"
```
## Save Locations
| Choice | Path | Scope |
|--------|------|-------|
| User | `$HOME/.baoyu-skills/baoyu-translate/EXTEND.md` | All projects |
| Project | `.baoyu-skills/baoyu-translate/EXTEND.md` | Current project |
## After Setup
1. Create directory if needed
2. Write EXTEND.md with selected values
3. Confirm: "Preferences saved to [path]"
4. Mention: "You can add custom glossary terms to EXTEND.md anytime. See the `glossary` section in the file for the format."
5. Continue with translation using saved preferences
## EXTEND.md Template
```yaml
target_language: [zh-CN/zh-TW/en/ja/...]
default_mode: [quick/normal/refined]
audience: [general/technical/academic/business/custom]
style: [storytelling/formal/technical/literal/academic/business/humorous/conversational/elegant]
# Custom glossary (optional) — add your own term translations here
# glossary:
# - from: "Term"
# to: "翻译"
# - from: "Another Term"
# to: "另一个翻译"
# note: "Usage context"
```
## Modifying Preferences Later
Users can edit EXTEND.md directly or delete it to trigger setup again.

View File

@@ -0,0 +1,21 @@
# English → Chinese Glossary
Terms where standard translation is non-obvious or easily mistranslated. Common terms with straightforward translations (e.g., Machine Learning → 机器学习) are omitted — the model already knows these.
| English | Chinese | Notes |
|---------|---------|-------|
| AI Agent | AI 智能体 | |
| Vibe Coding | 凭感觉编程 | |
| the Bitter Lesson | 苦涩的教训 | Rich Sutton's essay |
| Context Engineering | 上下文工程 | |
| AI Wrapper | AI 套壳 | |
| RLHF | 基于人类反馈的强化学习 | |
| Hallucination | 幻觉 | AI-specific meaning |
| Alignment | 对齐 | AI safety context |
| Guardrails | 护栏 | AI safety context |
| Agentic | 智能体化的 | |
| Grounding | 基础化/落地 | Context-dependent |
| Embedding | 嵌入/向量化 | Context-dependent |
| Moat | 护城河 | Business context |
| Flywheel | 飞轮效应 | |
| Boilerplate | 样板代码 | |

View File

@@ -0,0 +1,151 @@
# Translation Workflow Details
This file provides detailed guidelines for each workflow step. Steps are shared across modes:
- **Quick**: Translate only (no steps from this file)
- **Normal**: Step 1 (Analysis) → Translate
- **Refined**: Step 1 (Analysis) → Step 2 (Draft) → Step 3 (Review) → Step 4 (Revision) → Step 5 (Polish)
- **Normal → Upgrade**: After normal mode, user can continue with Step 3 → Step 4 → Step 5
All intermediate results are saved as files in the output directory.
## Step 1: Content Analysis
Before translating, analyze the source material. Save analysis to `01-analysis.md` in the output directory.
### 1.1 Content Summary
- What is this content about? What is the core argument?
- Author background, stance, and writing context
- Purpose and intended audience of the original
### 1.2 Terminology
- List technical terms, proper nouns, brand names, acronyms
- Cross-reference with loaded glossaries
- For terms not in glossary, determine standard translations
- Record in a terminology table
### 1.3 Tone & Style
- Formal or conversational? Humor, metaphor, cultural references?
- What register is appropriate for the translation given the target audience?
### 1.4 Translation Challenges
Identify what may cause difficulty in translation:
- **Comprehension gaps**: Terms or references that target readers may not understand — note what explanation is needed
- **Figurative language**: Metaphors, idioms, expressions that don't translate literally — note intended meaning and target-language approach (interpret / substitute / retain)
- **Structural challenges**: Long complex sentences, wordplay, puns, or humor that needs creative adaptation
**Save `01-analysis.md`** with:
```
## Content Summary
[Core argument, author, context, purpose]
## Terminology
[term → translation, ...]
## Tone & Style
[assessment]
## Translation Challenges
- [term/passage] → [challenge type] → [suggested approach]
- ...
```
## Step 2: Assemble Translation Prompt
Main agent reads `01-analysis.md` and assembles a complete translation prompt using [references/subagent-prompt-template.md](subagent-prompt-template.md). Inline the following from analysis:
- **Target style**: Resolved style preset + source voice assessment from §1.3
- **Content background**: Summary from §1.1
- **Glossary**: Merged glossary with analysis-extracted terms from §1.2
- **Translation challenges**: All challenges from §1.4
Save to `02-prompt.md`. This prompt is used by the subagent (chunked) or by the main agent itself (non-chunked).
## Step 3: Initial Draft
Save to `03-draft.md` in the output directory.
For chunked content, the subagent produces this draft (merged from chunk translations). For non-chunked content, the main agent produces it directly.
Translate the full content following `02-prompt.md`. Apply all **Translation principles** from SKILL.md.
## Step 4: Critical Review
The main agent critically reviews the draft against the source. Save review findings to `04-critique.md`. This step produces **diagnosis only** — no rewriting yet.
### 4.1 Accuracy
- Compare each paragraph against the original
- Verify facts, numbers, dates, proper nouns
- Flag content accidentally added, removed, or altered
- Check terminology consistency with glossary
### 4.2 Native Voice
- Flag sentences that read as "translated" rather than "written" — unnatural word order, calques, stiff phrasing
- For CJK targets: check for unnecessary connectives (因此/然而/此外), passive voice abuse (被/由/受到), noun pile-ups, over-nominalization
- Flag metaphors translated literally that sound unnatural in the target language
- Check emotional connotations are preserved, not flattened
- Note where sentence restructuring would improve readability
### 4.3 Notes & Adaptation
- Are translator's notes accurate, concise, and genuinely helpful?
- Flag missed comprehension challenges that need notes, and over-annotations on obvious terms
- Were translation strategies from `02-prompt.md` followed?
- Do cultural references work in the target language?
**Save `04-critique.md`** with:
```
## Accuracy
- [issue]: [location] — [description]
## Native Voice
- [issue]: [example] → [suggested fix]
## Notes & Adaptation
- [add/remove/revise]: [term/passage] — [reason]
## Summary
[Overall assessment: X critical issues, Y improvements]
```
## Step 5: Revision
Apply all findings from `04-critique.md` to produce a revised translation. Save to `05-revision.md`.
Read `03-draft.md` and `04-critique.md`, fix all accuracy issues, rewrite unnatural expressions, adjust notes, and improve flow.
## Step 6: Polish
Save final version to `translation.md`.
Final pass on `05-revision.md` for publication quality:
- Read the entire translation as a standalone piece — does it flow as native content?
- Smooth remaining rough transitions
- Ensure consistent narrative voice and style throughout
- Final terminology consistency check
- Verify formatting is preserved correctly
## Subagent Responsibility
Each subagent (one per chunk) is responsible **only** for producing the initial draft of its chunk (Step 3). The main agent assembles the shared prompt (Step 2), spawns all subagents in parallel, then takes over for critical review (Step 4), revision (Step 5), and polish (Step 6).
## Chunked Refined Translation
When content exceeds the chunk threshold and uses refined mode:
1. Main agent runs analysis (Step 1) on the **entire** document first → `01-analysis.md`
2. Main agent assembles translation prompt → `02-prompt.md`
3. Split into chunks → `chunks/`
4. Spawn one subagent per chunk in parallel (each reads `02-prompt.md` for shared context) → merge all results into `03-draft.md`
5. Main agent critically reviews the merged draft → `04-critique.md`
6. Main agent revises based on critique → `05-revision.md`
7. Main agent polishes → `translation.md`
8. Final cross-chunk consistency check: terminology, narrative flow, transitions at chunk boundaries

View File

@@ -0,0 +1,79 @@
# Subagent Translation Prompt Template
Two parts:
1. **`02-prompt.md`** — Shared context (saved to output directory). Contains background, glossary, challenges, and principles. No task-specific instructions.
2. **Subagent spawn prompt** — Task instructions passed when spawning each subagent. One subagent per chunk (or per source file if non-chunked).
The main agent reads `01-analysis.md` (if exists), inlines all relevant context into `02-prompt.md`, then spawns subagents in parallel with task instructions referencing that file.
Replace `{placeholders}` with actual values. Omit sections marked "if analysis exists" for quick mode.
---
## Part 1: `02-prompt.md` (shared context, saved as file)
```markdown
You are a professional translator. Your task is to translate markdown content from {source_lang} to {target_lang}.
## Target Audience & Style
**Audience**: {audience description}
**Target style**: {style description — e.g., "storytelling: engaging narrative flow, smooth transitions, vivid phrasing" or custom style from user}
**Source voice** (from analysis, if exists): {Brief description of the original author's voice — formal/conversational, humor, register, sentence rhythm.}
## Content Background
{Inlined from 01-analysis.md if analysis exists: content summary, core argument, author background, context.}
## Glossary
Apply these term translations consistently. First occurrence: include original in parentheses.
{Merged glossary — one per line: English → Translation}
## Translation Challenges
{Inlined from 01-analysis.md §1.4 if analysis exists. Comprehension gaps, figurative language, structural challenges with suggested approaches:}
- **{term/passage}**: {challenge type} → {suggested approach}
## Translation Principles
Rewrite the content into natural, engaging {target_lang} — not merely translate it. Every sentence should read as if a skilled native writer composed it from scratch.
- **Accuracy first**: Facts, data, and logic must match the original exactly
- **Natural flow**: Use idiomatic {target_lang} word order. Break long source sentences into shorter, natural ones. Interpret metaphors and idioms by intended meaning, not word-for-word
- **Terminology**: Use glossary translations consistently. Annotate with original in parentheses on first occurrence of specialized terms
- **Preserve format**: Keep all markdown formatting (headings, bold, italic, images, links, code blocks)
- **Proactive interpretation**: For jargon or concepts the target audience may lack context for, add concise explanations in **bold parentheses** `(**解释**)`. Keep annotations few — only where genuinely needed
```
---
## Part 2: Subagent spawn prompt (passed as Agent tool prompt)
### Chunked mode (one subagent per chunk, all spawned in parallel)
```
Read the translation instructions from: {output_dir}/02-prompt.md
You are translating chunk {NN} of {total_chunks}.
Context: {brief description of what this chunk covers and where it sits in the overall argument}
Translate this chunk:
1. Read `{output_dir}/chunks/chunk-{NN}.md`
2. Translate following the instructions in 02-prompt.md
3. Save translation to `{output_dir}/chunks/chunk-{NN}-draft.md`
```
### Non-chunked mode
```
Read the translation instructions from: {output_dir}/02-prompt.md
Translate the source file and save the result:
1. Read `{source_file_path}`
2. Save translation to `{output_path}`
```

View File

@@ -0,0 +1,25 @@
# Workflow Mechanics
Details for source materialization, output directory creation, and conflict resolution.
## Materialize Source
| Input Type | Action |
|------------|--------|
| File | Use as-is (no copy needed) |
| Inline text | Save to `translate/{slug}.md` |
| URL | Fetch content, save to `translate/{slug}.md` |
`{slug}`: 2-4 word kebab-case slug derived from content topic.
## Create Output Directory
Create a subdirectory next to the source file: `{source-dir}/{source-basename}-{target-lang}/`
Examples:
- `posts/article.md` → `posts/article-zh/`
- `translate/ai-future.md` → `translate/ai-future-zh/`
## Conflict Resolution
If the output directory already exists, rename the existing one to `{name}.backup-YYYYMMDD-HHMMSS/` before creating the new one. Never overwrite existing results.