Add baoyu-skills package
This commit is contained in:
@@ -0,0 +1,35 @@
|
||||
---
|
||||
name: codex-image2-fallback
|
||||
description: Fallback behavior when baoyu-image-gen lacks OpenAI API credentials but Codex/native image generation is available
|
||||
---
|
||||
|
||||
# Codex Image2 Fallback
|
||||
|
||||
When using `baoyu-image-gen` with `--provider openai --model gpt-image-2.5-flare`, the CLI can fail with:
|
||||
|
||||
```text
|
||||
OPENAI_API_KEY is required. Codex/ChatGPT desktop login does not automatically grant OpenAI Images API access to this script.
|
||||
```
|
||||
|
||||
This is expected. The `openai` provider uses the public OpenAI Images API and needs `OPENAI_API_KEY`. Codex / ChatGPT image2 entitlement is a separate runtime-native path.
|
||||
|
||||
## Practical fallback pattern
|
||||
|
||||
1. Try `baoyu-image-gen` when provider credentials are available.
|
||||
2. If it fails only because `OPENAI_API_KEY` is missing, do not leave the user waiting.
|
||||
3. Prefer a Codex/native raster backend in this order:
|
||||
- Codex runtime native `imagegen` skill/tool, if available.
|
||||
- `baoyu-image-gen --provider codex-cli` (preferred — wraps the bundled `scripts/codex-imagegen/main.ts`; the underlying repo-level package lives at `packages/baoyu-codex-imagegen/src/main.ts` for standalone callers), if `codex` CLI is installed/logged in.
|
||||
- Hermes native `image_generate`, if available.
|
||||
4. Be transparent about reference-image behavior:
|
||||
- If the fallback backend accepts references, pass the reference images.
|
||||
- If it does not, derive a concise identity-preserving prompt from the references and state that it is a text-description fallback, not strict reference-image editing.
|
||||
5. Return the generated media path or structured backend error promptly.
|
||||
|
||||
## User-facing wording
|
||||
|
||||
Use concise wording such as:
|
||||
|
||||
> The OpenAI API path needs `OPENAI_API_KEY`; Codex login is a separate image2 backend. I used the available Codex/native image backend instead. Reference images were [passed directly / reconstructed from visual traits].
|
||||
|
||||
Avoid implying that `baoyu-image-gen --provider openai` can use Codex OAuth without a dedicated provider implementation.
|
||||
@@ -0,0 +1,18 @@
|
||||
# Codex OAuth vs OpenAI API key for baoyu-image-gen
|
||||
|
||||
`baoyu-image-gen --provider openai` uses the standard OpenAI Images API and requires `OPENAI_API_KEY`. It calls OpenAI-compatible image endpoints such as `/images/generations` and `/images/edits`.
|
||||
|
||||
Codex / ChatGPT login is different. Codex image generation is driven by Codex OAuth and the Codex runtime's `image_gen` capability, not by the public OpenAI Images API key path. A Codex OAuth token is not a drop-in replacement for `OPENAI_API_KEY`, and setting `OPENAI_BASE_URL` to a Codex backend will not make baoyu-image-gen's existing `openai` provider work because the auth, route, and payload shape differ.
|
||||
|
||||
## What to use instead
|
||||
|
||||
- If running inside Codex and the native `imagegen` skill/tool is available, use it directly.
|
||||
- If running outside Codex but the `codex` CLI is installed and logged in, call `baoyu-image-gen --provider codex-cli` (preferred). It spawns the bundled `scripts/codex-imagegen/main.ts` and surfaces its retry/cache/log machinery through baoyu-image-gen's standard CLI + batch flow. Standalone callers outside this skill can run the same code at `packages/baoyu-codex-imagegen/src/main.ts`. Both invoke `codex exec` and the Codex `image_gen` tool; no `OPENAI_API_KEY` is required.
|
||||
- If running inside Hermes and a native `image_generate` tool is available, use that as a runtime-native fallback. Be explicit about whether reference images are passed directly or only reconstructed from extracted traits.
|
||||
- `baoyu-image-gen` already exposes a distinct `codex-cli` provider (wraps the bundled `scripts/codex-imagegen/`); do not modify the existing `openai` provider to add Codex OAuth.
|
||||
|
||||
## Reference-image prompting note
|
||||
|
||||
When using actual reference images for identity preservation, avoid long generic descriptions of the subject. Long descriptions can cause the model to synthesize a new similar-looking person/object. Prefer direct wording:
|
||||
|
||||
> Use the person/object in the reference image(s) as the same identity. Do not redesign it or create a similar-looking new subject. Only change scene, clothing, pose, lighting, rendering style, and composition.
|
||||
@@ -0,0 +1,392 @@
|
||||
---
|
||||
name: first-time-setup
|
||||
description: First-time setup and default model selection flow for baoyu-image-gen
|
||||
---
|
||||
|
||||
# First-Time Setup
|
||||
|
||||
## Overview
|
||||
|
||||
Triggered when:
|
||||
1. No EXTEND.md found → full setup (provider + model + preferences)
|
||||
2. EXTEND.md found but `default_model.[provider]` is null → model selection only
|
||||
|
||||
## Setup Flow
|
||||
|
||||
```
|
||||
No EXTEND.md found EXTEND.md found, model null
|
||||
│ │
|
||||
▼ ▼
|
||||
┌─────────────────────┐ ┌──────────────────────┐
|
||||
│ AskUserQuestion │ │ AskUserQuestion │
|
||||
│ (full setup) │ │ (model only) │
|
||||
└─────────────────────┘ └──────────────────────┘
|
||||
│ │
|
||||
▼ ▼
|
||||
┌─────────────────────┐ ┌──────────────────────┐
|
||||
│ Create EXTEND.md │ │ Update EXTEND.md │
|
||||
└─────────────────────┘ └──────────────────────┘
|
||||
│ │
|
||||
▼ ▼
|
||||
Continue Continue
|
||||
```
|
||||
|
||||
## Flow 1: No EXTEND.md (Full Setup)
|
||||
|
||||
**Language**: Use user's input language or saved language preference.
|
||||
|
||||
Use AskUserQuestion with ALL questions in ONE call:
|
||||
|
||||
### Question 1: Default Provider
|
||||
|
||||
```yaml
|
||||
header: "Provider"
|
||||
question: "Default image generation provider?"
|
||||
options:
|
||||
- label: "Google (Recommended)"
|
||||
description: "Gemini multimodal - high quality, reference images, flexible sizes"
|
||||
- label: "OpenAI"
|
||||
description: "GPT Image 2 - latest OpenAI image model, reference-image workflows"
|
||||
- label: "Azure OpenAI"
|
||||
description: "Azure-hosted GPT Image deployments with resource-specific routing"
|
||||
- label: "OpenRouter"
|
||||
description: "Router for Gemini/FLUX/OpenAI-compatible image models"
|
||||
- label: "DashScope"
|
||||
description: "Alibaba Cloud - Qwen-Image, strong Chinese/English text rendering"
|
||||
- label: "Z.AI"
|
||||
description: "GLM-image, strong poster and text-heavy image generation"
|
||||
- label: "MiniMax"
|
||||
description: "MiniMax image generation with subject-reference character workflows"
|
||||
- label: "Replicate"
|
||||
description: "Curated Replicate image families - nano-banana-2, Seedream, and Wan image models"
|
||||
- label: "Agnes"
|
||||
description: "Sapiens AI Agnes - optimized for high information density, complex layouts, reference-image support"
|
||||
```
|
||||
|
||||
### Question 2: Default Google Model
|
||||
|
||||
Only show if user selected Google or auto-detect (no explicit provider).
|
||||
|
||||
```yaml
|
||||
header: "Google Model"
|
||||
question: "Default Google image generation model?"
|
||||
options:
|
||||
- label: "gemini-3-pro-image (Recommended)"
|
||||
description: "Highest quality, best for production use"
|
||||
- label: "gemini-3.1-flash-image"
|
||||
description: "Fast generation, good quality, lower cost"
|
||||
- label: "gemini-3.1-flash-lite-image"
|
||||
description: "Cheapest and fastest; 1K output only"
|
||||
- label: "gemini-3-flash-preview"
|
||||
description: "Fast generation, balanced quality and speed"
|
||||
```
|
||||
|
||||
### Question 2b: Default OpenRouter Model
|
||||
|
||||
Only show if user selected OpenRouter.
|
||||
|
||||
```yaml
|
||||
header: "OpenRouter Model"
|
||||
question: "Default OpenRouter image generation model?"
|
||||
options:
|
||||
- label: "google/gemini-3.1-flash-image (Recommended)"
|
||||
description: "Best general-purpose OpenRouter image model with reference-image workflows"
|
||||
- label: "google/gemini-2.5-flash-image-preview"
|
||||
description: "Fast Gemini preview model on OpenRouter"
|
||||
- label: "black-forest-labs/flux.2-pro"
|
||||
description: "Strong text-to-image quality through OpenRouter"
|
||||
```
|
||||
|
||||
### Question 2c: Default Azure Deployment
|
||||
|
||||
Only show if user selected Azure OpenAI.
|
||||
|
||||
```yaml
|
||||
header: "Azure Deploy"
|
||||
question: "Default Azure image deployment name?"
|
||||
options:
|
||||
- label: "gpt-image-2.5-flare (Recommended)"
|
||||
description: "Use if your Azure deployment uses the GPT Image 2.5 Flare model name"
|
||||
- label: "gpt-image-2.5-sunburst"
|
||||
description: "Use if your Azure deployment uses the GPT Image 2.5 Sunburst model name"
|
||||
- label: "gpt-image-2"
|
||||
description: "Use if your Azure deployment uses the GPT Image 2 model name"
|
||||
- label: "gpt-image-1.5"
|
||||
description: "Previous GPT Image deployment name"
|
||||
- label: "gpt-image-1"
|
||||
description: "Earlier GPT Image deployment name"
|
||||
```
|
||||
|
||||
### Question 2d: Default MiniMax Model
|
||||
|
||||
Only show if user selected MiniMax.
|
||||
|
||||
```yaml
|
||||
header: "MiniMax Model"
|
||||
question: "Default MiniMax image generation model?"
|
||||
options:
|
||||
- label: "image-01 (Recommended)"
|
||||
description: "Best default, supports aspect ratios and custom width/height"
|
||||
- label: "image-01-live"
|
||||
description: "Faster variant, use aspect ratio instead of custom size"
|
||||
```
|
||||
|
||||
### Question 2e: Default Z.AI Model
|
||||
|
||||
Only show if user selected Z.AI.
|
||||
|
||||
```yaml
|
||||
header: "Z.AI Model"
|
||||
question: "Default Z.AI image generation model?"
|
||||
options:
|
||||
- label: "glm-image (Recommended)"
|
||||
description: "Best default for posters, diagrams, and text-heavy images"
|
||||
- label: "cogview-4-250304"
|
||||
description: "Legacy Z.AI image model on the same endpoint"
|
||||
```
|
||||
|
||||
### Question 3: Default Quality
|
||||
|
||||
```yaml
|
||||
header: "Quality"
|
||||
question: "Default image quality?"
|
||||
options:
|
||||
- label: "2k (Recommended)"
|
||||
description: "2048px - covers, illustrations, infographics"
|
||||
- label: "normal"
|
||||
description: "1024px - quick previews, drafts"
|
||||
```
|
||||
|
||||
### Question 4: Save Location
|
||||
|
||||
```yaml
|
||||
header: "Save"
|
||||
question: "Where to save preferences?"
|
||||
options:
|
||||
- label: "Project (Recommended)"
|
||||
description: ".baoyu-skills/ (this project only)"
|
||||
- label: "User"
|
||||
description: "~/.baoyu-skills/ (all projects)"
|
||||
```
|
||||
|
||||
### Save Locations
|
||||
|
||||
| Choice | Path | Scope |
|
||||
|--------|------|-------|
|
||||
| Project | `.baoyu-skills/baoyu-image-gen/EXTEND.md` | Current project |
|
||||
| User | `$HOME/.baoyu-skills/baoyu-image-gen/EXTEND.md` | All projects |
|
||||
|
||||
### EXTEND.md Template
|
||||
|
||||
```yaml
|
||||
---
|
||||
version: 1
|
||||
default_provider: [selected provider or null]
|
||||
default_quality: [selected quality]
|
||||
default_aspect_ratio: null
|
||||
default_image_size: null
|
||||
default_image_api_dialect: null
|
||||
default_model:
|
||||
google: [selected google model or null]
|
||||
openai: null
|
||||
azure: [selected azure deployment or null]
|
||||
openrouter: [selected openrouter model or null]
|
||||
dashscope: null
|
||||
zai: [selected Z.AI model or null]
|
||||
minimax: [selected minimax model or null]
|
||||
replicate: null
|
||||
agnes: null
|
||||
---
|
||||
```
|
||||
|
||||
If the user selects `OpenAI` but says their endpoint is only OpenAI-compatible and fronts another image model family, save `default_image_api_dialect: ratio-metadata` when they explicitly confirm the gateway expects aspect-ratio `size` plus metadata-based resolution. Otherwise leave it `null` / `openai-native`.
|
||||
|
||||
## Flow 2: EXTEND.md Exists, Model Null
|
||||
|
||||
When EXTEND.md exists but `default_model.[current_provider]` is null, ask ONLY the model question for the current provider.
|
||||
|
||||
### Google Model Selection
|
||||
|
||||
```yaml
|
||||
header: "Google Model"
|
||||
question: "Choose a default Google image generation model?"
|
||||
options:
|
||||
- label: "gemini-3-pro-image (Recommended)"
|
||||
description: "Highest quality, best for production use"
|
||||
- label: "gemini-3.1-flash-image"
|
||||
description: "Fast generation, good quality, lower cost"
|
||||
- label: "gemini-3.1-flash-lite-image"
|
||||
description: "Cheapest and fastest; 1K output only"
|
||||
- label: "gemini-3-flash-preview"
|
||||
description: "Fast generation, balanced quality and speed"
|
||||
```
|
||||
|
||||
### OpenAI Model Selection
|
||||
|
||||
```yaml
|
||||
header: "OpenAI Model"
|
||||
question: "Choose a default OpenAI image generation model?"
|
||||
options:
|
||||
- label: "gpt-image-2.5-flare (Recommended)"
|
||||
description: "Latest GPT Image model, fastest; flexible sizes up to 4K, high-fidelity image inputs"
|
||||
- label: "gpt-image-2.5-sunburst"
|
||||
description: "Most capable GPT Image model for complex scenes and precise edits; slower"
|
||||
- label: "gpt-image-2"
|
||||
description: "Previous GPT Image generation, flexible sizes up to 4K"
|
||||
- label: "gpt-image-1.5"
|
||||
description: "Previous GPT Image model"
|
||||
- label: "gpt-image-1"
|
||||
description: "Earlier GPT Image model"
|
||||
```
|
||||
|
||||
### Azure Deployment Selection
|
||||
|
||||
```yaml
|
||||
header: "Azure Deploy"
|
||||
question: "Choose a default Azure image deployment name?"
|
||||
options:
|
||||
- label: "gpt-image-2.5-flare (Recommended)"
|
||||
description: "Use when your Azure deployment name matches the GPT Image 2.5 Flare model"
|
||||
- label: "gpt-image-2.5-sunburst"
|
||||
description: "Use when your Azure deployment name matches the GPT Image 2.5 Sunburst model"
|
||||
- label: "gpt-image-2"
|
||||
description: "Use when your Azure deployment name matches the GPT Image 2 model"
|
||||
- label: "gpt-image-1.5"
|
||||
description: "Use when your Azure deployment name matches the GPT Image 1.5 model"
|
||||
- label: "gpt-image-1"
|
||||
description: "Use when your Azure deployment name matches GPT-image-1"
|
||||
```
|
||||
|
||||
Notes for Azure setup:
|
||||
|
||||
- In `baoyu-image-gen`, Azure `--model` / `default_model.azure` should be the Azure deployment name, not just the underlying model family.
|
||||
- If the deployment name is custom, save that exact deployment name in `default_model.azure`.
|
||||
|
||||
### OpenRouter Model Selection
|
||||
|
||||
```yaml
|
||||
header: "OpenRouter Model"
|
||||
question: "Choose a default OpenRouter image generation model?"
|
||||
options:
|
||||
- label: "google/gemini-3.1-flash-image (Recommended)"
|
||||
description: "Recommended for image output and reference-image edits"
|
||||
- label: "google/gemini-2.5-flash-image-preview"
|
||||
description: "Fast preview-oriented image generation"
|
||||
- label: "black-forest-labs/flux.2-pro"
|
||||
description: "High-quality text-to-image through OpenRouter"
|
||||
```
|
||||
|
||||
### DashScope Model Selection
|
||||
|
||||
```yaml
|
||||
header: "DashScope Model"
|
||||
question: "Choose a default DashScope image generation model?"
|
||||
options:
|
||||
- label: "qwen-image-2.0-pro (Recommended)"
|
||||
description: "Best-verified DashScope model for text rendering and custom sizes"
|
||||
- label: "qwen-image-3.0-pro"
|
||||
description: "Newest Qwen flagship; same sizing rules, not yet the built-in default"
|
||||
- label: "qwen-image-2.0"
|
||||
description: "Faster 2.0 variant with flexible output size"
|
||||
- label: "qwen-image-max"
|
||||
description: "Legacy Qwen model with five fixed output sizes"
|
||||
- label: "qwen-image-plus"
|
||||
description: "Legacy Qwen model, same current capability as qwen-image"
|
||||
- label: "wan2.7-image-pro"
|
||||
description: "Wan 2.7 Pro — supports up to 4K text-to-image and reference-image editing"
|
||||
- label: "wan2.7-image"
|
||||
description: "Wan 2.7 base — faster generation, up to 2K, supports reference-image editing"
|
||||
- label: "z-image-turbo"
|
||||
description: "Legacy DashScope model for compatibility"
|
||||
- label: "z-image-ultra"
|
||||
description: "Legacy DashScope model, higher quality but slower"
|
||||
```
|
||||
|
||||
Notes for DashScope setup:
|
||||
|
||||
- Prefer `qwen-image-2.0-pro` (or the newer `qwen-image-3.0-pro`) when the user needs custom `--size`, uncommon ratios like `21:9`, or strong Chinese/English text rendering.
|
||||
- `qwen-image-max` / `qwen-image-plus` / `qwen-image` only support five fixed sizes: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`.
|
||||
- `wan2.7-image-pro` and `wan2.7-image` are the only DashScope models that accept `--ref`. Pick one of these when the user wants reference-image editing or multi-image fusion via DashScope.
|
||||
- In `baoyu-image-gen`, `quality` is a compatibility preset. It is not a native DashScope parameter.
|
||||
|
||||
### Z.AI Model Selection
|
||||
|
||||
```yaml
|
||||
header: "Z.AI Model"
|
||||
question: "Choose a default Z.AI image generation model?"
|
||||
options:
|
||||
- label: "glm-image (Recommended)"
|
||||
description: "Current flagship image model with better text rendering and poster layouts"
|
||||
- label: "cogview-4-250304"
|
||||
description: "Legacy model on the sync image endpoint"
|
||||
```
|
||||
|
||||
Notes for Z.AI setup:
|
||||
|
||||
- Prefer `glm-image` for posters, diagrams, and Chinese/English text-heavy layouts.
|
||||
- In `baoyu-image-gen`, Z.AI currently exposes text-to-image only; reference images are not wired for this provider.
|
||||
- The sync Z.AI image API returns a downloadable image URL, which the runtime saves locally after download.
|
||||
|
||||
### Replicate Model Selection
|
||||
|
||||
```yaml
|
||||
header: "Replicate Model"
|
||||
question: "Choose a default Replicate image generation model?"
|
||||
options:
|
||||
- label: "google/nano-banana-2 (Recommended)"
|
||||
description: "Current default for general Replicate image generation in baoyu-image-gen"
|
||||
- label: "bytedance/seedream-4.5"
|
||||
description: "Replicate Seedream 4.5 with validated local size/ref guardrails"
|
||||
- label: "bytedance/seedream-5-lite"
|
||||
description: "Replicate Seedream 5 Lite with validated local size/ref guardrails"
|
||||
- label: "wan-video/wan-2.7-image-pro"
|
||||
description: "Replicate Wan 2.7 Image Pro with 4K text-to-image support"
|
||||
```
|
||||
|
||||
### MiniMax Model Selection
|
||||
|
||||
```yaml
|
||||
header: "MiniMax Model"
|
||||
question: "Choose a default MiniMax image generation model?"
|
||||
options:
|
||||
- label: "image-01 (Recommended)"
|
||||
description: "Best general-purpose MiniMax image model with custom width/height support"
|
||||
- label: "image-01-live"
|
||||
description: "Lower-latency MiniMax image model using aspect ratios"
|
||||
```
|
||||
|
||||
Notes for MiniMax setup:
|
||||
|
||||
- `image-01` is the safest default. It supports official `aspect_ratio` values and documented custom `width` / `height` output sizes.
|
||||
- `image-01-live` is useful when the user prefers faster generation and can work with aspect-ratio-based sizing.
|
||||
- MiniMax subject reference currently uses `subject_reference[].type = character`; docs recommend front-facing portrait references in JPG/JPEG/PNG under 10MB.
|
||||
|
||||
### Update EXTEND.md
|
||||
|
||||
After user selects a model:
|
||||
|
||||
1. Read existing EXTEND.md
|
||||
2. If `default_model:` section exists → update the provider-specific key
|
||||
3. If `default_model:` section missing → add the full section:
|
||||
|
||||
```yaml
|
||||
default_model:
|
||||
google: [value or null]
|
||||
openai: [value or null]
|
||||
azure: [value or null]
|
||||
openrouter: [value or null]
|
||||
dashscope: [value or null]
|
||||
zai: [value or null]
|
||||
minimax: [value or null]
|
||||
replicate: [value or null]
|
||||
agnes: [value or null]
|
||||
```
|
||||
|
||||
Only set the selected provider's model; leave others as their current value or null.
|
||||
|
||||
## After Setup
|
||||
|
||||
1. Create directory if needed
|
||||
2. Write/update EXTEND.md with frontmatter
|
||||
3. Confirm: "Preferences saved to [path]"
|
||||
4. Continue with image generation
|
||||
@@ -0,0 +1,149 @@
|
||||
---
|
||||
name: preferences-schema
|
||||
description: EXTEND.md YAML schema for baoyu-image-gen user preferences
|
||||
---
|
||||
|
||||
# Preferences Schema
|
||||
|
||||
## Full Schema
|
||||
|
||||
```yaml
|
||||
---
|
||||
version: 1
|
||||
|
||||
default_provider: null # google|openai|azure|openrouter|dashscope|zai|minimax|replicate|jimeng|seedream|codex-cli|agnes|null (null = auto-detect; codex-cli is never auto-detected — pin it here or via --provider)
|
||||
|
||||
default_quality: null # normal|2k|null (null = use default: 2k)
|
||||
|
||||
default_aspect_ratio: null # "16:9"|"1:1"|"4:3"|"3:4"|"2.35:1"|null
|
||||
|
||||
default_image_size: null # 1K|2K|4K|null (Google/OpenRouter, overrides quality)
|
||||
|
||||
default_image_api_dialect: null # openai-native|ratio-metadata|null (OpenAI-compatible gateways; null = use env/default)
|
||||
|
||||
default_model:
|
||||
google: null # e.g., "gemini-3-pro-image", "gemini-3.1-flash-image", "gemini-3.1-flash-lite-image"
|
||||
openai: null # e.g., "gpt-image-2.5-flare", "gpt-image-2.5-sunburst", "gpt-image-2", "gpt-image-1.5"
|
||||
azure: null # Azure deployment name, e.g., "gpt-image-2.5-flare" or "image-prod"
|
||||
openrouter: null # e.g., "google/gemini-3.1-flash-image"
|
||||
dashscope: null # e.g., "qwen-image-3.0-pro", "qwen-image-2.0-pro"
|
||||
zai: null # e.g., "glm-image"
|
||||
minimax: null # e.g., "image-01"
|
||||
replicate: null # e.g., "google/nano-banana-2"
|
||||
codex-cli: null # Logical label only — Codex image_gen has no user-selectable model. Default: "codex-image-gen"
|
||||
agnes: null # e.g., "agnes-image-2.5-flash"
|
||||
|
||||
batch:
|
||||
max_workers: 10
|
||||
provider_limits:
|
||||
replicate:
|
||||
concurrency: 5
|
||||
start_interval_ms: 700
|
||||
google:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
openai:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
azure:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
openrouter:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
dashscope:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
zai:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
minimax:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
codex-cli:
|
||||
concurrency: 1
|
||||
start_interval_ms: 2000
|
||||
agnes:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
---
|
||||
```
|
||||
|
||||
## Field Reference
|
||||
|
||||
| Field | Type | Default | Description |
|
||||
|-------|------|---------|-------------|
|
||||
| `version` | int | 1 | Schema version |
|
||||
| `default_provider` | string\|null | null | Default provider (null = auto-detect) |
|
||||
| `default_quality` | string\|null | null | Default quality (null = 2k) |
|
||||
| `default_aspect_ratio` | string\|null | null | Default aspect ratio |
|
||||
| `default_image_size` | string\|null | null | Google/OpenRouter image size (overrides quality) |
|
||||
| `default_image_api_dialect` | string\|null | null | OpenAI-compatible image dialect (`openai-native` or `ratio-metadata`) |
|
||||
| `default_model.google` | string\|null | null | Google default model |
|
||||
| `default_model.openai` | string\|null | null | OpenAI default model |
|
||||
| `default_model.azure` | string\|null | null | Azure default deployment name |
|
||||
| `default_model.openrouter` | string\|null | null | OpenRouter default model |
|
||||
| `default_model.dashscope` | string\|null | null | DashScope default model |
|
||||
| `default_model.zai` | string\|null | null | Z.AI default model |
|
||||
| `default_model.minimax` | string\|null | null | MiniMax default model |
|
||||
| `default_model.replicate` | string\|null | null | Replicate default model |
|
||||
| `default_model.codex-cli` | string\|null | null | Codex-CLI logical label (Codex image_gen has no user-selectable model) |
|
||||
| `default_model.agnes` | string\|null | null | Agnes default model |
|
||||
| `batch.max_workers` | int\|null | 10 | Batch worker cap |
|
||||
| `batch.provider_limits.<provider>.concurrency` | int\|null | provider default | Max simultaneous requests per provider |
|
||||
| `batch.provider_limits.<provider>.start_interval_ms` | int\|null | provider default | Minimum gap between request starts per provider |
|
||||
|
||||
## Examples
|
||||
|
||||
**Minimal**:
|
||||
```yaml
|
||||
---
|
||||
version: 1
|
||||
default_provider: google
|
||||
default_quality: 2k
|
||||
default_image_api_dialect: null
|
||||
---
|
||||
```
|
||||
|
||||
**Full**:
|
||||
```yaml
|
||||
---
|
||||
version: 1
|
||||
default_provider: google
|
||||
default_quality: 2k
|
||||
default_aspect_ratio: "16:9"
|
||||
default_image_size: 2K
|
||||
default_image_api_dialect: null
|
||||
default_model:
|
||||
google: "gemini-3-pro-image"
|
||||
openai: "gpt-image-2.5-flare"
|
||||
azure: "gpt-image-2.5-flare"
|
||||
openrouter: "google/gemini-3.1-flash-image"
|
||||
dashscope: "qwen-image-2.0-pro"
|
||||
zai: "glm-image"
|
||||
minimax: "image-01"
|
||||
replicate: "google/nano-banana-2"
|
||||
agnes: "agnes-image-2.5-flash"
|
||||
batch:
|
||||
max_workers: 10
|
||||
provider_limits:
|
||||
replicate:
|
||||
concurrency: 5
|
||||
start_interval_ms: 700
|
||||
azure:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
zai:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
openrouter:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
minimax:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
agnes:
|
||||
concurrency: 3
|
||||
start_interval_ms: 1100
|
||||
---
|
||||
```
|
||||
@@ -0,0 +1,55 @@
|
||||
# Sapiens AI Agnes Image
|
||||
|
||||
Read when the user picks `--provider agnes` or sets `default_model.agnes`. Default model is `agnes-image-2.5-flash`.
|
||||
|
||||
## Models
|
||||
|
||||
**`agnes-image-2.5-flash`** (only model)
|
||||
|
||||
- Text-to-image and image-to-image (with `--ref`) in a single `/images/generations` endpoint
|
||||
- Supports reference images as public URLs or Data URI (base64)
|
||||
- Optimized for high information density, complex layouts, and rich details
|
||||
- Size rules: both dimensions divisible by 32 (720px exception), long edge ≤ 2048, total pixels ≤ ~4M
|
||||
- Default size: `1024x1024`; custom `--size` supports arbitrary WxH within the above rules
|
||||
- `--ar` supported: computed as 2048-based size (long edge ≤ 2048, short edge proportional, both snapped to 32px); `1:1` special-cased to `1024x1024`
|
||||
|
||||
## Response Format
|
||||
|
||||
- The sync API always returns a URL
|
||||
- Default (`--response-format file`): downloads the image and saves as `.png`
|
||||
- Pass `--response-format url`: writes the URL string to `.txt` instead
|
||||
|
||||
## `--n` Behavior
|
||||
|
||||
The Agnes API returns a single image per request regardless of the `n` parameter. Passing `--n > 1` triggers a local error from `validateArgs` before any API call is made.
|
||||
|
||||
## Behavior Notes
|
||||
|
||||
- API key required: `AGNES_API_KEY`
|
||||
- Base URL: `https://apihub.agnes-ai.com/v1` (override with `AGNES_BASE_URL`)
|
||||
- Model override: `AGNES_IMAGE_MODEL` env
|
||||
- `response_format` is always embedded in `extra_body` (not at request top level)
|
||||
- Reference images: local files converted to Data URI base64 inline; remote URLs passed through
|
||||
- Rate limit defaults: concurrency=3, startIntervalMs=1100 (override via `BAOYU_IMAGE_GEN_AGNES_CONCURRENCY` / `BAOYU_IMAGE_GEN_AGNES_START_INTERVAL_MS`)
|
||||
- Timeout: 120s per request
|
||||
|
||||
## Size Resolution
|
||||
|
||||
- `--size <WxH>` wins over `--ar`
|
||||
- `--ar` maps to a concrete size using the algorithm: long edge ≤ 2048, short edge proportional, both dimensions snapped to 32px
|
||||
- `--ar 1:1` is special-cased to `1024x1024`
|
||||
|
||||
### Common `--ar` Results
|
||||
|
||||
| Aspect Ratio | Result |
|
||||
|--------------|--------|
|
||||
| `1:1` | `1024x1024` |
|
||||
| `16:9` | `2048x1152` |
|
||||
| `4:3` | `2048x1536` |
|
||||
| `3:2` | `2048x1376` |
|
||||
| `21:9` | `2048x896` |
|
||||
| Unlisted ratio | Computed on the fly (portrait mirror swaps width/height) |
|
||||
|
||||
## Official References
|
||||
|
||||
- [Agnes AIGC API Hub](https://apihub.agnes-ai.com)
|
||||
@@ -0,0 +1,81 @@
|
||||
# Codex CLI (`--provider codex-cli`)
|
||||
|
||||
Read when the user picks `--provider codex-cli`, sets `default_provider: codex-cli`, or asks for "Codex image generation without an OpenAI API key". This provider is a thin baoyu-image-gen wrapper around the bundled `scripts/codex-imagegen/main.ts` (synced from `packages/baoyu-codex-imagegen`), which spawns `codex exec --json --sandbox danger-full-access` and routes the request to Codex CLI's built-in `image_gen` tool. The Codex CLI uses the **user's Codex / ChatGPT subscription** — no `OPENAI_API_KEY` is read or sent.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
```bash
|
||||
npm install -g @openai/codex
|
||||
codex login # signs in with the user's OpenAI / Codex account
|
||||
codex --version # confirm >= 0.130
|
||||
```
|
||||
|
||||
`bun` is required for running the underlying wrapper (`scripts/codex-imagegen/main.ts`, carrying `#!/usr/bin/env bun`). If `bun` is missing from the runtime, `npx -y bun` works as a fallback.
|
||||
|
||||
## Selection
|
||||
|
||||
- **Never auto-selected.** `detectProvider` only picks `codex-cli` when it is pinned explicitly: pass `--provider codex-cli` or set `default_provider: codex-cli` in EXTEND.md.
|
||||
- Choose this provider when:
|
||||
- The user has a Codex subscription and explicitly does **not** want to manage an OpenAI API key.
|
||||
- You need Codex's specific `image_gen` behavior or quality.
|
||||
- Avoid this provider when latency matters — Codex CLI is typically 5–10× slower than direct OpenAI / Google API calls (except on cache hits).
|
||||
|
||||
## Supported flags
|
||||
|
||||
| Flag | Behavior |
|
||||
|------|----------|
|
||||
| `--prompt <text>` / `--promptfiles <files>` | Required. Written to a temp file and passed to the wrapper as `--prompt-file`. |
|
||||
| `--image <path>` | Required. Final output PNG location. |
|
||||
| `--ar <ratio>` | Mapped to wrapper's `--aspect`. Supported by Codex: `1:1` (default), `16:9`, `9:16`, `4:3`, `2.35:1`. |
|
||||
| `--ref <files...>` | Mapped to wrapper's repeated `--ref`. Codex's `image_gen` accepts reference images for style/composition guidance. |
|
||||
| `--n` | Must be `1`. `validateArgs` throws if `n > 1` because Codex `image_gen` returns a single image per call. |
|
||||
| `--imageApiDialect` | Not applicable. Throws if set to a non-default value. |
|
||||
| `--size`, `--imageSize`, `--quality` | Silently ignored — Codex picks pixel dimensions from the aspect ratio. |
|
||||
| `--model`, `-m` | Logical label only. The wrapper does not forward a model selector to Codex; the underlying engine is whichever model Codex's `image_gen` currently uses. Default label: `codex-image-gen`. |
|
||||
|
||||
## Environment variables
|
||||
|
||||
| Variable | Effect |
|
||||
|----------|--------|
|
||||
| `BAOYU_CODEX_IMAGEGEN_BIN` | Override the wrapper path. Default: bundled `scripts/codex-imagegen/main.ts` resolved relative to this skill's installed location. Accepts a `.ts` file (spawned with `bun`) or a legacy `.sh`/binary (spawned directly). |
|
||||
| `BAOYU_CODEX_IMAGEGEN_CACHE_DIR` | Enable the wrapper's idempotency cache. Disabled by default; set to e.g. `~/.cache/baoyu-codex-imagegen` for high-value reuse. |
|
||||
| `BAOYU_CODEX_IMAGEGEN_TIMEOUT_MS` | Per-attempt `codex exec` timeout in ms. Default: `300000` (5 min). Raise for slow networks or large prompts. |
|
||||
| `BAOYU_CODEX_IMAGEGEN_RETRIES` | Wrapper-side retry attempts on retryable errors. Default: `2` (3 total attempts). |
|
||||
| `BAOYU_CODEX_IMAGEGEN_LOG_FILE` | Append a structured JSONL diagnostic log. Useful when triaging timeouts or `agent_refused` errors. |
|
||||
| `BAOYU_IMAGE_GEN_CODEX_CLI_CONCURRENCY` | Batch-mode concurrency for the `codex-cli` provider. Default: `1` — Codex exec is a heavy single-process workflow; raising this rarely helps. |
|
||||
| `BAOYU_IMAGE_GEN_CODEX_CLI_START_INTERVAL_MS` | Batch-mode minimum start-gap. Default: `2000` ms. |
|
||||
|
||||
## Error model
|
||||
|
||||
The wrapper emits a single JSON line on stdout. On failure:
|
||||
|
||||
```json
|
||||
{"status":"error","path":"...","bytes":0,"error":"...","error_kind":"..."}
|
||||
```
|
||||
|
||||
The provider re-throws each wrapper error as `Invalid codex-cli result (<error_kind>): <message>`. The `"Invalid "` prefix triggers `isRetryableGenerationError` to mark it **non-retryable** in baoyu-image-gen's outer retry loop — the wrapper has already retried internally per `BAOYU_CODEX_IMAGEGEN_RETRIES`, so re-spawning Codex from main.ts would only multiply latency without changing the outcome.
|
||||
|
||||
`error_kind` values to expect:
|
||||
|
||||
| Kind | Cause | Action |
|
||||
|------|-------|--------|
|
||||
| `codex_not_installed` | `codex` not on `PATH` or unreadable | `npm install -g @openai/codex`, then `codex login`. |
|
||||
| `invalid_args` | Programmer error in the spawn invocation | Inspect provider source; usually a path-injection guard fired. |
|
||||
| `prompt_file_missing` | Temp prompt file vanished mid-call | Retry once; check `$TMPDIR` permissions. |
|
||||
| `spawn_failed` | OS / process-launch failure | Verify `bun` or `npx` is installed; check filesystem permissions. |
|
||||
| `timeout` | `codex exec` exceeded `--timeout` | Raise `BAOYU_CODEX_IMAGEGEN_TIMEOUT_MS`; check network. |
|
||||
| `no_image_gen_tool_use` | Codex agent answered without calling `image_gen` | Often transient — retry. If persistent, refine the prompt. |
|
||||
| `output_missing` / `invalid_png` | Agent reported success but file is absent or not a valid PNG | Retry; check disk space. |
|
||||
| `agent_refused` | Codex agent refused (policy or content) | Adjust the prompt; surface the refusal to the user. |
|
||||
| `lock_busy` | Another `codex-imagegen` invocation holds the file lock | Wait or set a distinct `--cache-dir` per concurrent caller. |
|
||||
|
||||
## Trade-offs
|
||||
|
||||
- Slow: 5–10× direct OpenAI API latency (except cache hits).
|
||||
- Subject to the same TOS as interactive `codex exec` use — programmatic invocation from baoyu-image-gen is the same usage class.
|
||||
- Stateful: requires `codex login` to be live; an expired session manifests as `codex_not_installed` or `agent_refused`.
|
||||
|
||||
## See also
|
||||
|
||||
- `references/codex-oauth-vs-openai-api-key.md` — why Codex OAuth is not interchangeable with `OPENAI_API_KEY`.
|
||||
- `references/codex-image2-fallback.md` — when to fall back to `codex-cli` from a failed `openai` provider call.
|
||||
@@ -0,0 +1,71 @@
|
||||
# DashScope (阿里通义万象)
|
||||
|
||||
Read when the user picks `--provider dashscope`, sets `default_model.dashscope`, or asks for Qwen-Image behavior. The SKILL.md only names the default — this file covers model families, sizing rules, and limits.
|
||||
|
||||
## Model Families
|
||||
|
||||
**`qwen-image-3.0*` / `qwen-image-2.0*`** — recommended modern family (identical sizing envelope). Members: `qwen-image-3.0-pro`, `qwen-image-2.0-pro`, `qwen-image-2.0-pro-2026-04-22`, `qwen-image-2.0-pro-2026-03-03`, `qwen-image-2.0`, `qwen-image-2.0-2026-03-03`.
|
||||
|
||||
- `qwen-image-3.0-pro` is the newest flagship; `qwen-image-2.0-pro` remains the built-in default until the 3.0 request shape is verified end to end
|
||||
|
||||
- Free-form `size` in `宽*高` format
|
||||
- Total pixels must be between `512*512` and `2048*2048`
|
||||
- Default ≈ `1024*1024`
|
||||
- Best choice for custom ratios (e.g. `21:9`) and text-heavy Chinese/English layouts
|
||||
|
||||
**Fixed-size family** — `qwen-image-max`, `qwen-image-max-2025-12-30`, `qwen-image-plus`, `qwen-image-plus-2026-01-09`, `qwen-image`.
|
||||
|
||||
- Only five sizes allowed: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`
|
||||
- Default is `1664*928`
|
||||
- `qwen-image` currently has the same capability as `qwen-image-plus`
|
||||
|
||||
**`wan2.7-image*`** — multimodal Wan 2.7 family. Members: `wan2.7-image-pro`, `wan2.7-image`.
|
||||
|
||||
- Free-form `size` in `宽*高` format, plus aspect-ratio inference
|
||||
- `wan2.7-image-pro` text-to-image (no `--ref`): total pixels in `[768*768, 4096*4096]`, ratio in `[1:8, 8:1]`
|
||||
- `wan2.7-image-pro` with reference images and `wan2.7-image` (all scenarios): total pixels in `[768*768, 2048*2048]`, ratio in `[1:8, 8:1]`
|
||||
- Default: `1024*1024` (`--quality normal`) or `2048*2048` (`--quality 2k`); 4K requires explicit `--size`
|
||||
- Supports up to 9 reference images in `--ref` (image editing / multi-image fusion)
|
||||
- Reference images are sent inline as base64 (or passed through if the path is an `http(s)://` URL)
|
||||
- API does NOT use `prompt_extend`; the skill omits it for this family
|
||||
- The Wan 2.7 API defaults `n` to **4** in non-collage mode and bills per generated image. baoyu-image-gen forces `n: 1` and rejects `--n > 1` to avoid silently paying for and discarding extra images.
|
||||
|
||||
**Legacy** — `z-image-turbo`, `z-image-ultra`, `wanx-v1`. Only use when the user explicitly asks for legacy behavior.
|
||||
|
||||
## Size Resolution
|
||||
|
||||
- `--size` wins over `--ar`
|
||||
- For `qwen-image-3.0*` / `qwen-image-2.0*`: prefer explicit `--size`; otherwise infer from `--ar` using the recommended table below
|
||||
- For `qwen-image-max/plus/image`: only use the five fixed sizes; if the requested ratio doesn't fit, switch to `qwen-image-2.0-pro` or `qwen-image-3.0-pro`
|
||||
- For `wan2.7-image*`: explicit `--size` is validated against the per-mode pixel/ratio limits; otherwise the size is derived from `--ar` and `--quality` (`normal` ≈ 1K, `2k` ≈ 2K). To request 4K with `wan2.7-image-pro` text-to-image, pass `--size` explicitly (e.g. `4096*4096`, `3840*2160`)
|
||||
- `--quality` is a baoyu-image-gen preset, not an official DashScope field. The mapping of `normal`/`2k` onto the `qwen-image-3.0*` / `qwen-image-2.0*` and `wan2.7-image*` tables is an implementation choice, not an API guarantee
|
||||
|
||||
### Recommended `qwen-image-3.0*` / `qwen-image-2.0*` sizes
|
||||
|
||||
| Ratio | `normal` | `2k` |
|
||||
|-------|----------|------|
|
||||
| `1:1` | `1024*1024` | `1536*1536` |
|
||||
| `2:3` | `768*1152` | `1024*1536` |
|
||||
| `3:2` | `1152*768` | `1536*1024` |
|
||||
| `3:4` | `960*1280` | `1080*1440` |
|
||||
| `4:3` | `1280*960` | `1440*1080` |
|
||||
| `9:16` | `720*1280` | `1080*1920` |
|
||||
| `16:9` | `1280*720` | `1920*1080` |
|
||||
| `21:9` | `1344*576` | `2048*872` |
|
||||
|
||||
## Reference Images
|
||||
|
||||
- Only `wan2.7-image-pro` and `wan2.7-image` accept `--ref`. Other DashScope models (qwen-image-3.0*, qwen-image-2.0*, qwen-image-max/plus/image, legacy) reject `--ref` and the user is steered to a different provider/model.
|
||||
- Up to 9 reference images per request. Local files are inlined as base64 data URLs; `http(s)://` URLs are forwarded as-is.
|
||||
- Supplying any `--ref` automatically clamps the wan2.7-image-pro pixel ceiling from 4K to 2K (the API only supports 4K for pure text-to-image with no image input).
|
||||
|
||||
## Not Exposed
|
||||
|
||||
DashScope APIs also support `negative_prompt`, `prompt_extend`, `watermark`, `thinking_mode`, `seed`, `bbox_list`, `enable_sequential`, and `color_palette`. `baoyu-image-gen` does not expose them as CLI flags today; the wan2.7 family relies on the API defaults (e.g. `thinking_mode=true`). The skill always sends `n=1` for wan2.7 — if you want grid/collage mode you currently need to call the API directly.
|
||||
|
||||
## Official References
|
||||
|
||||
- [Qwen-Image API](https://help.aliyun.com/zh/model-studio/qwen-image-api)
|
||||
- [Text-to-image guide](https://help.aliyun.com/zh/model-studio/text-to-image)
|
||||
- [Qwen-Image Edit API](https://help.aliyun.com/zh/model-studio/qwen-image-edit-api)
|
||||
- [Wan 2.7 image generation & editing API](https://help.aliyun.com/zh/model-studio/wan-image-generation-and-editing-api-reference)
|
||||
@@ -0,0 +1,29 @@
|
||||
# MiniMax
|
||||
|
||||
Read when the user picks `--provider minimax` or sets `default_model.minimax`. Default model is `image-01`.
|
||||
|
||||
## Models
|
||||
|
||||
**`image-01`** (recommended default)
|
||||
|
||||
- Supports text-to-image and subject-reference image generation
|
||||
- Supports official `aspect_ratio` values: `1:1`, `16:9`, `4:3`, `3:2`, `2:3`, `3:4`, `9:16`, `21:9`
|
||||
- Supports documented custom `width` / `height` via `--size <WxH>`
|
||||
- Both width and height must be in `[512, 2048]` and divisible by `8`
|
||||
|
||||
**`image-01-live`** — lower-latency variant
|
||||
|
||||
- Use `--ar` for sizing; MiniMax documents custom `width`/`height` only for `image-01`
|
||||
|
||||
## Subject Reference
|
||||
|
||||
- `--ref` files are sent as MiniMax `subject_reference`
|
||||
- `subject_reference[].type` is currently `character`
|
||||
- Official docs say `image_file` supports public URLs or Base64 Data URLs; baoyu-image-gen sends local refs as Data URLs
|
||||
- Recommended refs: front-facing portraits, JPG/JPEG/PNG, under 10MB
|
||||
|
||||
## Official References
|
||||
|
||||
- [Image Generation Guide](https://platform.minimaxi.com/docs/guides/image-generation)
|
||||
- [Text-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-t2i)
|
||||
- [Image-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-i2i)
|
||||
@@ -0,0 +1,19 @@
|
||||
# OpenRouter
|
||||
|
||||
Read when the user picks `--provider openrouter`. Default model is `google/gemini-3.1-flash-image`.
|
||||
|
||||
## Common Models
|
||||
|
||||
Use full OpenRouter model IDs:
|
||||
|
||||
- `google/gemini-3.1-flash-image` (recommended — supports image output and reference-image workflows)
|
||||
- `google/gemini-2.5-flash-image-preview`
|
||||
- `black-forest-labs/flux.2-pro`
|
||||
- Any other OpenRouter image-capable model ID
|
||||
|
||||
## Behavior Notes
|
||||
|
||||
- OpenRouter image generation uses `/chat/completions`, not the OpenAI `/images` endpoints
|
||||
- `--ref` requires a multimodal model that supports both image input and image output
|
||||
- `--imageSize` maps to `imageGenerationOptions.size`
|
||||
- `--size <WxH>` is converted to the nearest supported OpenRouter size, and the aspect ratio is inferred when possible
|
||||
@@ -0,0 +1,50 @@
|
||||
# Replicate
|
||||
|
||||
Read when the user picks `--provider replicate`. Replicate support is intentionally scoped to model families baoyu-image-gen can validate locally and save without dropping outputs.
|
||||
|
||||
## Supported Families
|
||||
|
||||
**`google/nano-banana*`** (default: `google/nano-banana-2`)
|
||||
|
||||
- Supports prompt-only and reference-image generation
|
||||
- Uses Replicate `aspect_ratio`, `resolution`, and `output_format`
|
||||
- `--size <WxH>` is accepted only as a shorthand for a documented `aspect_ratio` plus `1K` / `2K`
|
||||
|
||||
**`bytedance/seedream-4.5`**
|
||||
|
||||
- Supports prompt-only and reference-image generation
|
||||
- Uses Replicate `size`, `aspect_ratio`, and `image_input`
|
||||
- Local validation blocks unsupported `1K` requests before the API call
|
||||
|
||||
**`bytedance/seedream-5-lite`**
|
||||
|
||||
- Supports prompt-only and reference-image generation
|
||||
- Uses Replicate `size`, `aspect_ratio`, and `image_input`
|
||||
- Local validation currently accepts `2K` / `3K` only
|
||||
|
||||
**`wan-video/wan-2.7-image`**
|
||||
|
||||
- Supports prompt-only and reference-image generation
|
||||
- Uses Replicate `size` and `images`
|
||||
- Max output is 2K
|
||||
|
||||
**`wan-video/wan-2.7-image-pro`**
|
||||
|
||||
- Supports prompt-only and reference-image generation
|
||||
- Uses Replicate `size` and `images`
|
||||
- 4K is allowed only for text-to-image; local validation blocks `4K + --ref`
|
||||
|
||||
## Guardrails
|
||||
|
||||
- Replicate currently supports only single-output save semantics in this tool — keep `--n 1`
|
||||
- If a model is outside the compatibility list above, baoyu-image-gen treats it as prompt-only and rejects advanced local options instead of guessing a nano-banana-style schema
|
||||
|
||||
## Examples
|
||||
|
||||
```bash
|
||||
# Default model
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate
|
||||
|
||||
# Explicit model
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate --model google/nano-banana
|
||||
```
|
||||
@@ -0,0 +1,27 @@
|
||||
# Z.AI GLM-Image
|
||||
|
||||
Read when the user picks `--provider zai` or sets `default_model.zai`. Default model is `glm-image`.
|
||||
|
||||
## Models
|
||||
|
||||
**`glm-image`** (recommended default)
|
||||
|
||||
- Text-to-image only in baoyu-image-gen (no `--ref` support yet)
|
||||
- Native `quality` options are `hd` and `standard`; this skill maps `2k → hd` and `normal → standard`
|
||||
- Recommended sizes: `1280x1280`, `1568x1056`, `1056x1568`, `1472x1088`, `1088x1472`, `1728x960`, `960x1728`
|
||||
- Custom `--size` requires width/height in `[1024, 2048]`, divisible by `32`, total pixels ≤ `2^22`
|
||||
|
||||
**`cogview-4-250304`** (legacy family, same endpoint)
|
||||
|
||||
- Custom `--size` requires width/height in `[512, 2048]`, divisible by `16`, total pixels ≤ `2^21`
|
||||
|
||||
## Behavior Notes
|
||||
|
||||
- The sync API returns a temporary URL; baoyu-image-gen downloads it and writes locally
|
||||
- `--ref` is not supported for Z.AI in this skill yet
|
||||
- The sync API returns a single image, so `--n > 1` is rejected
|
||||
|
||||
## Official References
|
||||
|
||||
- [GLM-Image Guide](https://docs.z.ai/guides/image/glm-image)
|
||||
- [Generate Image API](https://docs.z.ai/api-reference/image/generate-image)
|
||||
137
baoyu-skills/skills/baoyu-image-gen/references/usage-examples.md
Normal file
137
baoyu-skills/skills/baoyu-image-gen/references/usage-examples.md
Normal file
@@ -0,0 +1,137 @@
|
||||
# Usage Examples
|
||||
|
||||
Extended CLI examples. SKILL.md shows the minimum set; read this file when the user asks about provider-specific invocation, batch generation, or less-common flags.
|
||||
|
||||
## Core Patterns
|
||||
|
||||
```bash
|
||||
# Basic text-to-image
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image cat.png
|
||||
|
||||
# With aspect ratio
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A landscape" --image out.png --ar 16:9
|
||||
|
||||
# High quality
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --quality 2k
|
||||
|
||||
# Prompt from files
|
||||
${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png
|
||||
|
||||
# With reference images (any provider family that supports refs)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --ref source.png
|
||||
```
|
||||
|
||||
## Per-Provider
|
||||
|
||||
```bash
|
||||
# OpenAI
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openai --model gpt-image-2.5-flare
|
||||
|
||||
# Azure OpenAI (model = deployment name)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider azure --model gpt-image-2.5-flare
|
||||
|
||||
# OpenAI GPT Image 2 custom 4K size
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic landscape" --image out.png --provider openai --model gpt-image-2.5-sunburst --size 3840x2160
|
||||
|
||||
# Google with explicit model
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider google --model gemini-3-pro-image --ref source.png
|
||||
|
||||
# OpenRouter (recommended default)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openrouter
|
||||
|
||||
# OpenRouter with reference
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider openrouter --model google/gemini-3.1-flash-image --ref source.png
|
||||
|
||||
# DashScope (default model)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "一只可爱的猫" --image out.png --provider dashscope
|
||||
|
||||
# DashScope Qwen-Image 2.0 Pro (custom size, Chinese text)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "为咖啡品牌设计一张 21:9 横幅海报,包含清晰中文标题" --image out.png --provider dashscope --model qwen-image-2.0-pro --size 2048x872
|
||||
|
||||
# DashScope legacy fixed-size
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "一张电影感海报" --image out.png --provider dashscope --model qwen-image-max --size 1664x928
|
||||
|
||||
# DashScope Wan 2.7 Image Pro (4K text-to-image)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "一间有着精致窗户的花店" --image out.png --provider dashscope --model wan2.7-image-pro --size 4096x4096
|
||||
|
||||
# DashScope Wan 2.7 Image with reference image (multi-image fusion)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "把图2的涂鸦喷绘在图1的汽车上" --image out.png --provider dashscope --model wan2.7-image-pro --ref car.webp paint.webp
|
||||
|
||||
# Z.AI GLM-image
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "一张带清晰中文标题的科技海报" --image out.png --provider zai
|
||||
|
||||
# Z.AI with custom size
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A science illustration with labels" --image out.png --provider zai --model glm-image --size 1472x1088
|
||||
|
||||
# MiniMax
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A fashion editorial portrait" --image out.jpg --provider minimax
|
||||
|
||||
# MiniMax with subject reference (character/portrait consistency)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A girl by the library window" --image out.jpg --provider minimax --model image-01 --ref portrait.png --ar 16:9
|
||||
|
||||
# Replicate (default: google/nano-banana-2)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate
|
||||
|
||||
# Replicate Seedream 4.5
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic portrait" --image out.png --provider replicate --model bytedance/seedream-4.5 --ar 3:2
|
||||
|
||||
# Replicate Wan 2.7 Image Pro
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A concept frame" --image out.png --provider replicate --model wan-video/wan-2.7-image-pro --size 2048x1152
|
||||
|
||||
# Codex CLI (uses Codex / ChatGPT subscription — no OPENAI_API_KEY; requires `codex` on PATH and `codex login`)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic portrait" --image out.png --provider codex-cli --ar 16:9
|
||||
|
||||
# Codex CLI with reference images (style/composition guidance)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "Match this color palette" --image out.png --provider codex-cli --ref source.png --ar 1:1
|
||||
|
||||
# Agnes (default model)
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A detailed infographic" --image out.png --provider agnes
|
||||
|
||||
# Agnes with aspect ratio and URL output
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic scene" --image out.txt --provider agnes --ar 16:9 --response-format url
|
||||
|
||||
# Agnes with reference image
|
||||
${BUN_X} {baseDir}/scripts/main.ts --prompt "Apply this style" --image out.png --provider agnes --ref source.png
|
||||
```
|
||||
|
||||
Notes on `codex-cli`:
|
||||
- Never auto-selected — pin via `--provider codex-cli` or `default_provider: codex-cli` in EXTEND.md.
|
||||
- Only `n=1` supported (Codex `image_gen` returns one image per call); `--size`, `--imageSize`, `--quality`, and `--imageApiDialect` are ignored or rejected.
|
||||
- Typically 5–10× slower than direct OpenAI / Google API calls (except on cache hits). Tune via `BAOYU_CODEX_IMAGEGEN_TIMEOUT_MS`, `BAOYU_CODEX_IMAGEGEN_RETRIES`, and `BAOYU_CODEX_IMAGEGEN_CACHE_DIR`.
|
||||
|
||||
## Batch Mode
|
||||
|
||||
```bash
|
||||
# Batch from saved prompt files
|
||||
${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json
|
||||
|
||||
# Batch with explicit worker count
|
||||
${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4 --json
|
||||
```
|
||||
|
||||
### Batch File Format
|
||||
|
||||
```json
|
||||
{
|
||||
"jobs": 4,
|
||||
"tasks": [
|
||||
{
|
||||
"id": "hero",
|
||||
"promptFiles": ["prompts/hero.md"],
|
||||
"image": "out/hero.png",
|
||||
"provider": "replicate",
|
||||
"model": "google/nano-banana-2",
|
||||
"ar": "16:9",
|
||||
"quality": "2k"
|
||||
},
|
||||
{
|
||||
"id": "diagram",
|
||||
"promptFiles": ["prompts/diagram.md"],
|
||||
"image": "out/diagram.png",
|
||||
"ref": ["references/original.png"]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Paths in `promptFiles`, `image`, and `ref` are resolved relative to the batch file's directory. `jobs` is optional (overridden by CLI `--jobs`). A top-level array without the `jobs` wrapper is also accepted.
|
||||
Reference in New Issue
Block a user