imagencn
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
Chineseimagencn - Multi-Cloud Text-to-Image Skill
imagencn - 多云平台文本转图像Skill
Overview
概述
imagencn — Image Generation, Cloud-Native: one CLI, every image cloud. The project started with China-friendly clouds and now covers international providers as well.
Generate images using Alibaba Cloud Bailian API. Default endpoint is China region.
Supports six platforms across ten model families:
- Alibaba Cloud Bailian (DashScope): Qwen-Image 2.0, Qwen-Image Edit, Qwen-Image legacy, Wan Series, Z-Image
- ByteDance Volcano Ark: Doubao-Seedream series (OpenAI-compatible REST)
- Tencent Hunyuan: Hunyuan Image 3.0 (OpenAI-compatible REST)
- Zhipu / BigModel: CogView-4 and GLM-Image (OpenAI-compatible REST)
- StepFun / 阶跃星辰: Step-2X and Step-Image-Edit (OpenAI-compatible REST)
- Google Gemini (international): Gemini 3 Pro Image (generateContent REST)
Cross-platform support: Windows, macOS, Linux
imagencn — 图像生成,云原生:一个CLI,对接所有图像云服务。该项目最初面向国内云平台开发,目前已覆盖国际服务商。
可通过阿里云百炼API生成图像。默认接入国内区域端点。
支持6大平台的10个模型系列:
- 阿里云百炼(DashScope):Qwen-Image 2.0、Qwen-Image Edit、Qwen-Image legacy、万相系列、Z-Image
- 字节跳动火山方舟(Volcano Ark):豆包-Seedream系列(兼容OpenAI REST接口)
- 腾讯混元(Hunyuan):混元图像3.0(兼容OpenAI REST接口)
- 智谱(Zhipu)/ BigModel:CogView-4和GLM-Image(兼容OpenAI REST接口)
- 阶跃星辰(StepFun):Step-2X和Step-Image-Edit(兼容OpenAI REST接口)
- Google Gemini(国际版):Gemini 3 Pro Image(generateContent REST接口)
跨平台支持:Windows、macOS、Linux
When to Use This Skill
使用场景
Automatically activate this skill when:
- User requests image generation with Chinese text or calligraphy
- Need photorealistic images or photography-style visuals
- Creating commercial posters, illustrations, or digital art
- User mentions any of these: Alibaba Cloud / Bailian / Qwen / Wan / DashScope, ByteDance / Volcano Ark / Seedream / Doubao, Tencent / Hunyuan, Google / Gemini / Nano Banana
- User wants an international (non-China) image provider — use the Gemini platform
- Any task where AI-generated image with strong Chinese support would be helpful
满足以下任一条件时自动激活该Skill:
- 用户请求生成包含中文文本或书法的图像
- 需要写实风格图像或摄影级视觉效果
- 制作商业海报、插画或数字艺术作品
- 用户提及以下任一平台:阿里云/百炼/Qwen/万相/DashScope、字节跳动/火山方舟/Seedream/豆包、腾讯/混元、Google/Gemini/Nano Banana
- 用户需要国际(非国内)图像生成服务商——使用Gemini平台
- 任何需要强中文支持的AI图像生成任务
Model Reference
模型参考
When the user wants to compare models, check pricing, or browse options before
choosing, open the local model reference page in their browser:
bash
open ~/.claude/skills/imagencn/docs/models.htmlThis page shows all 31 models across 6 platforms with pricing, resolution,
feature highlights, and a quick-reference guide. On Linux use ;
the file also works from with no server needed.
xdg-openfile://当用户需要对比模型、查看定价或选择模型时,在浏览器中打开本地模型参考页面:
bash
open ~/.claude/skills/imagencn/docs/models.html该页面展示了6大平台的31个模型,包含定价、分辨率、功能亮点及快速参考指南。Linux系统使用;该文件可通过协议直接打开,无需服务器。
xdg-openfile://Workflow
工作流程
Step 1 — Refine the prompt (interactive, never skip)
步骤1 — 优化提示词(交互式,不可跳过)
Users often give short, casual descriptions ("生成一只猫"). Before calling the
API, present 3 refined prompt options with different style directions.
Add, as appropriate:
- Subject details (shape, colour, material, expression, pose)
- Lighting (golden hour, studio, rim light, soft diffused, neon, cinematic)
- Composition (rule of thirds, shallow depth of field, wide shot, close-up)
- Style / medium (photorealistic, oil painting, watercolour, 3D render, vector)
- Mood / atmosphere (serene, dramatic, whimsical, dystopian, elegant)
- Quality keywords (8K, hyperdetailed, award-winning, professional photography)
- For Chinese text on images: text content, placement, font style, colour, size
Label the options clearly (e.g. A / B / C) with a one-line summary of each
direction. Let the user pick one, combine elements from multiple, or request
a new direction. Iterate until they confirm ("go", "generate", "ok", etc.),
then proceed to generation.
用户通常会给出简短、随意的描述(如"生成一只猫")。调用API前,提供3个不同风格方向的优化提示词选项,可根据需要添加以下元素:
- 主体细节(形状、颜色、材质、表情、姿态)
- 光线效果(黄金时段、影棚光、轮廓光、柔和漫射光、霓虹光、电影级光线)
- 构图方式(三分法、浅景深、广角镜头、特写)
- 风格/媒介(写实风格、油画、水彩、3D渲染、矢量图)
- 氛围/情绪(宁静、戏剧性、奇幻、反乌托邦、优雅)
- 画质关键词(8K、超精细、获奖作品、专业摄影)
- 图像中的中文文本:文本内容、位置、字体风格、颜色、尺寸
为每个选项添加清晰标签(如A/B/C)及一行风格方向总结。让用户选择其中一个、组合多个选项的元素,或要求新的方向。迭代直到用户确认(如"开始"、"生成"、"ok"等),再进入生成环节。
Step 2 — Pick a model
步骤2 — 选择模型
Choose based on the request (see Model Selection Guide below). Default to
if unsure. Mention your choice to the user.
qwen-image-2.0-pro根据请求选择模型(参考下方模型选择指南)。若不确定,默认使用。向用户说明你的选择。
qwen-image-2.0-proStep 3 — Pick a size
步骤3 — 选择尺寸
Native 2K for Qwen-Image 2.0, // for Wan2.7, or an aspect-ratio
preset (, , etc.).
1K2K4K16:91:1Qwen-Image 2.0默认原生2K;万相2.7支持//;或使用宽高比预设(、等)。
1K2K4K16:91:1Step 4 — Generate
步骤4 — 生成图像
Run with the confirmed prompt and output path.
scripts/generate_image.py运行,传入确认后的提示词和输出路径。
scripts/generate_image.pyStep 5 — Save
步骤5 — 保存图像
If the output path was implicit, save into the user's current working directory.
若输出路径未明确指定,保存到用户当前工作目录。
Models
模型详情
Qwen-Image 2.0 family - Latest Flagship (MultiModalConversation API)
Qwen-Image 2.0系列 - 最新旗舰模型(MultiModalConversation API)
| Model | Description |
|---|---|
| Default. Latest flagship, native 2K, strongest typography and detail |
| Latest snapshot (Jun 2026): generation + editing fusion, better text rendering and prompt adherence |
| Standard 2.0 tier, native 2K |
| Previous-gen flagship (Dec 2025) |
| qwen-image-max snapshot: improved realism, fewer AI artifacts |
| 模型 | 描述 |
|---|---|
| 默认模型。最新旗舰版,原生2K分辨率,文本渲染和细节表现最强 |
| 最新快照版本(2026年6月):融合生成与编辑功能,文本渲染和提示词遵循度更优 |
| 标准2.0版本,原生2K分辨率 |
| 上一代旗舰模型(2025年12月) |
| qwen-image-max快照版本:写实度提升,AI伪影减少 |
Qwen-Image Edit family - Image Editing (MultiModalConversation API)
Qwen-Image Edit系列 - 图像编辑模型(MultiModalConversation API)
Editing models require an input image via (local path or URL). Omit to match the input image dimensions.
--image--size| Model | Description |
|---|---|
| Flagship editing model, strongest instruction following |
| Latest max snapshot (Jan 2026) |
| Faster, lower-cost editing |
编辑模型需通过参数传入输入图像(本地路径或URL)。省略参数将匹配输入图像尺寸。
--image--size| 模型 | 描述 |
|---|---|
| 旗舰编辑模型,指令遵循度最强 |
| 最新max快照版本(2026年1月) |
| 更快、低成本的编辑模型 |
Qwen-Image legacy (ImageSynthesis API)
Qwen-Image legacy(ImageSynthesis API)
| Model | Description |
|---|---|
| Distilled accelerated version of qwen-image-max |
| qwen-image-plus snapshot (Jan 2026): faster high-quality generation |
| Base model |
| 模型 | 描述 |
|---|---|
| qwen-image-max的蒸馏加速版本 |
| qwen-image-plus快照版本(2026年1月):更快生成高质量图像 |
| 基础模型 |
Wan Series - Photorealistic Generation (ImageGeneration API)
万相系列 - 写实风格生成模型(ImageGeneration API)
| Model | Description |
|---|---|
| Latest. Up to 4K output, unified architecture (T2I + edit + multi-image) |
| Wan 2.7 standard, up to 2K |
| Wan 2.6, flexible sizing |
| High quality, up to 768x2700 |
| Speed-optimized |
| Professional tier |
| Fast execution |
| Professional tier |
| Earlier generation |
| 模型 | 描述 |
|---|---|
| 最新版本。最高支持4K输出,统一架构(文本转图像+编辑+多图像) |
| 万相2.7标准版,最高支持2K |
| 万相2.6,尺寸灵活 |
| 高质量,最高支持768x2700 |
| 速度优化版本 |
| 专业版 |
| 快速执行版本 |
| 专业版 |
| 早期版本 |
Z-Image - Lightweight & Fast (MultiModalConversation API)
Z-Image - 轻量快速模型(MultiModalConversation API)
| Model | Description |
|---|---|
| Fast, low-cost generation; bilingual (CN/EN) text rendering, high-fidelity portraits and product images. Pixel area 512x512 to 2048x2048 |
| 模型 | 描述 |
|---|---|
| 快速、低成本生成;支持中英双语文本渲染,人像和产品图像保真度高。像素范围512x512至2048x2048 |
Volcano Ark - ByteDance Seedream (OpenAI-compatible API)
火山方舟 - 字节跳动Seedream(兼容OpenAI API)
| Model | Description |
|---|---|
| Ark default. Latest, up to 3K, PNG/JPEG output, best text rendering |
| Seedream 4.5, up to 4K |
| Seedream 4.0, up to 4K, budget-friendly |
| 模型 | 描述 |
|---|---|
| 方舟默认模型。最新版本,最高支持3K,输出PNG/JPEG格式,文本渲染效果最佳 |
| Seedream 4.5,最高支持4K |
| Seedream 4.0,最高支持4K,性价比高 |
Tencent Hunyuan (OpenAI-compatible API)
腾讯混元(兼容OpenAI API)
| Model | Description |
|---|---|
| Hunyuan default. Flagship 3.0, strong composition awareness, handles complex Chinese prompts up to 8K chars |
| 模型 | 描述 |
|---|---|
| 混元默认模型。旗舰3.0版本,构图感知能力强,支持最长8000字符的复杂中文提示词 |
Zhipu / BigModel - CogView-4 & GLM-Image (OpenAI-compatible API)
智谱(Zhipu)/ BigModel - CogView-4 & GLM-Image(兼容OpenAI API)
| Model | Description |
|---|---|
| Zhipu default. Stable alias for latest CogView-4, native Chinese text rendering |
| CogView-4 fixed snapshot (Mar 2025), reproducible results |
| GLM-Image flagship, up to 2048x2048, hybrid autoregressive/diffusion |
| 模型 | 描述 |
|---|---|
| 智谱默认模型。CogView-4最新版本的稳定别名,原生支持中文文本渲染 |
| CogView-4固定快照版本(2025年3月),结果可复现 |
| GLM-Image旗舰模型,最高支持2048x2048,混合自回归/扩散架构 |
StepFun / 阶跃星辰 - Step-2X (OpenAI-compatible API)
阶跃星辰(StepFun)- Step-2X(兼容OpenAI API)
| Model | Description |
|---|---|
| StepFun default. High quality (0.1 RMB/image), up to 1024x1024 |
| Fast & cheap (0.02 RMB/image), supports negative prompts, 8 inference steps |
| 模型 | 描述 |
|---|---|
| 阶跃星辰默认模型。高质量(0.1元/张),最高支持1024x1024 |
| 快速低成本(0.02元/张),支持负提示词,8步推理 |
Google Gemini - International (generateContent API)
Google Gemini(国际版)(generateContent API)
| Model | Description |
|---|---|
| Gemini default. Google flagship image model, 512/1K/2K named sizes plus aspect-ratio presets |
| 模型 | 描述 |
|---|---|
| Gemini默认模型。Google旗舰图像模型,支持512/1K/2K命名尺寸及宽高比预设 |
Usage
使用方法
Basic Usage
基础用法
bash
undefinedbash
undefinedDefault model (qwen-image-2.0-pro, native 2K output)
默认模型(qwen-image-2.0-pro,原生2K输出)
python ~/.claude/skills/imagencn/scripts/generate_image.py "A cute cat" output.png
python ~/.claude/skills/imagencn/scripts/generate_image.py "一只可爱的猫" output.png
Photorealistic with Wan model (Wan2.7 supports 4K)
使用万相模型生成写实风格图像(万相2.7支持4K)
python ~/.claude/skills/imagencn/scripts/generate_image.py --model wan2.7-image-pro --size 4K "Realistic photo of mountains at sunset" photo.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --model wan2.7-image-pro --size 4K "日落时分的写实山脉照片" photo.png
Edit an existing image (requires --image; local path or URL)
编辑现有图像(需--image参数;本地路径或URL)
python ~/.claude/skills/imagencn/scripts/generate_image.py --model qwen-image-edit-max --image input.png "Change the background to a beach at sunset" edited.png
undefinedpython ~/.claude/skills/imagencn/scripts/generate_image.py --model qwen-image-edit-max --image input.png "将背景改为日落时分的海滩" edited.png
undefinedSize Options
尺寸选项
bash
undefinedbash
undefinedUse ratio preset
使用宽高比预设
python ~/.claude/skills/imagencn/scripts/generate_image.py --size 16:9 "Wide landscape" landscape.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --size 16:9 "宽幅风景" landscape.png
Use exact dimensions
使用精确尺寸
python ~/.claude/skills/imagencn/scripts/generate_image.py --size 1280*720 "Custom size" custom.png
undefinedpython ~/.claude/skills/imagencn/scripts/generate_image.py --size 1280*720 "自定义尺寸" custom.png
undefinedSize Presets
尺寸预设
Qwen-Image 2.0 (native 2K):
- -> 2048x2048 (default)
1:1 - -> 2688x1536
16:9 - -> 1536x2688
9:16 - -> 2304x1728
4:3 - -> 1728x2304
3:4 - -> 1024x1024
1K - -> 2048x2048
2K
Qwen-Image legacy:
- -> 1328x1328
1:1 - -> 1664x928
16:9 - -> 928x1664
9:16 - -> 1472x1104
4:3 - -> 1104x1472
3:4
Z-Image (pixel area 512x512 to 2048x2048):
- -> 1024x1024 (default)
1:1 - -> 1280x720
16:9 - -> 720x1280
9:16 - -> 1024x1536
2:3 - -> 1536x1024
3:2 - -> 1024x1024
1K
Wan Series (Wan2.7 also accepts //):
1K2K4K- -> 1024x1024
1:1 - -> 1280x1280
1:1-large - -> 1280x720
16:9 - -> 720x1280
9:16 - -> 1200x900
4:3 - -> 900x1200
3:4 - -> 1440x720
2:1
Volcano Ark (Seedream):
- -> 2048x2048
1:1 - -> 2848x1600
16:9 - -> 1600x2848
9:16 - -> 2304x1728
4:3 - -> 1728x2304
3:4 - -> 2496x1664
3:2 - -> 1664x2496
2:3 - /
1K/2K/3K(model-dependent max resolution)4K
Tencent Hunyuan (colon-separated format):
- -> 1024:1024
1:1 - -> 1920:1080
16:9 - -> 1080:1920
9:16 - -> 1600:1200
4:3 - -> 1200:1600
3:4
Zhipu (CogView-4 / GLM-Image):
- -> 1024x1024 (default)
1:1 - -> 1344x768
16:9 - -> 768x1344
9:16 - -> 1152x864
4:3 - -> 864x1152
3:4 - -> 1440x720
2:1 - -> 720x1440
1:2
StepFun (Step-2X):
- -> 1024x1024 (default)
1:1 - -> 512x512
1:1-small - -> 1280x800
16:9 - -> 800x1280
9:16
Google Gemini (named sizes + aspect ratios):
- /
512(default) /1K-> named output size2K - ,
1:1,16:9,9:16,4:3-> aspect ratio (no exact pixel sizes)3:4
Qwen-Image 2.0(原生2K):
- -> 2048x2048(默认)
1:1 - -> 2688x1536
16:9 - -> 1536x2688
9:16 - -> 2304x1728
4:3 - -> 1728x2304
3:4 - -> 1024x1024
1K - -> 2048x2048
2K
Qwen-Image legacy:
- -> 1328x1328
1:1 - -> 1664x928
16:9 - -> 928x1664
9:16 - -> 1472x1104
4:3 - -> 1104x1472
3:4
Z-Image(像素范围512x512至2048x2048):
- -> 1024x1024(默认)
1:1 - -> 1280x720
16:9 - -> 720x1280
9:16 - -> 1024x1536
2:3 - -> 1536x1024
3:2 - -> 1024x1024
1K
万相系列(万相2.7还支持//):
1K2K4K- -> 1024x1024
1:1 - -> 1280x1280
1:1-large - -> 1280x720
16:9 - -> 720x1280
9:16 - -> 1200x900
4:3 - -> 900x1200
3:4 - -> 1440x720
2:1
火山方舟(Seedream):
- -> 2048x2048
1:1 - -> 2848x1600
16:9 - -> 1600x2848
9:16 - -> 2304x1728
4:3 - -> 1728x2304
3:4 - -> 2496x1664
3:2 - -> 1664x2496
2:3 - /
1K/2K/3K(取决于模型支持的最大分辨率)4K
腾讯混元(冒号分隔格式):
- -> 1024:1024
1:1 - -> 1920:1080
16:9 - -> 1080:1920
9:16 - -> 1600:1200
4:3 - -> 1200:1600
3:4
智谱(CogView-4 / GLM-Image):
- -> 1024x1024(默认)
1:1 - -> 1344x768
16:9 - -> 768x1344
9:16 - -> 1152x864
4:3 - -> 864x1152
3:4 - -> 1440x720
2:1 - -> 720x1440
1:2
阶跃星辰(Step-2X):
- -> 1024x1024(默认)
1:1 - -> 512x512
1:1-small - -> 1280x800
16:9 - -> 800x1280
9:16
Google Gemini(命名尺寸+宽高比):
- /
512(默认) /1K-> 命名输出尺寸2K - ,
1:1,16:9,9:16,4:3-> 宽高比(无精确像素尺寸)3:4
Advanced Options
高级选项
bash
undefinedbash
undefinedWith negative prompt
使用负提示词
python ~/.claude/skills/imagencn/scripts/generate_image.py --negative "blurry, low quality" "High quality portrait" portrait.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --negative "模糊, 低画质" "高质量人像" portrait.png
Disable automatic prompt extension (DashScope only)
禁用自动提示词扩展(仅DashScope支持)
python ~/.claude/skills/imagencn/scripts/generate_image.py --no-extend "A photorealistic cat" cat.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --no-extend "写实风格的猫" cat.png
Set random seed for reproducibility
设置随机种子以复现结果
python ~/.claude/skills/imagencn/scripts/generate_image.py --seed 42 "A cat" cat.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --seed 42 "一只猫" cat.png
Guidance scale (Volcano Ark only)
设置引导尺度(仅火山方舟支持)
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform ark --guidance-scale 7.5 "Portrait" portrait.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform ark --guidance-scale 7.5 "人像" portrait.png
Disable watermark (Volcano Ark only)
禁用水印(仅火山方舟支持)
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform ark --no-watermark "Artwork" art.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform ark --no-watermark "艺术作品" art.png
Auto-enhance prompt on/off (Tencent Hunyuan only, --revise 0=off 1=on)
开启/关闭提示词自动优化(仅腾讯混元支持,--revise 0=关闭 1=开启)
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform hunyuan --revise 0 "A cat" cat.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform hunyuan --revise 0 "一只猫" cat.png
Add AI logo (Tencent Hunyuan only, --logo 0=no 1=yes)
添加AI标识(仅腾讯混元支持,--logo 0=不添加 1=添加)
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform hunyuan --logo 1 "Poster" poster.png
python ~/.claude/skills/imagencn/scripts/generate_image.py --platform hunyuan --logo 1 "海报" poster.png
Dry run (preview without making API call)
试运行(预览但不调用API)
python ~/.claude/skills/imagencn/scripts/generate_image.py --dry-run --platform ark "Test prompt"
python ~/.claude/skills/imagencn/scripts/generate_image.py --dry-run --platform ark "测试提示词"
List all models
列出所有模型
python ~/.claude/skills/imagencn/scripts/generate_image.py --list-models
undefinedpython ~/.claude/skills/imagencn/scripts/generate_image.py --list-models
undefinedRequirements
依赖要求
bash
pip install dashscope requestsbash
pip install dashscope requestsOptional: for coloured output and styled tables
可选:用于彩色输出和样式化表格
pip install rich
undefinedpip install rich
undefinedEnvironment Variables
环境变量
bash
undefinedbash
undefinedAlibaba Cloud Bailian (DashScope)
阿里云百炼(DashScope)
export DASHSCOPE_API_KEY="your_api_key" # Required
export DASHSCOPE_MODEL="wan2.7-image-pro" # Optional default model
export DASHSCOPE_API_BASE="cn" # Optional: cn, sg, us
export DASHSCOPE_API_KEY="your_api_key" # 必填
export DASHSCOPE_MODEL="wan2.7-image-pro" # 可选:默认模型
export DASHSCOPE_API_BASE="cn" # 可选:cn, sg, us
ByteDance Volcano Ark
字节跳动火山方舟
export ARK_API_KEY="your_api_key" # Required for Ark
export ARK_MODEL="doubao-seedream-5-0-260128" # Optional default model
export ARK_API_KEY="your_api_key" # 使用方舟必填
export ARK_MODEL="doubao-seedream-5-0-260128" # 可选:默认模型
Tencent Hunyuan (TokenHub)
腾讯混元(TokenHub)
export HUNYUAN_API_KEY="your_api_key" # Required for Hunyuan
export HUNYUAN_MODEL="hy-image-v3.0" # Optional default model
export HUNYUAN_API_KEY="your_api_key" # 使用混元必填
export HUNYUAN_MODEL="hy-image-v3.0" # 可选:默认模型
Zhipu / BigModel
智谱(Zhipu)/ BigModel
export ZHIPUAI_API_KEY="your_api_key" # Required for Zhipu
export ZHIPUAI_MODEL="cogview-4" # Optional default model
export ZHIPUAI_API_KEY="your_api_key" # 使用智谱必填
export ZHIPUAI_MODEL="cogview-4" # 可选:默认模型
StepFun / 阶跃星辰
阶跃星辰(StepFun)
export STEP_API_KEY="your_api_key" # Required for StepFun
export STEP_MODEL="step-2x-large" # Optional default model
export STEP_API_KEY="your_api_key" # 使用阶跃星辰必填
export STEP_MODEL="step-2x-large" # 可选:默认模型
Google Gemini (international)
Google Gemini(国际版)
export GEMINI_API_KEY="your_api_key" # Required for Gemini
export GEMINI_MODEL="gemini-3-pro-image-preview" # Optional default model
Get API Keys:
- DashScope: https://bailian.console.aliyun.com/
- Volcano Ark: https://console.volcengine.com/ark/region:ark+cn-beijing/apikey
- Tencent Hunyuan: https://console.cloud.tencent.com/tokenhub/apikey
- Zhipu: https://bigmodel.cn
- StepFun: https://platform.stepfun.com/interface-key
- Google Gemini: https://aistudio.google.com/export GEMINI_API_KEY="your_api_key" # 使用Gemini必填
export GEMINI_MODEL="gemini-3-pro-image-preview" # 可选:默认模型
获取API密钥:
- DashScope:https://bailian.console.aliyun.com/
- Volcano Ark:https://console.volcengine.com/ark/region:ark+cn-beijing/apikey
- Tencent Hunyuan:https://console.cloud.tencent.com/tokenhub/apikey
- Zhipu:https://bigmodel.cn
- StepFun:https://platform.stepfun.com/interface-key
- Google Gemini:https://aistudio.google.com/Config File (Optional)
配置文件(可选)
Create for personal defaults, or in a project
directory for per-project overrides. API keys stay in environment variables for
security.
~/.imagencn.json.imagencn.jsonjson
{
"platform": "ark",
"model": "doubao-seedream-5-0-260128",
"size": "2K"
}All keys are optional. Priority (highest first):
- CLI arguments (,
--platform,--model)--size - Project config (in current directory)
.imagencn.json - User config ()
~/.imagencn.json - Environment variables (,
DASHSCOPE_MODEL,ARK_MODEL,HUNYUAN_MODEL,ZHIPUAI_MODEL,STEP_MODEL)GEMINI_MODEL - Built-in defaults
创建设置个人默认值,或在项目目录中创建设置项目级覆盖配置。API密钥需保存在环境变量中以保证安全。
~/.imagencn.json.imagencn.jsonjson
{
"platform": "ark",
"model": "doubao-seedream-5-0-260128",
"size": "2K"
}所有配置项均为可选。优先级(从高到低):
- CLI参数(,
--platform,--model)--size - 项目配置(当前目录下的)
.imagencn.json - 用户配置()
~/.imagencn.json - 环境变量(,
DASHSCOPE_MODEL,ARK_MODEL,HUNYUAN_MODEL,ZHIPUAI_MODEL,STEP_MODEL)GEMINI_MODEL - 内置默认值
API Endpoints
API端点
| Region | Alias | URL |
|---|---|---|
| China (default) | | |
| Singapore | | |
| Virginia | | |
bash
undefined| 区域 | 别名 | URL |
|---|---|---|
| 国内(默认) | | |
| 新加坡 | | |
| 弗吉尼亚 | | |
bash
undefinedSwitch to Singapore endpoint
切换到新加坡端点
export DASHSCOPE_API_BASE="sg"
export DASHSCOPE_API_BASE="sg"
Or use full URL
或使用完整URL
export DASHSCOPE_API_BASE="https://dashscope-intl.aliyuncs.com/api/v1"
undefinedexport DASHSCOPE_API_BASE="https://dashscope-intl.aliyuncs.com/api/v1"
undefinedModel Selection Guide
模型选择指南
Quick Pick — You Only Need Eight
快速选择 — 只需记住8个模型
| What you want | Model | Platform |
|---|---|---|
| Default / general (posters, text) | | DashScope |
| Photorealistic (portraits, landscapes) | | DashScope |
| Edit an image | | DashScope |
| Cheap & fast | | DashScope |
| Photo + text combo | | Volcano Ark |
| Complex Chinese composition | | Tencent Hunyuan |
| Chinese text in images | | Zhipu |
| Ultra-cheap volume gen | | StepFun |
| International (non-China) | | Google Gemini |
All other models are legacy/snapshot variants.
| 需求 | 模型 | 平台 |
|---|---|---|
| 默认/通用场景(海报、文本) | | DashScope |
| 写实风格(人像、风景) | | DashScope |
| 图像编辑 | | DashScope |
| 低成本快速生成 | | DashScope |
| 照片+文本组合 | | Volcano Ark |
| 复杂中文构图 | | Tencent Hunyuan |
| 图像中的中文文本 | | Zhipu |
| 超低成本批量生成 | | StepFun |
| 国际场景(非国内) | | Google Gemini |
其他模型均为旧版/快照变体。
Full Reference
完整参考
| Use Case | Recommended Model |
|---|---|
| General high-quality (default) | |
| Chinese text/calligraphy | |
| English text on images | |
| Posters with typography | |
| Photorealistic photos (4K) | |
| Photorealistic photos (2K) | |
| Portrait photography | |
| Image editing (best quality) | |
| Image editing (fast, low-cost) | |
| Fast, low-cost generation | |
| High-fidelity portraits / product shots (fast) | |
| Fast photorealistic (Wan) | |
| Lower-cost text rendering | |
| ByteDance best quality | |
| Budget-friendly 4K (ByteDance) | |
| Complex Chinese prompts (Tencent) | |
| 使用场景 | 推荐模型 |
|---|---|
| 通用高质量生成(默认) | |
| 中文文本/书法 | |
| 图像中的英文文本 | |
| 带排版的海报 | |
| 写实风格照片(4K) | |
| 写实风格照片(2K) | |
| 人像摄影 | |
| 图像编辑(最佳质量) | |
| 图像编辑(快速低成本) | |
| 快速低成本生成 | |
| 高保真人像/产品拍摄(快速) | |
| 快速写实风格生成(万相) | |
| 低成本文本渲染 | |
| 字节跳动最佳质量 | |
| 高性价比4K生成(字节跳动) | |
| 复杂中文提示词(腾讯) | |
Platform Quick Comparison
平台快速对比
| Feature | DashScope | Ark | Hunyuan | Zhipu | StepFun | Gemini |
|---|---|---|---|---|---|---|
| Best for | Text, variety | Photo+text | Complex CN | CN text in image | Ultra-cheap | International |
| Max res | 4K | 4K | 2K | 2K | 1K | 2K |
| SDK | | None | None | None | None | None |
| Price | Varies | ~0.22 | ~0.20 | ~0.06 | ~0.02 | ~$0.13 |
| Env var | | | | | | |
| 特性 | DashScope | Ark | Hunyuan | Zhipu | StepFun | Gemini |
|---|---|---|---|---|---|---|
| 最佳适用场景 | 文本、多样化风格 | 照片+文本 | 复杂中文 | 图像中的中文文本 | 超低成本 | 国际场景 |
| 最大分辨率 | 4K | 4K | 2K | 2K | 1K | 2K |
| SDK | | 无 | 无 | 无 | 无 | 无 |
| 价格 | 浮动 | ~0.22元 | ~0.20元 | ~0.06元 | ~0.02元 | ~$0.13 |
| 环境变量 | | | | | | |
Examples
示例
Volcano Ark (ByteDance)
火山方舟(字节跳动)
bash
undefinedbash
undefinedDefault Ark model (Seedream 5.0)
默认方舟模型(Seedream 5.0)
ARK_API_KEY="xxx" python scripts/generate_image.py
--platform ark
"A vibrant close-up editorial portrait, Vogue magazine cover style"
portrait.png
--platform ark
"A vibrant close-up editorial portrait, Vogue magazine cover style"
portrait.png
ARK_API_KEY="xxx" python scripts/generate_image.py
--platform ark
"充满活力的特写人像,Vogue杂志封面风格"
portrait.png
--platform ark
"充满活力的特写人像,Vogue杂志封面风格"
portrait.png
With 4K output
4K输出
ARK_API_KEY="xxx" python scripts/generate_image.py
--platform ark --model doubao-seedream-4-5-251128 --size 4K
"Breathtaking mountain sunset, golden hour, professional photography"
landscape.png
--platform ark --model doubao-seedream-4-5-251128 --size 4K
"Breathtaking mountain sunset, golden hour, professional photography"
landscape.png
undefinedARK_API_KEY="xxx" python scripts/generate_image.py
--platform ark --model doubao-seedream-4-5-251128 --size 4K
"令人惊叹的山脉日落,黄金时段,专业摄影"
landscape.png
--platform ark --model doubao-seedream-4-5-251128 --size 4K
"令人惊叹的山脉日落,黄金时段,专业摄影"
landscape.png
undefinedTencent Hunyuan
腾讯混元
bash
undefinedbash
undefinedDefault Hunyuan model (Image 3.0)
默认混元模型(图像3.0)
HUNYUAN_API_KEY="xxx" python scripts/generate_image.py
--platform hunyuan
"An astronaut riding a horse on the moon, cinematic lighting, 8K detail"
scifi.png
--platform hunyuan
"An astronaut riding a horse on the moon, cinematic lighting, 8K detail"
scifi.png
HUNYUAN_API_KEY="xxx" python scripts/generate_image.py
--platform hunyuan
"宇航员在月球上骑马,电影级光线,8K细节"
scifi.png
--platform hunyuan
"宇航员在月球上骑马,电影级光线,8K细节"
scifi.png
With prompt auto-enhance disabled
禁用提示词自动优化
HUNYUAN_API_KEY="xxx" python scripts/generate_image.py
--platform hunyuan --revise 0
"A cute orange cat napping in sunlight, oil painting style"
cat.png
--platform hunyuan --revise 0
"A cute orange cat napping in sunlight, oil painting style"
cat.png
undefinedHUNYUAN_API_KEY="xxx" python scripts/generate_image.py
--platform hunyuan --revise 0
"一只可爱的橘猫在阳光下打盹,油画风格"
cat.png
--platform hunyuan --revise 0
"一只可爱的橘猫在阳光下打盹,油画风格"
cat.png
undefinedGoogle Gemini (international)
Google Gemini(国际版)
bash
undefinedbash
undefinedDefault Gemini model (Gemini 3 Pro Image)
默认Gemini模型(Gemini 3 Pro Image)
GEMINI_API_KEY="xxx" python scripts/generate_image.py
--platform gemini --size 2K
"A serene Japanese garden with koi pond, soft morning light"
garden.png
--platform gemini --size 2K
"A serene Japanese garden with koi pond, soft morning light"
garden.png
undefinedGEMINI_API_KEY="xxx" python scripts/generate_image.py
--platform gemini --size 2K
"宁静的日式花园,带有锦鲤池,柔和的晨光"
garden.png
--platform gemini --size 2K
"宁静的日式花园,带有锦鲤池,柔和的晨光"
garden.png
undefinedChinese New Year Poster (DashScope)
春节海报(DashScope)
bash
python ~/.claude/skills/imagencn/scripts/generate_image.py \
"A beautiful Chinese New Year poster with red background, golden text, fireworks and firecrackers" \
new_year_poster.pngbash
python ~/.claude/skills/imagencn/scripts/generate_image.py \
"精美的春节海报,红色背景,金色文字,烟花和鞭炮" \
new_year_poster.pngPhotorealistic Landscape (4K)
写实风格风景(4K)
bash
python ~/.claude/skills/imagencn/scripts/generate_image.py \
--model wan2.7-image-pro \
--size 4K \
"Breathtaking sunset over mountain range, golden hour, professional photography" \
landscape.pngbash
python ~/.claude/skills/imagencn/scripts/generate_image.py \
--model wan2.7-image-pro \
--size 4K \
"令人惊叹的山脉日落,黄金时段,专业摄影" \
landscape.pngProduct Shot
产品拍摄
bash
python ~/.claude/skills/imagencn/scripts/generate_image.py \
--model wan2.7-image \
--size 2K \
"Professional product photography of a coffee cup on marble surface, studio lighting" \
product.pngbash
python ~/.claude/skills/imagencn/scripts/generate_image.py \
--model wan2.7-image \
--size 2K \
"大理石桌面上的咖啡杯专业产品摄影,影棚光线" \
product.png