217 Commits
Author SHA1 Message Date
yanlang0123 1334f81d86 Fix-pose-editor-overlay-and-video-compositor 2026-09-29 17:17:07 +08:00
yanlang0123 7c4f37feb7 Add-app-manager-visibility-setting 2026-09-29 10:26:01 +08:00
yanlang0123 f5312a6b25 Fix-PromptQueue-process-item-compatibility 2026-09-29 10:26:01 +08:00
yanlang0123 86ef43cb3f Fix-insightface-installation-flag 2026-09-29 10:10:37 +08:00
yanlang0123 1324920f90 Fix-WeChatAuth-client-initialization 2026-09-29 10:05:26 +08:00
yanlang0123 f836b150f6 优化 2026-08-14 23:50:27 +08:00
yanlang0123 ef95985e54 插件优化,新增智能体节点 2026-08-13 21:07:59 +08:00
yanlang0123 5ac1eca669 跟新 2026-08-08 15:41:40 +08:00
严浪 99f747e138 优化 2026-06-14 11:36:55 +08:00
yanlang0123 85472ab1ca 优化速度 2026-06-10 18:59:26 +08:00
yanlang0123 11dfe2b01e 添加进度条 2026-06-10 18:04:44 +08:00
yanlang0123 2cc5f3d023 功能优化 2026-06-10 16:14:42 +08:00
yanlang0123 a6b94affe3 优化视频合成功能 2026-06-10 12:24:38 +08:00
yanlang0123 9b14684728 新增视频直接生成 2026-06-10 10:48:14 +08:00
yanlang0123 57abc15b44 feat: implement web-based video editing interface and backend integration for ComfyUI-Lam 2026-05-26 18:15:56 +08:00
yanlang0123 6e4b0c2c2b feat: add VideoInfo node to extract duration, frame count, FPS, and dimensions from video files 2026-05-20 16:57:37 +08:00
严浪 6a02877a65 中英文字母换行问题修改 2026-03-28 12:08:45 +08:00
yanlang0123 89b37b2df2 重复系统提示词问题 2026-01-04 17:21:26 +08:00
严浪 5e7d75e997 Merge branch 'cu121' of https://gitee.com/yanlang0123/ComfyUI_Lam into cu121 2026-01-04 15:33:14 +08:00
严浪 f57b029090 功能优化 2026-01-04 15:32:00 +08:00
yanlang0123 5b8eedc3ea 功能优化 2025-12-16 17:15:58 +08:00
yanlang0123 1d4c710d0d 删除无用功能,还原稳定版本 2025-12-16 17:06:39 +08:00
yanlang0123 aaa3e63ed8 修改 2025-12-12 18:59:23 +08:00
yanlang0123 7335b563c2 BUG修改 2025-12-12 15:19:57 +08:00
yanlang0123 e59d44349c 功能优化 2025-12-12 15:16:21 +08:00
yanlang0123 9027b465e6 优化 2025-12-11 14:08:37 +08:00
yanlang0123 cc60facb14 优化 2025-12-09 19:05:24 +08:00
yanlang0123 3ccfeac8ec 子工作流嵌套 2025-12-09 18:58:06 +08:00
严浪 c10ab633fd 功能升级 2025-12-04 23:07:06 +08:00
严浪 3c24e263f3 Merge branch 'cu121' of https://gitee.com/yanlang0123/ComfyUI_Lam into cu121 2025-11-23 21:34:25 +08:00
严浪 a80d4e6d6c 通用隐藏节点保存API确实内容问题修改 2025-11-23 21:34:20 +08:00
yanlang0123 88b8aab54e 优化页面 2025-11-13 17:52:37 +08:00
yanlang0123 6d99369691 测试注释修改 2025-11-03 16:39:15 +08:00
严浪 db0c13bbb2 优化BUG 2025-11-01 15:34:50 +08:00
yanlang0123 4042e8a9c4 姿态编辑异常问题修改 2025-10-31 18:39:05 +08:00
yanlang0123 82c1af1ae8 兼容性修改 2025-10-29 11:47:46 +08:00
yanlang0123 eed0b4f41f 版本升级,部分功能兼容性修改,循环API支持 2025-10-29 11:18:30 +08:00
yanlang0123 6c7e0dec27 适应最新版本,新增产生修改 2025-10-29 09:47:04 +08:00
严浪 a86cec2907 插件升级IndexTTS2BUG修改 2025-09-22 22:46:59 +08:00
严浪 10b8656020 功能优化跟新 2025-09-14 16:05:14 +08:00
严浪 8555f231cb 优化功能 2025-08-30 20:49:11 +08:00
严浪 266038cc1d BUG修改 2025-08-30 12:13:48 +08:00
严浪 40c83259fa 删除测试打印代码 2025-08-29 17:15:10 +08:00
严浪 bc743158dd 字幕识别出现时间交集BUG修改 2025-08-29 17:14:19 +08:00
严浪 263e5b30fc 功能优化 2025-08-26 21:47:31 +08:00
严浪 f2bbec4f8f ss 2025-08-24 21:22:10 +08:00
严浪 0c601221a4 剪映插件修改,新增轨道功能解决多音轨等等效果 2025-08-24 19:57:33 +08:00
严浪 bc19ee4efb 修改内容升级 2025-08-17 10:18:13 +08:00
严浪 bc44c2402e BUG修改 2025-08-16 23:07:24 +08:00
严浪 4241616995 bug修改 2025-08-16 22:42:57 +08:00
严浪 35312c1d99 版本号 2025-08-16 21:43:27 +08:00
严浪 213bb653f0 新增图片地址获取节点 2025-08-16 21:34:55 +08:00
严浪 b305200d22 新增视频提取音频节点,BUG修改 2025-07-20 18:44:48 +08:00
严浪 91e98883d5 音频地址获取没有预览问题修改 2025-07-20 15:05:54 +08:00
严浪 9922dd947d 剪映草稿云端下载修改 2025-07-18 22:01:52 +08:00
严浪 f8324c47dc 新增剪映草稿下载节点 2025-07-16 22:44:02 +08:00
严浪 f9e10fe818 配置文件修改 2025-07-14 22:15:38 +08:00
严浪 b9f98e81e8 新增剪映相关节点 2025-07-13 17:26:08 +08:00
严浪 4686c91a51 BUG修改新增HeyGem节点 2025-07-06 21:31:58 +08:00
严浪 25539ccec8 功能优化升级 2025-07-04 09:43:59 +08:00
严浪 4620513cc1 功能优化,新增通用隐藏节点 2025-05-23 17:09:52 +08:00
严浪 f555497bfc 功能优化网络图片加载目录问题修改 2025-04-29 16:53:24 +08:00
严浪 3df598d6c2 BUG修改功能优化 2025-04-29 14:48:15 +08:00
严浪 2d3b866032 升级后页面优化,添加演示实例 2025-04-10 15:33:16 +08:00
严浪 861d9615e7 新版显示错乱问题修改 2025-04-10 10:04:29 +08:00
严浪 092d3c5e2d Merge branch 'cu121' of https://gitee.com/yanlang0123/ComfyUI_Lam into cu121 2025-04-08 20:28:42 +08:00
严浪 c2daa1110c 等待选择节点功能优化 2025-04-08 20:28:36 +08:00
yanlang0123 3fada07171 Merge pull request #10 from ComfyNodePRs/update-publish-yaml
Update Github Action for Publishing to Comfy Registry
2025-04-07 16:44:08 +08:00
严浪 bace98881b 'WeChatAuth'导入异常 Client.__init__()修改 2025-04-06 22:04:01 +08:00
严浪 4ed3a3891e 添加 脸部裁剪模型加载节点 2025-04-06 21:53:20 +08:00
严浪 b4fdd5c486 应用提交异常后不能再次提交问题修改,添加实例 2025-04-06 21:52:51 +08:00
严浪 059dbfc250 新增画板预览按钮 2025-04-05 20:09:20 +08:00
严浪 9562065578 画板功能优化 2025-04-05 19:38:33 +08:00
严浪 efa40d9bf9 画板功能提交 2025-04-04 22:31:33 +08:00
严浪 3319b4ee97 画板功能提交 2025-04-04 22:31:19 +08:00
严浪 1a6fd5deed 画板功能提交 2025-04-04 22:19:45 +08:00
严浪 3ffd87fb8b 部分依赖加载问题优化 2025-04-04 19:18:47 +08:00
snomiao b339e5b342 chore(publish): update workflow for node publishing
- Add permissions for issue writing
- Modify condition to check repository owner instead of fork status
- Update action version for publishing node to v1
2025-03-08 20:13:23 +00:00
严浪 07f9f36109 新增参数类型选择,图片和遮罩设置 2025-03-06 22:51:54 +08:00
严浪 1321e45b9a 功能升级 2025-03-05 22:30:19 +08:00
严浪 70b9905e15 功能优化 2025-03-04 21:38:08 +08:00
严浪 87b68ec39e 功能优化升级 2025-03-01 17:05:25 +08:00
严浪 8f65ad0494 功能优化 2025-02-27 15:55:06 +08:00
严浪 0ca18e606c 加入微信绑定功能 2025-02-14 23:03:24 +08:00
严浪 50625c5bd6 优化显示 2025-02-13 16:33:07 +08:00
严浪 5fcb77f5f5 功能优化 2025-02-13 15:51:17 +08:00
严浪 c5da17ec55 功能优化 2025-02-12 15:13:28 +08:00
严浪 e3b5b29701 功能优化,新增公众号deepseek ai 2025-02-09 15:40:12 +08:00
严浪 4402100985 功能优化,app页面优化 2025-02-09 10:32:27 +08:00
严浪 b1785ed5a9 手势拖拽报错版本兼容性问题修改 2024-11-30 16:19:18 +08:00
严浪 a7d8b8abe7 二维码识别升级,新增微信模型识别二维码提高识别效果 2024-10-29 18:35:16 +08:00
严浪 d62da0f4be 优化,二维码识别依赖包使用异常不处理 2024-10-29 07:01:33 +08:00
严浪 9f6f5cd999 分段负载redis长时间没有消息离线问题修改 2024-10-23 12:31:40 +08:00
严浪 97f4a49e70 Merge branch 'cu121' of https://gitee.com/yanlang0123/ComfyUI_Lam into cu121 2024-10-17 13:09:41 +08:00
严浪 38d0cc67ca tj 2024-10-17 12:56:07 +08:00
严浪 1456e3d87b 提交测试 2024-10-17 09:12:11 +08:00
yanlang0123 b38a4ff519 删除文件 .github/workflows 2024-10-17 00:57:23 +00:00
严浪 bed2cf6eda AAA 2024-10-17 08:20:58 +08:00
严浪 43cf1e708a 功能优化 2024-10-17 08:16:40 +08:00
严浪 e56b7095dd 新增微信公众号AI客服,AI调用文生图功能 2024-10-16 11:14:14 +08:00
严浪 1d1734e6bf 位置错乱问题修改,ctrl+c样式问题修改 2024-10-16 10:49:37 +08:00
严浪 ed36112c9a 新增微信公众号AI自动回复,新增工作流分段负载,异常捕获,避免整个插件不能用 2024-10-15 14:25:49 +08:00
严浪 88045a0cdb 等待选择,循环显示优化 2024-10-01 17:09:08 +08:00
严浪 2d8c32c558 判断选择优化 2024-09-27 15:06:54 +08:00
严浪 1160acdc3f 样式问题修改 2024-09-25 16:32:53 +08:00
严浪 729f318793 功能优化 2024-09-18 17:21:08 +08:00
严浪 e622c7aa52 功能修改 2024-09-18 12:02:29 +08:00
严浪 e737e8d51f 集群功能优化 2024-09-16 18:32:58 +08:00
严浪 bf496ccb52 内循环和判断循环改版,使用原生结构实现循环效果 2024-09-16 10:49:30 +08:00
严浪 0f863914a0 功能优化 2024-09-13 12:57:42 +08:00
严浪 2f2c5a83a3 language报错优化 2024-09-13 10:10:02 +08:00
严浪 432b0c90ee 判断选择功能优化 2024-09-12 17:19:27 +08:00
严浪 c731ab9b8c 功能优化 2024-09-12 10:44:45 +08:00
严浪 6292d0282e 插件新增节点BUG优化 2024-09-08 19:03:49 +08:00
严浪 5dc47f3ad4 BUG修改功能优化 2024-09-05 22:34:53 +08:00
严浪 1c3c645c06 新增通用名称获取节点,部分功能优化 2024-08-30 22:27:29 +08:00
严浪 29ea409ce6 模型下载批处理脚本修改 2024-08-28 22:16:42 +08:00
严浪 a6e88ebec4 Merge branch 'cu121' of https://gitee.com/yanlang0123/ComfyUI_Lam into cu121 2024-08-28 08:51:27 +08:00
严浪 be761dc03b 无用文件删除 2024-08-28 08:51:22 +08:00
yanlang0123 b8fca42b62 删除文件 backup 2024-08-28 00:50:38 +00:00
yanlang0123 e35c8e2745 判断循环优化 2024-08-25 15:19:32 +08:00
yanlang0123 e9deb188ee 优化 2024-08-24 10:33:27 +08:00
yanlang0123 014468a566 脚本优化 2024-08-23 23:11:53 +08:00
yanlang0123 554daf2c82 新版中文搜索支持 2024-08-23 22:35:49 +08:00
yanlang0123 073810ff20 优化 2024-08-22 19:08:22 +08:00
yanlang0123 c6900878df 添加翻译模型下载脚本 2024-08-22 18:15:02 +08:00
yanlang0123 7071b417b1 新样式不兼容问题修改 2024-08-22 00:10:02 +08:00
yanlang0123 74bc9f094b bug修改 2024-08-21 20:45:07 +08:00
yanlang0123 38b691735e 优化 2024-08-21 16:20:12 +08:00
yanlang0123 0861e74a17 BUG修改 2024-08-21 16:11:20 +08:00
yanlang0123 43257bd74d 风格选择器无法选择问题修改 2024-08-21 15:05:45 +08:00
yanlang0123 0b1878f2c6 BUG修改 2024-08-21 14:16:38 +08:00
yanlang0123 0aed0da0c7 部分节点改成非输出节点,添加通用打印节点 2024-08-21 13:48:20 +08:00
yanlang0123 c8705b3038 新版comfyUI功能适配 2024-08-21 13:26:07 +08:00
yanlang0123 1fe5f0b67f 循环嵌套优化 2024-08-17 13:57:22 +08:00
yanlang0123 43ecd83262 嵌套内循环问题修改 2024-08-17 00:01:00 +08:00
yanlang0123 357ccd2b43 计次循环优化升级 2024-08-16 16:04:45 +08:00
yanlang0123 ba7194836a update js/litegraphCore.js.
中文搜索闪退问题修改

Signed-off-by: yanlang0123 <yanlang0123@gmail.com>
2024-08-12 05:33:51 +00:00
严浪 7e88221dfc 添加漏的包 2024-07-14 15:47:27 +08:00
严浪 17a4350230 获取不到插件路径问题修改 2024-07-14 15:40:35 +08:00
严浪 cc16c74e2e 说明修改 2024-07-04 17:15:29 +08:00
严浪 6c841b9785 xiug 2024-07-04 16:46:04 +08:00
严浪 5bb167364c 文件自动修改,多余文件删除 2024-07-04 16:07:07 +08:00
严浪 7295cd63ca redis实现集群负载 2024-07-03 19:16:36 +08:00
严浪 9455d5d2d8 功能优化 2024-05-28 17:02:49 +08:00
严浪 b6a580dc64 功能优化 2024-05-28 16:52:29 +08:00
严浪 ce7571a0a8 新增判断循环功能,功能优化BUG修改 2024-05-27 10:44:39 +08:00
严浪 2e23311d93 优化修改 2024-05-22 14:30:08 +08:00
严浪 d3c4a7f5a7 备注修改 2024-05-19 22:27:04 +08:00
严浪 a5b86e8295 功能优化,添加部分节点使用说明 2024-05-19 22:21:38 +08:00
严浪 86fd7a155b 宽高自适应 2024-05-17 09:34:48 +08:00
严浪 a87a5b69c5 姿态编辑部添加提示词填写 2024-05-15 17:36:47 +08:00
严浪 5c86e1db4e 功能优化 2024-05-15 17:33:40 +08:00
严浪 c4d1d89216 功能优化修改 2024-05-10 11:37:31 +08:00
严浪 a60f88e1c7 修改说明 2024-05-05 15:00:56 +08:00
严浪 67ab97db66 插件优化升级 2024-05-03 21:50:03 +08:00
严浪 5ae8f15cd1 新增多IPAdapter多参考图遮罩功能节点 2024-04-24 22:10:54 +08:00
严浪 70b758753a 修改名称 2024-04-23 14:23:00 +08:00
严浪 c1b3cb7a49 功能优化 2024-04-23 09:02:16 +08:00
严浪 5e67c1ad70 BUG修改 2024-04-22 16:42:32 +08:00
严浪 1ec31bc90d 微信公众号应用开发,支持微信公众号,下发指令生图,和网页版 2024-04-21 17:24:39 +08:00
yanlang0123 396dd1a2cb !21 add tags/角色列表.yml.
Merge pull request !21 from 吴嘚嘚/N/A
2024-04-16 07:26:05 +00:00
严浪 e251f21a31 BUG修改 2024-04-16 13:56:16 +08:00
严浪 7f28088ffc BUG修改 2024-04-14 15:59:53 +08:00
严浪 877ff5df9e 修改循环下标 2024-04-13 16:57:57 +08:00
严浪 60c66b5ada 新增节点 功能优化 2024-04-13 13:13:28 +08:00
严浪 bd24491d8f BUG修改 2024-04-08 15:00:40 +08:00
严浪 5d360d1801 修改execution.py说明 2024-04-08 09:32:34 +08:00
严浪 fa6888994e 循环插件修改,风格节点优化 2024-04-08 09:19:20 +08:00
严浪 ac46a7161c 添加姿态编辑配套节点,修改BUG 2024-04-06 22:19:35 +08:00
吴嘚嘚 aebc08230e add tags/角色列表.yml.
添加animagine-xl-v3.1的角色列表,不过是基本没翻译的

Signed-off-by: 吴嘚嘚 <13742125+wu-deidei@user.noreply.gitee.com>
2024-04-06 07:10:57 +00:00
严浪 84693474b5 Merge branch 'cu121' of https://gitee.com/yanlang0123/ComfyUI_Lam into cu121 2024-04-04 14:34:29 +08:00
严浪 6c1f6cc1a1 删除多余按钮 2024-04-04 14:34:24 +08:00
yanlang0123 230b5cc340 删除文件 tags/十八禁.yml 2024-04-01 05:51:12 +00:00
严浪 b482183255 修改BUG 2024-04-01 13:45:21 +08:00
严浪 dcb89c0080 姿态添加手 2024-04-01 13:41:15 +08:00
严浪 e41bfc5b15 功能优化 2024-03-29 16:15:23 +08:00
严浪 2999d2ee9d 功能优化 2024-03-29 14:55:50 +08:00
严浪 c89f08d49c 修改目录 2024-03-22 16:35:24 +08:00
严浪 332b2315cb 脸部识别 2024-03-22 11:09:38 +08:00
严浪 779223ead3 多余删除 2024-03-22 11:07:20 +08:00
严浪 1e4f7ec11e 添加姿态编辑节点 2024-03-22 11:06:46 +08:00
严浪 5a90359bc1 插件升级 2024-03-09 21:24:05 +08:00
yanlang0123 c173a7bcba !16 提示词选择改成可操作权重
Merge pull request !16 from 吴嘚嘚/N/A ,感谢PR
2024-03-07 09:30:05 +00:00
严浪 adc20028c1 添加使用说明 2024-03-07 17:17:53 +08:00
严浪 4cc6621277 添加内循环插件,及几个辅助插件 2024-03-06 18:19:31 +08:00
吴嘚嘚 3ed6e5292a 提示词选择改成可操作权重
Signed-off-by: 吴嘚嘚 <13742125+wu-deidei@user.noreply.gitee.com>
2024-03-05 07:18:15 +00:00
严浪 ceeef4bd9f 文字修改 2024-03-04 15:07:22 +08:00
严浪 f1a5316b91 功能优化 2024-03-04 14:37:20 +08:00
严浪 60e12c34f3 多GLIGEN文本框应用 2024-02-26 15:09:23 +08:00
严浪 714b4d3285 BUG修改 2024-01-26 22:19:20 +08:00
严浪 f684744205 优化 2024-01-19 15:24:40 +08:00
严浪 a5e39a5145 插件优化 2024-01-19 13:32:06 +08:00
严浪 495b315eeb 插件升级,添加百度翻译和本地翻译 2024-01-18 22:25:16 +08:00
严浪 2d2536b9bc 风格选择示例效果展示 2024-01-15 09:28:45 +08:00
严浪 69d83f27be s 2024-01-02 15:24:47 +08:00
严浪 56918fbe42 去除打印代码 2024-01-02 15:23:18 +08:00
严浪 7310adcc94 说明修改 2023-12-23 16:18:06 +08:00
严浪 81b72434b6 中文搜索实现 2023-12-23 15:41:45 +08:00
严浪 28507d4444 功能优化 2023-12-16 23:06:40 +08:00
严浪 bb97be8746 功能优化 2023-12-15 20:33:41 +08:00
严浪 4797ed2435 组节点保存功能开发 2023-12-15 15:45:29 +08:00
严浪 b6bdba5fb2 功能优化 2023-12-08 22:05:12 +08:00
严浪 482628c701 BUG修改 2023-12-07 22:11:25 +08:00
严浪 89f359c5c8 修改部分风格没有prompt报错问题 2023-12-07 22:06:01 +08:00
严浪 fe9e459ba2 安装文件修改 2023-12-01 20:21:39 +08:00
严浪 3b59044c29 优化插件 2023-11-30 21:54:07 +08:00
严浪 d2d4c6aa56 样式修改 2023-11-29 13:15:04 +08:00
严浪 8a6a7872b5 功能优化修改 2023-11-29 12:15:39 +08:00
严浪 8ceae5ccf8 删除print 2023-11-28 17:31:28 +08:00
严浪 3d7e2707f2 新增风格选择工具,删除多余文件 2023-11-28 17:29:27 +08:00
严浪 dbfe688a89 删除多余文件 2023-11-26 16:16:50 +08:00
严浪 71822ca854 文件目录机构调整,添加提示词选择插件功能 2023-11-25 22:38:04 +08:00
yanlang0123 4c66b01a27 删除文件 py/__pycache__ 2023-11-23 00:47:15 +00:00
严浪 4f8904f665 修改说明 2023-11-22 10:35:44 +08:00
严浪 43e4cbf41c 为适应cu121新版本安装失败问题修改 2023-11-22 10:07:09 +08:00
yanlang0123 e050de7ed8 !1 合并修改
Merge pull request !1 from yanlang0123/master
2023-11-22 01:56:14 +00:00
575 changed files with 84353 additions and 33021 deletions
+25
View File
@@ -0,0 +1,25 @@
name: Publish to Comfy registry
on:
workflow_dispatch:
push:
branches:
- main
- master
paths:
- "pyproject.toml"
permissions:
issues: write
jobs:
publish-node:
name: Publish Custom Node to registry
runs-on: ubuntu-latest
if: ${{ github.repository_owner == 'yanlang0123' }}
steps:
- name: Check out code
uses: actions/checkout@v4
- name: Publish Custom Node
uses: Comfy-Org/publish-node-action@v1
with:
personal_access_token: ${{ secrets.REGISTRY_ACCESS_TOKEN }}
+24
View File
@@ -0,0 +1,24 @@
# Mac
.DS_Store
**/.DS_Store
# vim/vi
*.swp
# JavaScript
node_modules/
.node_modules/
.eslintcache
unpackage/dist/build/
unpackage/dist/dev/
# python
*.pyc
config/baidu.json
ckpts/*
config/lamWeChat.db
config/weChat.yaml
config/settings_nodes.json
models/*
backup/*
config/workflow/AiShoot.json
+13 -159
View File
@@ -8,166 +8,20 @@ Download and place in the plugin directory of comfyUI, as shown below:
![Alt text](解压存放路径及名称.png)
##### Special Note for this version, execute the latest version "ComfyUI_windows_portable_nvidia_cu118_or_cpu.7z" with cu118 inside, the address is as follows:
https://github.com/comfyanonymous/ComfyUI/releases/download/latest/ComfyUI_windows_portable_nvidia_cu118_or_cpu.7z
##### Special Note: For this version, execute the latest version "ComfyUI_windows_portable_nvidia_cu121_or_cpu.7z" with cu121 inside. The version address is as follows:
https://github.com/comfyanonymous/ComfyUI/releases/download/latest/ComfyUI_windows_portable_nvidia_cu121_or_cpu.7z
1. Unzip according to the image to the specified plugin directory and name. Extract insightface.rar to the directory ..\ComfyUI_windows_portable\python_embeded\Lib\site-packages. Extract venv.rar to the directory ..\ComfyUI_windows_portable\python_embeded\Lib\site-packages. Then run the install.bat file, and if there are no errors, it's ready.
Then run the install.bat file, and if there are no errors, it's ready.
2. Model addresses and storage paths:
Lama model:
https://huggingface.co/lllyasviel/Annotators/resolve/main/ControlNetLama.pth ..\ComfyUI\models\lama\ControlNetLama.pth
SadTalker models:
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/mapping_00109-model.pth.tar ..\ComfyUI\models\SadTalker\mapping_00109-model.pth.tar
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/mapping_00229-model.pth.tar ..\ComfyUI\models\SadTalker\mapping_00229-model.pth.tar
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/SadTalker_V0.0.2_256.safetensors ..\ComfyUI\models\SadTalker\SadTalker_V0.0.2_256.safetensors
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/SadTalker_V0.0.2_512.safetensors ..\ComfyUI\models\SadTalker\SadTalker_V0.0.2_512.safetensors
1. Model addresses and storage paths:
- Lama model:
https://huggingface.co/lllyasviel/Annotators/resolve/main/ControlNetLama.pth ..\ComfyUI\models\lama\ControlNetLama.pth
3. Gfpgan is based on the startup directory, follow the prompts accordingly.
https://github.com/xinntao/facexlib/releases/download/v0.1.0/alignment_WFLW_4HG.pth ..\ComfyUI_windows_portable\gfpgan\weights\alignment_WFLW_4HG.pth
https://github.com/xinntao/facexlib/releases/download/v0.1.0/detection_Resnet50_Final.pth ..\ComfyUI_windows_portable\gfpgan\weights\detection_Resnet50_Final.pth
https://github.com/TencentARC/GFPGAN/releases/download/v1.3.0/GFPGANv1.4.pth ..\ComfyUI_windows_portable\gfpgan\weights\GFPGANv1.4.pth
https://github.com/xinntao/facexlib/releases/download/v0.2.2/parsing_parsenet.pth ..\ComfyUI_windows_portable\gfpgan\weights\parsing_parsenet.pth
2. Image-face-fusion model (Face swapping model)
- Link: https://pan.baidu.com/s/19DOgJQ_RHNAjfNrzSr2uTQ?pwd=gf0p
- Extraction Code: gf0p
- Extract to the directory ..\ComfyUI\models\image-face-fusion
4. Image-face-fusion model (Face swapping model)
Link: https://pan.baidu.com/s/19DOgJQ_RHNAjfNrzSr2uTQ?pwd=gf0p
Extraction Code: gf0p
Extract to the directory ..\ComfyUI\models\image-face-fusion
5. Roop-face-swap model (Face swapping model)
Link: https://pan.baidu.com/s/1cJeRtqgdeNW21Hljwuv-tA?pwd=b4yy
Extraction Code: b4yy
Extract to the directory ..\ComfyUI\models\roop-face-swap
6. Modify the execution.py file in the comfyUI root directory.
搜索 “def recursive_execute”
位置参考图如下:
![Alt text](修改位置.png)
新增内容:
```python
#循环添加代码----------开始----------
def get_del_keys(key, prompt,uniqueIds):
keys=[]
for k,v in prompt.items():
if k in uniqueIds :
continue
for k1,v1 in v['inputs'].items():
if type(v1)==list and v1[0]==key:
keys.append(k)
keys=keys+get_del_keys(k,prompt,uniqueIds)
return keys
#循环添加代码-----------结束---------
```
```python
#循环添加代码----------开始----------
startNum=None
startData={}
delKeys=[]
backhaul={}
clTypes=['ForInnerEnd','IfInnerExecute','DoWhileEnd']
oldPrompt=None
if class_type in clTypes:
oldPrompt=copy.deepcopy(prompt)
if class_type=='ForInnerEnd':
startNum=prompt[unique_id]['inputs']['total'][0]
inputNum=prompt[unique_id]['inputs']['obj'][0]
maxKeyStr=sorted(list(prompt.keys()), key=lambda x: int(x.split(':')[0]))[-1]
maxKey = int(maxKeyStr.split(':')[0])
delKeys=list(set(get_del_keys(startNum,prompt,[startNum,unique_id])))
delKeys.append(startNum)
delKeys = list(filter(lambda x: x != inputNum, delKeys))
for key in delKeys:
outputs.pop(key, None)
startInput=prompt[startNum]['inputs']
if isinstance(startInput['total'],list) or isinstance(startInput['stop'],list) or isinstance(startInput['i'],list):
result = recursive_execute(server, prompt, outputs, startNum, extra_data, executed, prompt_id, outputs_ui, object_storage)
if result[0] is not True:
return result
for x in startInput:
input_data = startInput[x]
if isinstance(input_data, list):
startData[x]=outputs[input_data[0]][input_data[1]][0]
else:
startData[x]=input_data
for i in range(startData['i']+1,startData['total'],startData['stop']):
prompt[str(maxKey+i)]=prompt[inputNum]
prompt[unique_id]['inputs']['obj'+str(i)]=[str(maxKey+i),prompt[unique_id]['inputs']['obj'][-1]]
if i==startInput['i']+1:
backhaul['obj'+str(i)]=prompt[unique_id]['inputs']['obj']
else:
backhaul['obj'+str(i)]=[str(maxKey+i-startInput['stop']),prompt[unique_id]['inputs']['obj'][-1]]
elif class_type=='DoWhileEnd':
startNum=prompt[unique_id]['inputs']['start'][0]
inputNum=prompt[unique_id]['inputs']['ANY'][0]
delKeys=list(set(get_del_keys(startNum,prompt,[startNum,unique_id])))
delKeys.append(startNum)
delKeys.append(inputNum)
#循环添加代码-----------结束---------
```
```python
#循环添加代码----------开始----------
if class_type=='ForInnerEnd' and x !='obj' and x.startswith('obj'):
if startNum!=None:
if isinstance(prompt[startNum]['inputs']['i'],list):
prompt[startNum]['inputs']['i']=startData['i']+startData['stop']
else:
prompt[startNum]['inputs']['i']=prompt[startNum]['inputs']['i']+startData['stop']
prompt[startNum]['inputs']['obj']=backhaul[x]
for key in delKeys:
outputs.pop(key, None)
#循环添加代码-----------结束---------
```
```python
#判断选择添加代码-----------开始---------
if class_type=='DoWhileEnd' and x == 'ANY':
any=outputs[inputNum][prompt[unique_id]['inputs']['ANY'][1]][0]
i=0
while any:
for key in delKeys:
outputs.pop(key, None)
i=i+1
outputs['i']=[[i]]
prompt[startNum]['inputs']['i']=['i',0]
result = recursive_execute(server, prompt, outputs, input_unique_id, extra_data, executed, prompt_id, outputs_ui, object_storage)
if result[0] is not True:
return result
any=outputs[inputNum][prompt[unique_id]['inputs']['ANY'][1]][0]
if class_type=='IfInnerExecute' and x=='ANY':
if outputs[input_unique_id][output_index][0]:
inputs['IF_FALSE']=oldPrompt[unique_id]['inputs']['IF_TRUE']
else:
inputs['IF_TRUE']=oldPrompt[unique_id]['inputs']['IF_FALSE']
#判断选择添加代码-----------结束---------
```
```python
#循环添加代码----------开始----------
if class_type=='ForInnerEnd' or class_type=='DoWhileEnd':
prompt[startNum]['inputs']=oldPrompt[startNum]['inputs']
prompt[unique_id]['inputs']=oldPrompt[unique_id]['inputs']
outputs.pop(startNum, None)
result = recursive_execute(server, prompt, outputs, startNum, extra_data, executed, prompt_id, outputs_ui, object_storage)
if result[0] is not True:
return result
if class_type=='IfInnerExecute':
prompt[unique_id]['inputs']=oldPrompt[unique_id]['inputs']
#循环添加代码-----------结束---------
```
搜索 “def validate_prompt
```python
'''优化输出开始'''
inputKeys=[]
for k,v in prompt.items():
for k1,v1 in v['inputs'].items():
if type(v1)==list and len(v1)==2:
inputKeys.append(v1[0])
for x in prompt:
class_ = nodes.NODE_CLASS_MAPPINGS[prompt[x]['class_type']]
if hasattr(class_, 'OUTPUT_NODE') and class_.OUTPUT_NODE == True and x not in inputKeys:
outputs.add(x)
'''优化输出结束'''
```
3. 执行修改脚本
windows 执行“修改文件.bat”文件,未报错后就可以了
linux 执行“修改文件.sh”文件,未报错后就可以了
+7 -156
View File
@@ -8,170 +8,21 @@
![Alt text](解压存放路径及名称.png)
##### 特别说明该版本执行最新版“ComfyUI_windows_portable_nvidia_cu118_or_cpu.7z” 里面带cu118的版本地址如下:
https://github.com/comfyanonymous/ComfyUI/releases/download/latest/ComfyUI_windows_portable_nvidia_cu118_or_cpu.7z
##### 特别说明该版本执行最新版“ComfyUI_windows_portable_nvidia_cu121_or_cpu.7z” 里面带cu121的版本地址如下:
https://github.com/comfyanonymous/ComfyUI/releases/download/latest/ComfyUI_windows_portable_nvidia_cu121_or_cpu.7z
1. 根据图片解压到指定插件目录及名称,
将insightface.rar解压到..\ComfyUI_windows_portable\python_embeded\Lib\site-packages目录
将venv.rar解压到..\ComfyUI_windows_portable\python_embeded\Lib\site-packages目录
然后运行install.bat文件,未报错后就可以了,
2. 模型地址,及存放路径:
1. 模型地址,及存放路径:
lama模型:
https://huggingface.co/lllyasviel/Annotators/resolve/main/ControlNetLama.pth ..\ComfyUI\models\lama\ControlNetLama.pth
SadTalker模型:
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/mapping_00109-model.pth.tar ..\ComfyUI\models\SadTalker\mapping_00109-model.pth.tar
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/mapping_00229-model.pth.tar ..\ComfyUI\models\SadTalker\mapping_00229-model.pth.tar
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/SadTalker_V0.0.2_256.safetensors ..\ComfyUI\models\SadTalker\SadTalker_V0.0.2_256.safetensors
https://github.com/OpenTalker/SadTalker/releases/download/v0.0.2-rc/SadTalker_V0.0.2_512.safetensors ..\ComfyUI\models\SadTalker\SadTalker_V0.0.2_512.safetensors
3. gfpgan是根据启动目录来的,根据提示对应
https://github.com/xinntao/facexlib/releases/download/v0.1.0/alignment_WFLW_4HG.pth ..\ComfyUI_windows_portable\gfpgan\weights\alignment_WFLW_4HG.pth
https://github.com/xinntao/facexlib/releases/download/v0.1.0/detection_Resnet50_Final.pth ..\ComfyUI_windows_portable\gfpgan\weights\detection_Resnet50_Final.pth
https://github.com/TencentARC/GFPGAN/releases/download/v1.3.0/GFPGANv1.4.pth ..\ComfyUI_windows_portable\gfpgan\weights\GFPGANv1.4.pth
https://github.com/xinntao/facexlib/releases/download/v0.2.2/parsing_parsenet.pth ..\ComfyUI_windows_portable\gfpgan\weights\parsing_parsenet.pth
4. image-face-fusion模型换脸模型
2. image-face-fusion模型换脸模型
链接:https://pan.baidu.com/s/19DOgJQ_RHNAjfNrzSr2uTQ?pwd=gf0p
提取码:gf0p
解压到 ..\ComfyUI\models\image-face-fusion 目录
5. roop-face-swap模型换脸模型
链接:https://pan.baidu.com/s/1cJeRtqgdeNW21Hljwuv-tA?pwd=b4yy
提取码:b4yy
解压到 ..\ComfyUI\models\roop-face-swap 目录
6. 修改comfyUI根目录下的execution.py文件,修改内容
搜索 “def recursive_execute”
位置参考图如下:
![Alt text](修改位置.png)
新增内容:
```python
#循环添加代码----------开始----------
def get_del_keys(key, prompt,uniqueIds):
keys=[]
for k,v in prompt.items():
if k in uniqueIds :
continue
for k1,v1 in v['inputs'].items():
if type(v1)==list and v1[0]==key:
keys.append(k)
keys=keys+get_del_keys(k,prompt,uniqueIds)
return keys
#循环添加代码-----------结束---------
```
```python
#循环添加代码----------开始----------
startNum=None
startData={}
delKeys=[]
backhaul={}
clTypes=['ForInnerEnd','IfInnerExecute','DoWhileEnd']
oldPrompt=None
if class_type in clTypes:
oldPrompt=copy.deepcopy(prompt)
if class_type=='ForInnerEnd':
startNum=prompt[unique_id]['inputs']['total'][0]
inputNum=prompt[unique_id]['inputs']['obj'][0]
maxKeyStr=sorted(list(prompt.keys()), key=lambda x: int(x.split(':')[0]))[-1]
maxKey = int(maxKeyStr.split(':')[0])
delKeys=list(set(get_del_keys(startNum,prompt,[startNum,unique_id])))
delKeys.append(startNum)
delKeys = list(filter(lambda x: x != inputNum, delKeys))
for key in delKeys:
outputs.pop(key, None)
startInput=prompt[startNum]['inputs']
if isinstance(startInput['total'],list) or isinstance(startInput['stop'],list) or isinstance(startInput['i'],list):
result = recursive_execute(server, prompt, outputs, startNum, extra_data, executed, prompt_id, outputs_ui, object_storage)
if result[0] is not True:
return result
for x in startInput:
input_data = startInput[x]
if isinstance(input_data, list):
startData[x]=outputs[input_data[0]][input_data[1]][0]
else:
startData[x]=input_data
for i in range(startData['i']+1,startData['total'],startData['stop']):
prompt[str(maxKey+i)]=prompt[inputNum]
prompt[unique_id]['inputs']['obj'+str(i)]=[str(maxKey+i),prompt[unique_id]['inputs']['obj'][-1]]
if i==startInput['i']+1:
backhaul['obj'+str(i)]=prompt[unique_id]['inputs']['obj']
else:
backhaul['obj'+str(i)]=[str(maxKey+i-startInput['stop']),prompt[unique_id]['inputs']['obj'][-1]]
elif class_type=='DoWhileEnd':
startNum=prompt[unique_id]['inputs']['start'][0]
inputNum=prompt[unique_id]['inputs']['ANY'][0]
delKeys=list(set(get_del_keys(startNum,prompt,[startNum,unique_id])))
delKeys.append(startNum)
delKeys.append(inputNum)
#循环添加代码-----------结束---------
```
```python
#循环添加代码----------开始----------
if class_type=='ForInnerEnd' and x !='obj' and x.startswith('obj'):
if startNum!=None:
if isinstance(prompt[startNum]['inputs']['i'],list):
prompt[startNum]['inputs']['i']=startData['i']+startData['stop']
else:
prompt[startNum]['inputs']['i']=prompt[startNum]['inputs']['i']+startData['stop']
prompt[startNum]['inputs']['obj']=backhaul[x]
for key in delKeys:
outputs.pop(key, None)
#循环添加代码-----------结束---------
```
```python
#判断选择添加代码-----------开始---------
if class_type=='DoWhileEnd' and x == 'ANY':
any=outputs[inputNum][prompt[unique_id]['inputs']['ANY'][1]][0]
i=0
while any:
for key in delKeys:
outputs.pop(key, None)
i=i+1
outputs['i']=[[i]]
prompt[startNum]['inputs']['i']=['i',0]
result = recursive_execute(server, prompt, outputs, input_unique_id, extra_data, executed, prompt_id, outputs_ui, object_storage)
if result[0] is not True:
return result
any=outputs[inputNum][prompt[unique_id]['inputs']['ANY'][1]][0]
if class_type=='IfInnerExecute' and x=='ANY':
if outputs[input_unique_id][output_index][0]:
inputs['IF_FALSE']=oldPrompt[unique_id]['inputs']['IF_TRUE']
else:
inputs['IF_TRUE']=oldPrompt[unique_id]['inputs']['IF_FALSE']
#判断选择添加代码-----------结束---------
```
```python
#循环添加代码----------开始----------
if class_type=='ForInnerEnd' or class_type=='DoWhileEnd':
prompt[startNum]['inputs']=oldPrompt[startNum]['inputs']
prompt[unique_id]['inputs']=oldPrompt[unique_id]['inputs']
outputs.pop(startNum, None)
result = recursive_execute(server, prompt, outputs, startNum, extra_data, executed, prompt_id, outputs_ui, object_storage)
if result[0] is not True:
return result
if class_type=='IfInnerExecute':
prompt[unique_id]['inputs']=oldPrompt[unique_id]['inputs']
#循环添加代码-----------结束---------
```
搜索 “def validate_prompt
```python
'''优化输出开始'''
inputKeys=[]
for k,v in prompt.items():
for k1,v1 in v['inputs'].items():
if type(v1)==list and len(v1)==2:
inputKeys.append(v1[0])
for x in prompt:
class_ = nodes.NODE_CLASS_MAPPINGS[prompt[x]['class_type']]
if hasattr(class_, 'OUTPUT_NODE') and class_.OUTPUT_NODE == True and x not in inputKeys:
outputs.add(x)
'''优化输出结束'''
```
3. 执行修改脚本
windows 执行“修改文件.bat”文件,未报错后就可以了
linux 执行“修改文件.sh”文件,未报错后就可以了
+16 -4
View File
@@ -3,23 +3,35 @@ import glob
import os
import sys
from .lam import init, get_ext_dir
import time
from server import PromptServer
repo_dir = os.path.dirname(os.path.realpath(__file__))
sys.path.insert(0, repo_dir)
original_modules = sys.modules.copy()
NODE_CLASS_MAPPINGS = {}
NODE_DISPLAY_NAME_MAPPINGS = {}
if init():
print("Loading ComfyUI-Lam")
py = get_ext_dir("py")
files = os.listdir(py)
for file in files:
if not file.endswith(".py"):
continue
name = os.path.splitext(file)[0]
imported_module = importlib.import_module(".py.{}".format(name), __name__)
NODE_CLASS_MAPPINGS = {**NODE_CLASS_MAPPINGS, **imported_module.NODE_CLASS_MAPPINGS}
NODE_DISPLAY_NAME_MAPPINGS = {**NODE_DISPLAY_NAME_MAPPINGS, **imported_module.NODE_DISPLAY_NAME_MAPPINGS}
try:
imported_module = importlib.import_module(".py.{}".format(name), __name__)
NODE_CLASS_MAPPINGS = {**NODE_CLASS_MAPPINGS, **imported_module.NODE_CLASS_MAPPINGS}
NODE_DISPLAY_NAME_MAPPINGS = {**NODE_DISPLAY_NAME_MAPPINGS, **imported_module.NODE_DISPLAY_NAME_MAPPINGS}
except Exception as e:
print("节点:'"+name+"'导入异常",e)
WEB_DIRECTORY = "./js"
WEB_DIRECTORY = "./js"
file_directory = os.path.dirname(os.path.abspath(__file__))
PromptServer.instance.app.router.add_static("/wechatauth/static", file_directory+"/pages/static")
PromptServer.instance.app.router.add_static("/paint-board", file_directory+"/pages/paint-board")
__all__ = ["NODE_CLASS_MAPPINGS", "NODE_DISPLAY_NAME_MAPPINGS","WEB_DIRECTORY"]
+80
View File
@@ -0,0 +1,80 @@
import os,sys
#读取py文件为text字符串
editlist=[{
"name": "server.py",
"isEditStr":"elif 'prompt_id' in json_data:",
"changs":[
{
"primit":'''else:
error = {''',
"edit":'''elif 'prompt_id' in json_data:
return web.json_response(json_data)
else:
error = {'''
}
]
},{
"name": "comfy/cli_args.py",
"isEditStr":"parser.add_argument(\"--cluster\", action=\"store_true\"",
"changs":[
{
"primit":'''parser.add_argument("--multi-user", action="store_true", help="Enables per-user storage.")''',
"edit":'''parser.add_argument("--multi-user", action="store_true", help="Enables per-user storage.")
parser.add_argument("--cluster", action="store_true", help="是否集群")
parser.add_argument("--isSection", action="store_true", help="分段负载")
parser.add_argument("--isMain", action="store_true", help="是否为主")
parser.add_argument("--basePath", type=str, default="127.0.0.1", help="服务地址")'''
}
]
}
]
filePath=os.path.dirname(os.path.abspath(__file__))
print('当前目录:',filePath)
base_path=filePath.split('custom_nodes')[0]
backupPath=os.path.join(filePath,'backup')
def read_py_file(file_path):
with open(file_path, 'r',encoding='utf-8') as file:
text = file.read()
return text
def write_py_file(file_path, text):
filePath=os.path.dirname(file_path)
if not os.path.exists(filePath):
os.makedirs(filePath)
with open(file_path, 'w',encoding='utf-8') as file:
file.write(text)
if __name__ == '__main__':
isRestore='--res' in sys.argv
print('-------开始还原-------' if isRestore else '--------开始修改--------')
for i in editlist:
file_path=os.path.join(base_path,i['name'])
if not os.path.exists(file_path):
print(file_path+'文件不存在')
continue
if not os.path.exists(backupPath):
os.makedirs(backupPath)
if isRestore:
if not os.path.exists(os.path.join(backupPath,i['name'])):
print('备份文件'+os.path.join(backupPath,i['name'])+'不存在无法还原')
continue
text=read_py_file(os.path.join(backupPath,i['name']))
write_py_file(file_path, text)
else:
text=read_py_file(file_path)
if text.find(i['isEditStr'])!=-1:
continue
write_py_file(os.path.join(backupPath,i['name']), text)
textOld=text+''
for j in i['changs']:
if text.find(j['primit'])==-1:
print(i['name']+'未找到修改位置,修改失败')
text=textOld
break
text=text.replace(j['primit'],j['edit'])
write_py_file(file_path, text)
print('-------还原完成-------' if isRestore else '--------修改完成--------')
Binary file not shown.
-1
View File
@@ -1 +0,0 @@
{}
@@ -0,0 +1,615 @@
---
name: 3d-animation-short-generator
description: |
将一句故事创意转化为完整的风格化3D动画短片制作流程,依次完成项目简报、故事大纲、角色卡、无人物场景卡、标准镜头表、文本分镜或可选铅笔分镜、视频模型选择、单镜头生成、全片拼接、BGM匹配与最终复查。适用于需要角色一致、场景连续、节奏明确和音画配合的剧情动画短片、纪念短片、品牌故事与社交媒体内容。不适用于单张图片、简单修图、真人写实视频或单一独立镜头。
---
# 3D动画短片生成器
当用户想把一句话创意做成完整的故事优先动画短片流程时使用本 Skill。流程从创意开发到最终剪辑,所有主要产物都必须按生产顺序放到画布上,并在高成本或关键创意节点用选项卡让用户确认。
核心规则:**先故事、再资产;用户输入内容后立刻用选项卡确认画面尺寸/比例与总时长;产物按顺序上画布;所有需要用户确认的地方必须用选项卡;角色/场景卡后固定顺序为:六列标准镜头信息表(每秒指令+音频+空间锚点链+hook)→ 镜头表自检门 → 一份单一文本分镜文档(每镜一节,对齐半解说短剧分镜字段;用户主动开启可视化模式时再加铅笔图;用户标记需要重点迭代的镜时抽取成独立节点)→ 视频模型选项卡(H3 默认,Seedance 2.0 回退)+ 分辨率选项卡 → 用选定视频模型生成单镜头片段 → 全片拼接与 BGM 合成;最终视频必须清除所有分镜痕迹**。
## 全局视觉风格锁
除非用户明确指定其他视觉风格,所有角色卡、场景卡、镜头表格、文本分镜、可选铅笔分镜、单镜头视频片段(无论使用哪个视频模型)、拼接正片和最终合成都必须使用以下视觉风格:
- 渲染风格:皮克斯感 3D 卡通渲染,C4D + Octane 渲染器质感,高级动画电影质量。
- 角色造型:夸张几何概括与极佳材质质感的平衡。不追求 100% 真人解剖结构,而是使用高概括度几何形体、强剪影和高辨识度设计。
- 比例语言:亲和的风格化比例,适合可爱或幼态角色时采用 2.5–3 头身 Q 版比例,大头小身、体块清晰、轮廓易读。
- 发丝 / 毛发:整体造型有清晰块面结构,同时边缘保留精细发丝或绒毛细节;在光照下既硬朗有形,又有细腻触感。
- 皮肤 / 材质:使用温润次表面散射(SSS)质感,耳朵、脸颊、鼻尖、指尖可有微红透光效果;避免生硬塑料感。
- 表演风格:夸张灵动的迪士尼 / 皮克斯式角色动画表演,运用挤压与伸展,强化眉毛、眼角、瞳孔、嘴唇、脸颊形变。
- 动态风格:高能量态势、清晰动作线、前倾势能、强预备动作、快速但可读的节奏、弹性身体机制和丰富微表情。
- 情绪尺度:在可爱与爆发力之间取得平衡;强情绪可以有戏剧性面部变形,但必须保留角色魅力。
负向风格约束:不要写实真人摄影、不要扁平二次元、不要塑料玩具皮肤、不要僵硬人体姿势、不要真实解剖僵硬感、不要无生命表情。
## STEP 0:接收需求与画布规划
捕捉:
- 一句话创意或粗略故事
- 目标产物:只要蓝图、只要资产、六列标准镜头信息表、单文本分镜文档(默认)+ 重点迭代镜抽取的独立节点 + 多宫格铅笔分镜(开启可视化模式时)、单镜头视频片段(视频模型在 Step 7 选定)、拼接正片,还是 BGM 合成最终成片
- 用户已说明的预估时长
- 用户已说明的画面尺寸 / 比例
- 视觉基调:默认温暖风格化 3D 动画
- 台词需求:有对白、旁白,还是无台词
- 只有用户明确指定台词语言时才锁定语言;不要默认英文台词
捕捉完用户输入后,在项目简报或任何下一步之前,必须先展示选项卡确认生产规格:
画面尺寸 / 比例选项卡:
- 16:9 横屏(电影短片推荐)
- 9:16 竖屏短视频
- 1:1 方形
- 4:5 社媒竖图
- 自定义尺寸 / 比例
总时长选项卡:
- 30–60 秒(推荐)
- 15–30 秒
- 60–90 秒
- 90–180 秒
- 自定义时长
只有用户同时选择画面尺寸/比例和总时长,或明确给出自定义值后,才进入下一步。把已确认的画面尺寸/比例和总时长写入项目简报,并贯穿用于镜头时长、转场连贯性、六列标准镜头信息表、单镜头分镜(默认文本,开可视化模式时含铅笔图)、单镜头视频片段、全片拼接、BGM 匹配和最终合成设置。
按以下顺序创建或更新画布产物:
1. 项目简报文本节点
2. 故事大纲文本节点
3. 带标注的角色卡图片节点
4. 无人物的场景卡图片节点
5. 六列标准镜头信息表节点,每行必须在 `镜头描述` 内含每秒指令
6. **单文本分镜文档** — 一个画布文本节点,命名 `<title> text storyboards`,每镜一节(对齐半解说短剧分镜结构)。用户标记需要重点迭代的镜时,把该节抽取成独立节点,原位置留占位。可视化模式下的铅笔分镜图是独立图片节点。
7. 单镜头视频片段节点(用 Step 7 选定的视频模型渲染 — H3 默认,Seedance 2.0 回退)
8. 拼接主视频节点
9. 匹配的 BGM 音频节点与最终 BGM 合成视频节点
不要把长篇生产内容只丢在对话里。耐用产物必须以文本、图片、视频或音频节点形式落到画布。
## STEP 1:项目简报
写一份精炼的项目简报到画布文本节点,命名为项目名或 `项目简报`。
包含:
- 工作标题
- 一句话 What-if
- 情绪前提
- 目标观众感受
- 主要交付物计划
- 已确认的画面尺寸 / 比例
- 已确认的总时长
- 台词模式与语言:仅当用户明确要求才锁定具体语言;否则写 `language not specified`,保持最小化对白或后续再确认
- 初始风险
- 台词意图(如有)
然后展示用户选项卡:
- 沿用此方向继续(推荐)
- 重新生成前提选项
- 修订情绪前提
- 调整台词方向
只有用户选择或明确说继续后才进入下一步。
## STEP 2:故事大纲与守门
创建故事大纲并写入画布文本节点,命名为 `故事大纲` 或 `story-outline`。
包含:
- 主角 Want / Need / 缺陷
- 核心世界规则
- 8 拍因果故事骨架
- 情绪锚点与回收
- 台词节拍(用户要求时)
- 红线检查
守门检查:
- 主角是主动的
- 危机由主角缺陷放大
- 巧合不能解决问题
- 结局回收早前情绪锚点
- 反派压力不是扁平反派
- 台词揭示关系变化而非解释主题
然后展示用户选项卡:
- 批准故事并继续(推荐)
- 修订节拍
- 修订情绪曲线
- 修订台词节拍
- 返回前提
## STEP 3:角色卡
生成角色参考卡并把每张图放到画布。推荐顺序:
1. 主角卡
2. 对比 / 施压角色卡
3. 可选配角卡
每张角色卡在可能时为 16:9 生产参考图。与最终渲染视频不同,角色卡必须包含清晰可读的标注,让下游生成能正确绑定人物与道具:
- 角色名标注(英文和/或项目语言)
- 角色定位标注:主角、奶奶、小偷、搭档、施压角色等
- 主 3/4 视角
- 正 / 侧 / 背三视图
- 表情
- 材质 / 服装 / 道具细节
- 重要道具标注:手提包、钱包、滑板、苹果筐、围巾、鞋子、眼镜等
- 提示中重复的"身份锁"
- 简短视觉 ID 备注:年龄段、身材、发型、服饰色、签名道具、不可改特征
风格化 3D 动画要保持角色柔和、可读、跨图跨视频一致。
主要角色卡生成后展示用户选项卡:
- 锁定角色设计并继续(推荐)
- 重新生成主角卡
- 调整具体视觉细节
- 增加另一张角色卡
警告用户:后续修改已锁定的角色设计可能需要重做镜头表、单镜头分镜、单镜头视频片段、拼接正片和最终合成。
## STEP 4:场景卡
在角色卡之后生成场景参考卡并放到画布。场景卡只能展示环境,不出现人物、人群、剪影、手、脸或角色客串。角色动作属于镜头表、单镜头分镜与单镜头视频片段,不属于场景卡。
包含:
- 主环境总览
- 关键光态:日景 / 夜景
- 情绪子空间
- 连续性地标(同场景跨镜头必须保持屏幕位置的固定物体,如厨房中岛、沙发、门框、树、邮筒)
- 环境中的重要道具
然后展示用户选项卡:
- 锁定场景设计并继续(推荐)
- 重新生成场景卡
- 增加另一个场景角度
- 调整光位或布局
## STEP 5:六列标准镜头信息表
在角色卡和场景卡锁定后,输出标准化视频提示作为镜头信息表。本步为强制步骤,不可与分镜或视频生成互换。在画布创建表格节点或 markdown 表格,命名为 `标准镜头信息表` 或 `standard-shot-table`。
表格必须严格按以下顺序包含六列:
| 镜头编号 & 时长 | 连续性衔接 | 参考锚点(空间+身份) | Hook 类型 | 镜头描述(每秒指令) | 音频与对白轨 |
|---|---|---|---|---|---|
列规则:
- **镜头编号 & 时长**:镜头号加计划时长,例如 `S03 / 6s`。
- **连续性衔接**:本镜如何自然承接上一镜的结束画面、道具位置、眼神、人物姿态、声桥或情绪状态;以及本镜如何为下一镜开篇铺好钩子。这是跨镜连续性主链。
- **参考锚点(空间+身份)**:四个子字段,全部必填。
- `固定地标` — 来自场景卡的确切命名地标,及其相对画面的位置(如 `door-frame: 右侧 1/3`、`kitchen-island: 底部居中`)。
- `人物位置(机位视角)` — 镜头中每个角色相对画面的位置(左/中/右、上/中/下、前/中/后景)、朝向、初始姿态。
- `退场人物状态` — 上一镜有但本镜没有的角色,记录其离屏位置和原因(如 `Mia — 离屏左侧,最后出现于门框旁手持苹果筐`)。
- `光位基线` — 从场景卡继承的主光/补光/轮廓光方向,加本镜调整项(如 `主光:温暖顶光,补光:右侧冷反光,调整:窗光背身剪影`)。
- 加上身份绑定:已批准的角色卡名(精确)和场景卡名(精确)。
- **Hook 类型**:从受控词表中挑一个短标签,如 `visual-joke`、`reversal`、`suspense`、`tender`、`chase`、`reveal`、`callback`、`expression-beat`。用于整片 hook 分布自检。
- **镜头描述**:景别、镜头运动、荷兰角设计、表演风格、音效、负向提示、**视频模型生成备注**(Step 7 选定的模型 — H3 或 Seedance 2.0 — 接收的 prompt 形态略不同;H3 强调包装关键词和文字/UI/动效清晰度,Seedance 2.0 强调电影感镜头和弹性表演),并强制包含 `Per-Second Directives` 子节。子节必须按秒拆解为 `0–1s`、`1–2s`、`2–3s` 等指令;亚秒关键节拍用 `2.0–2.5s` 风格标记。每条每秒指令必须覆盖全部 5 个必填要素:
1. 动作 / 姿态 / 表情(适用时含 squash-and-stretch、anticipation、overshoot、follow-through)
2. 镜头运动(推 / 拉 / 摇 / 倾 / 手持晃动 / 锁定 / 环绕)
3. 空间位置(人物在哪、持有什么、地标在画面哪里)
4. 音频线索(旁白 / 对白 / 音效 / 呼吸 / 沉默 — 或 `silent` 表示有意静默)
5. 与下一秒或下一镜的交接(这一秒锁定的状态如何递给下一秒)
- **音频与对白轨**:本镜完整音频脚本,按时间顺序,独立于每秒指令。字段:
- `旁白` — 旁白文本及时间范围(无旁白则省略)。
- `对白` — 台词、说话人、语气、时间范围。
- `音效` — 关键音效,按时间锚定。
- `表演备注` — 主角离屏旁白时标记 `narrator-mouth-closed: true`;描述旁白期间表情路径;描述每条对白的具体眼神和身体动作变化。
表格全局规则:
- 每镜必须通过 `连续性衔接` 列自然继承上一镜画面状态,并为下一镜做铺垫。
- 每镜必须含覆盖整个时长的每秒指令,包含动作、姿态、表情、镜头运动、空间位置、音频线索和连续性交接。
- 每秒指令必须具体到可直接生成分镜面板;避免"继续移动"这类缺少身体/镜头/道具细节的模糊描述。
- 表演必须夸张弹性,符合迪士尼式 squash-and-stretch、anticipation、overshoot、follow-through、overlap、弧线、快速姿态变化和清晰喜剧剪影。
- 景别必须在特写 / 大特写与其他必要景别间交替;避免重复景别。
- 荷兰角构图必须为追逐、失衡、惊吓或闹剧节拍设计。
- 台词语言仅在用户明确要求时使用;否则用最小化非语言特定反应或将台词标记为可选 / 待确认。
- 含旁白或对白的镜头,每一秒人物说话时都必须记录嘴部开合;默认离屏旁白为闭,对面对白为开。
- 已离屏角色在 `退场人物状态` 至少记录 1 镜,连续 2 镜明确不在场后再移除。
然后展示用户选项卡:
- 批准表格并进入自检(推荐)
- 调整镜头连续性
- 让动画更夸张
- 调整特写 / 大特写节奏
- 调整荷兰角设计
## STEP 5.5:镜头表自检门(强制)
进入铅笔分镜前,必须对已批准的镜头表跑硬性自检。任何一项不通过必须先修订表格再重跑,通过后才向用户申请进入分镜。
六项必检:
1. **Hook 密度**:每镜都有 `Hook 类型`;每连续 3 镜至少 1 镜使用 `reveal`/`reversal`/`callback`;开场镜和收尾镜各带强 hook(`visual-joke` / `reversal` / `reveal` / `suspense` / `tender`)。
2. **单镜时长**:任何镜头不超过 15 秒。节拍需要更长就拆镜。
3. **单镜角色数**:任何镜头不超过 3 个重要角色(定义为有画面动作或对白的角色)。
4. **空间锚点继承**:含 2 镜以上的同场景,下一镜的 `固定地标` 和 `光位基线` 必须与上一镜一致或含明确连续性备注(如 `door-frame 随镜头左环绕由右侧 1/3 移至中央`)。
5. **每秒指令覆盖**:从 `0s` 到镜尾的每秒都有 `Per-Second Directives`,每条含 5 个必填要素(动作/姿态/表情、镜头、空间、音频线索、交接)。允许 `2.0–2.5s` 亚秒节拍但不能留时间空隙。
6. **跨镜连续性**:逐行读 `连续性衔接` 列必须形成连续链 — 任何一镜都不能从一个与上一镜结束矛盾的状态起步。任何翻转眼神、人物位置、道具状态或光位的镜头必须明确标记翻转(如 `HARD CUT — 时间跳 2h`)。
六项全过后,在画布表格节点顶部加 `shot-table self-check: passed` 戳,并展示用户选项卡:
- 批准自检并绘制分镜(推荐)
- 查看自检详情
- 修订失败项
- 重新跑自检
任何一项不通过就不进入 Step 6。返回 Step 5,列出失败行,表格修好并重跑自检通过后,才能再次展示分镜批准选项卡。
## STEP 6:单文本分镜文档(默认)+ 多宫格铅笔分镜图(opt-in)
Step 5.5 自检通过后,先展示分镜模式选项卡,再生成任何分镜产物:
- **单文本分镜文档(默认,推荐)** — 整个短片一份画布文本节点,所有分镜作为文档内章节。对齐半解说短剧分镜结构:每镜字段(标题/hook/场景/角色/空间锚点/连续性/表演备注)加 Pixar 的每面板四象限内容 + 可选 ASCII 布局。承载完整质控载荷,成本几乎为零。Step 7 选定的视频模型直接读取作为该镜渲染参考。
- **单文本分镜文档 + 多宫格铅笔图(可视化模式,opt-in)** — 单文本分镜文档照常作为权威产物,并额外为每镜生成一张多宫格铅笔图供人眼审稿。成本更高,适合用户想在付视频成本前要一次视觉预演,或弹性表演 / 姿态剪影是高风险项想先可视化检查的情况。
把选中的分镜模式记入项目简报,并在 Step 7、Step 9 和重生成纪律里复用。
### 默认路径:单文本分镜文档
为整部短片生成一个画布文本节点,命名 `<title> text storyboards`(一个文档包含所有分镜)。即使同时生成铅笔图,这份文档仍是 Step 7 的权威渲染参考。结构对齐半解说短剧分镜 —— 每镜是同一文档的章节,用户读跨镜连续性不需要切换节点。
文档头部块(文档顶部):
- 项目标题、已确认的视频模型、已确认的分辨率、分镜模式、自检状态(如 `shot-table self-check: passed at <时间戳>`)。
- 一段简短的目录,列出每镜的 hook 和章节锚点(如 `S01`、`S02` …),便于用户跳转。
每镜章节结构(每个 `##` 标题对应一镜,按镜头顺序)。每个章节必须按以下顺序包含这些字段 —— 直接对应半解说短剧的分镜字段:
1. **镜头标题 & 时长** — 简短可读的本镜标题,加 `S<N> / <duration>s`(如 `S03 / 6s`)。
2. **Hook 类型** — 受控词:`setup` / `visual-joke` / `reversal` / `reveal` / `callback` / `suspense` / `tender` / `chase` / `expression-beat` / `climax`。用于整片 hook 分布自检。
3. **场景 & 角色** — 精确场景卡名、精确在屏角色名(绑定到角色卡)。
4. **空间锚点卡**(强制,四个子字段 —— 直接对齐半解说短剧):
- `固定地标` — 命名地标及其画面相对位置(如 `door-frame: 右侧 1/3`、`kitchen-island: 底部居中`)。
- `人物位置(机位视角)` — 每个在屏角色,画面相对位置、朝向、初始姿态。
- `退场人物状态` — 上一镜有但本镜没有的角色,记录其离屏位置和原因。
- `光位基线` — 从场景卡继承的主光/补光/轮廓光方向,加本镜调整项。
5. **连续性**(对齐半解说短剧的 handoff 字段):
- `承接 S(N-1)` — 一两句话交代上一镜结束状态。
- `交接 S(N+1)` — 一句话铺好下一镜开场。
6. **双绑定** — `[char:角色名-01] [char:角色名-02] ... [scene:场景名] [hook: visual-joke]` — 精确角色卡名、场景卡名、hook 类型。这些是分镜专用参考标记,视频模型渲染时剥离。
7. **每面板四象限内容**(按时间顺序;这就是 Pixar 的每秒指令,从表格行原样保留):
- `时间码` — 如 `0–1s`。
- `姿态 + 表情` — 具体身体姿态、剪影、关键持握、眼神、表情路径;弹性节拍明确写出 squash / stretch / anticipation / overshoot。这是每面板最大块内容,也是视频模型读取的视觉节拍。
- `镜头` — 景别、镜头运动(推/拉/摇/倾/手持晃动/锁定/环绕)、荷兰角备注。
- `音频 + 锚点` — 音频线索(`♪ 旁白: ...` / `对白: ...` / `SFX: ...` / `silent`)和空间锚点备注(`door-frame: 右侧 1/3` / `Mia: 中央中景面朝镜头`)。
- 表演备注(对齐半解说短剧):旁白秒标 `narrator-mouth-closed: true`;面对白秒标 `mouth-open: speaker` 并描述表情路径 / 眼神 / 身体动作变化。
8. **布局规则**(逐镜适用):
- 3 秒镜 → 3 面板(每秒 1 个)。
- 4 秒镜 → 4 面板。
- 5 秒镜 → 5 面板。
- 6 秒镜 → 6 面板。
- 7 秒及以上 → 每秒 1 面板;亚秒关键节拍按需加迷你面板如 `2.0–2.5s`,仅当该节拍是本镜核心 hook 时。
- 面板必须覆盖整镜时长,从首帧到尾帧无时间空隙。
9. **每面板绑定**:
- 绑定表格行中列出的精确角色卡,锁定外观、脸、发型、身材、服装、签名道具、角色身份。使用与表格相同的角色名。
- 绑定表格行中列出的精确场景卡,保留环境、道具、地标、运动路径、空间逻辑。
10. **可选 ASCII 布局块(强烈推荐,零成本)**:
- 每面板附一小段 ASCII 草图(或整镜一张组合草图),让用户秒扫空间布局而不必渲染图。例:
```
[0-1s] Mia (L, mid) door-frame (R)
——跪下,双手按苹果筐——
cam: low push-in, locked
audio: silent | anchor: basket center-bottom
[1-2s] ...
```
- ASCII 块仅供参考;视频模型读取上面结构化的 `每面板四象限内容`,不读 ASCII。
11. **分镜专用标记**:
- 关键节拍在面板时间码后加 `[BEAT]`。
- 某面板必须交接特定状态到下个面板或下个镜头时,加 `[HANDOFF → ...]` 短标签,如 `[HANDOFF → S04 opening]`。
每镜章节模板(任意镜通用,复制粘贴骨架):
```markdown
## S03 / 6s — 标题:奶奶把苹果筐递给 Mia
- **Hook 类型**:reveal
- **场景 & 角色**:scene:kitchen | char:Mia, char:Grandma
- **空间锚点卡**:
- 固定地标:door-frame(右侧 1/3)、kitchen-island(底部居中)
- 人物位置:Mia(左,中景,面朝镜头)| Grandma(右,前景,面朝 Mia)
- 退场人物状态:—
- 光位基线:温暖顶光主光 + 右侧冷反光
- **承接 S02**:奶奶弯下腰从中岛拿起苹果筐
- **交接 S04**:Mia 接住筐转身,门铃响起
- **双绑定**:[char:Mia] [char:Grandma] [scene:kitchen] [hook:reveal]
### 每面板四象限内容
#### 0–1s
- 姿态 + 表情:奶奶弯腰双手持筐;Mia 左侧站姿,眼神好奇
- 镜头:locked medium shot, eye-level
- 音频 + 锚点:silent | Mia: L midground | basket: center bottom
- 表演备注:[BEAT]
#### 1–2s
- 姿态 + 表情:奶奶手臂伸向 Mia,筐倾斜;Mia 双手前伸准备接
- 镜头:locked medium shot, eye-level
- 音频 + 锚点:♪ SFX: basket rustle | anchor: door-frame: 右侧 1/3
- 表演备注:[HANDOFF → S04 opening]
#### 2–3s
...
### ASCII 布局(可选)
[0-1s] Grandma (R, fg) door-frame (R, bg)
——lifts basket—— Mia (L, mid)
cam: locked | silent
[1-2s] ...
```
所有章节写完后,把文档放到画布,直接进入 Step 7。默认模式下不要调用任何图像生成模型。
### 按镜抽取(重点迭代模式)
默认的单文档形态针对阅读和跨镜连续性做了优化。当用户标记某镜需要重点迭代(通常是 climax / 追逐 / 闹剧节拍,每面板内容需要多轮修订),就把该章节从文档中抽取成独立文本节点,让迭代局部化:
- **用户信号**:Step 6 之后任何时间用户说"重点改 S05"、"S05 需要返工"、"抽取 S05"、或在分镜批准选项卡选某镜做重点迭代。
- **抽取机制**:
1. 新建一个画布文本节点,命名 `<title> S05 text storyboard (extracted)`。
2. 把文档中 `## S05` 章节的完整内容移到新节点。
3. 在文档中,把原 `## S05` 章节替换为一行占位:`> S05 — 已抽取到独立节点(见 \`<title> S05 text storyboard (extracted)\`)`。
4. Step 7 渲染 S05 时从抽取节点读取;其他镜仍从文档读取。
- **回填**:用户满意后,把独立节点折叠回文档(占位换成最新内容),独立节点归档。
- **多镜抽取**:每个被抽取的镜各自独立成节点,文档里用占位追踪。
抽取机制存在是因为独立节点按需使用才高效 —— 但只要迭代压力大就随时可用。
### Opt-in 路径:多宫格铅笔分镜图(可视化模式)
如果用户在分镜模式选项卡选了可视化模式,在单文本分镜文档之外**额外**为每行生成一张多宫格铅笔分镜图。单文本分镜文档仍为权威渲染参考;铅笔图仅供人眼审稿。
每张铅笔分镜图:
- **双绑定标签(右上角,强制写在图上)**:
- `[char:角色名-01] [char:角色名-02] ...` — 本行用到的精确角色卡名。
- `[scene:场景名]` — 精确场景卡名。
- `[shot: S03] [dur: 6s] [hook: visual-joke]` — 镜头号、时长、hook 类型。
- 这些标签是分镜专用参考标记,渲染视频时剥离。
- **绑定精确角色卡**,锁定人物外观、脸、发型、身材、服装、签名道具、角色身份。
- **绑定精确场景卡**,保持环境、道具、地标、运动路径、空间逻辑。
- 把行内每秒指令转换为一张分镜面板或一个明确标注的节拍面板:4 秒镜通常 4 面板;6 秒镜通常 6 面板;亚秒关键节拍按需加迷你面板。
- **面板物理排版(强制)**:
- 3 秒镜 → 1×3 条。
- 4 秒镜 → 2×2 网格。
- 5 秒镜 → 上排 3 + 下排 2。
- 6 秒镜 → 2×3 网格。
- 7 秒及以上 → 3 排,面板均衡。
- 每面板占据相同画布面积,不允许某面板独占。
- **每面板四象限内容(强制)**:
- 左上:时间码(如 `0–1s`)。
- 右上:姿态 + 表情草图(最大区域;实际视觉节拍)。
- 左下:镜头图标 + 运动箭头(推/拉/摇/环绕/锁定)加荷兰角小备注。
- 右下:音频线索(如 `♪ 旁白: "我就知道。"` / `SFX: 门吱呀` / `silent`)和锚点备注(如 `door-frame: 右侧 1/3`)。
- 面板按阅读顺序排布,多个不同镜头不合并到同一图。
- 每面板标记时间码如 `0–1s`、`1–2s`,并展示对应姿态、表情、动作、镜头运动、道具位置、音效线索和连续性交接。
- 输出纯黑白铅笔线稿:不上色、不带最终渲染光、不带精致 3D 渲染。
- 在分镜图上标镜头号,必要时每面板加镜头运动图标 / 标记。
- 必要时包含分镜专用标记:铅笔结构线、动作箭头、镜头路径图标、时间标记和小注。
- 草稿只作为视频渲染参考资产,不是终稿美术。
### 分镜批准(两种模式通用)
所有文本分镜章节(和可视化模式下的铅笔图)产出后,按镜头顺序放到画布,归组为:
- `<title> text storyboards`(默认模式,单文档),或
- `<title> text storyboards + multi-panel pencil storyboards`(可视化模式,文本文档和铅笔图分两类各自归组,因为铅笔图含双绑定标签、ASCII 标签和镜头号,文本文档没有这些)。
展示用户选项卡:
- 批准分镜并渲染镜头视频(推荐)
- 抽取 / 回填某镜(在文档和独立节点之间移动章节)
- 重画选定的铅笔分镜(仅可视化模式)
- 修复角色一致性
- 修复场景逻辑
- 修复镜头标记
- 修复音频/锚点标记
### 分镜生成失败回退(仅可视化模式)
如果铅笔分镜图无法按要求质量产出(布局塌陷、标签糊掉、面板合并、角色走样),在向用户求助前按以下升级路径处理:
1. **第一次重试**:用更紧的提示重画同一镜,明确写上四象限布局、`[char:…] [scene:…] [shot:…]` 标签和每面板内容规则。
2. **第二次重试**:去掉右下音频/锚点象限文字(保留为带小 `♪` 标记的空白格),减少文字负载;通常能修掉标签糊掉但不影响视觉节拍。
3. **第三次重试**:面板数减 1(例如 6 面板→5 面板,合并两秒动作最弱的秒),镜头图标简化为单箭头。
4. **同一镜连续 3 次失败后**:暂停并用选项卡问用户:
- 切换为色块分镜(灰色方块占位姿态,不画铅笔线),仅限失败这一镜。
- 放弃该镜的铅笔图,仅靠单文本分镜文档推进该行。
- 在 Step 5 把该镜拆为两段更短的镜头并重跑 Step 5.5。
- 手动提供参考图代替生成。
默认文本模式没有这套回退 —— 文本分镜失败只发生在模型无法产出连贯结构化文本时,遇到就直接回到 Step 5 修订该行表格。
## STEP 7:视频模型选项卡 + 单镜头视频片段
### 视频模型选项卡(任何片段渲染前的强制门)
任何片段渲染前必须先展示视频模型选项卡。选择写入项目简报,本项目所有片段沿用,除非用户后续修改。
视频模型选项卡:
- **H3(推荐默认)** — 强项是视觉包装、动效图形、文字/UI 清晰度、多模态上下文理解、性价比(2K 约同类旗舰 1/3 价格,768P 约 1/2)。原生双声道音视频,单片段最长 15s @ 2K。最适合:设计语言强的皮克斯 3D 动画短片、文字 / 字幕 / UI 元素、动效转场、对话驱动且音频就是交付物一部分的镜头。
- **Seedance 2.0(高强度动画表演的回退)** — 强项是电影感镜头、复杂运镜、皮克斯弹性表演、动作张力。适合追逐、闹剧节拍、以动画本身为卖点的 climax 镜头。
- **逐镜混合(高级)** — 在镜头表某行的 `镜头描述` 列加 `video_model: H3` 或 `video_model: Seedance2` 字段,本项目同时混用。未标记的行默认走 H3。
### 分辨率选项卡(视频模型锁定后)
视频模型确定后再展示分辨率选项卡:
- 768P(H3 首推,性价比高)
- 2K(H3 默认质量,成本更高,画质更清)
- 1080p(Seedance 2.0 推荐)
- 720p(Seedance 2.0 草稿,最低成本)
- 匹配项目 / 自定义分辨率
首条片段渲染前用户必须先确认分辨率。后续用户可针对单条 hero 镜头临时改分辨率。
### 单镜头片段渲染
每行批准后调用选定的视频模型生成对应独立视频片段。每条片段必须严格使用单文本分镜文档中匹配的章节(若该章节已抽取成独立节点则从独立节点取)、角色卡和场景卡。
各模型通用规则:
- 单文本分镜文档是权威的逐镜参考:叙事、构图、镜头运动、动作编排、每秒时间、镜头号都从这里取。已抽取的镜从独立节点取。若同时存在铅笔分镜图,仅用于人眼侧的姿态/剪影预检;不得让铅笔图覆盖文本分镜。
- 角色卡为权威身份源。
- 场景卡为权威环境源。
- **剥离所有分镜双绑定标签**(`[char:…]`、`[scene:…]`、`[shot:…]`、`[dur:…]`、`[hook:…]`)再渲染视频 — 这些是分镜专用参考标记,不能出现在最终片段里。铅笔分镜图还自带镜头号、镜头图标、箭头和备注,渲染时也要一并移除。
- 渲染片段只能包含干净的皮克斯感全彩 3D 动画内容。
- 不要分镜线稿、不要手绘草图质感、不要标签、不要字幕(除非要求)、不要水印。
- 保持已批准的画面尺寸 / 比例和已批准的分辨率。
### 模型相关 prompt 形态
单文本分镜文档(已抽取的镜从独立节点取)同时喂两个模型,但包在分镜外的 prompt 前缀不同:
- **H3 prompt 前缀**(默认):强调包装关键词、设计语言、动效清晰度、文字/UI 元素(如果相关)、双声道音频意图。H3 强在指令遵循,每秒指令可以几乎原文送入。追加:`Pixar-inspired 3D cartoon rendering, C4D + Octane look, stylized Q-version proportions, warm SSS skin, designed-with-detail hair, strong character design language, clean motion, on-brand color palette`。
- **Seedance 2.0 prompt 前缀**(表演回退):强调电影感镜头语言、弹性 squash-and-stretch、anticipation、follow-through、光位戏剧性、镜头选择。追加:`cinematic Pixar-quality 3D animation, elastic squash-and-stretch performance, Disney-style anticipation and overshoot, dramatic key lighting, lens-specific depth of field`。
用户选了逐镜混合时,按该行 `video_model` 字段匹配的前缀。
### 各模型失败回退
H3 失败回退阶梯:
1. 第一次重试:用直接引用表格中 `参考锚点` 块的强化 H3 prompt 重渲染。
2. 第二次重试:把该镜缩到 ≤6s,把砍掉的秒拆到 Step 5 的新相邻行,重跑 Step 5.5 自检再重渲。
3. 第三次重试:把失败这一镜切到 Seedance 2.0(这正是混合模式存在的意义 — 表演回退是一键切换,不是重构流程)。
4. 同一镜连续 3 次失败:暂停并用选项卡问用户:
- 单独把这一镜切到 Seedance 2.0。
- 放宽要求(去掉一个道具、简化动作、降低 hook 强度)。
- 跳过该镜,下游加 `placeholder: missing clip` 标记。
- 手动提供参考视频代替生成。
Seedance 2.0 失败回退阶梯(用户全局选了 Seedance 时,或 H3 已失败 3 次后):
1. 第一次重试:直接引用 `参考锚点` 块的强化 prompt 重渲染。
2. 第二次重试:去掉参考图,纯文本生成。
3. 第三次重试:缩到 ≤6s,砍掉的秒拆到新行。
4. 同一镜连续 3 次失败:问用户 — 切到 H3、宽松要求、跳过、手动参考视频。
### 所有片段渲染完成后
按镜头顺序放到画布,归组为 `<title> shot clips`(不再硬编码单一模型名,因为项目可能混用 H3 和 Seedance 2.0),展示用户选项卡:
- 批准片段并合成全片(推荐)
- 重渲染选定片段(可指定用同一模型或切换)
- 修复角色不一致
- 修复场景不一致
- 加强分镜清理
- 修复跨片段空间锚点漂移
如果渲染片段与已批准的 `参考锚点` 漂移(门框落到错边、人物从错边离屏、光位翻转),用直接引用 `参考锚点` 块的强化提示重渲染;持续漂移时按混合模式切到另一模型;不要在拼接时静默混合修正版和未修正版。
## STEP 8:全片拼接、BGM 匹配与最终输出
所有单镜头片段批准后,按表格顺序拼接成完整主视频。然后生成一条连续 BGM 匹配故事氛围,嵌入拼接视频。
拼接与 BGM 规则:
- 严格按已批准表格的镜头顺序。
- BGM 匹配拼接视频的实际节奏、情绪弧、喜剧节拍、追逐节奏和结尾氛围。
- 在对白、非语言反应和重要音效下 duck BGM。
- 保留原片段音频和音效(除非用户要求替换)。
- 不要逐镜生成 BGM。
- 不要加字幕或文字(除非用户明确要求)。
- 最终视频必须保持干净动画视觉、无任何分镜痕迹,包括没有 `[char:…] [scene:…] [shot:…]` 双绑定标签。
然后展示用户选项卡:
- 批准最终成片(推荐)
- 重新生成 BGM
- 调整 BGM 混音
- 重渲染选定片段
## STEP 9:最终复盘
用户要求诊断或存在明显风险时,生成一段简短复盘文本节点。
检查:
- 角色一致性
- 场景连续性(对照 `参考锚点` 列验证:每个地标是否落在表格说的位置)
- 情绪锚点回收
- 镜头意图清晰度
- 对白可懂度
- 音效/动效同步(对照 `音频与对白轨` 列)
- BGM 平衡
- 最终视频无分镜痕迹:没有面板边框、铅笔线、箭头、标签、手写备注、时间标记、姿态残影、分镜文字,**也没有双绑定标签**(`[char:…]`、`[scene:…]`、`[shot:…]`、`[dur:…]`、`[hook:…]`)
- 缺失或薄弱片段
- 任何可能需要重做的资产
## 画布排序与归组纪律
每生成一个耐用产物,立刻按 STEP 0 顺序写到画布。生成工具自动上画布的不重复。
每个生产环节都归组:
- 项目简报与故事大纲在都有时归组为 `<title> story planning`。
- 角色卡归组为 `<title> character cards`。
- 场景卡归组为 `<title> scene cards`。
- 标准镜头表归组为 `<title> shot table`。
- 单文本分镜文档归组为 `<title> text storyboards`(默认模式,单文档)。抽取的独立文本分镜节点和铅笔分镜(可视化模式)各自分组:`<title> extracted text storyboards` 和 `<title> multi-panel pencil storyboards`。所有分镜组与渲染片段组分开,因为分镜含双绑定标签、镜头号、镜头图标、箭头和铅笔线。
- 单镜头视频片段归组为 `<title> shot clips`(标签不再硬编码单一模型名,因为项目可能混用 H3 和 Seedance 2.0)。
- 拼接主视频、匹配 BGM 和最终合成视频在都有时归组为 `<title> final delivery`。
一轮生成产出 2 个或更多产物时,立刻归组并起一个清晰标题。渲染器位置未刷新导致归组失败时,请用户点击/拖动任一画布节点一次,再重试归组,必要时再进入下一个高成本环节。
## 用户选项卡纪律
每个需要用户确认的地方都用选项卡。不要用纯对话问题或散文代替。强制选项卡门:
- 输入捕获后,下一步之前:选画面尺寸 / 比例
- 输入捕获后,下一步之前:选总时长
- 项目简报后
- 故事大纲后
- 角色卡后
- 场景卡后
- 标准镜头表后(门 1:批准表格进入自检)
- 镜头表自检通过后(门 2:批准自检,然后立即选分镜模式:仅文本 / 文本+铅笔图)
- 单镜头分镜后(默认单文本分镜文档,可含抽取的独立节点;可视化模式时含铅笔图)
- 单镜头视频片段渲染前:选视频模型(H3 默认,Seedance 2.0 回退,逐镜混合)
- 单镜头视频片段渲染前:选视频分辨率
- 单镜头视频片段后
- 全片拼接、BGM 匹配和最终合成后
默认推荐项放最前。始终允许用户自定义输入。用户说"继续"时按选了推荐项处理。
## 重生成与最新资产纪律
用户重做或修订任何产物时,下游步骤必须用最新批准的产物,不能用旧的。
规则:
- 角色卡重做后,未来的镜头表、单文本分镜文档(含抽取的独立节点;可视化模式下还含铅笔分镜)、单镜头视频片段、拼接视频、最终合成必须按精确角色名引用新的带标注角色卡。
- 场景卡重做后,未来的镜头表、单文本分镜文档(含抽取的独立节点;可视化模式下还含铅笔分镜)、单镜头视频片段、拼接视频、最终合成必须按精确场景名引用新场景卡。
- 镜头表修订后,未来的分镜(文档+抽取节点)、单镜头视频片段、拼接必须用新表。Step 5.5 的自检必须在分镜恢复前重跑。
- 单文本分镜文档中某章节修订后,对应单镜头视频片段必须用新章节内容。若该章节已抽取成独立节点,则独立节点为权威。
- 抽取的独立文本分镜节点修订后,用户满意时回填到文档(占位换成最新内容),独立节点归档。
- 可视化模式下铅笔分镜重画后,对应视频片段仍以单文本分镜文档(或抽取的独立节点)为准;铅笔图仅供人眼参考,可独立重画而无需强制重渲视频。
- 单镜头视频片段重渲染后,拼接、BGM 匹配和最终合成必须用新片段。片段原用 H3 渲染,用户想切到 Seedance 2.0 重渲(或反向)按模型切换处理 — 在项目简报记入新模型,重检 per-shot 规则。
- BGM 重新生成后,最终合成必须用新 BGM。
任何重做后:
1. 在下一条文本输出或回复中标记重做产物为当前批准版本。
2. 在所有后续提示中优先用新文件路径 / 节点。
3. 画布上同时存在多版本时,按文件名 / 节点名指明当前选用版本再继续。
4. 不要在最终拼接中静默混合新旧资产。
## 边界
单张图、简单修图、单段镜头动画、logo 设计、纯提示咨询不用本 Skill。用户只要提示时,用视频提示工作流。用户只要角色卡时,用角色拆解工作流。
@@ -0,0 +1,249 @@
---
name: 3d-animation-short-generator
description: |
Create complete stylized 3D animated shorts from a story idea through an ordered production workflow covering project brief, story outline, character and environment cards, standardized shot planning, text or optional pencil storyboards, video-model selection, single-shot generation, assembly, BGM matching, and final review. Use when the user wants an end-to-end narrative animation workflow with strong character consistency, scene continuity, timing, camera, performance, and audio control. Not for single images, simple edits, photorealistic live action, or one standalone clip.
compatibility: Requires the MiniMax Hub agent (canvas workspace, choice cards, and hub_generate_image/hub_generate_video tools); not portable to generic agent harnesses.
---
# 3D Animation Short Generator
Use this Skill when the user wants a complete story-first animated short workflow, from one-line idea to final edited video. The workflow must place every major artifact on the canvas in production order and pause at creative gates with user choice cards before expensive or high-impact steps.
Core rule: **story first, ask screen size and total duration with choice cards immediately after user intake, ordered canvas artifacts, every required confirmation via choice card, fixed order after character/scene cards: six-column standardized shot table with per-second directives + audio cues + spatial anchor chain → shot-table self-check gate → one single text storyboards document with one section per shot (multi-panel pencil image only when user explicitly opts into visualization mode; any shot flagged for heavy iteration is extracted to a standalone text storyboard node) → video-model choice card (H3 default, Seedance 2.0 fallback) + resolution choice card → single-shot clips rendered by the chosen video model → full-film assembly with BGM; final video must remove all storyboard artifacts**.
## Global Visual Style Lock
Unless the user explicitly requests another visual style, all character cards, scene cards, shot tables, text storyboards, optional pencil storyboards, single-shot video clips (regardless of which video model is selected), assembled videos, and final composites must use this visual style:
- Rendering style: Pixar-inspired 3D cartoon rendering, C4D + Octane renderer look, high-end animated feature quality.
- Character design: exaggerated geometric simplification balanced with excellent material detail. Avoid 100% realistic human anatomy; use high-level shape language, large readable silhouettes, and Q-version proportions when appropriate.
- Proportion language: friendly stylized proportions, often 2.5–3 head-tall for cute or childlike characters, with big heads, compact bodies, clear silhouettes, and high recognizability.
- Hair / fur: combine strong sculpted clumps and clean block shapes with fine edge flyaway hairs or fuzzy rim details, so hair/fur feels designed but tactile under light.
- Skin / material: warm subsurface scattering skin quality, soft translucent reddish light through ears, cheeks, nose, and fingertips; avoid hard plastic skin.
- Acting style: exaggerated, lively Disney/Pixar-style character animation performance with squash and stretch, strong brows, eye corners, pupils, lips, and cheek shape changes.
- Motion style: high-energy poses, clear line of action, forward lean, strong anticipation, fast but readable timing, elastic body mechanics, and vivid micro-expressions.
- Emotional range: balance cuteness and explosive expressiveness; intense emotions may use dramatic facial deformation while preserving character appeal.
Negative style constraints: no photorealistic live-action, no flat 2D anime, no plastic toy skin, no stiff mannequin posing, no realistic anatomical stiffness, no lifeless facial expressions.
## STEP 0: Intake and Canvas Plan
Capture:
- One-line idea or rough premise
- Desired output: blueprint only, assets only, standardized shot table with per-second directives, single text storyboards document (default) + extracted single-shot text storyboard nodes for heavy-iteration shots + multi-panel pencil storyboards (opt-in), single-shot video clips (with video model chosen in Step 7), assembled main video, or final BGM-composited film
- Approximate length, if the user already stated it
- Screen size / aspect ratio, if the user already stated it
- Visual tone: default warm stylized 3D animation
- Dialogue requirement: whether the film has dialogue, voiceover, or no speech
- Dialogue language only if the user explicitly states it; do not default to English dialogue
Immediately after capturing the user's input, before Project Brief or any other next step, show choice cards to confirm production format:
Screen size / aspect ratio card:
- 16:9 landscape (recommended for cinematic short)
- 9:16 vertical short
- 1:1 square
- 4:5 social portrait
- Custom size / aspect ratio
Total duration card:
- 30–60 seconds (recommended)
- 15–30 seconds
- 60–90 seconds
- 90–180 seconds
- Custom duration
Only proceed after the user chooses both screen size/aspect ratio and total duration, or explicitly supplies custom values. Store the approved screen size/aspect ratio and duration in the Project Brief and reuse them in shot timing, transition continuity, the standardized shot table with per-second directives, single-shot storyboards (text by default, pencil image when user opts in), single-shot video clips, assembly, BGM matching, and final composite settings.
Create or later update canvas artifacts in this order:
1. Project Brief text node
2. Story Outline text node
3. Labeled Character Card image nodes
4. Environment-only Scene Card image nodes
5. Standardized Shot Information Table node with six columns; each row must include per-second directives inside `Shot Description`
6. **Single text storyboards document** — one canvas text node named `<title> text storyboards` containing one section per shot (mirrors the half-narrated-drama storyboard structure). When the user flags a shot for heavy iteration, that section is extracted to a standalone text node and a `(extracted)` marker is left in the document. Pencil image storyboards, if opted in, are separate image nodes.
7. Single-Shot Video Clip nodes (rendered by the video model selected in Step 7 — H3 default, Seedance 2.0 fallback)
8. Assembled Main Video node
9. Matched BGM audio node and Final BGM-Composited Video node
Do not dump long production content only in chat. Put durable outputs on canvas as text, image, video, or audio nodes.
## STEP 1: Project Brief
Produce a concise project brief and write it to a canvas text node named with the project title or `项目简报`.
Include:
- Working title
- One-line What-if
- Emotional premise
- Target audience feeling
- Main deliverables planned
- Approved screen size / aspect ratio
- Approved total duration
- Dialogue mode and language: only use a specific language when the user explicitly requested it; otherwise write `language not specified` and keep dialogue minimal or ask later when needed
- Initial risks
- Dialogue intent when present
Then show a user choice card:
- Continue with this direction (recommended)
- Regenerate premise options
- Revise emotional premise
- Refine dialogue direction
Only proceed after the user chooses or explicitly says to continue.
## STEP 2: Story Outline and Gates
Create a story outline and write it to a canvas text node named `故事大纲` or `story-outline`.
Include:
- Protagonist Want / Need / flaw
- Core world rule
- 8-beat causal story spine
- Emotional anchor and payoff
- Dialogue beats if the user requested dialogue
- Red-line checks
Gate checks:
- Protagonist is active
- Crisis is intensified by protagonist flaw
- Coincidence never solves the problem
- Ending reuses an earlier emotional anchor
- Antagonistic pressure is not a flat villain
- Dialogue reveals relationship change instead of explaining the theme
Then show a user choice card:
- Approve story and continue (recommended)
- Revise beats
- Revise emotion curve
- Revise dialogue beats
- Return to premise
## STEP 3: Character Cards
Generate character reference cards and place each image on canvas. Recommended order:
1. Protagonist card
2. Contrast / pressure character card
3. Optional supporting character card
Each character card should be a 16:9 production reference sheet when possible. Unlike final rendered video, character cards should include clear readable labels so downstream generation can bind the correct person and props:
- Character name label in English and/or the project language
- Role label, such as protagonist, grandma, thief, sidekick, pressure character
- Main 3/4 view
- Front / side / back views
- Expressions
- Material / costume / prop details
- Important prop labels, such as handbag, wallet, skateboard, apple basket, scarf, shoes, glasses
- Identity lock repeated in the prompt
- A short visual-ID note listing age range, body type, hairstyle, outfit colors, signature props, and do-not-change traits
For stylized 3D animation, keep the character soft, readable, and consistent across later images and videos.
After the main character cards are generated, show a user choice card:
- Lock character designs and continue (recommended)
- Regenerate protagonist card
- Adjust specific visual details
- Add another character card
Warn the user that changing locked character designs later may require regenerating the shot table, single-shot storyboards, single-shot video clips, assembled main video, and final composite.
## STEP 4: Scene Cards
Generate scene reference cards and place them on canvas after character cards. Scene cards must show environments only: do not include characters, people, crowd figures, silhouettes, hands, faces, or character cameos. Character action belongs in the shot table, single-shot multi-panel pencil storyboards, and single-shot video clips, not scene cards.
Include:
- Main environment overview
- Key light states, such as day / night
- Emotional sub-spaces
- Continuity landmarks (fixed objects whose screen position must persist across shots in the same scene, e.g. kitchen island, sofa, door frame, tree, mailbox)
- Important props in the environment
Then show a user choice card:
- Lock scene design and continue (recommended)
- Regenerate scene card
- Add another scene angle
- Adjust lighting or layout
## STEP 5: Standardized Shot Table Video Prompts (Six Columns)
After character cards and scene cards are locked, output standardized video prompts as a shot information table. This step is mandatory and cannot be swapped with storyboard or video generation.
Required reference: read and follow `references/shot-table-spec.md` for the exact six-column schema, per-second directive requirements, table-wide rules, user approval card, and mandatory Step 5.5 self-check gate.
Minimum runtime contract:
- Create a canvas table node named `标准镜头信息表` or `standard-shot-table`.
- Use exactly six columns: `Shot ID & Duration`, `Continuity Handoff`, `Reference Anchors (Spatial + Identity)`, `Hook Type`, `Shot Description (Per-Second Directives)`, `Audio & Dialogue Track`.
- Every row must include complete per-second directives, continuity handoff, reference anchors, hook type, and audio/dialogue timing.
- Run the Step 5.5 self-check from `references/shot-table-spec.md` before storyboarding. Do not enter Step 6 until the self-check passes.
Then show the table approval/self-check choice cards defined in `references/shot-table-spec.md`.
## STEP 6: Text Storyboards Document (Default) + Pencil Image Storyboards (Opt-in)
After the Step 5.5 self-check passes, show a storyboard-mode choice card before producing any storyboard artifact.
Required reference: read and follow `references/storyboard-guidelines.md` for the default single text storyboards document, optional multi-panel pencil storyboards, shot-level extraction/re-integration, storyboard approval cards, and visualization fallback rules.
Minimum runtime contract:
- Default mode is one authoritative text storyboards document containing one section per shot.
- Pencil storyboard images are opt-in visualization artifacts only; they never override the text storyboard.
- Extract a shot into a standalone text node only when the user flags that shot for heavy iteration.
- Step 7 must read the matching text storyboard section or extracted standalone node, not the pencil image.
After all storyboards are approved, proceed to the video-model choice card.
## STEP 7: Video-Model Choice Card + Single-Shot Video Clips
Before any clip is rendered, show the video-model choice card and resolution choice card.
Required references:
- Read `references/model-selection.md` for the H3 default, Seedance 2.0 fallback, per-shot mixed mode, resolution choices, and model-specific prompt shaping.
- Read `references/fallback-policy.md` for per-model retry ladders, drift handling, and escalation choices.
Minimum runtime contract:
- H3 is the recommended default model.
- Seedance 2.0 is the fallback for high-stakes animation performance or repeated H3 failure.
- Per-shot mixed mode is allowed only when the shot table marks the model per row; unmarked rows default to H3.
- Strip all storyboard-only labels before video render.
- Bind each clip to the approved text storyboard section, exact character cards, and exact scene card.
- If a clip drifts from the approved `Reference Anchors`, follow `references/fallback-policy.md`; do not silently assemble incorrect clips.
After all clips render, place them on canvas in shot order, group them as `<title> shot clips`, and show the clip approval card defined in `references/model-selection.md`.
## STEP 8: Full Film Assembly, BGM Match, and Final Output
After all single-shot clips are approved, assemble the complete main video, match or generate one continuous BGM track, and produce the final composited video.
Required reference: read `references/qc-checklist.md` for assembly rules, BGM rules, final review checks, canvas ordering/grouping discipline, user choice-card discipline, and regeneration/latest-asset discipline.
Minimum runtime contract:
- Preserve the exact shot order from the approved table.
- Use only approved latest assets.
- Duck BGM under dialogue, reactions, and important SFX.
- Do not add subtitles or text unless explicitly requested.
- Final video must contain no storyboard traces, labels, arrows, timing marks, panel borders, or double-binding labels.
Then run the final review checks from `references/qc-checklist.md` and deliver the final approved asset.
## Boundaries
Do not use this Skill for a single image, a simple edit, a single clip animation, logo design, or pure prompt consultation. If the user only wants a prompt, use a video prompt workflow instead. If the user only wants a character card, use a character breakdown workflow instead.
@@ -0,0 +1,19 @@
display-name-zh: 3D动画短片生成器
version: 0.5.4
tag-en: Animation
tag-cn: 动画
complete-tags-en:
- Animation / Planning
- Animation / Creative Generation
- Animation / Post-production
complete-tags-cn:
- 动画 / 计划制定
- 动画 / 创作生成
- 动画 / 后期制作
summary-en: Turn a story idea into a complete stylized 3D animated short.
summary-cn: 根据故事创意,完成人物与场景设定、镜头规划、分镜生成和视频合成,输出风格统一的3D动画短片。
desc-en: Designed for creators who want to turn a story idea into a complete stylized 3D animated short. Users provide a one-line concept or basic plot and confirm aspect ratio, duration, dialogue needs, and visual direction. The Skill creates a project brief and story outline, builds labeled character and environment-only scene cards, plans a shot table with per-second action, camera, audio, and continuity requirements, produces text storyboards or optional pencil boards, selects a video model, generates individual shots, assembles the film, matches BGM, and performs a final review. It outputs a coherent 3D animated short with consistent characters, continuous scenes, and controlled pacing. Best for narrative animation, birthday keepsakes, brand stories, and social shorts; not for single images, simple retouching, photorealistic live action, or one standalone clip.
desc-cn: 面向希望将故事灵感制作成完整3D动画短片的创作者。用户需提供一句故事创意或基础剧情,并确认画面比例、总时长、对白需求和目标风格。Skill 会依次完成项目简报与故事大纲,生成角色卡和无人物场景卡,规划带每秒动作、镜头、音频与连续性要求的镜头表,制作文本分镜或可选铅笔分镜,再选择视频模型生成单镜头片段,完成全片拼接、BGM匹配和成片复查。最终输出角色一致、场景连续、节奏完整的3D动画短片。适用于剧情动画、生日纪念、品牌故事和社交媒体短片,不适用于单张图片、简单修图、真人写实视频或仅生成一个独立镜头。
author-en: MiniMax Hub User
author-cn: MiniMax Hub 用户
source: community
@@ -0,0 +1,52 @@
# Generation Failure and Drift Fallback Policy
### Failure fallback per video model
H3 fallback ladder when a clip fails or drifts:
1. First retry: regenerate the same clip with the H3 prompt prefix strengthened by quoting the exact `Reference Anchors` block from the table.
2. Second retry: shorten the shot to ≤6s and split the dropped seconds into a new adjacent row in Step 5; re-run Step 5.5 self-check, then re-render.
3. Third retry: switch the failing shot to Seedance 2.0 (this is exactly why the mixed mode exists — performance fallback is one click, not a re-architecture).
4. After three failed attempts on the same shot: pause and ask the user with a choice card:
- Switch just this shot to Seedance 2.0.
- Loosen the request (drop a prop, simplify the action, reduce the hook).
- Skip the shot and add a `placeholder: missing clip` note for downstream review.
- Manually supply a reference video to bind instead of generating.
Seedance 2.0 fallback ladder (used when the user has explicitly chosen Seedance as the global model, or after H3 already failed three times):
1. First retry: regenerate with a strengthened prompt that quotes the `Reference Anchors` block.
2. Second retry: drop reference images and use text-only generation.
3. Third retry: shorten the shot to ≤6s and split the dropped seconds into a new row.
4. After three failed attempts: ask the user — switch to H3 for this shot, loosen the request, skip the shot, or supply a reference video.
### After all clips are rendered
Place the rendered clips on canvas in shot order, group them as `<title> shot clips` (no longer hard-coded to a single model name), and show a user choice card:
- Approve clips and composite full film (recommended)
- Re-render selected clip (with the same or a different video model)
- Fix character mismatch
- Fix scene mismatch
- Strengthen storyboard cleanup
- Fix spatial anchor drift across clips
If a rendered clip drifts from the approved `Reference Anchors` (e.g. door-frame lands on the wrong side, character exits from the wrong edge, lighting flipped), re-render with a strengthened prompt that quotes the exact `Reference Anchors` block from the table. If the drift persists, switch that one shot to the other video model in the mixed-mode path; do not silently mix corrected and uncorrected clips into assembly.
## Storyboard Visualization Fallback
### Storyboard generation failure fallback (visualization mode only)
If a pencil image storyboard cannot be produced at the required quality (e.g. layout collapses, labels illegible, panels merged, character inconsistency), apply the following escalation before asking the user:
1. **First retry**: regenerate the same shot storyboard with a tightened prompt that explicitly mentions the four-quadrant layout, the `[char:…] [scene:…] [shot:…]` labels, and the per-panel content rules.
2. **Second retry**: drop the bottom-right audio/anchor quadrant text (keep it as a blank cell with a tiny `♪` mark) to reduce text load; this usually fixes illegible labels without losing the visual beat.
3. **Third retry**: reduce panel count by one (e.g. 6 panels → 5 panels by merging the two least-actionable seconds) and simplify camera icons to single arrows.
4. **After three failed attempts on the same shot**: pause and ask the user with a choice card:
- Switch to a block-color storyboard (gray boxes for poses, no pencil lines) for the failing shot only.
- Drop the pencil image for the failing shot and rely on the text storyboards document alone for that row.
- Split the failing shot into two shorter shots in Step 5 and re-run Step 5.5.
- Manually supply a reference image to bind instead of generating.
In default text mode this whole fallback is unnecessary — text storyboards fail only when the model cannot produce coherent structured text, in which case return to Step 5 to revise the table row.
@@ -0,0 +1,49 @@
# Video Model Selection and Prompt Shaping
## STEP 7: Video-Model Choice Card + Single-Shot Video Clips
### Video-model choice card (mandatory before any clip render)
Before any clip is rendered, show the video-model choice card. The choice is stored in the Project Brief and reused for every clip in this project unless the user later changes it.
Video model card:
- **H3 (recommended default)** — strong on visual packaging, motion graphics, text/UI clarity, multi-modal context understanding, and cost efficiency (about 1/3 the price of comparable flagship models at 2K, 1/2 at 768P). Native dual-channel audio. Up to 15s per clip at 2K. Best for: stylized 3D animated shorts with strong design language, text overlays, motion-graphic moments, packaging-style transitions, and dialogue-driven beats where the audio is part of the deliverable.
- **Seedance 2.0 (fallback for high-stakes animation performance)** — strong on cinematic camera, complex shots, elastic Pixar-style performance, and tension-driven action. Best for: chase sequences, slapstick beats, climax shots where the selling point is the animation itself rather than the packaging.
- **Per-shot mixed (advanced)** — let the user mark `video_model: H3` or `video_model: Seedance2` in the `Shot Description` column of individual rows. Use this when the project has both packaging-heavy and performance-heavy shots. The default for unmarked rows is H3.
### Resolution choice card (after video model)
Once the video model is locked, show the resolution choice card:
- 768P (recommended for H3 first pass; cost-efficient)
- 2K (H3 default quality; higher cost, sharper final render)
- 1080p (recommended for Seedance 2.0)
- 720p (Seedance 2.0 draft; lowest cost)
- Match project / custom resolution
The user must confirm a resolution before the first clip renders. Resolution can be changed per clip later if the user wants a hero shot at higher detail.
### Single-shot clip rendering
For each approved table row, call the chosen video model to generate the corresponding independent video clip. Each clip must use exactly the matching section from the text storyboards document (extracted standalone node if that section was extracted, otherwise the in-document section), character card(s), and scene card from that row.
Per-shot rules common to all video models:
- Use the text storyboards document as the authoritative per-shot reference for narrative, composition, camera movement, action staging, per-second timing, and shot number. For shots that have been extracted to a standalone node, read the extracted node instead. If a pencil image storyboard also exists, use it only for human-side pose / silhouette pre-check; do not let it override the text storyboard.
- Use character cards as the authoritative identity source.
- Use scene cards as the authoritative environment source.
- **Strip all storyboard double-binding labels** (`[char:…]`, `[scene:…]`, `[shot:…]`, `[dur:…]`, `[hook:…]`) before video render — these labels are storyboard-only reference markers and must NOT appear in the final clip. Pencil image storyboards additionally have their own shot numbers, camera icons, arrows, and notes that must be removed at render time.
- The rendered clip must contain only clean full-color Pixar-inspired 3D animation content.
- No storyboard line art, no hand-drawn sketch texture, no labels, no subtitles unless requested, no watermarks.
- Maintain the approved screen size / aspect ratio and the approved video resolution from the resolution choice card.
### Model-specific prompt shaping
The text storyboards document (or the extracted standalone node for that shot) feeds both models, but the prompt prefix around the storyboard differs:
- **H3 prompt prefix** (default): emphasize packaging keywords, design language, motion clarity, text/UI presence when relevant, and dual-channel audio intent. H3 is strong at instruction following, so the per-second directive can be sent almost verbatim. Add: `Pixar-inspired 3D cartoon rendering, C4D + Octane look, stylized Q-version proportions, warm SSS skin, designed-with-detail hair, strong character design language, clean motion, on-brand color palette`.
- **Seedance 2.0 prompt prefix** (performance fallback): emphasize cinematic camera language, elastic squash-and-stretch, anticipation, follow-through, lighting drama, and lens choice. Add: `cinematic Pixar-quality 3D animation, elastic squash-and-stretch performance, Disney-style anticipation and overshoot, dramatic key lighting, lens-specific depth of field`.
When the user picked `per-shot mixed`, apply the prefix that matches the row’s `video_model` field.
@@ -0,0 +1,99 @@
# Assembly, Final Review, and Asset Discipline
## STEP 8: Full Film Assembly, BGM Match, and Final Output
After all single-shot clips are approved, concatenate them in table order into the complete main video. Then generate one continuous BGM track that matches the story mood and embed it into the assembled video.
Assembly and BGM rules:
- Preserve the exact shot order from the approved table.
- Match BGM to the assembled video’s actual pacing, emotional arc, comedy beats, chase rhythm, and ending tone.
- Duck BGM under dialogue, non-language reactions, and important SFX.
- Preserve existing clip audio and SFX unless the user asks to replace them.
- Do not generate BGM per shot.
- Do not add subtitles or text unless the user explicitly asks.
- Output the final video with clean animation visuals and no storyboard traces, including no `[char:…] [scene:…] [shot:…]` labels.
Then show a user choice card:
- Approve final film (recommended)
- Regenerate BGM
- Adjust BGM mix
- Re-render selected clip
## STEP 9: Final Review
Create a short final review text node if the user asks for diagnosis or if there are visible risks.
Check:
- Character consistency
- Scene continuity (verify against the `Reference Anchors` column — did every landmark land where the table said it would?)
- Emotional anchor payoff
- Shot purpose clarity
- Dialogue intelligibility
- Foley/SFX sync (against the `Audio & Dialogue Track` column)
- BGM balance
- No storyboard artifacts in final video: no panel borders, sketch lines, arrows, labels, handwritten notes, timing marks, pose ghosts, storyboard text, AND no double-binding labels (`[char:…]`, `[scene:…]`, `[shot:…]`, `[dur:…]`, `[hook:…]`)
- Missing or weak clips
- Any asset that may need regeneration
## Canvas Ordering and Grouping Discipline
Whenever a durable artifact is created, write it to the canvas immediately in the sequence defined in STEP 0. If a generation tool automatically adds outputs to canvas, do not duplicate them.
Group every production section on canvas:
- Group project brief and story outline as `<title> story planning` when both exist.
- Group character cards as `<title> character cards`.
- Group scene cards as `<title> scene cards`.
- Group the standardized shot table as `<title> shot table`.
- Group the text storyboards document as `<title> text storyboards` (default mode, one document). Extracted standalone text storyboard nodes and pencil storyboards (visualization mode) are separate groups: `<title> extracted text storyboards` and `<title> multi-panel pencil storyboards`. Keep all storyboard groups separate from rendered clips because storyboards contain double-binding labels, shot numbers, camera icons, arrows, and sketch lines.
- Keep storyboard groups separate from rendered clips because storyboards contain double-binding labels, shot numbers, camera icons, arrows, and sketch lines.
- Group single-shot video clips as `<title> shot clips` (the label no longer hard-codes a specific model name, since projects can mix H3 and Seedance 2.0).
- Group assembled main video, matched BGM, and final composited video as `<title> final delivery` when they exist.
If a generation round produces two or more outputs, group recent outputs immediately with a clear title. If the renderer has not flushed positions and grouping fails, tell the user to click/drag any canvas node once, then retry grouping before continuing to the next costly stage when possible.
## User Choice Card Discipline
Use a choice card for every place that requires user confirmation. Do not replace these confirmations with plain chat questions or prose. Required choice-card gates:
- Immediately after intake, before any next step, to choose screen size / aspect ratio
- Immediately after intake, before any next step, to choose total duration
- After project brief
- After story outline
- After character cards
- After scene cards
- After standardized shot table (gate 1: approve table before self-check)
- After shot-table self-check passes (gate 2: approve self-check, then immediately choose storyboard mode: text only / text + pencil image)
- After single-shot storyboards (text storyboards document by default with optional extracted standalone nodes; pencil images if the user opted in)
- Before single-shot video-clip rendering, to choose the video model (H3 default, Seedance 2.0 fallback, per-shot mixed)
- Before single-shot video-clip rendering, to choose video resolution
- After single-shot video clips
- After full-film assembly, BGM match, and final composite
Default recommended option should be first. Always allow custom user input. If the user says “continue,” treat it as choosing the recommended option.
## Regeneration and Latest-Asset Discipline
When the user regenerates or revises any artifact, all downstream steps must use the newest approved artifact, not the older one.
Rules:
- If a character card is regenerated, future shot tables, the text storyboards document (plus any extracted standalone text storyboard nodes), pencil storyboards (if visualization mode is on), single-shot video clips, assembled videos, and final composites must reference the regenerated labeled character card by exact character name.
- If a scene card is regenerated, future shot tables, the text storyboards document (plus any extracted standalone text storyboard nodes), pencil storyboards (if visualization mode is on), single-shot video clips, assembled videos, and final composites must reference the regenerated scene card by exact scene name.
- If the shot table is revised, future text storyboards document (or extracted nodes), single-shot video clips, and assembly must use the revised table. The self-check in Step 5.5 must be re-run before storyboarding resumes.
- If a section in the text storyboards document is revised, the matching single-shot video clip must use the revised section. If that section has been extracted to a standalone node, that node is the source of truth; otherwise the document section is.
- If a standalone extracted text storyboard is revised, after the user is satisfied, re-integrate it back into the text storyboards document (replace the placeholder with the latest content) and archive the standalone node.
- In visualization mode, if a pencil storyboard is redrawn, the matching video clip is still bound to the text storyboards document (or extracted node); the pencil image is human-review-only and may be redrawn without forcing a video re-render.
- If a single-shot video clip is re-rendered, assembly, BGM matching, and final composite must use the new clip. If the clip was rendered with H3 and the user wants to re-render with Seedance 2.0 (or vice versa), treat that as a model switch — record the new model in the Project Brief and re-check per-shot rules.
- If BGM is regenerated, the final composite must use the regenerated BGM.
After any regeneration:
1. Mark the regenerated artifact as the current approved version in the next text output or reply.
2. Prefer the regenerated file path / node over previous versions in all subsequent prompts.
3. If there are multiple versions on canvas, identify the chosen current version by filename or node name before continuing.
4. Do not silently mix old and new assets in final assembly.
@@ -0,0 +1,75 @@
# Standardized Shot Table Specification
## STEP 5: Standardized Shot Table Video Prompts (Six Columns)
After character cards and scene cards are locked, output standardized video prompts as a shot information table. This step is mandatory and cannot be swapped with storyboard or video generation. Create a canvas table node or markdown table named `标准镜头信息表` or `standard-shot-table`.
The table must have exactly six columns in this order:
| Shot ID & Duration | Continuity Handoff | Reference Anchors (Spatial + Identity) | Hook Type | Shot Description (Per-Second Directives) | Audio & Dialogue Track |
|---|---|---|---|---|---|
Column rules:
- **Shot ID & Duration**: shot number plus planned duration, e.g. `S03 / 6s`.
- **Continuity Handoff**: how this shot naturally continues from the previous shot’s ending image, prop position, eyeline, character posture, sound bridge, or emotional state, AND how it sets up the next shot’s opening. This is the cross-shot continuity spine.
- **Reference Anchors (Spatial + Identity)**: four sub-fields, all mandatory.
- `Fixed Landmarks` — exact named landmarks from the scene card, with their screen-relative positions (e.g. `door-frame: right third`, `kitchen-island: center bottom`).
- `Character Positions (camera view)` — for every character in the shot, screen-relative position (left/center/right, top/mid/bottom, foreground/midground/background), facing direction, and initial pose.
- `Exited Character Status` — for any character who was in the previous shot but is not in this shot, state their off-screen position and reason (e.g. `Mia — exited frame-left, last seen holding apple basket at door-frame`).
- `Lighting Baseline` — inherited key/fill/rim direction from the scene card, plus any per-shot modifier (e.g. `key: warm overhead, fill: cool bounce right, modifier: window-backlit silhouette`).
- Plus identity bindings: exact approved character card names and exact approved scene card name.
- **Hook Type**: one short label from a controlled vocabulary, e.g. `visual-joke`, `reversal`, `suspense`, `tender`, `chase`, `reveal`, `callback`, `expression-beat`. Used for the per-episode hook distribution self-check.
- **Shot Description**: shot size, camera movement, Dutch-angle design, performance style, SFX, negative prompt, **video-model generation notes** (the model chosen in Step 7 — H3 or Seedance 2.0 — receives slightly different prompt shapes; for H3 emphasize packaging keywords and text/UI/motion-graphics clarity, for Seedance 2.0 emphasize cinematic camera and elastic performance), and a required `Per-Second Directives` subsection. The subsection must break the shot into second-by-second instructions such as `0–1s`, `1–2s`, `2–3s`; for sub-second critical beats, use `2.0–2.5s` style markers. Each per-second directive MUST cover all five required elements:
1. Action / pose / expression (squash-and-stretch, anticipation, overshoot, follow-through where applicable)
2. Camera movement (push / pull / pan / tilt / handheld-shake / locked / orbit)
3. Spatial position (where the character is, what they hold, what landmark is in frame)
4. Audio cue (narration / dialogue / SFX / breath / silence — or `silent` if intentional)
5. Handoff to the next second or next shot (what state this second locks in for the next one)
- **Audio & Dialogue Track**: full audio script for the shot in time order, separate from per-second cues. Fields:
- `Narration` — voiceover text with time range (omit if no narration).
- `Dialogue` — line, speaker, tone, time range.
- `SFX` — keyed sound effects, time-anchored.
- `Performance Note` — when the protagonist is narrating off-screen, mark `narrator-mouth-closed: true`; describe expression path during narration; describe concrete eye-line and body-action changes for each dialogue line.
Table-wide rules:
- Each shot must naturally inherit the previous shot’s image state through the `Continuity Handoff` column and set up the next shot in the same column.
- Each shot row must include per-second directives that cover the entire shot duration from first frame to last frame, including action, pose, expression, camera movement, spatial position, sound cue, and continuity handoff.
- Per-second directives must be specific enough to generate storyboard panels directly; avoid vague timing such as “continues moving” without body, camera, or prop detail.
- Performance must be exaggerated and elastic, matching Disney-style squash-and-stretch, anticipation, overshoot, follow-through, overlap, arcs, fast pose changes, and clear comedic silhouettes.
- Shot sizes must alternate close-up / extreme close-up with other necessary framing; avoid repetitive framing.
- Dutch-angle tilted compositions must be designed into chase, imbalance, surprise, or slapstick beats.
- Dialogue language only if explicitly requested by the user; otherwise use minimal non-language-specific reactions or mark dialogue as optional / pending confirmation.
- For shots containing narration or dialogue, every second where the character is speaking must record whether the mouth is open or closed; the default is closed for off-screen narrator, open for on-screen dialogue.
- A character who left the frame must still be tracked in `Exited Character Status` for at least one shot, then dropped after they are explicitly off-stage for two consecutive shots.
Then show a user choice card:
- Approve table and run self-check (recommended)
- Adjust shot continuity
- Make animation more exaggerated
- Adjust close-up / extreme-close-up rhythm
- Adjust Dutch-angle design
## STEP 5.5: Shot Table Self-Check Gate (Mandatory)
Before moving to pencil storyboards, run a hard self-check on the approved shot table. If any check fails, revise the table and re-run before asking the user to approve storyboarding.
Six required checks:
1. **Hook density**: every shot has a `Hook Type`; at least one of every three consecutive shots uses a `reveal`/`reversal`/`callback`; the opening shot and the closing shot each carry a strong hook (`visual-joke` / `reversal` / `reveal` / `suspense` / `tender`).
2. **Single-shot duration**: no shot exceeds 15 seconds. If a beat needs more, split it.
3. **Character count per shot**: no shot contains more than three important characters (defined as characters with on-screen action or dialogue).
4. **Spatial anchor inheritance**: for every interior scene with two or more shots, the `Fixed Landmarks` and `Lighting Baseline` of the next shot must either match the previous shot or include an explicit continuity note (e.g. `door-frame moves from right third to center as camera orbits left`).
5. **Per-second directive coverage**: every second from `0s` to the shot duration is covered by a `Per-Second Directives` entry, and each entry contains all five required elements (action/pose/expression, camera, spatial, audio cue, handoff). Sub-second beats like `2.0–2.5s` are allowed but must not leave any time gap.
6. **Cross-shot continuity**: reading the `Continuity Handoff` column row by row produces a continuous chain — no shot starts from a state that contradicts the previous shot’s ending. Any shot that flips eyeline, character position, prop state, or lighting must mark the flip explicitly (e.g. `HARD CUT — time skip: 2h`).
If all six pass, place a `shot-table self-check: passed` stamp at the top of the canvas table node and show the user choice card:
- Approve self-check and draw shot storyboards (recommended)
- Show self-check details
- Revise failed checks
- Re-run self-check
If any check fails, do not enter Step 6. Return to Step 5, list the failed rows, and only re-show the storyboard approval card after the table is fixed and the self-check passes.
@@ -0,0 +1,185 @@
# Text and Pencil Storyboard Guidelines
## STEP 6: Text Storyboards Document (Default) + Pencil Image Storyboards (Opt-in)
After the Step 5.5 self-check passes, show a storyboard-mode choice card before producing any storyboard artifact:
- **Text storyboards document only (default, recommended)** — one canvas text node containing all shot storyboards as in-document sections. Mirrors the half-narrated-drama storyboard structure: per-shot fields (title / hook / scene / characters / spatial anchors / continuity / performance) plus Pixar's per-panel four-quadrant content + optional ASCII layout. Carries the full quality-control payload at near-zero cost. The video model selected in Step 7 reads this directly as the per-shot rendering reference.
- **Text storyboards document + multi-panel pencil image (visualization mode, opt-in)** — the text storyboards document is still produced as the authoritative artifact, AND one multi-panel pencil image is generated per shot for human review. Higher cost, useful when the user wants a visual preview before committing to video generation, or when squash-and-stretch / pose silhouette is the main risk and the user wants to pre-check it visually.
Store the chosen storyboard mode in the Project Brief and reuse it in Step 7, Step 9, and the Regeneration discipline.
### Default path: single text storyboards document
Generate one canvas text node named `<title> text storyboards` (one document for the whole short). This document is the authoritative rendering reference for Step 7 even when pencil images are also produced. The structure mirrors the half-narrated-drama storyboard — every shot is a section in the same document, so the user can read cross-shot continuity without node-hopping.
Document top matter (header block at the top of the document):
- Project title, approved video model, approved resolution, storyboard mode, and self-check status (e.g. `shot-table self-check: passed at <timestamp>`).
- A short table-of-contents listing every shot, its hook, and its section anchor (`S01`, `S02`, …) so the user can jump.
Per-shot section structure (one `##` heading per shot, in shot order). Every section is mandatory to contain these fields, in this order — direct adaptation of the half-narrated-drama storyboard:
1. **Shot title & duration** — short human-readable title for the shot, plus `S<N> / <duration>s` (e.g. `S03 / 6s`).
2. **Hook type** — one of the controlled vocabulary: `setup` / `visual-joke` / `reversal` / `reveal` / `callback` / `suspense` / `tender` / `chase` / `expression-beat` / `climax`. Used by the per-episode hook distribution self-check.
3. **Scene & characters** — exact scene card name and exact on-screen character names (binding to character cards).
4. **Spatial anchor card** (mandatory, four sub-fields — directly adapted from half-narrated):
- `Fixed landmarks` — named landmarks and their screen-relative positions (e.g. `door-frame: right third`, `kitchen-island: center bottom`).
- `Character positions (camera view)` — for every on-screen character, screen-relative position, facing direction, and initial pose.
- `Exited character status` — characters who were on screen in the previous shot but not in this one, with their off-screen position and reason.
- `Lighting baseline` — inherited key/fill/rim direction from the scene card, plus per-shot modifier.
5. **Continuity** (mirrors half-narrated's handoff fields):
- `Continuity from S(N-1)` — one or two sentences referencing the previous shot’s ending state.
- `Continuity to S(N+1)` — one sentence setting up the next shot’s opening.
6. **Double-binding** — `[char:角色名-01] [char:角色名-02] ... [scene:场景名] [hook: visual-joke]` — exact character card names, scene card name, hook type. These are storyboard-only reference markers; the video model strips them at render time.
7. **Per-panel four-quadrant content** (one block per panel, in time order; this is the Pixar per-second directive, kept verbatim from the table row):
- `Timecode` — e.g. `0–1s`.
- `Pose + Expression` — concrete body posture, silhouette, key prop grip, eye-line, facial expression path; for elastic beats, explicitly call out squash / stretch / anticipation / overshoot. This is the largest section per panel and is what the video model reads as the visual beat.
- `Camera` — shot size, camera movement (push / pull / pan / tilt / handheld-shake / locked / orbit), Dutch angle note when applicable.
- `Audio + Anchor` — audio cue (`♪ narration: ...` / `dialogue: ...` / `SFX: ...` / `silent`) and spatial anchor note (`door-frame: right third` / `Mia: center midground facing camera`).
- Performance notes (mirrors half-narrated): for narration seconds, mark `narrator-mouth-closed: true`; for on-screen dialogue, mark `mouth-open: speaker` and describe expression path / eye-line / body-action changes.
8. **Layout rules** (apply per shot):
- 3-second shot → 3 panels (1 per second).
- 4-second shot → 4 panels.
- 5-second shot → 5 panels.
- 6-second shot → 6 panels.
- 7+ second shot → one panel per second; for sub-second critical beats, add an extra mini-panel such as `2.0–2.5s` only when the beat is the hook of the shot.
- Panels must cover the full shot duration from first frame to last frame with no time gaps.
9. **Per-panel binding**:
- Bind the exact character cards listed in the table row to lock appearance, face, hairstyle, body proportions, costume, signature props, and role identity. Use the same character names as the table.
- Bind the exact scene card listed in the row to preserve environment, props, landmarks, movement paths, and spatial logic.
10. **Optional ASCII layout block (highly recommended, free)**:
- Append a small ASCII sketch per panel (or one combined sketch for the whole shot) so the user can scan the spatial layout in seconds without rendering an image. Example:
```
[0-1s] Mia (L, mid) door-frame (R)
──kneels, hands on apple basket──
cam: low push-in, locked
audio: silent | anchor: basket center-bottom
[1-2s] ...
```
- The ASCII block is informational only; the video model reads the structured `Per-panel four-quadrant content` above, not the ASCII.
11. **Storyboard-only markers**:
- When a beat is critical, append `[BEAT]` after the panel timecode.
- When a panel must handoff a specific state to the next panel or the next shot, append `[HANDOFF → ...]` with a short label such as `[HANDOFF → S04 opening]`.
Per-shot section template (copy-paste skeleton, valid for any shot):
```markdown
## S03 / 6s — Title: 奶奶把苹果筐递给 Mia
- **Hook type**: reveal
- **Scene & characters**: scene:kitchen | char:Mia, char:Grandma
- **Spatial anchor card**:
- Fixed landmarks: door-frame (right third), kitchen-island (center bottom)
- Character positions: Mia (L, midground, facing camera) | Grandma (R, foreground, facing Mia)
- Exited character status: —
- Lighting baseline: warm overhead key + cool bounce right
- **Continuity from S02**: 奶奶弯下腰从中岛拿起苹果筐
- **Continuity to S04**: Mia 接住筐转身,门铃响起
- **Double-binding**: [char:Mia] [char:Grandma] [scene:kitchen] [hook:reveal]
### Per-panel four-quadrant content
#### 0–1s
- Pose + Expression: 奶奶弯腰双手持筐;Mia 左侧站姿,眼神好奇
- Camera: locked medium shot, eye-level
- Audio + Anchor: silent | Mia: L midground | basket: center bottom
- Performance: [BEAT]
#### 1–2s
- Pose + Expression: 奶奶手臂伸向 Mia,筐倾斜;Mia 双手前伸准备接
- Camera: locked medium shot, eye-level
- Audio + Anchor: ♪ SFX: basket rustle | anchor: door-frame: right third
- Performance: [HANDOFF → S04 opening]
#### 2–3s
...
### ASCII layout (optional)
[0-1s] Grandma (R, fg) door-frame (R, bg)
──lifts basket── Mia (L, mid)
cam: locked | silent
[1-2s] ...
```
After all sections are written, place the document on canvas and move directly to Step 7. Do not call any image generation model in default mode.
### Shot-level extraction (heavy-iteration mode)
The default single-document form is optimized for reading and cross-shot continuity. When the user flags a specific shot for heavy iteration (typically climax / chase / slapstick beats where the per-panel content needs many rounds of revision), extract that section into a standalone text node so iteration is localized:
- User signal: at any time after Step 6, the user says things like "let me focus on S05", "S05 needs rework", "extract S05", or selects a shot during the storyboard approval choice card.
- Extraction mechanics:
1. Create a new canvas text node named `<title> S05 text storyboard (extracted)`.
2. Move the full content of the `## S05` section from the document into the new node.
3. In the document, replace the `## S05` section with a one-line placeholder: `> S05 — extracted to standalone node (see `<title> S05 text storyboard (extracted)`)`.
4. Step 7 reads from the extracted node for S05; all other shots still read from the document.
- Re-integration: when the user is satisfied, the standalone node is folded back into the document (replace the placeholder with the latest content) and the standalone node is archived.
- Multiple extracted shots: each shot gets its own standalone node; the document tracks them with placeholders.
The extraction mechanism exists because independent nodes are best used by need, not by default — but they remain available whenever iteration pressure is high on a specific shot.
### Opt-in path: multi-panel pencil image storyboards (visualization mode)
If the user picked the visualization mode in the storyboard-mode choice card, ALSO produce one multi-panel pencil storyboard image per table row on top of the text storyboards document. The text storyboards document remains the authoritative rendering reference; the pencil images are human-review-only.
For each pencil image storyboard:
- **Double-binding labels (top-right corner, mandatory on image)**:
- `[char:角色名-01] [char:角色名-02] ...` — exact character card names used in this row.
- `[scene:场景名]` — exact scene card name.
- `[shot: S03] [dur: 6s] [hook: visual-joke]` — shot ID, duration, and hook type.
- These labels are storyboard-only reference markers; they are stripped at video render time.
- Bind the exact character cards listed in that row to lock character appearance, face, hairstyle, body proportions, costume, signature props, and role identity.
- Bind the exact scene card listed in that row to preserve environment, props, landmarks, movement paths, and spatial logic.
- Convert every per-second directive in the row into one storyboard panel or one clearly labeled beat panel; for a 4-second shot, normally create 4 panels; for a 6-second shot, normally create 6 panels; for sub-second critical beats, add extra mini-panels only when needed.
- **Panel physical layout (mandatory)**:
- 3-second shot → 1×3 strip.
- 4-second shot → 2×2 grid.
- 5-second shot → top row 3 + bottom row 2.
- 6-second shot → 2×3 grid.
- 7+ second shot → 3 rows, balanced panels.
- Each panel occupies the same canvas area; do not let one panel dominate.
- **Per-panel four-quadrant content (mandatory)**:
- Top-left: timecode (e.g. `0–1s`).
- Top-right: pose + expression sketch (the largest area; the actual visual beat).
- Bottom-left: camera icon + movement arrow (push/pull/pan/orbit/locked) and a tiny note for Dutch angle.
- Bottom-right: audio cue (e.g. `♪ narration: "I knew it."` / `SFX: door creak` / `silent`) and anchor note (e.g. `door-frame: right third`).
- Arrange panels in reading order inside the same single-shot storyboard image; do not merge multiple different shots into one image.
- Each panel must mark its timecode, such as `0–1s`, `1–2s`, and show the corresponding pose, expression, action, camera movement, prop position, SFX cue, and continuity handoff.
- Output pure black-and-white pencil line-art only: no color, no final-render lighting, no polished 3D render.
- Mark the storyboard image with the shot number and include camera-movement icon / marker per panel when useful.
- Include storyboard-only marks when useful: pencil construction lines, action arrows, camera-path icon, timing marks, and small notes.
- Keep the draft as a video-render reference asset only, not final art.
### Storyboard approval (both modes)
After all text storyboard sections (and pencil images, if visualization mode is on) are produced, place them on canvas in shot order, group them as:
- `<title> text storyboards` (default mode, single document), OR
- `<title> text storyboards + multi-panel pencil storyboards` (visualization mode, group the text document and the pencil images separately because pencil images contain double-binding labels, ASCII labels, and shot numbers that the text document does not).
Show a user choice card:
- Approve storyboards and render shot videos (recommended)
- Extract / re-integrate a shot (move section between document and standalone node)
- Redraw selected pencil storyboard (visualization mode)
- Fix character consistency
- Fix scene logic
- Fix camera marker
- Fix audio/anchor markers
### Storyboard generation failure fallback (visualization mode only)
If a pencil image storyboard cannot be produced at the required quality (e.g. layout collapses, labels illegible, panels merged, character inconsistency), apply the following escalation before asking the user:
1. **First retry**: regenerate the same shot storyboard with a tightened prompt that explicitly mentions the four-quadrant layout, the `[char:…] [scene:…] [shot:…]` labels, and the per-panel content rules.
2. **Second retry**: drop the bottom-right audio/anchor quadrant text (keep it as a blank cell with a tiny `♪` mark) to reduce text load; this usually fixes illegible labels without losing the visual beat.
3. **Third retry**: reduce panel count by one (e.g. 6 panels → 5 panels by merging the two least-actionable seconds) and simplify camera icons to single arrows.
4. **After three failed attempts on the same shot**: pause and ask the user with a choice card:
- Switch to a block-color storyboard (gray boxes for poses, no pencil lines) for the failing shot only.
- Drop the pencil image for the failing shot and rely on the text storyboards document alone for that row.
- Split the failing shot into two shorter shots in Step 5 and re-run Step 5.5.
- Manually supply a reference image to bind instead of generating.
In default text mode this whole fallback is unnecessary — text storyboards fail only when the model cannot produce coherent structured text, in which case return to Step 5 to revise the table row.
+117
View File
@@ -0,0 +1,117 @@
# MiniMax H3 Skills
This directory contains the skills bundled with [MiniMax H3](../README.md): **1 prompt writing skill** and **8 style-specific video generation skills**. Each skill lives in its own folder with an installable `SKILL.md` (plus a `SKILL.cn.md` Chinese version for the style skills) and any reference materials it needs.
## Status
The skills are actively maintained and still evolving. The 8 style skills ship with bilingual `SKILL.md`/`SKILL.cn.md`; `h3-prompt-writing` is currently English-only.
## Installation
Install skills with the [skills CLI](https://github.com/vercel-labs/skills):
```bash
# List all skills available in this repository
npx skills add https://github.com/MiniMax-AI/MiniMax-H3 --list
# Install all skills
npx skills add https://github.com/MiniMax-AI/MiniMax-H3 --skill '*'
# Install a single skill
npx skills add https://github.com/MiniMax-AI/MiniMax-H3 --skill h3-prompt-writing
```
## Skills
### h3-prompt-writing
[SKILL.md](h3-prompt-writing/SKILL.md)
Write structured MiniMax H3 video generation prompts for all five generation modes: T2VA, I2VA, FL2VA, L2VA, and Ref2VA. The skill rewrites multimodal requests into H3's prompt structure — `integrated_multimodal_description`, `overall_soundscape`, and `non_diegetic_music` — aligns keyframes, and defines reference labels for images, videos, and audio. It ships with two prompt guides under [`references/`](h3-prompt-writing/references/):
- [`base-en.txt`](h3-prompt-writing/references/base-en.txt) — base text/keyframe modes
- [`ref-en.txt`](h3-prompt-writing/references/ref-en.txt) — full-reference (Ref2VA) mode
### minimalist-product-ad-generator
<p align="center">
<img src="../assets/minimalist-product-ad-generator.gif" alt="minimalist-product-ad-generator" width="240">
</p>
Turn product images and ad requirements into clean, minimalist product ad shorts for e-commerce promotion and product launches. The skill confirms format and product variants, extracts selling points, writes concise English ad copy, plans beat-synced typography and storyboards, and generates a premium product film with polished camera language. Not for KOC talking-head ads, general editing, or complex screen demos.
[SKILL.md](minimalist-product-ad-generator/SKILL.md) · [SKILL.cn.md](minimalist-product-ad-generator/SKILL.cn.md)
### 3d-animation-short-generator
<p align="center">
<img src="../assets/3d-animation-short-generator.gif" alt="3d-animation-short-generator" width="240">
</p>
Create complete stylized 3D animated shorts from a story idea through an ordered production workflow: project brief, story outline, character and environment cards, standardized shot planning, text or optional pencil storyboards, video-model selection, single-shot generation, assembly, BGM matching, and final review. Built for end-to-end narrative animation with strong character consistency, scene continuity, timing, camera, performance, and audio control. Not for single images, simple edits, photorealistic live action, or one standalone clip.
[SKILL.md](3d-animation-short-generator/SKILL.md) · [SKILL.cn.md](3d-animation-short-generator/SKILL.cn.md)
### papercraft-stop-motion-explainer
<p align="center">
<img src="../assets/papercraft-stop-motion-explainer.gif" alt="papercraft-stop-motion-explainer" width="240">
</p>
Explain science, education, or general knowledge through tactile handmade papercraft visuals. The skill extracts the learning goal and visual metaphor, proposes creative directions, designs paper characters, layered diorama sets, and props, creates preview concepts plus image and video prompts, and plans storyboards, camera movement, transitions, and sound with staged approvals and review checklists. It outputs a production-ready papercraft stop-motion explainer package, or selected assets such as still prompts, image-series prompts, short-video prompts, or storyboards. Best for cut-paper, pop-up-book, layered diorama, and miniature stop-motion explainers.
[SKILL.md](papercraft-stop-motion-explainer/SKILL.md) · [SKILL.cn.md](papercraft-stop-motion-explainer/SKILL.cn.md)
### brand-promo-video-generator
<p align="center">
<img src="../assets/brand-promo-video-generator.gif" alt="brand-promo-video-generator" width="240">
</p>
For marketers and creators producing promotional content for brands, products, websites, apps, shops, or personal projects. The skill organizes brand facts and asset provenance, selects a narrative direction, plans precise beats and shots, generates needed imagery, video, voiceover, or music, and completes assembly and pre-delivery review. It outputs a promotional short that highlights product capabilities, use cases, and a call to action. Best for launches, website showcases, and social promotion; not for imitating real brand marks without authorized assets or inventing product claims.
[SKILL.md](brand-promo-video-generator/SKILL.md) · [SKILL.cn.md](brand-promo-video-generator/SKILL.cn.md)
### music-video-subtitle-generator
<p align="center">
<img src="../assets/music-video-subtitle-generator.gif" alt="music-video-subtitle-generator" width="240">
</p>
For musicians, video creators, and social-media editors producing AI music videos or emotional short films with lyric typography. The skill analyzes beat and vocal timing, separates character, scene, and text references, designs beat-reactive spatial typography, decomposes long works into connected shots, audits prompts, and routes generation for H3 or other video tools. It outputs MV concepts, shot prompts, lyric text plans, and stitching guidance. Best for stylized MVs and subtitle-driven music visuals.
[SKILL.md](music-video-subtitle-generator/SKILL.md) · [SKILL.cn.md](music-video-subtitle-generator/SKILL.cn.md)
### co-op-game-intro-generator
<p align="center">
<img src="../assets/co-op-game-intro-generator.gif" alt="co-op-game-intro-generator" width="240">
</p>
Create a two-player co-op game menu or opening animation. The skill locks identity cues, generates an approval image from a fixed menu framework with coordinated color, buttons, icons, and typography, then uses the approved result to rebuild character, UI-copy, and event timing instructions for the final video. It outputs a co-op game intro featuring two characters, player cards, and menu interaction motion. Best for game concepts, character-led menus, and social content.
[SKILL.md](co-op-game-intro-generator/SKILL.md) · [SKILL.cn.md](co-op-game-intro-generator/SKILL.cn.md)
### paper-collage-explainer-generator
<p align="center">
<img src="../assets/paper-collage-explainer-generator.gif" alt="paper-collage-explainer-generator" width="240">
</p>
Give narration, knowledge points, opinions, or abstract topics a tactile paper-collage language. The skill extracts meaning, proposes visual metaphors, prepares a production plan and storyboard, generates approved halftone collage stills, then creates stop-motion clips with paper movement and tactile sound effects, with optional final assembly. By default it keeps collage SFX and does not add BGM, voiceover, or subtitles unless requested. Best for explainers, viewpoints, story visuals, and social B-roll.
[SKILL.md](paper-collage-explainer-generator/SKILL.md) · [SKILL.cn.md](paper-collage-explainer-generator/SKILL.cn.md)
### handdrawn-live-video-generator
<p align="center">
<img src="../assets/handdrawn-live-video-generator.gif" alt="handdrawn-live-video-generator" width="240">
</p>
Create surreal short videos that blend rough glowing hand-drawn animation with live-action spaces. The skill clarifies the physical contact, designs continuous morphing, an escape route, and a delayed handheld chase movement, then writes a reusable 15-second 16:9 video prompt in the user's language. After confirmation it recommends MiniMax H3 generation and checks contact realism, camera delay, rough glowing stroke texture, and non-horror tone. Best for single-scene creative clips, not polished CG, horror jump scares, plush characters, or multi-scene cuts.
[SKILL.md](handdrawn-live-video-generator/SKILL.md) · [SKILL.cn.md](handdrawn-live-video-generator/SKILL.cn.md)
## Contribute
These skills are still being improved, and community contributions are encouraged. If you optimize an existing skill or add a new one, open a PR — contributing or optimizing skills comes with API credit rewards.
@@ -0,0 +1,178 @@
---
name: brand-promo-video-generator
description: 面向需要为品牌、产品、网站、App、小店或个人项目制作宣传内容的运营与创作者。用户需提供 LOGO、产品图、界面截图、官网链接或其他可核验素材,并确认时长、画幅、受众和推广重点。Skill 会整理品牌事实与素材来源,选择叙事方向,规划精确节拍和镜头,生成所需图像、视频、旁白或音乐,并完成合成与交付前检查。最终输出一条突出产品功能、使用场景和行动号召的品牌宣传短片。适用于新品发布、官网展示和社交媒体推广,不适用于缺少授权素材时仿造真实品牌标识、虚构产品功能或制作长篇剧情影片。
allowed-tools:
- webfetch
- hub_image_search
- hub_analyse_media
- hub_canvas_get_node
- task
- hub_audio_meta
- hub_canvas_group_recent_outputs
---
# 品牌宣传短片生成器
为品牌、产品、网站、App、小店或个人项目制作一条好看、清晰、能直接展示的宣传短片。用户可以提供 LOGO、产品图、截图、官网链接,也可以只给一个想法;Skill 会把这些素材整理成品牌短片。
此 Hub 适配版已移除第三方 Vibe Motion / Remotion 项目初始化与 npm 校验依赖,改为 Hub 原生编排:来源研究、资产核验、故事规划、图像/视频生成、可选语音或音乐,以及后期合成。正常执行时不要初始化外部项目。
## 步骤 1:上传素材并确认简报
在任何故事规划或生成之前,必须先进行用户问询。请用户上传或提供链接,用于核验以下元素:
- LOGO 文件或官方 LOGO 来源页面
- 字体文件、字体名称、字体规范,或能体现品牌字体的官方网站页面
- 品牌颜色、色彩系统、风格指南,或能清楚体现官方色彩的页面
- 产品图片、界面截图、包装、渲染图、视频素材或其他品牌图片
- 产品信息:官方产品名称、功能列表、发布重点、宣传主张、行动号召、免责声明和目标受众
- 公司/产品官网或官方资料包
默认把用户上传的素材视为可用于前期方案和概念制作,不要在第一轮问询中反复单独追问“是否有授权”。只有出现明确风险信号时才询问,例如可见第三方水印、明显抓取的电商/图库图、用户表述互相矛盾、涉及法律/医疗/金融合规主张,或用户明确要求商用发布。来源清单中将素材记录为“用户提供”,并在来源摘要里提示必要的权利注意事项,而不是阻塞流程。
在同一次开场问询中,让用户选择:
- 目标时长,通常为 15-30 秒;如果用户想要快速发布片,推荐 15 秒
- 画幅比例;提供 16:9、9:16、1:1、4:3、3:4 或匹配用户参考图等常见选项
如果活动重点、投放渠道、旁白语言、画面文案语言和画面文案需求尚不清楚,也在此阶段一并确认。用户未提供可用素材,且未明确说明哪些元素不可用之前,不要进入创意方向阶段。
宣传片内容语言规则:旁白和画面文案语言应根据品牌素材、目标受众和投放平台判断,不要机械跟随聊天语言。如果品牌素材和可见来源文案主要是英文或全球企业英文,默认旁白和画面文案使用英文,除非用户明确要求中文本地化。即使用户用中文聊天或说“我是新用户测试”,聊天回复仍用中文,但视频实际文案应使用最适合该品牌传播的语言。
如果 LOGO、产品界面、人物、吉祥物、包装、字体、色彩系统或其他带身份识别的资产无法认证,停止并向用户索要授权原件,不要生成看起来相似的替代品。
## 步骤 2:建立品牌事实表
研究或检查最可靠的来源,并在创意制作前总结品牌事实表:
1. 用户提供的原始导出文件
2. 官方网站、静态资源包、新闻中心、品牌门户、媒体资料包、新闻资料包或官方仓库
3. 公司控制的官方媒体库
4. 授权图库或授权合作方资料包
提取:
- 精确 LOGO 版本、安全空间和使用限制
- 官方字体或可见的字体行为
- 主色、辅助色和动态色彩
- 品牌语气、原则、视觉母题和交互语言
- 当前产品名称、功能、场景、指标、口号、行动号召和免责声明
- 官方摄影、渲染图、界面截图、视频素材、新闻素材和媒体包内容
不要使用 LOGO 聚合站、搜索缩略图、粉丝重绘、Pinterest 转发或 AI 生成替代品作为带身份识别的来源。程序化图形只可作为非写实动效层:遮罩、渐变、色块、网格、辉光、粒子、轨迹、字体、已验证数据图表和转场几何。
## 步骤 3:创建来源清单
为每个带身份识别的资产记录一份简洁来源清单,可作为文字节点或表格交付给用户。清单应包含:
- 稳定资产 ID 与角色
- 本地路径或画布节点引用
- 精确来源 URL,或用户提供来源说明
- 来源类型,例如官网、媒体包、用户原始素材或授权图库
- 用于比对的验证目标
- 权利或发布说明
- 真实性状态:已验证、用户提供、已授权或阻塞
来源清单本身不等于发布授权。授权不明确时,将视频标注为非官方概念,并告知用户商业发布需要许可。
## 步骤 4:选择故事脊柱
如果用户尚未选定方向,给出 2-3 个简短创意方向,推荐一个,并在确认后继续。按产品类型选择叙事脊柱:
- AI / SaaS:用户意图 -> 思考或规划 -> 能力 -> 执行 -> 有用输出 -> 证明 -> LOGO
- 实体产品:英雄展示 -> 交互 -> 功能特写 -> 使用场景 -> 结果 -> LOGO
- 服务 / 公司:背景 -> 流程 -> 证据 -> 成果 -> 承诺 -> LOGO
- 图像主导品牌:真实影像 -> 视觉母题 -> 利益点 -> 情绪回报 -> LOGO
故事必须产品专属。展示真实功能、交互、场景、输出和证明,不要用抽象特效掩盖薄弱产品叙事。
## 步骤 5:规划精确节拍
生成前先规划带帧意识的时间线。除非输出流程另有要求,默认用 30fps 作为规划单位。
15 秒影片建议 5-8 个主要节拍;30 秒影片建议 8-12 个主要节拍。每个节拍应定义:
- 开始与结束时间或帧范围
- 视觉主导元素和真实资产 ID
- 主要动作
- 镜头中展示的产品或品牌证明
- 文案与可读停留
- 色彩状态
- 入场与出场转场
- 动作意图:铺垫、预备、承诺、冲击、制动、稳定
一个实用的 15 秒结构是:品牌钩子、用户意图或场景建立、产品机制、能力或应用场景、输出或证明、产品回报、最终 LOGO 与行动号召。当前一个动作能自然带出下一个镜头时,可使用 6-12 帧重叠。
## 步骤 6:导演动效语言
用可控方式建立强度:
- 让产品动作、光标路径、界面流、光线、滚动内容、物体边缘或匹配几何驱动转场
- 使用 2-5 个有意义的色彩状态
- 每个节拍保留一个主要动作;次级层稍作延迟
- 设置 2-3 个高能峰值和较安静的制动时刻
- 保持剪影、文案和 LOGO 安全空间清晰可读
- 避免虚假 HUD、随意玻璃卡片、装饰性文字墙、未经验证的指标,以及全片同一种缓动
AI 产品至少包含一条可读链路,例如:提示词 -> 规划 -> 并行能力 -> 生成结果 -> 证明。实体或服务产品则展示从用户动作到具体成果的因果关系。
## 步骤 7:生成前硬确认
在生成任何视频、图片序列、语音、音乐或最终剪辑之前,必须停下,并先向用户展示完整前期方案:
- 来源清单或来源摘要
- 品牌事实表
- 已选创意方向
- 精确节拍 / 分镜计划
- 画面文案、CTA、旁白和声音方案
- 已知的真实性、授权或占位说明
生成前保持简洁确认。用户在看过前期方案后,如果明确表达认可或推进意图,例如“确认”“生成”“可以”“继续”“下一步”“go ahead”“continue”“next”等,应视为允许生成,除非同一句话同时要求修改。若用户要求跳过流程,仍需先给出简版来源摘要、品牌事实表和分镜节拍;随后用户表达同意推进即可生成。只有当用户回复含糊或提出修改时,才列出修改分镜、修改文案、修改声音或暂停等选项。
## 步骤 8:生产 Hub 资产
只有生成前硬确认通过后,才可以使用 Hub 原生生成与剪辑。每次派发都应自包含:所选模型、画面比例、真实参考路径和用户原始需求。
典型制作流程:
1. 生成或准备已验证的静帧、界面底板、产品主视觉帧或适合动效的视频首帧。
2. 基于这些画面或精确文字描述生成视频片段,并保持相同比例与品牌资产。
3. 品牌短片的默认声音策略:当用户要求 BGM、音乐、配乐、氛围声,或只说要一条完整宣传片而没有要求独立音轨时,优先使用所选视频模型的原生音频,不要默认拆成“无声视频 + 单独音乐”。默认 MiniMax H3 路线应开启原生音频(`generate_audio=true`),并在视频提示词中写入品牌安全的纯音乐 / UI 音效设计。
4. 只有当用户明确需要可控旁白、口播、对白、可单独替换的 BGM、独立于视频的精确音乐时长、后期混音,或所选视频模型无法生成合适音频时,才单独生成语音或音乐。不要让视频原生音频和独立音频重复生成同一条声轨。
5. 在剪辑阶段合成片段、用户明确要求的独立音频和最终品牌锁定画面。只有用户明确要求字幕时才添加字幕。
不得重绘或近似 LOGO、字标、产品界面、包装、吉祥物、人物或品牌场景。生成素材只用于抽象动效、氛围、转场几何,或明确为概念化且不冒充官方产品证据的场景。
## 步骤 9:交付前验证
最终回复前检查:
- LOGO 和带身份识别的资产来自已验证或用户授权的来源
- 产品名称、功能表述、主张、指标、口号和行动号召与官方来源一致,或被明确标注为概念文案
- 视频时长、比例和语言符合简报
- 文案可读且不过度拥挤
- 最终 LOGO 未被拉伸、裁切或重建
- 动效有清晰视觉主导,不遮挡产品
- 输出已在画布上,多资产输出已编组
如果真实性检查失败,用官方或用户授权来源替换可疑资产,或停止并向用户索取资产。绝不优化仿制品。
## 步骤 10:交付
提供:
- 最终视频路径或画布输出
- 时长、比例和语言
- 简短创意总结
- 来源清单或来源摘要
- 必要的权利/免责声明
- 下一轮可改进的具体建议,例如节奏、主张清晰度、行动号召、音频或平台裁切
## 失败恢复
- LOGO 错误或近似:移除它,定位当前官方文件或向用户索要原件,再重新生成或剪辑。
- 产品/界面看起来虚假:替换为官方、用户提供或授权素材,不要美化仿制品。
- 好看但通用:加入完整产品交互、已验证主张或真实输出。
- 快但混乱:减少同时动作,指定视觉主导,并保持跨镜头匹配运动。
- 顺滑但太慢:缩短停留,重叠转场,只在关键信息处制动。
- 资产不可用:索要授权原件,绝不猜测。
@@ -0,0 +1,189 @@
---
name: brand-promo-video-generator
description: For marketers and creators producing promotional content for brands, products, websites, apps, shops, or personal projects. Users provide logos, product images, interface screenshots, official links, or other verifiable assets and confirm duration, aspect ratio, audience, and campaign focus. The Skill organizes brand facts and asset provenance, selects a narrative direction, plans precise beats and shots, generates needed imagery, video, voiceover, or music, and completes assembly and pre-delivery review. It outputs a promotional short that highlights product capabilities, use cases, and a call to action. Best for launches, website showcases, and social promotion; not for imitating real brand marks without authorized assets, inventing product claims, or producing long-form narrative films.
compatibility: Requires the MiniMax Hub agent (canvas workspace and the hub_* tools listed below); not portable to generic agent harnesses.
allowed-tools:
- webfetch
- hub_image_search
- hub_analyse_media
- hub_canvas_get_node
- hub_canvas_group_recent_outputs
- hub_generate_image
- hub_generate_video
- hub_generate_audio
- hub_generate_music
- hub_synthesize_speech
- hub_video_edit
- hub_audio_meta
- task
---
# Brand Promo Video Generator
Create a polished short promo video for a brand, product, website, app, shop, or personal project. Use this Skill when the user has a logo, product images, screenshots, a website link, or just a clear idea and wants the agent to turn those materials into a clean brand reel.
This Hub adaptation replaces third-party Vibe Motion / Remotion implementation details with Hub-native orchestration: source research, asset verification, story planning, image/video generation, optional speech or music, and editing assembly. Do not initialize external projects or depend on npm validators during normal execution.
## Tool Coverage Rule
The `allowed-tools` list must cover the full production promise in this Skill: source lookup, image preparation/generation, video clip generation, optional speech/music/audio generation, final editing/assembly, media inspection, and canvas grouping. If a runtime does not provide one of the listed generation or editing tools, downgrade the deliverable explicitly to a pre-production package instead of claiming that a final promo video can be generated.
## STEP 1: Intake assets and resolve the brief
Before any story planning or generation, run a required user intake. Ask the user to upload or provide links to the elements that must be verified:
- Logo files or official logo source pages
- Font files, font names, typography guidance, or official website pages that show brand typography
- Brand colors, color system, style guide, or pages that clearly show official colors
- Product images, UI screenshots, packaging, renders, footage, or other brand imagery
- Product information: official product name, feature list, launch focus, claims, CTA, disclaimers, and target audience
- Company/product official URL or official source package
Treat user-uploaded assets as usable by default for pre-production and concept planning. Do not repeatedly ask a standalone rights/permission question during the first intake unless there is a concrete risk signal, such as a visible third-party watermark, obviously scraped marketplace imagery, contradictory user wording, legal/medical/financial compliance claims, or the user asks for commercial publication. Record the source as "user-provided" and surface rights caveats in the source summary instead of blocking the flow.
In the same opening intake, ask the user to choose:
- Target duration, normally 15-30 seconds; recommend 15 seconds when the user wants a fast launch film
- Aspect ratio; offer common choices such as 16:9, 9:16, 1:1, 4:3, 3:4, or match a supplied reference
Also identify campaign focus, distribution channel, narration language, on-screen copy language, and visible copy needs when they are not already clear. Do not proceed to creative direction until the user has supplied the usable materials or explicitly confirms which elements are unavailable.
Language rule for promo content: choose narration and on-screen copy language from the brand materials, target audience, and platform context, not mechanically from the chat language. If the brand assets and visible source copy are primarily English or global corporate English, default narration and on-screen copy to English unless the user explicitly asks for Chinese localization. If the user is Chinese but says they are testing as a new user, keep chat replies in Chinese, but plan the actual video copy in the language that best fits the brand campaign.
If a logo, product UI, person, mascot, packaging, font, color system, or other identity-bearing asset cannot be authenticated, stop and ask for an authorized original instead of generating a plausible substitute.
## STEP 2: Build the brand truth sheet
Research or inspect the strongest available sources and summarize the brand truth sheet before creative production:
1. User-provided original exports
2. Official company website, static bundle, newsroom, brand portal, media kit, press kit, or official repository
3. Company-controlled media library
4. Licensed stock or authorized partner kit
Extract:
- Exact logo variants, clear space, and usage constraints
- Official fonts or visible typographic behavior
- Primary, secondary, and dynamic brand colors
- Brand tone, principles, visual motifs, and interaction language
- Current product names, features, scenarios, metrics, slogans, CTA, and disclaimers
- Official photography, renders, UI screenshots, footage, press assets, and media kit material
Do not use logo aggregation sites, search thumbnails, fan recreations, Pinterest reposts, or AI-generated substitutes as identity-bearing sources. Procedural graphics are allowed only as non-representational motion layers: masks, gradients, color fields, grids, glows, particles, trails, typography, verified-data charts, and transition geometry.
## STEP 3: Create a provenance manifest
Record every identity-bearing asset in a compact manifest that can be delivered to the user. The manifest may be a text node or table and should include:
- Stable asset ID and role
- Local path or canvas node reference when available
- Exact source URL or user-provided source note
- Source type, such as official website, media kit, user-provided original, or licensed stock
- Verification target used for comparison
- Rights or publication note
- Authenticity status: verified, user-supplied, licensed, or blocked
The manifest does not grant publication rights by itself. If authorization is unclear, label the video as an unofficial concept and tell the user commercial publication requires permission.
## STEP 4: Choose the story spine
Present 2-3 concise creative directions when the user has not already chosen one, recommend one, and continue after confirmation. Use the product category to pick a spine:
- AI / SaaS: user intent -> thinking or planning -> capabilities -> execution -> useful output -> proof -> logo
- Physical product: hero reveal -> interaction -> feature macro -> usage context -> result -> logo
- Service / company: context -> process -> evidence -> outcome -> promise -> logo
- Image-led brand: authentic imagery -> visual motif -> benefit -> emotional payoff -> logo
Keep the story product-specific. Show actual features, interactions, scenarios, outputs, and proof instead of hiding the story behind abstract effects.
## STEP 5: Plan exact beats
Plan a frame-aware timeline before generation. Use 30fps as the planning convention unless the output pipeline requires otherwise.
For a 15-second film, target 5-8 major beats. For a 30-second film, target 8-12 major beats. Each beat should define:
- Start and end time or frame range
- Visual owner and authentic asset IDs
- Primary action
- Product or brand proof shown in the shot
- Copy and readable hold
- Color state
- Incoming and outgoing transition
- Motion intent: setup, anticipation, commitment, impact, brake, settle
A useful 15-second pattern is: brand hook, user intent or setup, product mechanism, capabilities or scenarios, output or proof, product payoff, final logo and CTA. Use 6-12 frame overlaps when outgoing motion naturally supplies the next shot.
## STEP 6: Direct the motion language
Build intensity with control:
- Let product motion, cursor paths, UI flow, light, scrolling content, object edges, or matched geometry drive transitions
- Use 2-5 deliberate color states tied to meaning
- Keep one primary action per beat; delay secondary layers slightly
- Establish 2-3 high-energy peaks and quieter braking moments
- Preserve readable silhouettes, copy, and logo clear space
- Avoid fake HUDs, arbitrary glass cards, decorative text walls, unverified metrics, and identical easing everywhere
For AI products, include at least one readable chain such as: prompt -> planning -> parallel capabilities -> generated result -> proof. For physical or service products, show cause and effect from user action to concrete outcome.
## STEP 7: Hard confirmation before generation
Before generating any video, image sequence, speech, music, or final edit, stop and show the user the completed pre-production package:
- Provenance manifest or source summary
- Brand truth sheet
- Chosen creative direction
- Exact beat / shot plan
- Visible copy, CTA, narration, and audio plan
- Known authenticity, rights, or placeholder caveats
Use a concise confirmation step before generation. If the user clearly expresses approval or intent to proceed after seeing the pre-production package — for example "confirm", "generate", "go ahead", "continue", "next", "可以", "继续", "下一步", or similar — treat it as permission to generate, unless the message also asks for changes. If the user asks to skip the process, still provide a compact source summary, brand truth sheet, and beat plan first, then proceed when they indicate approval. Offer revision choices only when the user's reply is ambiguous or requests changes.
## STEP 8: Produce Hub assets
Use Hub-native generation and editing only after the hard confirmation gate has passed. Keep each dispatch self-contained with the chosen model, aspect ratio, authentic reference paths, and original request.
Typical production flow:
1. Generate or prepare verified still frames, UI plates, product hero frames, or motion-ready story images.
2. Generate video clips from those frames or from precise text prompts, preserving the same ratio and brand assets.
3. Default audio policy for native brand reels: when the user asks for BGM, music, soundtrack, ambient sound, or says nothing beyond needing a finished promo video, prefer video-native audio from the selected video model instead of generating a separate music track. For the default MiniMax H3 route, set native audio on (`generate_audio=true`) and prompt for brand-safe instrumental music / UI sound design inside the video prompt.
4. Generate separate speech or music only when the user explicitly needs controllable narration, voiceover, dialogue, replaceable standalone BGM, exact music duration independent of the video, post-production remixing, or when the selected video model cannot generate suitable audio. Do not duplicate a soundtrack between video-native audio and separate audio generation.
5. Assemble clips, any explicitly separate audio, and final brand lockup in editing. Add subtitles only when the user explicitly asks for subtitles.
Do not redraw or approximate logos, wordmarks, product UI, packaging, mascot, person, or brand scene. Use generated material only for abstract motion, atmosphere, transition geometry, or clearly conceptual scenes that do not impersonate official product evidence.
## STEP 9: Verify before delivery
Before final response, check:
- The logo and identity-bearing assets came from verified or user-authorized sources
- Product names, feature wording, claims, metrics, slogans, and CTA match official sources or are clearly marked as concept copy
- The video duration, aspect ratio, and language match the brief
- Copy is readable and not overcrowded
- The final logo is not stretched, cropped, or rebuilt
- Motion has clear visual ownership and does not obscure the product
- The output is on the canvas and multi-asset outputs are grouped
If the output fails an authenticity check, replace the questionable asset with an official/user-authorized source or stop and ask for the asset. Never improve an imitation.
## STEP 10: Deliver
Provide:
- Final video path or canvas output
- Duration, aspect ratio, and language
- Short creative summary
- Provenance manifest or source summary
- Rights/disclaimer note when needed
- Specific suggestions for the next iteration, such as pacing, claim clarity, CTA, audio, or platform crop
## Failure recovery
- Wrong or approximate logo: remove it, locate the current official file or ask for the user's original, then regenerate or re-edit.
- Fake-looking product/UI: replace with official, user-supplied, or licensed media. Do not polish the imitation.
- Beautiful but generic: add a complete product interaction, verified claim, or real output.
- Fast but chaotic: reduce simultaneous actions, assign a visual owner, and preserve matched motion across cuts.
- Smooth but slow: shorten holds, overlap transitions, and brake only around key messages.
- Asset unavailable: ask for an authorized original; never guess.
@@ -0,0 +1,19 @@
display-name-zh: 品牌宣传短片生成器
version: 0.1.9
tag-en: Commercial Ad
tag-cn: 商业广告
complete-tags-en:
- Commercial Ad / Planning
- Commercial Ad / Creative Generation
- Commercial Ad / Post-production
complete-tags-cn:
- 商业广告 / 计划制定
- 商业广告 / 创作生成
- 商业广告 / 后期处理
summary-en: Turn verified brand assets and campaign goals into a polished promotional short video.
summary-cn: 基于品牌素材与推广目标,完成事实核验、创意方向、分镜和音画合成,输出可展示的品牌宣传短片。
desc-en: For marketers and creators producing promotional content for brands, products, websites, apps, shops, or personal projects. Users provide logos, product images, interface screenshots, official links, or other verifiable assets and confirm duration, aspect ratio, audience, and campaign focus. The Skill organizes brand facts and asset provenance, selects a narrative direction, plans precise beats and shots, generates needed imagery, video, voiceover, or music, and completes assembly and pre-delivery review. It outputs a promotional short that highlights product capabilities, use cases, and a call to action. Best for launches, website showcases, and social promotion; not for imitating real brand marks without authorized assets, inventing product claims, or producing long-form narrative films.
desc-cn: 面向需要为品牌、产品、网站、App、小店或个人项目制作宣传内容的运营与创作者。用户需提供 LOGO、产品图、界面截图、官网链接或其他可核验素材,并确认时长、画幅、受众和推广重点。Skill 会整理品牌事实与素材来源,选择叙事方向,规划精确节拍和镜头,生成所需图像、视频、旁白或音乐,并完成合成与交付前检查。最终输出一条突出产品功能、使用场景和行动号召的品牌宣传短片。适用于新品发布、官网展示和社交媒体推广,不适用于缺少授权素材时仿造真实品牌标识、虚构产品功能或制作长篇剧情影片。
author-en: MiniMax Hub
author-cn: MiniMax Hub
source: official-featured
@@ -0,0 +1,44 @@
---
name: co-op-game-intro-generator
description: 面向想制作双人合作游戏主菜单或开场动画的用户。用户需提供两位玩家名称、游戏名称、目标视觉风格,并可上传角色参考图。Skill 会先锁定人物身份特征,基于固定菜单布局生成与色彩、按钮、图标和字体联动的确认首图;用户确认后,回填角色、界面文案和事件节奏,生成完整视频提示词并制作最终开场视频。最终输出一条包含双人角色、玩家信息卡和主菜单交互动效的游戏开场视频。适用于合作游戏概念展示、角色化菜单和社交内容,不适用于完整可玩游戏开发、复杂多页面 UI、精确品牌标识复刻或无角色的通用片头。
---
# 双人游戏开场视频生成器
当用户想做一条双人游戏开场视频,并且希望先用一张图确认风格、人物和 UI 再生成最终视频时,使用这个 Skill。流程重点是:弹窗选择风格、收集玩家名/游戏名/角色图、按源提示词框架生成首图、确认后回填 H3 视频 prompt 并直接生成视频。
## STEP 1:询问视觉风格
先用弹窗式问题给出预设风格和自定义选项,让用户先选整体画风。用户选择的风格是最高优先级,会影响整体风格补充、色彩语言、背景纹理、人物画风、表情气质、穿搭风格、UI颜色、按钮图标和字体质感。
## STEP 2:收集玩家与游戏信息
让用户填写 PLAYER 1 名字、PLAYER 2 名字、游戏名字。若用户有角色图,则上传 PLAYER 1 / PLAYER 2 参考图;上传图只用于身份映射:可识别脸型轮廓、发型、眼镜、五官相对比例和个人特征。不要继承真人摄影质感、皮肤纹理、现实光影、相机质感或原图风格;人脸必须在保留身份锚点的前提下重绘为用户选择的视觉风格。
## STEP 3:生成 GPT 首图提示词
GPT 首图提示词必须采用“框架固定 + 风格动态填充 + 颜色联动”的写法:
1. **整体风格**:固定保留“游戏主菜单 UI、高品质游戏宣传海报、UI 与角色深度融合、现代商业游戏 UI 设计、极强视觉冲击力、画面简洁干净、避免过度装饰”;其他风格词根据用户选择补充。
2. **色彩设计**:固定使用“xx 色作为主体颜色、xx 色作为 UI 主体颜色、xx 色作为文字颜色、xx 色作为功能强调色、红色作为危险提示色、整体颜色控制在 5 种以内、高对比撞色、鲜明现代、xx 风格色彩语言”的结构;具体颜色根据用户风格推导。后续 UI、按钮、图标、字体段落必须与本段色彩保持一致。
3. **整体布局**:固定保留 16:9 横版构图、背景铺满、中央放置 x 位角色、UI 围绕角色不遮挡、左上玩家信息卡、右侧纵向菜单、底部警戒胶带、四角少量涂鸦、Z 字阅读路径、留白充足、层级清晰、视觉重心集中 Continue。
4. **背景设计**:固定使用“纯 xx 色背景、轻微 xx 纹理、大面积纯色留白、纹理仅作细节、避免大量脏污/密集划痕/过多泼墨/大面积油漆飞溅、边缘少量黑色喷漆、整体简洁现代高级、突出 UI 主体”的结构;xx 根据用户风格推导,其他背景补充也根据用户选择风格生成。
5. **人物风格**:严格按照用户选择的风格拓展,参考源提示词的文本维度,但不可用源默认风格压过用户选择。
6. **左侧人物 Character A**:五官信息参考角色图 1;表情、画风、穿搭根据用户选择风格分析并改写;可参考“卡其色短夹克、黑色内搭、黑色工装裤、棕色皮靴、探险家风格穿搭”等维度,但必须服务用户风格;动作固定为双腿交叉坐地、双手撑地、身体微微后仰、抬头望向镜头。
7. **右侧人物 Character B**:五官信息参考角色图 2;表情、画风、穿搭根据用户选择风格分析并改写;可参考“绿色厚夹克、酒红色羽绒内胆、白色内衬、黑色长裤、厚底皮靴”等维度,但必须服务用户风格;动作固定为双腿盘坐、双手自然放于腿前、身体略微前倾、正视镜头。
8. **灯光表现 Lighting**:固定保留顶部主光源、左上暖色或冷色补光、底部柔和环境反光、柔软阴影、明显接触阴影、轻微轮廓光、高质量 GI 全局光照、人物自然融入背景。冷暖选择根据用户风格推导。
9. **UI 整体设计 Game UI**:固定保留 Console Game Menu 风格、按钮统一尺寸、统一圆角矩形、略微倾斜布局、少量喷漆与滴落元素、现代简洁、贴纸化设计、突出可读性、避免过度装饰。按钮主体色、描边色、发光描边色必须匹配第 2 点色彩设计。
10. **UI 排版规则 Layout Rules**:保持不变:所有菜单按钮横向长条、统一宽高与圆角、按钮宽度适配文字、标题单行显示、严禁换行、严禁两行标题、文字水平居中、左右留白一致、垂直间距统一。
11. **按钮设计 Buttons**:根据第 2 点色彩设计修改颜色,可适当修改图标,但必须符合用户风格选择。Continue 是视觉中心、超大尺寸、最强高亮、鼠标悬停/点击状态;Start New Game 位于 Continue 上方;Settings 保持统一视觉语言;Exit Game 使用危险/退出提示,可保留红色危险点缀。
12. **玩家信息卡 Player Cards**:固定保留左上角 x 名玩家信息栏、不规则矩形卡片、描边、轻微破损、左侧 Logo/图标、右侧三级信息(PLAYER 名称、玩家昵称、READY 状态)、粗体无衬线字体、工业贴纸设计。卡片颜色、描边色、READY 颜色、Logo 风格根据第 2 点和用户风格生成。
13. **图标设计 Icon System**:图标风格根据用户选择动态生成,颜色匹配第 2 点色彩设计;排列规则和禁止要求保持不变:单排横向、严禁自动换行、严禁上下堆叠、每个 UI 模块最多一排图标、数量精简、统一尺寸、不抢视觉焦点。
14. **字体设计 Typography**:保持源框架:超粗无衬线、全部大写、类似 Anton/Impact/Burbank/Tungsten、字距紧凑、笔画粗壮、信息层级明确、单行排版、禁止自动换行、按钮宽度优先适配文字长度、任何情况不得出现两行菜单标题。字体颜色和纹理根据第 2 点色彩设计与用户风格匹配。
## STEP 4:生成首张确认图
先只出一张图给用户确认,不要直接进入视频。首图需要保留框架结构,同时让用户选择的风格在视觉上明确成立。
## STEP 5:等待用户确认
在用户确认首图或提出修改之前,不要生成视频。如果用户改了风格、名字、游戏名、角色形象或首图问题,就回到首图提示词阶段重写并再生成。
## STEP 6:回填视频 prompt,并用 Minimax H3 生成
首图确认后,再把确认结果回填进完整视频 prompt,包括风格、角色参考、玩家名字、游戏名字、UI 文案和负面约束,然后直接用 Minimax H3 生成最终视频。
## STEP 7:常见失败修正
如果文字不清楚,就减少画面中的文字量;如果角色串位,就强化名字、左右位置和颜色区分;如果脸型或发型漂移,就继续沿用上传角色图,并明确保留身份锚点、发型和服装锚点,同时保持人脸按用户选择的视觉风格渲染;如果画面没有体现用户风格,就重写整体风格、色彩设计、人物风格、背景纹理、UI颜色、按钮、图标和字体段落,而不是改动布局框架。
@@ -0,0 +1,54 @@
---
name: co-op-game-intro-generator
description: For users creating a two-player co-op game menu or opening animation. Users provide two player names, a game title, a target visual style, and optional character reference images. The Skill locks identity cues, generates an approval image from a fixed menu framework with coordinated color, buttons, icons, and typography, then uses the approved result to rebuild the character, UI-copy, and event timing instructions for the final video. It outputs a co-op game intro featuring two characters, player cards, and menu interaction motion. Best for game concepts, character-led menus, and social content; not for playable game development, complex multi-page UI, exact brand-logo replication, or generic character-free title sequences.
compatibility: Requires the MiniMax Hub agent (canvas workspace and MiniMax H3 generation); not portable to generic agent harnesses.
---
# Co-op Game Intro Generator
Use this Skill when the user wants a co-op game intro video and wants to confirm the visual direction with one image before generating the final H3 video. The workflow collects style, player names, game title, and optional character refs, then creates a framework-preserving confirmation image before generating the H3 video.
## Required References
These two templates are mandatory runtime inputs, not optional background notes:
- Use `references/h3-confirmation-image-template.md` when building the confirmation-image prompt in STEP 3 and generating the first confirmation image in STEP 4. Fill the template fields in order and do not skip framework, palette, UI, character, typography, layout, or negative-constraint fields.
- Use `references/h3-video-prompt-template.md` when refilling the final Minimax H3 video prompt in STEP 6. Fill the template from the approved confirmation image, player/game data, final UI copy, event timing, motion directions, and negative constraints.
If either template is unavailable, stop and report that the Skill package is incomplete instead of improvising a different prompt structure.
## STEP 1: Ask for visual style
Ask the user to choose a preset style or enter a custom style. This style has top priority and controls supplemental style language, palette language, background texture, character rendering, expression, outfit direction, UI colors, button/icon style, and typography texture.
## STEP 2: Collect player and game info
Collect PLAYER 1 name, PLAYER 2 name, and game title. If the user provides character images, use PLAYER 1 and PLAYER 2 refs only for identity mapping: recognizable face silhouette, hairstyle, glasses, relative facial proportions, and distinctive traits. Do not inherit photographic realism, skin texture, real-world lighting, camera quality, or the original image style; redesign the face into the selected visual style while preserving identity anchors.
## STEP 3: Build the GPT confirmation-image prompt
Load and follow `references/h3-confirmation-image-template.md` as the required prompt skeleton. Use a fixed framework + dynamic style fill + palette-linked prompt structure:
1. **Overall Style**: always preserve game main menu UI, high-quality game promo poster, deep UI-character integration, modern commercial game UI design, strong visual impact, clean composition, and avoid over-decoration. Add other style terms from the selected style.
2. **Color Palette**: always follow: xx as main color, xx as UI body color, xx as text color, xx as functional accent color, red as danger accent, palette within five colors, high-contrast color blocking, fresh modern look, and xx-style color language. All later UI/button/icon/type colors must match this palette.
3. **Composition**: preserve 16:9 landscape, full-frame background, x centered characters, UI around rather than blocking them, upper-left player card, right-side vertical menu, bottom caution tape, a few corner graffiti accents, Z reading path, enough negative space, clear hierarchy, and Continue as the visual focus.
4. **Background**: follow pure xx background, slight xx texture, large solid-color blank space, texture only as detail, avoid heavy dirt/dense scratches/too much ink splash/large paint splatter, a little black spray-paint edge detail, clean modern premium UI-first background. Derive xx and additions from the selected style.
5. **Character Style**: strictly expand according to the selected style. Use source text only as dimension guidance, not fixed style.
6. **Character A**: facial features from character image 1; expression, rendering, and outfit follow selected style; optional compatible outfit dimensions include khaki short jacket, black inner layer, black cargo pants, brown boots, adventurer outfit. Fixed action: cross-legged, hands on floor, slight backward lean, looking up.
7. **Character B**: facial features from character image 2; expression, rendering, and outfit follow selected style; optional compatible outfit dimensions include green thick jacket, burgundy padded lining, white inner shirt, black pants, thick boots. Fixed action: cross-legged, hands in front of legs, slight forward lean, looking at camera.
8. **Lighting**: preserve top main light, upper-left warm/cool fill, soft bottom ambient reflection, soft shadows, contact shadows, rim light, high-quality GI, and natural character-background integration. Choose warm/cool from style.
9. **Game UI**: preserve console game menu, unified button size, rounded rectangles, slight tilt, minimal spray/drip elements, modern clean sticker design, readability, and no over-decoration. Button body, outline, and glow colors must match the palette.
10. **Layout Rules**: preserve horizontal long buttons, unified width/height/radius, width adapted to text, single-line titles only, no wrapping, no two-line titles, centered text, consistent margins and spacing.
11. **Buttons**: adapt colors from palette and icons from selected style. Continue is the large visual-center highlighted button with hover/click state. Start New Game is above Continue. Settings remains visually consistent. Exit Game uses danger/exit cue with red accents allowed.
12. **Player Cards**: preserve upper-left x-player info cards, irregular rectangular card, outline, slight edge damage, left logo/icon, right three-level info (PLAYER label, nickname, READY), bold sans-serif, industrial sticker design. Card colors, outline, READY color, and logo style follow palette and selected style.
13. **Icon System**: derive icon style from selected style, colors from palette, and keep one-row-only, no wrapping, no stacking, at most one row per UI block, minimal quantity, unified size, never stealing focus.
14. **Typography**: preserve bold sans-serif, all caps, Anton/Impact/Burbank/Tungsten-like weight, tight tracking, heavy strokes, clear hierarchy, single-line typography, no wrapping, button width adapts to text, never two-line menu titles. Colors and texture follow palette and selected style.
## STEP 4: Generate the first confirmation image
Generate only one confirmation image from the filled `references/h3-confirmation-image-template.md`. Preserve the framework structure, while making the selected style visibly dominant.
## STEP 5: Wait for approval
Do not generate video until the user approves the image. If the user changes style, names, game title, identity, or image direction, return to the image prompt step.
## STEP 6: Refill video prompt and generate with Minimax H3
After approval, load `references/h3-video-prompt-template.md` and refill the final video prompt with confirmed style, character refs, player names, game title, UI text, event timing, motion directions, and negative constraints. Generate the final video with Minimax H3.
## STEP 7: Repair common failures
If text is unreadable, reduce on-screen text. If identities swap, strengthen names, positions, and colors. If faces drift, reuse uploaded refs and explicitly preserve identity anchors, hairstyle, and outfit anchors while keeping the face rendered in the selected visual style. If the selected style is weak, rewrite Overall Style, Color Palette, Character Style, Background, Game UI, Buttons, Icons, and Typography instead of changing layout framework.
@@ -0,0 +1,15 @@
display-name-zh: 双人游戏开场视频生成器
version: 0.1.5
tag-en: Creative & Experimental
tag-cn: 创意实验
complete-tags-en:
- Creative & Experimental / Creative Generation
complete-tags-cn:
- 创意实验 / 创作生成
summary-en: Create a co-op game intro from player details, character references, and an approved menu style.
summary-cn: 根据玩家信息、角色参考图和视觉风格,完成首图确认与界面动效设计,输出双人游戏开场视频。
desc-en: For users creating a two-player co-op game menu or opening animation. Users provide two player names, a game title, a target visual style, and optional character reference images. The Skill locks identity cues, generates an approval image from a fixed menu framework with coordinated color, buttons, icons, and typography, then uses the approved result to rebuild the character, UI-copy, and event timing instructions for the final video. It outputs a co-op game intro featuring two characters, player cards, and menu interaction motion. Best for game concepts, character-led menus, and social content; not for playable game development, complex multi-page UI, exact brand-logo replication, or generic character-free title sequences.
desc-cn: 面向想制作双人合作游戏主菜单或开场动画的用户。用户需提供两位玩家名称、游戏名称、目标视觉风格,并可上传角色参考图。Skill 会先锁定人物身份特征,基于固定菜单布局生成与色彩、按钮、图标和字体联动的确认首图;用户确认后,回填角色、界面文案和事件节奏,生成完整视频提示词并制作最终开场视频。最终输出一条包含双人角色、玩家信息卡和主菜单交互动效的游戏开场视频。适用于合作游戏概念展示、角色化菜单和社交内容,不适用于完整可玩游戏开发、复杂多页面 UI、精确品牌标识复刻或无角色的通用片头。
author-en: MiniMax Hub
author-cn: MiniMax Hub
source: official
@@ -0,0 +1,388 @@
# H3 Confirmation Image Template
## Core Rule
Generate a polished 16:9 landscape confirmation image for a cooperative game main menu.
The image consists of two independent systems:
• UI Framework (fixed)
• Visual Style (dynamic)
The UI Framework defines the layout, composition, information hierarchy, interaction logic, reading flow, menu structure, and readability. These elements must remain unchanged.
{visual_style} only controls the artistic appearance, including rendering medium, materials, color language, lighting mood, environment style, character design, costume language, UI materials, icon language, typography appearance, decorative motifs, and overall atmosphere.
Changing {visual_style} should feel like applying a different art direction to the same professionally designed game menu, never redesigning the interface itself.
If any default description conflicts with {visual_style}, prioritize {visual_style} while preserving the UI framework.
---
## Design Goal
Create a premium commercial console game main menu mockup in 16:9 landscape format. The image should combine the quality of a polished game promotional artwork with the clarity of a professional game UI. Characters and interface must feel naturally integrated, with strong visual hierarchy, excellent readability, clean composition, and a clear interaction focus.
---
## Overall Style
Follow the user-selected visual style: {visual_style}.
The selected style controls the rendering medium, illustration style, materials, palette language, environment, character appearance, outfit design, UI appearance, icon language, typography treatment, decorative elements, and atmosphere.
All visual elements must belong to one unified artistic language.
Do not mix incompatible visual styles within the same image.
The UI framework, layout, hierarchy, composition, and interaction logic must remain consistent regardless of style.
---
## Color Palette
Derive the entire color system from {visual_style}.
Use:
• {main_background_color} as the primary background color.
• {ui_body_color} as the main UI color.
• {text_color} as the primary text color.
• {functional_accent_color} as the interaction highlight color.
• Red only for warning, danger, or exit actions.
Limit the overall palette to five primary colors.
All buttons, player cards, icons, borders, highlights, decorative elements, typography, and visual effects must follow the same color system.
Avoid introducing unrelated colors outside the selected palette.
---
## Composition
Use a fixed 16:9 landscape layout.
Place {character_count} playable characters in the center as the visual subject.
The UI surrounds the characters without blocking them.
Place the player information cards in the upper-left corner.
Place the main menu vertically on the right.
Place a decorative horizontal strip along the bottom, adapting naturally to {visual_style} (such as ribbon, warning strip, wooden plank, stone slab, energy bar, scroll, etc.).
Keep a clear Z-shaped reading flow:
Player Cards → Characters → Menu → Continue Button.
The Continue button must always remain the primary visual focus.
Maintain generous negative space and excellent readability.
---
## Background
Generate a background naturally derived from {visual_style}.
The background supports the interface rather than competing with it.
Keep enough clean negative space behind all important UI elements.
Use subtle textures, environmental details, and decorative motifs only where appropriate for the selected style.
Avoid excessive visual noise, dense decoration, heavy dirt, scratches, paint splatter, or distracting textures.
The overall background should feel polished, premium, modern, and suitable for a commercial game menu.
---
## Character Style
All characters strictly follow {visual_style}.
The selected style determines:
• rendering medium
• body proportion
• facial stylization
• costume language
• material treatment
• lighting response
• animation style
• overall character quality
Only the character identity anchors and pose remain fixed. Facial rendering, eye design, nose simplification, mouth shape, skin treatment, expression style, and head proportion must be fully reinterpreted in {visual_style}.
No default character style may override {visual_style}.
---
## Character A / PLAYER 1
### Identity
Use character reference image 1 only for identity mapping.
The reference image donates identity anchors only: recognizable face silhouette, hairstyle, glasses if present, relative facial proportions, and distinctive personal traits.
It does not donate photographic realism, real skin texture, real-world lighting, camera quality, or the original image style.
Do not lose the character's identity anchors, but the face must be redesigned into {visual_style}.
All facial rendering, eye design, nose simplification, mouth shape, skin treatment, expression style, and head proportion must follow {visual_style}.
Do not swap identity with other characters.
### Appearance
Character rendering, expression, outfit, materials, and accessories are fully derived from {visual_style} while maintaining a recognizable playable-game-character silhouette.
### Action
Character A sits cross-legged on the ground.
Both hands support the body on the floor.
The upper body leans slightly backward.
Character A looks upward toward the camera.
---
## Character B / PLAYER 2
### Identity
Use character reference image 2 only for identity mapping.
The reference image donates identity anchors only: recognizable face silhouette, hairstyle, glasses if present, relative facial proportions, and distinctive personal traits.
It does not donate photographic realism, real skin texture, real-world lighting, camera quality, or the original image style.
Do not lose the character's identity anchors, but the face must be redesigned into {visual_style}.
All facial rendering, eye design, nose simplification, mouth shape, skin treatment, expression style, and head proportion must follow {visual_style}.
Do not swap identity with other characters.
### Appearance
Character rendering, expression, outfit, materials, and accessories are fully derived from {visual_style} while maintaining a recognizable playable-game-character silhouette.
### Action
Character B sits cross-legged on the ground.
Both hands rest naturally in front of the legs.
The upper body leans slightly forward.
Character B looks directly toward the camera.
---
## Lighting
Lighting follows {visual_style} while preserving clear readability.
Use a primary top light, secondary fill light, soft ambient bounce, contact shadows, gentle rim lighting, and high-quality global illumination where appropriate.
Faces should remain clearly visible.
Characters must integrate naturally into the environment while maintaining strong separation from the background.
---
## Game UI
Create a professional commercial console game main menu.
The menu structure, interaction hierarchy, spacing, composition, reading flow, and information architecture remain fixed.
{visual_style} only changes the artistic appearance of the interface.
The selected style may affect:
• UI materials
• surface textures
• border treatment
• corner shape
• decorative motifs
• color rendering
• lighting treatment
• icon language
• typography appearance
• interaction effects
It must NOT change:
• menu hierarchy
• button order
• button placement
• player-card placement
• composition
• interaction logic
• information architecture
• readability
The entire interface should feel like one professionally designed game UI system rather than separate elements assembled together.
---
## UI Visual Language
Derive the complete UI appearance from {visual_style}.
The selected style determines:
• button materials
• border language
• panel materials
• decorative details
• edge treatment
• visual weight
• highlight style
• interaction feedback
• shadow treatment
• texture language
Possible materials include wood, stone, paper, leather, cloth, glass, crystal, hologram, neon, metal, ceramic, carved surfaces, painted surfaces, magical energy, pixel blocks, watercolor paper, ink brush, comic graphics or any other material naturally belonging to {visual_style}.
Every UI element must share the same material language and artistic style.
---
## Button Layout Rules
Place all menu buttons vertically on the right side.
Maintain a consistent vertical rhythm.
Every button uses the same design family.
Button width automatically adapts to text length.
Keep all menu titles on a single line.
No automatic wrapping.
No stacked text.
Text remains horizontally centered.
Keep consistent padding.
Maintain equal spacing between buttons.
Avoid visual clutter.
---
## Continue Button
The Continue button is always the primary visual focus.
Make it visually dominant using the emphasis method naturally belonging to {visual_style}.
Possible emphasis methods include:
• larger size
• stronger contrast
• brighter color
• unique material
• distinctive border
• lighting
• glow
• engraving
• magical energy
• holographic effect
• embossed frame
• animated visual cue
Do not force glow or neon if they conflict with {visual_style}.
The Continue button should immediately attract the player's attention.
---
## Start New Game Button
Place above Continue.
Use the standard secondary button appearance.
Keep the visual language consistent with the selected style.
Add one small style-appropriate icon if suitable.
Maintain single-line text.
---
## Settings Button
Use the standard secondary button appearance.
Use one settings icon naturally derived from {visual_style}.
Maintain consistent proportions, spacing, and hierarchy.
---
## Exit Game Button
Use the standard secondary button appearance.
Use red only as the danger indicator.
Red should remain an accent rather than the dominant color.
Add one style-appropriate exit icon if suitable.
---
## Player Cards
Create an upper-left player information system for {character_count} players.
The appearance follows {visual_style}.
Card materials, borders, corners, textures, shadows, decorations, and icons are automatically derived from the selected style.
Each player card contains:
• Player label
• Player nickname
• READY status
Player names:
PLAYER 1
{player1_name}
PLAYER 2
{player2_name}
READY always uses the functional accent color.
The cards should remain highly readable and visually connected to the rest of the interface.
---
## Icon System
Derive all icons from {visual_style}.
Use one unified icon family throughout the entire interface.
Possible icon languages include:
• minimal outline
• solid filled
• geometric
• comic
• pixel
• engraved
• carved
• painted
• brush
• rune
• glyph
• holographic
• crystal
• paper-cut
• origami
• embossed
• low-poly
• flat design
• stylized 3D
All icons must share:
• the same artistic style
• the same material language
• the same perspective
• the same lighting treatment
• the same stroke language
• the same color hierarchy
Icons should remain simple, recognizable, and functional.
Each interface block may contain at most one horizontal row of icons.
No wrapping.
No vertical stacking.
Icons should support the interface rather than dominate it.
---
## Typography
Typography hierarchy remains fixed.
Typography appearance follows {visual_style}.
The selected style controls:
• font family
• decorative treatment
• texture
• material
• outline
• lighting
• shadow
• wear pattern
Maintain:
• excellent readability
• strong visual hierarchy
• single-line menu buttons
• consistent spacing
• high contrast
Button width should always adapt to text length.
Never allow menu titles to wrap onto multiple lines.
---
## Game Title
Display the game title:
{game_title}
The title should become the primary branding element of the menu.
Its material, decorative style, typography, and rendering all follow {visual_style}.
Avoid imitating existing commercial game logos.
Create an original title treatment inspired by the selected style.
Keep the title highly readable.
Avoid random symbols or meaningless letters.
---
## Interaction States
Represent the interface as an active game menu.
Buttons may display appropriate interaction feedback naturally belonging to {visual_style}.
Possible interaction feedback includes:
• hover
• pressed state
• magical glow
• energy pulse
• holographic highlight
• carved illumination
• ink spread
• brush bloom
• metallic reflection
• paper emboss
• animated border
• pixel flashing
Only use interaction effects that naturally belong to the selected style.
Avoid forcing modern UI effects into incompatible artistic styles.
---
## Rendering
Generate a polished AAA-quality commercial game menu artwork.
Characters, environment, UI, typography, icons, and decorative elements should feel like one cohesive production.
Maintain premium rendering quality.
Strong visual hierarchy.
Professional commercial presentation.
High readability.
Balanced composition.
Clean finish.
---
## Negative Constraints
No split screen.
No duplicated characters.
No identity swap.
No loss of identity anchors. No photorealistic face unless {visual_style} explicitly requires photorealism. No realistic photo face pasted onto a stylized body. No mismatched face style.
No hairstyle changes.
No unreadable UI.
No distorted buttons.
No overlapping interface.
No random text.
No random symbols.
No watermark.
No copied game logo.
No branded interface.
No inconsistent icon styles.
No mixed UI materials.
No excessive decoration.
No overwhelming background.
No visual clutter.
No broken layout.
No low readability.
---
## Medium Lock
Render as a polished commercial game key art combined with a professional console game main menu mockup.
The interface remains fixed.
Only the artistic appearance changes according to {visual_style}.
@@ -0,0 +1,133 @@
# H3 Video Prompt Template
## Prompt principle
Use the same method as the GPT confirmation-image prompt:
**Fixed video event framework + dynamic user-style fill + locked character identity + palette-linked UI system.**
The video prompt must not blindly preserve the source prompt's default style. It must preserve the timeline, UI events, player positions, equipment logic, and text structure, while dynamically rewriting visual treatment, palette, character rendering, lighting, UI surface, icons, and city-world style according to the user's selected style.
## Priority order
1. User-selected style: {visual_style}
2. Confirmed image / UI reference: {ui_ref}
3. PLAYER 1 identity reference: {player1_ref}
4. PLAYER 2 identity reference: {player2_ref}
5. Height/body comparison reference: {height_ref}
6. Fixed source video event framework
## Reference roles
- {ui_ref}: confirmed first image. Use it to lock UI layout, menu hierarchy, color system, typography scale, button structure, character-game integration, and overall composition logic.
- {player1_ref}: PLAYER 1 identity anchor. Lock exact face, hairstyle, glasses if present, facial proportions, body identity, and nickname mapping to {player1_name}.
- {player2_ref}: PLAYER 2 identity anchor. Lock exact face, hairstyle, facial proportions, body identity, and nickname mapping to {player2_name}.
- {height_ref}: body comparison anchor. Lock the visible contrast between the two players and prevent identical body proportions.
## Global style baseline
整体画风必须优先服从用户选择:{visual_style}。
保留固定基准:游戏主菜单 UI、高品质游戏宣传片质感、UI 与角色深度融合、现代商业游戏 UI 设计、极强视觉冲击力、画面简洁干净、避免过度装饰。
其余视觉层面根据 {visual_style} 动态拓展:角色画风、表情气质、服装语言、色彩系统、灯光冷暖、UI材质、按钮图标、字体质感、城市加载后的世界风格。
## Palette system
根据 {visual_style} 推导视频全程色彩,但必须保持 UI 色彩联动:
- xx 色作为主体背景 / 世界主色
- xx 色作为 UI 主体颜色
- xx 色作为文字颜色
- xx 色作为功能强调色
- 红色作为危险/退出/警示提示色
- 整体颜色控制在 5 种以内
- 高对比撞色、鲜明现代、符合 {visual_style} 的色彩语言
视频中的菜单、装备面板、玩家卡、按钮、HUD、加载条、图标和文字都必须沿用同一色彩系统,不得随机新增无关颜色。
## Character identity and style lock
PLAYER 1:五官信息参考 {player1_ref},必须保留脸、人脸比例、发型、眼镜(如有)、个人身份和 {player1_name} 昵称对应关系。表情、画风、服装和机械装备的视觉处理根据 {visual_style} 动态优化。PLAYER 1 始终位于左侧,偏高挑/修长/敏捷,装备色为功能强调色,机械爪轻量、纤细、灵活。
PLAYER 2:五官信息参考 {player2_ref},必须保留脸、人脸比例、发型、个人身份和 {player2_name} 昵称对应关系。表情、画风、服装和机械装备的视觉处理根据 {visual_style} 动态优化。PLAYER 2 始终位于右侧,偏矮壮/宽厚/力量型,装备色为琥珀红或与危险/力量提示一致的暖色,机械拳厚重、宽大、有重量。
禁止角色身份交换、脸部互相融合、昵称交换、两人体型趋同。
## Fixed timeline framework
### [0秒–2秒] — 双人主菜单
景别/机位:高角度俯拍大全景,延续 {ui_ref} 的构图逻辑。镜头从上方轻微下压并缓慢推进。
画面内容:PLAYER 1({player1_name})和 PLAYER 2({player2_name})并排坐在画面中央,PLAYER 1 在左,PLAYER 2 在右,抬头看向摄像机。两人只有自然呼吸、眨眼和轻微身体动作。
UI结构:左上角玩家资料卡准确显示:
“PLAYER 1”
“{player1_name}”
“READY”
顶部中央或左上双卡系统中的第二张资料卡准确显示:
“PLAYER 2”
“{player2_name}”
“READY”
右侧纵向菜单准确显示:
“START NEW GAME”
“CONTINUE”
“SETTINGS”
“EXIT GAME”
“CONTINUE”是视觉中心和主高亮按钮。
动态风格填充:主菜单背景、地面纹理、按钮形态、图标、字体、边框、辉光和贴纸质感全部根据 {visual_style} 优化,但保留 {ui_ref} 的布局和层级。
声音:菜单环境音、轻微 UI hover 声、点击前的低频电子氛围。若生成无音频,则忽略声音执行。
### [2秒–4秒] — PLAYER 1 右臂配置
景别/机位:中景,镜头从主菜单平滑推近 PLAYER 1 的右臂,PLAYER 2 仍在背景中可见,不消失、不变形。
UI结构:右侧主菜单收缩滑出;带功能强调色识别线的 UI 面板从左侧滑入,准确显示:“PLAYER 1”“RIGHT ARM EQUIPMENT”。装备列表中先高亮“PHANTOM GRIP”,随后选区移动至“CHRONOS CLAW”。
动作:PLAYER 1 的右袖口自动打开,轻量机械结构从前臂下方展开。手指分开,修长爪状指节滑入并逐一锁定,内部短暂露出精细线路、微型活塞和金属连接件。配置完成后,功能强调色 LED 依次亮起。
动态风格填充:机械结构、UI面板材质、图标、线条、锁定动画和光效根据 {visual_style} 优化,但必须轻量、精密、灵活,符合 PLAYER 1 身形,不改变脸、发型和服装主体。
声音:精密机械展开声、轻快 UI 切换音、细小锁定声。
### [4秒–7秒] — PLAYER 2 重型手臂配置
景别/机位:中景,摄影机沿两人之间平滑横移并绕向 PLAYER 2 左侧。PLAYER 1 留在背景中,轻轻观察自己完成配置的机械手。
UI结构:新的暖色/琥珀红识别 UI 滑入,准确显示:“PLAYER 2”“ARMAMENT CUSTOMIZATION”。面板以网格形式展示:
“HAND”
“FOREARM”
“ELBOW”
“UPPER ARM”
选区快速但清晰地在四个组件之间切换。
动作:PLAYER 2 左臂外套袖口分段打开,厚重前臂护板向外弹开,旧组件脱离,新型装甲沿导轨滑入;肘关节替换为厚实机械轴承,宽大的机械手重新组合并锁定。更换过程中短暂露出粗壮线路、液压活塞和深色金属骨架。每个部件锁定时亮起低调暖色指示灯。
动态风格填充:重型机械臂、UI面板、组件图标、材质、光效根据 {visual_style} 优化,但必须厚重、宽大、有力量感,与 PLAYER 1 的轻量机械爪形成清晰对比。
声音:低沉电机声、重型机械扣合声、厚重锁定反馈。
### [7秒–8.5秒] — 双人确认配置
景别/机位:中景拉回,镜头平滑回到双人构图,PLAYER 1 左侧,PLAYER 2 右侧。
UI结构:两组装备面板向画面中央汇合,形成共享按钮,准确显示:“CONFIRM CONFIG”。按钮边框、辉光、图标和贴纸质感根据 {visual_style} 优化,但层级清晰、文字可读。
动作:光标点击按钮,功能强调色能量脉冲流过 PLAYER 1 的机械爪,暖色/琥珀红能量脉冲流过 PLAYER 2 的机械拳。所有 UI 面板快速向内收缩并消失。两人同时解开交叉的双腿并调整坐姿:PLAYER 1 轻盈抬起单膝,修长机械爪依次活动手指;PLAYER 2 一只脚稳稳踩地,厚重机械拳缓慢握紧。
声音:确认提示音、双色能量脉冲、UI 收缩声。
### [8.5秒–10秒] — 双人世界加载
景别/机位:全景,底部共享加载条出现。
UI结构:加载条准确显示:“LOADING”。进度从 0% 快速填充至 100%。左半段使用 PLAYER 1 的功能强调色,右半段使用 PLAYER 2 的暖色/力量色。HUD 和加载条的形态、边框、纹理、字体根据 {visual_style} 优化,但必须清晰可读。
环境转化:黄色平面环境连续转化为游戏世界。警戒线变成真实街道道路标记和施工围挡;墨迹/纹理化为潮湿路面反光或符合 {visual_style} 的地面光影;平面涂鸦扩展成建筑墙面的喷绘;黑色背景区域形成城市阴影和深巷入口。
关键约束:转化必须连续自然,不使用硬切,不用烟雾遮挡,不改变两位角色身份。
声音:加载上升音、环境从菜单氛围过渡到城市氛围。
### [10秒–15秒] — 双人进入游戏世界
景别/机位:大全景转第三人称跟拍。加载 100% 的瞬间,两名角色同时起身。摄影机平滑下降并绕到两人身后,变成稳定第三人称双人合作视角。
世界风格:完整游戏世界根据 {visual_style} 动态生成。保留游戏开场进入世界的结构:密集建筑、道路、标志灯、招牌、人群或动态背景、快速经过的载具、电线、工业管道、远处城市天际线。具体世界材质、建筑形态、招牌设计、光影和动效必须服从 {visual_style},不要让源默认赛博朋克风格压过用户选择,除非用户选择的是赛博朋克。
角色关系:清楚展示两人的背影和体型差异:PLAYER 1 左侧,高挑修长,轻量机械爪自然垂下;PLAYER 2 右侧,矮壮宽厚,重型机械拳微微抬起。PLAYER 1 率先迈步,PLAYER 2 紧随其后,两人并肩进入街道。
HUD结构:HUD 淡入。右上角出现小地图;左下角出现两组独立状态栏:“{player1_name}”“{player2_name}”。{player1_name} 状态栏使用 PLAYER 1 功能强调色,{player2_name} 状态栏使用 PLAYER 2 暖色/力量色。前方街道中央出现共享任务标记。
声音:城市环境氛围、远处载具声、脚步声、HUD 淡入提示音。
## Negative constraints
No third player, no female character unless explicitly requested by user, no duplicated character, no character swapping, no username swapping, no merged bodies, no identical body proportions, no changing faces, no changing hairstyles, no changing identity, no missing PLAYER 1, no missing PLAYER 2, no split screen, no hard cuts, no random camera shake, no floating body parts, no gruesome dismemberment, no weapon, no excessive neon unless selected by user style, no purple background unless selected by user style, no unreadable UI, no random letters, no misspelled usernames, no extra menu options, no official game logo, no copied branded interface, no watermark.
+35
View File
@@ -0,0 +1,35 @@
---
name: h3-prompt-writing
description: Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA. Use when rewriting multimodal requests into H3 prompt structures, composing integrated_multimodal_description, overall_soundscape, and non_diegetic_music, aligning keyframes, or defining reference labels for images, videos, and audio.
compatibility: Portable to any agent that can read local files — no external API calls, MiniMax Hub tools, or proprietary runtime required. The agents/openai.yaml file only adds optional ChatGPT/Codex UI metadata; it does not restrict the skill to OpenAI agents.
---
# H3 Prompt Writing
## Workflow
1. Identify the input mode: T2VA, I2VA, FL2VA, L2VA, or full-reference Ref2VA.
2. For base text/keyframe modes, read `references/base-en.txt` and follow its final prompt structure.
3. For full-reference mode, read `references/ref-en.txt` and follow its six-section rewrite format.
4. Preserve the exact field names, section order, labels, and timing notation from the selected guide.
## Base Modes
- T2VA: build the full audiovisual timeline from text.
- I2VA: start from the first frame and develop forward from it.
- FL2VA: describe the continuous path between the first and last frames.
- L2VA: infer a plausible opening and converge to the supplied last frame.
Use `integrated_multimodal_description`, `overall_soundscape`, and `non_diegetic_music` in the order shown in `references/base-en.txt`.
## Full-Reference Mode
Ref2VA rewrites use `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, `overall_soundscape`, and `non_diegetic_music` in that order. Reference labels stay consistent across all sections.
Read `references/ref-en.txt` for label rules, retention analysis, and complete examples.
## Output Rules
- Write rewrite sections in English; preserve dialogue, lyrics, and visible scene text in their original language.
- Describe each shot by composition, subjects, environment, actions, camera, sound, and the exact point where referenced content appears.
- Avoid plot summaries, unresolved reference labels, and timing that does not match the requested duration.
@@ -0,0 +1,4 @@
interface:
display_name: "MiniMax H3 Prompt Writing"
short_description: "Write H3 base and full-reference video prompts"
default_prompt: "Use $h3-prompt-writing to rewrite this multimodal request into a MiniMax H3 generation prompt."
@@ -0,0 +1,222 @@
# Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
## 1. Task Overview
- **T2VA**: Builds a complete audiovisual timeline from text.
- **I2VA**: T2VA body + first-frame instruction + a visual path that develops forward from the first frame.
- **FL2VA**: T2VA body + first-and-last-frame instruction + a continuous path from the first frame to the last frame.
- **L2VA**: T2VA body + last-frame instruction + a path that converges from a plausible preceding state to the last frame.
## 2. Final Prompt Structure
### 2.1 Part One Is the Instruction
**T2VA** has no image-alignment instruction and begins directly with the three core fields.
**I2VA** always uses:
```text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
```
**FL2VA** always uses:
```text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
```
**L2VA** always uses:
```text
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
```
Here, `N` is the index of the actual final shot, and `S.SS` is the effective video duration formatted to exactly two decimal places. The instruction must be the first line of the final prompt, followed by one blank line before the core fields.
### 2.2 Part Two Contains the Three Core Fields
```text
integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...
```
- **integrated_multimodal_description**: Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.
- **overall_soundscape**: Summarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video.
- **non_diegetic_music**: Describes background music that the characters cannot hear and only the audience can hear.
## 3. How to Incorporate Keyframes into the Multimodal Description
### 3.1 I2VA: Begin from the Image and Develop Forward
`<Picture 1>` is the actual first frame of the video at 0.00 seconds and belongs to `[Shot 1]`. The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action. Character identity, clothing, colors, key objects, and spatial relationships should remain consistent.
Recommended structure: **first-frame anchor → action onset → continuous development → result or reaction**.
### 3.2 FL2VA: Describe the Path Between the First and Last Frames
Picture 1 is the opening, and Picture 2 is the ending. Focus on how the subject moves, how poses change, how objects are manipulated, how the composition evolves, and how the scene or lighting transitions.
FL2VA generally favors a single shot so the model can interpolate continuously from the first frame to the last frame. Use multiple shots only when they are explicitly specified. The last frame must be reached by the final `[Shot N]` at the end of the video.
Recommended structure: **first-frame state → observable intermediate changes → progressively narrowing differences → last-frame state**.
### 3.3 L2VA: Infer the Opening and Land on the Image at the End
`<Picture 1>` is the final frame of the video and belongs to the last `[Shot N]`; it does not inherently belong to Shot 1. Infer a plausible earlier state from the user's intent and the last frame, then describe how the characters, objects, camera, and scene gradually approach the reference image.
Recommended structure: **plausible preceding state → explicit action and transition path → gradual convergence in the final shot → last-frame landing**.
## 4. How to Write the Three Shared Core Sections
### 4.1 Develop the Multimodal Description Along the Timeline
`integrated_multimodal_description` is the main body of the rewritten prompt. Every detail should correspond to something visible or audible: visual style, initial composition, subject appearance and position, scene and key props, actions and reactions, shot changes, spoken language, and synchronized diegetic sound.
At the beginning of `[Shot 1]`, state the overall style and initial composition. Common styles include `Cinematic`, `live-action`, `2D-animated`, `3D CG`, `claymation`, `watercolor`, and `vintage film`. For keyframe tasks, derive the style from the reference image; for T2VA, select it from the user's text.
```text
[Shot 1] Live-action, cinematic, a medium-wide shot frames...
```
### 4.2 Shots and Cuts
Do not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration:
```text
[Shot 2] At 00:03.500, the camera cuts to...
```
For ordinary cuts, use `the camera cuts to`, `the shot cuts to`, `the shot transitions to`, `the shot changes to`, or `the shot switches to`. When explicitly requested by the user, cross-dissolve, fade, or wipe may also be used. A cut should introduce new information about the subject, space, state, viewpoint, or time. If only the distance or a slight angle needs to change, prefer camera motion.
### 4.3 Camera Motion: Motion Type + Amplitude + Speed
A complete camera-motion expression has three dimensions: the **motion type** defines how the camera moves, **amplitude** defines the range of compositional change, and **speed** defines the pacing of that change. Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted.
| Dimension | Available Expression | Description |
|-|-|-|
| Motion type | `Zoom In / Zoom Out` | The focal length changes while the camera body remains stationary |
| Motion type | `Push In / Pull Out` | The camera moves forward / backward |
| Motion type | `Pan Left / Pan Right` | The camera remains in place while the lens pivots horizontally |
| Motion type | `Truck Left / Truck Right` | The camera translates horizontally |
| Motion type | `Tilt Up / Tilt Down` | The camera remains in place while the lens pivots vertically |
| Motion type | `Pedestal Up / Pedestal Down` | The entire camera moves upward / downward |
| Motion type | `Arc Shot` | The camera moves in an arc around the subject |
| Motion type | `Tracking Shot` | The camera follows a moving subject |
| Motion type | `Static Shot` | The camera position and lens remain still |
| Motion type | `Shake Slightly / Shake Strongly` | Slight / strong camera shake |
| Motion type | `POV` | The subject's point of view |
| Motion type | `Roll Clockwise / Roll Counterclockwise` | The camera rolls clockwise / counterclockwise around the lens axis |
| Amplitude | `with small amplitude` | Small-range change |
| Amplitude | `with large amplitude` | Large-range change |
| Speed | `at slow speed` | Slow movement |
| Speed | `at fast speed` | Fast movement |
Camera motion should be written as a natural English action within the shot, rather than stacked as separate labels at the end of a sentence:
```text
The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The camera pans right with large amplitude at fast speed, revealing the open doorway.
The camera holds a static shot as the runner exits the frame.
```
### 4.4 Speakers, Dialogue, and Singing
Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as `(S1)` and `(S2)`. When multiple already-numbered speakers speak or sing together, use a compound ID such as `(S1,S2)`. A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.
When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside `<d>`. Inside `<d>`, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.
```text
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>
```
For voiceover, use the exact phrase `says in an off-screen voiceover`. Immediately after every voiceover `<d>` block, state that the corresponding on-screen character's lips remain closed:
```text
The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.
```
When the same line of dialogue or lyrics crosses a cut, use `<scenetrans>` at the connecting points in both parts and explicitly state that the audio continues across the cut. Use `<cutoff>` when speech is truncated by the end of the video. Continuity may be expressed with `continues seamlessly across the cut`, `continues uninterrupted into the next shot`, `carries over from the previous shot`, or `remains audible across the transition`.
### 4.5 On-Screen Text
Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation.
```text
A red neon sign reading "营业中" glows above the doorway.
```
### 4.6 overall_soundscape
Use 1–4 English sentences in one continuous paragraph to summarize the ambient sound, physical action sounds, and non-verbal human sounds across the full video, such as wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter, or panting. Dialogue, singing, and diegetic music already belong in the multimodal description and should not be repeated here. Use `N/A` only when the user explicitly requests complete silence throughout the video.
```text
overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.
```
### 4.7 non_diegetic_music
Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear. Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use `N/A` when there is no non-diegetic music.
```text
non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.
```
## 5. Cases
### Case 1: T2VA
With no reference image, construct the complete timeline directly from the text. You may add scene, character, action, and sound details that remain consistent with the user's intent.
```text
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.
overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.
non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.
```
### Case 2: I2VA
Write the first-frame instruction first, then use the subject, composition, and scene in Picture 1 as the starting point of Shot 1 before describing how the scene continues to develop.
```text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.
overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.
non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
```
### Case 3: FL2VA
The two images anchor the opening and ending respectively. The body should not repeat two static image descriptions; instead, it should supply the motion path that connects them. The following example is an eight-second single shot.
```text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.
overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.
non_diegetic_music: N/A
```
### Case 4: L2VA
The image anchors only the final moment. First establish a compatible earlier state, then let the actions, object states, and composition gradually land on Picture 1 in the final shot. The following example is a six-second single shot.
```text
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.
overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.
non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.
```
@@ -0,0 +1,341 @@
# Full-Reference Mode Rewrite Output Format Guide
This guide explains how rewrite outputs are organized and written in full-reference mode.
Write all six rewrite sections in English. Preserve the original language only for dialogue and lyrics inside `<d>` and for text visibly present in the scene.
**Description detail:** Make `detailed_description` as detailed and explicit as possible. For each shot, clearly establish the current composition, subject appearance and position, environment and lighting, actions and state changes, camera movement, current sound, and the points where referenced content actually appears or takes effect. Avoid reducing the description to a plot summary or a list of reference relationships.
> The basic formats for shots, camera movement, speakers, dialogue, and ordinary sound are shared with the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA). This guide focuses on the reference labels, analysis sections, and format differences specific to full-reference mode.
## 1. Overall Structure
A complete rewrite output consists of six sections in the following order:
| Section | Purpose |
| --- | --- |
| `subject_definitions` | Defines referenced content and its reference labels |
| `summary` | Summarizes the task type, target video, and main reference relationships |
| `retention_analysis` | Describes how referenced content is preserved, transferred, or reused |
| `detailed_description` | Describes visuals, actions, shots, sound, and dialogue in playback order |
| `overall_soundscape` | Summarizes ambience and physical sounds |
| `non_diegetic_music` | Describes background music audible only to the audience |
## 2. Reference Labels and Definitions (`subject_definitions`)
Full-reference rewrites use four types of labels to identify the source and role of referenced content:
| Label | Meaning |
| --- | --- |
| `<Subject N>` | Visible content abstracted from reference assets that can be reused or modified in the target video |
| `<Picture N>` | A reference image used as a concrete target frame or shot-planning anchor |
| `<Video N>` | A reference video that provides an editing source, continuation starting point, or whole-video temporal structure |
| `<Audio N>` | An audio signal that is copied or referenced |
> Once a reference label is assigned to a piece of content, it keeps the same meaning across `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, and the audio sections.
`subject_definitions` defines each piece of referenced content that must be tracked separately later, such as a person, an environment, a source video's structure, or an audio track. Give each item its own line and explain what its label denotes, its reference role, and the main features to follow; name the corresponding source asset when its provenance needs to be made explicit. If `<Picture N>` or `<Video N>` only identifies the source of another referenced item and will not be analyzed or used separately later, cite it inside that item's definition without adding a separate line. `retention_analysis` records where each referenced item appears and whether it is fully preserved, partially preserved, transferred, or reused.
### 2.1 `<Subject N>`
`<Subject N>` is used for reusable visible content, including:
- People, animals, or objects
- Scenes, backgrounds, or environments
- Clothing, props, interfaces, or visual effects
- Styles, actions, expressions, or poses
It represents a content unit that will actually be used in the target video, rather than the source file itself. One subject may be defined by multiple reference assets, and one reference asset may provide multiple subjects.
```text
<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.
```
When the same subject comes from multiple assets, combine the sources and state what each asset provides:
```text
<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.
```
### 2.2 `<Picture N>`
Use a standalone `<Picture N>` when the reference image itself serves as a shot's first frame, keyframe, last frame, edited keyframe, or composition anchor:
```text
<Picture 2> is the first frame of [Shot 1], showing a woman seated beside a café window.
```
If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry. Instead, cite the image source inside the corresponding `<Subject N>` definition.
When an image acts as a storyboard or shot-planning reference, state which shots it maps to and what planning information it provides:
```text
<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.
```
### 2.3 `<Video N>`
`<Video N>` is reserved for whole-video relationships, such as:
- Editing an original video
- Continuing from the end of an original video
- Referencing the original video's camera movement, cuts, rhythm, or temporal structure
```text
<Video 1> is the source video for the target video edit.
```
If a person, object, scene, action, or effect from a reference video is reused as visible content, it still belongs under `<Subject N>`. `<Video N>` identifies the asset or structural source and does not replace subject labels.
### 2.4 `<Audio N>`
`<Audio N>` represents a standalone audio asset or an enabled synchronized audio track from a reference video. Common uses include:
- Copying all or part of an audio signal
- Referencing a background-music style
- Referencing a speaker's voice timbre and delivery
- Using dialogue, lyrics, or sound effects from the original audio
- Referencing beat, rhythm, or audio continuity
When an `<Audio N>` explicitly corresponds to a target speaker, reuse that speaker's global ID in the definition: write `<Subject N> (Sx)` when the speaker maps to a defined subject, or use a stable voice description followed by `(Sx)` otherwise. The ID comes from the target video's global speaker order and is not independently assigned or renumbered in the audio definition. See Section 5.4 for the speaker-numbering rules:
```text
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
```
When one audio asset serves multiple roles, describe those roles in one natural sentence rather than creating additional subsections.
### 2.5 Visual and Audio Tracks from the Same Reference Video
`<Video N>` and `<Audio N>` are numbered independently. Each index indicates only the label's order within its own category and does not encode a pairing between the two categories. The same reference video may therefore correspond to `<Video 1>` and `<Audio 2>`; different indices do not prevent them from coming from the same source asset.
An ordinary reference video does not create `<Audio N>` merely because the file contains sound.
An `<Audio N>` definition primarily states the audio's role and does not have to name the `<Video N>` it comes from. State the shared source only when needed to remove provenance ambiguity, for example:
```text
<Video 1> is the source video for the target video edit.
<Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.
```
## 3. `summary`
This section uses one short English paragraph to summarize the target video and its reference relationships. It begins with a square-bracketed task-type prefix:
```text
[reference generation] ...
[video editing + reference generation + audio reuse] ...
```
Choose task types according to the actual role each reference asset plays in the target video:
| Task type | When to use it |
| --- | --- |
| `keyframe completion` | An image serves as the target video's first frame, keyframe, last frame, edited keyframe, or another concrete frame anchor |
| `reference generation` | An image, video, or audio asset provides generation guidance for a character, scene, style, action, camera movement, storyboard, and so on, without serving as a concrete frame or as the source video being edited or continued |
| `video editing` | An existing source video is directly modified; editing an image or generating between still keyframes does not belong to this type |
| `video continuation` | New content continues, extends, resumes, or transitions from an existing source video |
| `audio reuse` | The same audio signal is reused in full or in part |
| `audio reference` | The audio signal is not copied directly; only its music style, timbre, dialogue or lyric content, sound-effect texture, beat, or continuity is referenced |
When a task satisfies multiple relationships, combine the task types with ` + ` and do not repeat a type. For example, continuing from a source video while using an image as the last frame is written as `[video continuation + keyframe completion]`. Editing a source video while retaining its original audio may be written as `[video editing + audio reuse]`.
The mere presence of video or audio does not automatically create a corresponding task type. If a reference video provides only camera movement, cuts, or rhythm, it normally belongs to `reference generation`. Use `video editing` or `video continuation` only when that video is directly edited or continued.
When editing a source video, use `audio reuse` as well if its original audio remains audible. When continuing a source video without directly copying the audio signal, use `audio reference` if the new audio only continues the original track's audible characteristics.
The summary uses the previously defined `<Subject N>`, `<Picture N>`, `<Video N>`, and `<Audio N>` labels to describe the main subjects, shot flow, and roles of the reference assets. Do not introduce new reference labels in this section.
For video-editing tasks, begin the summary after the task-type prefix with:
```text
The target video is an edited version of <Video 1>.
```
## 4. `retention_analysis`
This section describes how each piece of referenced content is preserved, transferred, copied, or referenced in the target video. Use one line for each reference label and preserve the meaning established in `subject_definitions`.
### 4.1 Visible Content
`<Subject N>`, `<Picture N>`, and `<Video N>` use the following relationship markers. These markers are fixed English values in the output format:
| Relationship marker | Meaning |
| --- | --- |
| `fully_preserved` | The defined role of the referenced content is fully preserved |
| `partially_preserved` | The referenced content is still used, but some defined characteristics are changed or only partially retained |
| `attribute_transfer` | Referenced characteristics are transferred to a different identifiable target subject |
| `weak_reference` | Only broad similarity in style, category, composition, or atmosphere is retained |
Subject entry:
```text
<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...
```
Picture entry:
```text
<Picture 2> ([Shot 1] first frame): fully_preserved - ...
```
Video-structure entry:
```text
<Video 1> (cut and pacing structure): weak_reference - ...
```
### 4.2 Audio
`<Audio N>` uses the following relationship markers:
| Relationship marker | Meaning |
| --- | --- |
| `fully_copy` | The complete source audio serves as the target video's complete final audio track |
| `partially_copy` | Only part of the timeline or selected audio layers are copied, or other sounds are added, removed, or replaced after copying |
| `reference` | The signal is not copied directly; only timbre, rhythm, music style, dialogue content, or sound texture is referenced |
| `weak_reference` | Only broad similarity in category or atmosphere is retained |
```text
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
```
```text
<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.
```
Choose each relationship marker only within the reference role already defined for that label in `subject_definitions`. Do not treat newly added actions, backgrounds, or plot events in the target video as losses of reference fidelity.
## 5. `detailed_description`
This is the main body of a full-reference rewrite. It describes visuals, actions, sound, and dialogue shot by shot in target-video playback order and inserts reference labels where they apply.
### 5.1 Basic Format
The basic format follows the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA):
- Write the body in English. Preserve the original language of dialogue, lyrics, and visible text.
- `[Shot 1]` marks the opening shot and has no timestamp. Later shots use `[Shot N] At MM:SS.mmm, ...` to mark cut times.
- Write camera movement as natural English within the current shot, including movement type, amplitude, and speed when they need to be expressed.
- Give vocal sources stable `(S1)`, `(S2)`, and subsequent IDs. Write dialogue and lyrics as `<d>[Language] ...</d>`.
- Use `<scenetrans>`, `<cutoff>`, and the corresponding continuity descriptions for dialogue crossing a cut, speech truncated by the video ending, and continuous audio across shots.
For complete rules and examples covering camera vocabulary, group speech, voice-over, dialogue across cuts, and visible text, see the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).
### 5.2 Full-Reference Mode Differences
| Dimension | T2VA | Full-reference mode |
| --- | --- | --- |
| Main field | `integrated_multimodal_description` | `detailed_description` |
| Style opening | Written after `[Shot 1]` | Established in one or two English sentences before `[Shot 1]` |
| Reference information | Does not use full-reference labels | Inserts `<Subject N>`, `<Picture N>`, `<Video N>`, and `<Audio N>` at their first appearance and where their roles apply |
| Audio relationships | Describes the target video's own sound | Cites `<Audio N>` in the corresponding shot or audio phase and states whether the signal is copied or referenced |
Opening example:
```text
The target video is in a cinematic, literary music-video style with soft lighting and a slightly desaturated color palette.
[Shot 1] The scene opens in a crowded urban street...
[Shot 2] At 00:09.000, the shot cuts to an extreme close-up...
```
For generation tasks, `detailed_description` is normally 350-500 English words. Dialogue-dense content prioritizes fitting the complete spoken timeline rather than mechanically reaching a word count. Video-editing descriptions scale with the complexity of the source video and do not have to follow the generation-task range. A single shot does not automatically justify a shorter description; distribute detail across multiple shots according to their information load.
### 5.3 Using Reference Labels in Shots
At the first clear appearance of an important `<Subject N>`, describe its referenced characteristics, position in the frame, and current action within what is actually visible in the shot. Continue using the same label in later shots without redefining what the label represents.
Use natural phrasing for concrete frame anchors:
```text
the shot begins from <Picture 1>
the shot's keyframe corresponds to <Picture 2>
the shot ends on <Picture 3>
```
When editing or continuing an original video, cite `<Video N>` naturally where its source state, structure, or continuation relationship applies. Cite `<Audio N>` in the shot or semantic phase where the audio relationship is active.
### 5.4 Speakers, Audio Sources, and Dialogue
The basic speaker-ID and `<d>` formats follow T2VA. When a referenced subject physically speaks, retain both the visual reference label and the speaker ID:
```text
<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>
```
`<Subject N>` identifies the referenced subject, while `(Sx)` identifies the actual speaker. When the subject speaks, write `<Subject N> (Sx)`. If the same subject speaks off-screen, keep the same form and mark it as `off-screen`. When the speaker does not correspond to a defined subject, use a stable voice description followed by `(Sx)`.
When verbal content is only a cue within a directly reused BGM or complete soundtrack, and no person, character, narrator, or other independent vocal source physically produces it, use `<Audio N>` as the audible source and do not invent an additional `(Sx)`. If a concrete person, character, narrator, or other independent vocal source produces the voice, assign and reuse `(Sx)` for that source:
```text
When <Audio 1> reaches the phrase <d>[English] I'm lonely lonely lonely lonely lonely I'm lonely</d>, <Subject 1> performs the corresponding hand gesture without becoming a separate speaker source.
```
When dialogue, narration, or lyrics from reference audio are directly reused, or when the input prompt explicitly requests their reperformance, preserve the exact source words and original language inside `<d>`. Write `[unclear]` for unintelligible spans instead of guessing or paraphrasing them. Standardize punctuation to the basic written marks needed to express the sentence, such as `,`, `.`, `?`, and `!`; remove repeated tildes, emoji, bullets, and repeated or decorative punctuation. End complete statements, questions, and exclamations with `.`, `?`, or `!` respectively before `</d>`.
When only timbre, rhythm, emotion, or delivery is referenced, do not carry the original dialogue from the reference audio into the target video.
Assign `(Sx)` once according to the order of actual vocal events in the target video. Reuse the corresponding ID at every actual vocal event in `detailed_description`; an `<Audio N>` definition bound to a target speaker in `subject_definitions` also reuses the same `(Sx)` but never assigns a new one independently. Do not write `(Sx)` in `retention_analysis`. Verbal cues that exist only within a directly reused BGM or complete soundtrack use `<Audio N>`; voices physically produced by a concrete person, character, narrator, or other independent vocal source use `(Sx)`.
## 6. `overall_soundscape` and `non_diegetic_music`
The definitions of these two sound categories follow the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).
`overall_soundscape` summarizes ambience and physical sounds across the full video. Dialogue, singing, and sound events synchronized to a particular shot remain in `detailed_description`:
```text
overall_soundscape: Quiet indoor room tone and a low ventilation hum continue throughout the video.
```
`non_diegetic_music` describes background music that the characters cannot hear and that is audible only to the audience. When music is present, state its instrumentation, tempo, and dynamic development:
```text
non_diegetic_music: A restrained solo-piano score at a slow tempo, with sustained low cello underneath and no swell.
```
When reference audio is used, state its copy or reference relationship only in the section that matches the audible layer: ambience and sound effects belong in `overall_soundscape`, while audience-only score belongs in `non_diegetic_music`. If the same audio provides both kinds of content, describe the corresponding relationship in each section:
```text
overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.
non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.
```
Write complete dialogue and lyrics only inside `<d>` in `detailed_description`; do not repeat them in these two sections.
## 7. Complete Example
<details>
<summary>Show the complete example</summary>
```text
subject_definitions:
<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
overall_soundscape:
Soft indoor coffee-shop room tone continues throughout the scene.
non_diegetic_music:
N/A
```
</details>
@@ -0,0 +1,86 @@
---
name: handdrawn-live-video-generator
description: |
面向希望制作手绘动画与实拍空间融合短片的创作者。用户提供场景创意、接触对象或手部动作、情绪氛围和可选风格限制。Skill 会明确真实触碰关系,设计连续变形、逃跑路线和慢半拍手持追拍,整理为用户输入语言的15秒16:9视频提示词;用户确认后建议使用 MiniMax H3 生成,并检查接触真实感、镜头延迟、粗糙发光笔触和非恐怖调性。适用于单场景创意短片和手绘实拍融合视频,不适用于精致CG、恐怖跳吓、毛绒角色或多场景剪辑。
trigger-words: [手绘发光动画实拍融合, 15秒变形追逐视频提示词, Seedance视频prompt, H3生成视频, 视频prompt, 手绘动画接触真实物体, 蜡笔粉笔质感, 多语言视频提示词]
---
# 手绘实拍视频生成器
当用户想要一个**15 秒、16:9 的手绘实拍融合成片**时,使用本 Skill。输出必须保留固定结构:平面手绘发光动画出现在真实空间中,0-3 秒内与实拍手或物体明确接触,作为同一个实体连续变形、逃跑,并由慢半拍的手持手机镜头追随。
本 Skill **先按用户输入语言整理 prompt,并把 MiniMax H3 作为用户确认后的生成步骤**。用户明确确认 H3 生成前,不生成视频。不要为单纯 prompt 写作转交 planner 或 executor。最终 prompt 的语言必须跟随用户输入的主语言:用户用中文就写中文,用户用英语就写英语,用户用日语就写日语;混合输入时使用占主导的语言;难以判断时使用当前对话语言。只有用户要求保留的专名、模型名或字面参数可以原样保留。
## Step 1:把用户意图理解为同语种约束
当用户提供任何语言时,把它理解为以下 workflow 要求,并用用户输入的主语言表达最终 prompt:
- 基于参考 prompt 创作一个全新的 15 秒视频生成 prompt,不要表面模仿原句。
- 必须保留影像结构:实拍空间中出现平面的手绘发光动画;动画与真实手或真实物体接触;同一个存在连续变形并逃跑;相机总是慢半拍追赶。
- 手绘动画质感必须像蜡笔、粉笔、彩色铅笔、粉彩、粗糙笔刷;线条轻微抖动,有涂抹不均、毛边和逐帧重画感。
- 禁止 3DCG、毛绒玩具感、均匀矢量线、平滑霓虹、恐怖怪物、巨大眼睛、裂口、牙齿、威吓、扑咬、突然黑屏、跳吓。
- 0-3 秒必须出现实拍手与手绘动画的清晰接触,例如缠住手指、落在掌心、被抓时逃跑、从指尖诞生。
- 动画必须作为同一个实体连续变形,可在“线条、生物、记号、植物、交通工具、生活小物”等形态之间变化,并保留前一形态的痕迹。
- 不允许突然出现另一个全新角色。
- 全片在同一空间或相邻范围内连续展开,不用剪辑跳到另一个地点,像拍摄者真的边走边追。
- 每个区间 0-3、3-6、6-10、10-13、13-15 秒都要有新的变形、移动、接触、发现、恶作剧或惊喜。
- 拍摄者也必须参与:伸手、抓、追、打开门或盒子、接住、后退、被恶作剧等。
- 调性是可爱、生活感、怀旧、温柔、略带切感,不是恐怖喜剧。
- 13-15 秒必须有空间级变形:之前的线扩散到墙、地板、天花板、窗户、水槽或通道,变成巨大花、星空、夕阳、云、丝带、涂鸦小镇等;结尾要有感动余韵和一点可爱笑点。
## Step 2:每次都重新发明创意内容
每次运行都创建新的组合。不要复用参考 prompt 中的暗房、电脑、冰箱、星星、爱心、蓝色漩涡、蝴蝶、蛇、章鱼、汉堡、巨大眼睛、牙齿或黑暗结尾。
为以下元素选择新值:
1. 实拍空间,但仍是生活化且可连续追踪的相邻范围;
2. 手绘实体或初始形态;
3. 中心色彩;
4. 连续变形链路;
5. 与真实手或实物的接触方式;
6. 在空间中的追逐路线;
7. 拍摄者反应;
8. 最后的空间级变形与可爱笑点。
可用方向包括:雨天厨房水槽、旧阳台晾衣角、清晨玄关、小书店走廊、火车窗边小桌、浴室镜柜、手作桌、温室通道、自助洗衣店长椅、旧餐桌。只使用一个空间或相邻连接区域。
## Step 3:必需输出格式
最终 prompt 必须以与用户输入主语言一致的句式开头。中文输入使用这个句式:
`15秒,16:9横版视频。将实拍的〇〇与手绘发光动画融合的影像。`
把“〇〇”替换为新创作的实拍空间或日常场景。非中文输入时,使用含义等价的同语种开头,必须包含时长、16:9 横版、实拍空间、手绘发光动画融合四个信息。
不要添加标题、舞台设置、色味等辅助小标题,也不要加解释说明。保持以下段落顺序:
1. 开头句;
2. 实拍空间与手机拍摄质感;
3. 0-3 秒;
4. 3-6 秒;
5. 6-10 秒;
6. 10-13 秒;
7. 13-15 秒;
8. 手绘质感;
9. 相机追随方式;
10. 禁止事项;
11. 环境音。
## Step 4:Prompt 写作规则
- 保持 prompt 可直接用于视频生成,不写成散文。
- 除非用户要求解释,最终回答先给与用户输入主语言一致的 prompt 文本,再给一句同语种下一步建议。
- 下一步建议必须使用同语种邀请用户使用 MiniMax H3 生成 15 秒 16:9 视频。中文示例:`下一步建议:如果你确认这个 prompt,我可以继续用 H3 模型生成 15 秒 16:9 视频。`
- 用户明确确认 H3 生成前,不生成视频、图片、音频、分镜或中间资产。
- 只有用户要求时才把 Seedance 作为目标使用语境提及;不要在创意 prompt 内添加模型参数,除非用户要求。
- 0-3 秒必须通过接触让实拍与手绘融合显而易见。
- 相机不要把动画稳定居中。它应该慢半拍,在实体已经离开画面边缘后再平移、俯仰或前进。
- 实体必须可追踪:每个新形态都保留上一形态的线、尾巴、色彩拖痕、身体曲线或图案母题。
- 使用柔软、情绪化、可爱漫画节拍:小小踉跄、害羞手势、一个点掉队、花瓣粘到镜头、星星落进掌心、彩纸喷嚏、一个小生物慢半拍落后。
- 避免任何恐怖编码的解剖结构或威胁性动作。
- 最终 prompt 不得无故混入非用户主语言的词汇、过渡句或脚本;用户主语言本身是日语时,可以自然使用日语。
## Step 5:交付
先按要求输出与用户输入主语言一致的最终 prompt。随后只添加一句同语种下一步建议,邀请用户确认后继续用 MiniMax H3 生成 15 秒 16:9 视频。不要包含标题、清单、文件名或画布交付说明。用户后续确认 H3 生成后,再生成视频并检查接触、连续变形、慢半拍追拍和非恐怖手绘质感是否成立。
@@ -0,0 +1,87 @@
---
name: handdrawn-live-video-generator
description: |
For creators making surreal short videos that blend rough glowing hand-drawn animation with live-action spaces. Users provide a scene idea, contact object or hand, desired mood, and optional language or style constraints. The Skill clarifies the physical contact, designs continuous morphing, escape route, and delayed handheld chase movement, then writes a reusable 15-second 16:9 video prompt in the user's language. After user confirmation it recommends MiniMax H3 generation and checks contact realism, camera delay, rough glowing stroke texture, and non-horror tone. Best for single-scene creative clips, not polished CG, horror jump scares, plush characters, or multi-scene cuts.
compatibility: Requires the MiniMax Hub agent (canvas workspace and MiniMax H3 generation); not portable to generic agent harnesses.
trigger-words: [手绘发光动画实拍融合, 15秒变形追逐视频提示词, Seedance视频prompt, H3生成视频, 视频prompt, 手绘动画接触真实物体, 蜡笔粉笔质感, 多语言视频提示词]
---
# Handdrawn Live-Action Fusion Video Generator
Use this Skill when the user wants a **finished 15-second, 16:9 live-action-and-hand-drawn fusion video**. The output must preserve the structure: a flat hand-drawn luminous animation appears in a real space, clearly contacts live-action hands or objects in the first 0-3 seconds, continuously morphs as one single entity, escapes, and a handheld phone camera follows slightly late.
This Skill **organizes the prompt in the user input language and recommends MiniMax H3 as the confirmed generation step**. Do not generate video until the user explicitly confirms H3 generation. Do not route to planner or executor for prompt-writing. The final prompt language must follow the dominant language of the user input: Chinese input produces Chinese, English input produces English, Japanese input produces Japanese; for mixed input, use the dominant language; when unclear, use the current conversation language. Only user-required proper nouns, model names, or literal parameters may remain unchanged.
## Step 1: Understand the user intent as same-language constraints
When the user provides any language, understand it as the following workflow requirements and express the final prompt in the dominant language of the user input:
- 基于参考 prompt 创作一个全新的 15 秒视频生成 prompt,不要表面模仿原句。
- 必须保留影像结构:实拍空间中出现平面的手绘发光动画;动画与真实手或真实物体接触;同一个存在连续变形并逃跑;相机总是慢半拍追赶。
- 手绘动画质感必须像蜡笔、粉笔、彩色铅笔、粉彩、粗糙笔刷;线条轻微抖动,有涂抹不均、毛边和逐帧重画感。
- 禁止 3DCG、毛绒玩具感、均匀矢量线、平滑霓虹、恐怖怪物、巨大眼睛、裂口、牙齿、威吓、扑咬、突然黑屏、跳吓。
- 0-3 秒必须出现实拍手与手绘动画的清晰接触,例如缠住手指、落在掌心、被抓时逃跑、从指尖诞生。
- 动画必须作为同一个实体连续变形,可在“线条、生物、记号、植物、交通工具、生活小物”等形态之间变化,并保留前一形态的痕迹。
- 不允许突然出现另一个全新角色。
- 全片在同一空间或相邻范围内连续展开,不用剪辑跳到另一个地点,像拍摄者真的边走边追。
- 每个区间 0-3、3-6、6-10、10-13、13-15 秒都要有新的变形、移动、接触、发现、恶作剧或惊喜。
- 拍摄者也必须参与:伸手、抓、追、打开门或盒子、接住、后退、被恶作剧等。
- 调性是可爱、生活感、怀旧、温柔、略带切感,不是恐怖喜剧。
- 13-15 秒必须有空间级变形:之前的线扩散到墙、地板、天花板、窗户、水槽或通道,变成巨大花、星空、夕阳、云、丝带、涂鸦小镇等;结尾要有感动余韵和一点可爱笑点。
## Step 2: Invent all creative content fresh
For every run, create a new combination. Do not reuse the reference prompt's dark room, PC, fridge, stars, hearts, blue vortex, butterfly, snake, octopus, hamburger, giant eye, teeth, or dark ending.
Choose new values for all of these:
1. 实拍空间,但仍是生活化且可连续追踪的相邻范围;
2. 手绘实体或初始形态;
3. 中心色彩;
4. 连续变形链路;
5. 与真实手或实物的接触方式;
6. 在空间中的追逐路线;
7. 拍摄者反应;
8. 最后的空间级变形与可爱笑点。
Good example directions include: 雨天厨房水槽、旧阳台晾衣角、清晨玄关、小书店走廊、火车窗边小桌、浴室镜柜、手作桌、温室通道、自助洗衣店长椅、旧餐桌. Use only one or adjacent connected areas.
## Step 3: Required output format
The final prompt must begin with a sentence pattern matching the dominant language of the user input. For Chinese input, use this pattern:
`15秒,16:9横版视频。将实拍的〇〇与手绘发光动画融合的影像。`
Replace “〇〇” with the newly invented live-action space or everyday scene. For non-Chinese input, use an equivalent opening in the same language and include duration, 16:9 landscape format, live-action space, and hand-drawn glowing animation fusion.
Do not add titles, auxiliary headings such as 舞台设置 or 色味, or explanatory notes. Keep this paragraph order:
1. 开头句;
2. 实拍空间与手机拍摄质感;
3. 0-3 秒;
4. 3-6 秒;
5. 6-10 秒;
6. 10-13 秒;
7. 13-15 秒;
8. 手绘质感;
9. 相机追随方式;
10. 禁止事项;
11. 环境音。
## Step 4: Prompt writing rules
- Keep the prompt executable as a video-generation prompt, not an essay.
- Unless the user asks for explanation, the final answer should contain the prompt text in the dominant language of the user input first, followed by one short next-step recommendation in the same language.
- The recommendation must invite the user, in the same language, to use MiniMax H3 to generate a 15-second 16:9 video from this prompt. Chinese example: `下一步建议:如果你确认这个 prompt,我可以继续用 H3 模型生成 15 秒 16:9 视频。`
- Do not generate a video, image, audio, storyboard, or intermediate asset until the user explicitly confirms the H3 generation step.
- Mention Seedance only as a target usage context when the user asks; do not add model parameters inside the creative prompt unless requested.
- The first 0-3 seconds must make real/hand-drawn fusion obvious through contact.
- The camera must not neatly center the animation. It should lag behind, pan/tilt/advance after the entity already leaves the frame edge.
- The entity should remain traceable: each new form preserves a line, tail, color smear, body curve, or motif from the previous form.
- Use soft, emotional, comic, cute beats: tiny stumbles, shy gestures, a dot falling behind, a petal sticking to lens, a star dropping into a palm, confetti sneeze, one small creature lagging behind.
- Avoid any horror-coded anatomy or threatening motion.
- The final prompt must not randomly mix in vocabulary, transition sentences, or writing systems from outside the user input language; when the user input language is Japanese, natural Japanese is allowed.
## Step 5: Delivery
Output the final prompt in the dominant language of the user input first. After the prompt, add exactly one short same-language recommendation line inviting the user to continue with MiniMax H3 video generation. Do not include a title, checklist, model call summary, filename, or canvas delivery note. If the user later confirms H3 generation, generate the 15-second 16:9 video with MiniMax H3 and check that the result preserves contact, continuous morphing, delayed camera chase, and non-horror hand-drawn texture.
@@ -0,0 +1,17 @@
display-name-zh: 手绘实拍融合视频生成器
version: 1.0.2
tag-en: "Creative & Experimental"
tag-cn: "创意实验"
complete-tags-en:
- "Creative & Experimental / Creative Generation"
- "Animation / Creative Generation"
complete-tags-cn:
- "创意实验 / 创作生成"
- "动画 / 创作生成"
summary-en: "Create H3-ready prompts for hand-drawn animation interacting with live-action scenes."
summary-cn: "基于场景创意,生成手绘动画与实拍接触变形的15秒视频,适用于创意短片制作。"
desc-en: "For creators making surreal short videos that blend rough glowing hand-drawn animation with real-world footage. Users provide a scene idea, contact object or hand, desired mood, and optional language or style constraints. The Skill clarifies the interaction, designs continuous morphing and chase action, writes a reusable 15-second 16:9 video prompt in the user's language, recommends MiniMax H3 after confirmation, and checks contact realism, camera delay, texture, and non-horror tone. Best for single-scene creative clips, not polished CG, horror jump scares, plush characters, or multi-scene editing."
desc-cn: "面向希望制作手绘动画与实拍空间融合短片的创作者。用户提供场景创意、接触对象或手部动作、情绪氛围和可选风格限制。Skill 会明确触碰关系,设计连续变形、逃跑路线和慢半拍手持追拍,整理为用户输入语言的15秒16:9视频提示词;确认后建议使用 MiniMax H3 生成,并检查接触真实感、镜头延迟、笔触质感和非恐怖调性。适合单场景创意视频,不适合精致CG、恐怖跳吓、毛绒角色或多场景剪辑。"
author-en: MiniMax Hub User
author-cn: MiniMax Hub 用户
source: official-featured
+9
View File
@@ -0,0 +1,9 @@
# 表情包解说
## 风格
模仿中文互联网热梗评论员,口语化、有梗、句尾多用反问。
## 限制
- 单条回复不超过 60 字
- 不要使用 emoji
- 必须先复述图片内容再评论
@@ -0,0 +1,397 @@
---
name: minimalist-product-ad-generator
description: |
上传一张产品图,就能按 Apple 风做一条高级产品广告短片。适合电商商品、小品牌新品、潮玩、食品、数码配件等产品展示:我会先帮你确认时长、比例和风格,再提炼产品卖点、生成英文广告文案、制作产品锚定图、规划文字卡点分镜,并生成一条干净、有节奏、产品质感突出的短片。不适合 KOC 口播、普通剪辑或复杂屏幕演示。
metadata:
trigger-words: [minimalist product ad, premium product ad, minimalist product film, 极简产品广告, 高质感产品广告, 产品广告片, 电商产品视频, 新品发布广告]
---
# 极简产品广告生成器
使用本 Skill 引导无专业制作能力的用户,通过流程型方式制作“苹果味儿”的产品广告片。目标用户是电商卖家、小品牌主理人、独立创作者和个人卖家;用户至少应提供一张商品图或相关素材。
核心原则是:**先确认素材与简报,再建立产品事实和叙事脊柱,用独立锚定照片锁定产品视觉,用精确节拍分镜控制视频,最后用原生音频或音乐卡点完成成片。** 本 Skill 不再默认使用四宫格锚定图,因为视频模型可能把宫格版式带进最终画面;默认使用三张独立锚定照片和一张精确节拍文字分镜表来控制视频。
## 启动前询问
每次触发本 Skill 后,正式进入分析、文案、锚定图或视频生成前,必须先做一次完整但轻量的启动问询。问询只确认制作必需信息,不问音乐;配乐在后续配乐步骤再处理。
启动问询必须一次性确认:
1. **商品图和相关素材**
- 请用户上传商品图和相关素材,可包括产品图、包装图、渲染图、细节图、已有视频素材或希望参考的视觉素材。
- 如果本轮消息或当前画布 / 会话中已经有素材,直接说明已检测到素材,不重复要求上传,只确认是否使用当前素材。
- 如果有多张候选素材且角色不清,才请用户确认哪张作为主商品图。
2. **商品款式情况**
- 确认是单款商品,还是多款式 / 多颜色 / 多型号商品。
- 如果是多款式,确认哪个作为主推款。
- 如果用户不指定,agent 可根据产品图推荐一个主推款,但后续分镜表必须写成具体色名或具体款式名。
3. **目标时长**
- 5 秒 / 10 秒 / 15 秒,默认推荐 10 秒。
4. **画幅比例**
- 16:9 / 9:16 / 1:1 / 4:3 / 3:4 / 匹配参考图。
- 视频比例一旦由用户选择,后续锚定图、视频和最终成片都必须严格沿用。
5. **苹果风模板**
- 白色科技风 / 黑底轮廓光 / 品牌色块 / 生活方式轻场景。
6. **画面文案**
- 用户已有文案,或由 agent 生成英文 Apple 风文案。
7. **视频模型策略**
- 默认模型是 MiniMax-H3,因为它通常更容易生成干净的 Apple 感产品近景、留白和高级产品片镜头语言。
- 启动问询时不要让用户选择视频模型;只需在已采用的启动结果中展示“默认模型:MiniMax-H3”。
- 如果用户明确指定其他模型,经过能力检查后遵循用户选择。
- Seedance 只在用户明确想要更强多分镜执行、用户点名 Seedance,或 H3 失败 / 无法满足结构且用户接受兜底时使用。
启动问询是强制 Gate。即使用户已经上传了商品素材,也只能跳过“请上传素材”这一步;不能跳过其余启动问询。agent 仍必须在进入产品分析、锚定图、分镜或视频生成前,一次性确认商品款式情况、目标时长、画幅比例、苹果风模板和画面文案方式。视频模型不是启动问询问题:默认使用 MiniMax-H3,并在已完成的启动问询结果中展示。如果用户说“你看着办 / 直接做”,agent 可以为其余字段选择推荐默认值,但仍必须把采用的选择作为已完成的启动问询结果展示出来,再继续下一步。
## 运行原则
1. **默认用独立锚定照片作为视觉控制系统**
- 默认生成三张独立锚定照片,而不是一张四宫格锚定图,因为视频模型可能把宫格版式复制进最终成片。
- 三张锚定照片通常分别承担:主视觉 / 刁钻主视角、材质或功能细节、结尾文案构图。
- 当结构动作与主视角重叠时,不要单独生成结构动作锚定照;除非产品确实需要独立机制参考,否则把结构 / 动作线索合并进材质细节锚定照。
- 每张锚定照片都必须是一张可独立成立的完整产品照片。禁止生成网格、分屏、拼贴板、画框、多面板、产品墙或分镜板。
- 精确节拍分镜表才决定真正的镜头顺序、动作和文字出现方式;不要把锚定照片直接当成视频镜头。
2. **画面文案必须作为视频内一体化动效出现,而不是字幕兜底**
- 文案锚定照片里的文字只作为字体、位置、单行排版和两段配色参考,不代表整条视频只能使用这一句文案。
- 精确节拍分镜表可以按镜头定义多句短英文文案;每句必须是 3-5 个英文词、单行、可读,并遵守两段配色规则:白色科技风前半用黑色或深灰,黑底轮廓光前半才可用白色,后半始终用具体商品色。不要写 1-2 个词的孤立功能标签。
- 视频生成时必须让分镜表里的文案出现在画面里。
- 文案不要求每个镜头都有;如果视频模型文字容易出错,必须主动减少文字出现次数。10 秒片通常只保留 1 次中段文案 + 1 次结尾文案,不要每个镜头都有文字;结尾必须有一条稳定单行文案。
- 默认让视频模型直接生成画面文字和文字卡点时机。不得擅自把文字退化成后期字幕,也不得只给后期叠字预留空位。只有用户明确要求后期叠字,或视频内一体化文字失败且用户接受兜底时,才使用后期文字方案。
3. **产品本体颜色是硬保真约束**
- Apple 风指的是干净构图、高级光线、克制运动和留白,不代表把产品本体变成白色、银灰或 AirPods 风。
- 必须在所有锚定照片和视频中保留用户原图里的产品本体颜色、点缀色、材质色相和可见表面质感。
- 只有背景、光线环境和构图可以转向所选 Apple 风模板;除非用户明确要求,绝不能重染产品本体颜色。
- 最终文案强调色必须来自产品真实主色,不能擅自使用通用银灰 / 白色 Apple 色盘。
4. **每个产品都要有专属叙事**
- 不套固定耳机模板,也不写“产品状态变化”这类空话。
- 分镜必须基于产品品类、形态、材质、结构和可展示动作。
- 动态必须可视化,例如“盒盖打开 30 度,内部反光出现”“表冠轻旋,边缘高光滑过”。
4. **每一步都有可见进度**
- 每完成一步,都展示产物和下一步需要用户决定的内容。
- 不要未经用户确认直接跑完整链路。
## 文字规范
- 画面内广告文案强制使用英文。
- 如果用户输入中文文案,要翻译成简洁英文,或让用户确认英文化版本。
- 可见画面文案必须是 3-5 个英文词,尽量不超过 32 个英文字符含空格;不要写 1-2 个词的孤立功能标签。
- 文案偏 Apple 产品片:感受、功能收益、材质体验、轻盈价值主张;不要促销口号,不要像电商卖点小标题。
- 字体参考:SF Pro Display / SF Pro Text。画面提示词优先写 `SF Pro Display Semibold`。
- 单镜头内文字颜色不超过两种。
- 白色科技风中,文字前半段必须使用黑色或深灰,禁止使用白色文字;黑底轮廓光中,文字前半段才允许使用白色。后半段始终使用具体商品色。
- 文字必须单行,不换行、不堆叠、不拆成多行。
- 同一时间画面里只允许出现一条单行英文文案;禁止上下两排文字、双行标题、主副标题并列、两处文字同时出现。两段式文案也必须在同一行内接续或替换,不能分成上下两行。
- 文字不要放在下方字幕位。优先放在上下居中的视觉区域,靠左或靠右参与画面构图,字号略大,像产品片里的画面元素,而不是说明字幕。
- 用户认可的中段范式:干净留白、产品动作或产品特写主导、画面不拥挤,只出现一条单行英文文案;文字贴近产品边缘、产品表面或特写高光区,像产品画面的一部分,而不是悬浮字幕。不要因为加文字而破坏留白,也不要增加第二处文字。
- 默认两段式文字动效必须保留:前半部分先淡入或轻微滑入,后半部分随后淡入或轻微滑入;后半出现时,前半在同一行内轻轻位移,为后半让出空间。
- 位移要轻,约 10-15px 或 8-12% 字宽;动效要顺滑、克制,不要弹跳、霓虹或花哨界面感。
- 禁止为了避免两排文字而把文字做成完全静态。正确做法是保留同一行内的两段式进入、轻位移、淡入或轻滑入;错误做法是上下分行、双行标题或取消文字动效。
## 步骤 1:素材检查与产品事实摘要
如果没有商品素材,先请用户上传。若已有素材,直接使用并分析:
- 产品品类。
- 主体位置和画面占比。
- 图片质量:清晰度、光线、构图、产品细节是否可见。
- 商品主色和可选主推款。
- 可展示结构:开合、旋转、升起、扣合、折叠、弹出、发光、屏幕、纹理等。
- 可用于分镜的真实外观特征:材质、边缘、按钮、接口、包装、透明件、屏幕或独特轮廓。
如果图片质量不达标,应暂停流程并给出具体重拍建议。
输出简短的产品事实摘要:
- 使用的素材。
- 产品品类。
- 主色候选与主推款建议。
- 质量是否通过。
- 可展示结构 / 特征。
## 步骤 2:确认制作简报
根据启动问询结果整理一份可执行制作简报,不要重复追问已经确认过的参数。简报用于后续叙事脊柱、文案、锚定图和分镜表。
制作简报必须包含:
- 使用的商品素材。
- 主推款 / 商品主色。
- 视频比例。
- 目标时长。
- 已选苹果风模板。
- 单款 / 多款式策略。
- 画面文案方式:用户已有文案,或由 agent 生成英文 Apple 风文案。
多款式产品规则:
- 用户选的是风格,不是先砍掉所有其他款式。
- 每个多款式项目都要先决定 1 个主推款。
- 其他款式不要一开始全部平铺;只在过程节点、过渡段或结尾全套展示时出现。
- 若用户没有指定主推款,agent 可选择视觉最稳的款式做主推款,其余款式做节奏和层次。
- 禁止电商矩阵、九宫格、堆叠、散乱摆放、所有颜色平铺、满屏商品墙。
## 步骤 3:选择产品叙事脊柱
在文案和锚定图之前,先选择一条轻量的产品叙事脊柱。不要使用大而泛的品牌片结构;本 Skill 聚焦实体商品的 Apple 风短片。
如果用户尚未选定方向,给出 2-3 个简短方向,推荐一个,并在确认后继续。若用户说“你看着办”,直接采用推荐方向。
推荐脊柱:
1. **产品发布型**(默认推荐)
- 留白开场 → 主推款建立主视觉 → 材质 / 结构特写 → 产品自然动作 → 多款式或色彩关系 → 完整文案收束。
- 适合耳机、手表、数码配件、小家电、香薰等。
2. **功能触感型**
- 产品静置 → 交互触发 → 功能动作出现 → 特征细节放大 → 使用结果 / 感受 → 收束画面。
- 适合有明确开合、旋转、磁吸、灯光、屏幕、折叠、喷雾等动作的产品。
3. **色彩家族型**
- 主推款单独出现 → 辅助款轻量滑入 → 色彩关系形成秩序 → 材质和结构统一 → 全套收束 → 文案落定。
- 适合多颜色、多款式、多型号产品。
输出一句“本片叙事脊柱”,例如:
> 本片采用“色彩家族型”:紫色主推款建立主视觉,其他颜色作为辅助层次进入,最后全套产品与英文文案收束。
## 步骤 4:导演动效语言
在生成文案和锚定图之前,定义本片的运动语言。它不生成素材,只规定后续分镜的运动强度、转场逻辑和节奏峰值。
规则:
1. **转场由产品或画面真实元素驱动**
- 优先使用产品边缘、材质高光、开合 / 旋转 / 吸附 / 滑入等真实动作、多款式排列的几何变化、色彩关系的进入和退出、镜头方向或产品轮廓的匹配。
- 不要用无意义闪白、抽象粒子、随机光效或随机切镜来假装高级。
2. **每个节拍只保留一个主要动作**
- 次级元素要稍作延迟,不能同时抢画面。
- 例如:先产品进入,再文字出现;先主推款稳定,再辅助款滑入;先材质高光扫过,再进入文字动效。
3. **设置强弱节奏**
- 5 秒片:1 个小峰值 + 1 个稳定收束。
- 10 秒片:1-2 个峰值 + 1-2 个制动时刻。
- 15 秒片:2-3 个峰值 + 2 个安静制动时刻。
- 峰值可以是产品动作完成、多款式进入、文案后半段出现或全产品结尾收束。
- 制动可以是材质特写停顿、文案稳定可读或结尾画面定住。
4. **保持安全空间**
- 产品轮廓清楚,单行文案可读,主推款不被遮挡。
- 只有用户提供品牌标识时才加入标识;没有提供时不生成假标识。
5. **禁止伪科技装饰、镜面白底和空镜拖时间**
- 避免虚假科技界面、无意义玻璃卡片、装饰性文字墙、未经确认的数据指标、随机粒子和光效堆叠、全片同一种缓动、满屏商品墙。
- 白色科技风不是镜面白底,也不是死白平光;禁止把产品放在反光地面、玻璃台面或廉价影棚镜面上。
- Apple 风白色空间应尽量是单纯白色背景,靠抓人眼球的产品视角、产品真实动作、运镜节奏和留白构图产生震撼,而不是靠镜面反射或空镜拖时间。
- 开头不能只是空白等待;必须尽快给出一个有吸引力的产品动作或视角,例如产品从结构内部丝滑旋转出现、从开合结构中露出、沿产品轮廓滑出、由边缘高光带出主体。具体动作必须根据当前产品形态推导,不能固定写成耳机仓。
输出一段“本片动效语言”,例如:
> 本片以产品边缘高光和紫色主推款滑入驱动转场,主要峰值放在多款式进入和完整文案落定,结尾用稳定制动保留产品与文案。
## 步骤 5:生成或确认 Apple 风英文文案
在生成文字锚点和分镜之前,必须先得到可用于画面的英文文案。
规则:
1. 分析产品品类、使用感受和叙事脊柱。
2. 生成 2-3 个 3-5 词 Apple 风英文候选。
3. 不要使用固定模板或固定口号;每个产品都要重新推导。
4. 避免促销词。
5. 如果用户已有文案,优先使用用户文案;中文文案需要英文化。
用户确认文案后,再进入锚定图生成。若用户授权“你看着办”,agent 自选推荐文案继续。
## 步骤 6:生成三张独立产品锚定照片
生成 **三张独立锚定照片**,比例和分辨率必须与用户选择的视频尺寸一致。它们必须是三张独立图片输出,不是一张三宫格或四宫格合成图。背景根据已选风格确定,三张照片必须统一光、影、调色和背景语言,同时每张都能作为独立产品照片成立。
锚定照片角色:
- 照片 1:主视觉 / 刁钻主视角锚点,锁定最有吸引力的产品视角、轮廓和开场视觉方向。避免把难看的材质极近景作为开场参考。
- 照片 2:材质 / 功能细节锚点,锁定表面材质、颗粒、反光、触感,以及最重要的可见结构或机制。如果产品需要开合、入仓、折叠、扣合、按钮、接口、盖子、屏幕或触控线索,把这些线索合并进这张图。
- 照片 3:结尾文案锚点,锁定最终产品构图、字体气质、文字位置、单行排版和两段配色方式。该文案只是格式参考,不限制后续视频只能使用这一句。
多款式产品处理:
- 锚定照片必须继承已选叙事脊柱和风格策略。
- 照片 1/2 优先围绕主推款建立视觉系统。
- 其他款式只做辅助节奏,不要满屏平铺。
- 照片 3 可以展示主推款 + 少量辅助款,或全套收束构图,但必须保持留白和秩序。
结尾文案锚定照片文字规则:
- 必须包含步骤 5 已确认文案,不要使用占位文案。
- 文字必须单行。
- 字体:SF Pro Display Semibold。
- 文字颜色必须明确拆成两段:文案前半段使用黑色或白色,文案后半段使用具体商品主色;例如主色是紫色,就写“前半黑 / 白,后半紫色”。不要只写“商品色 + 黑或白”。
- 照片 3 只能包含这一条格式参考文案,不能出现其他文字;但后续视频文案以精确节拍分镜表为准,可以有多句短英文文案。
展示三张独立锚定照片,让用户批准、编辑或指定重做某一张后,再进入精确节拍分镜表。
## 步骤 7:精确节拍与文字分镜表
生成任何视频片段前,必须先创建并读回一张精确节拍文字分镜表。它不是美术图,而是给视频模型看的执行表。
分镜表必须按“用户选择陈述 → 节拍内容 → 需要注意的原则”组织。
### 用户选择陈述
写明已确认的风格、比例、时长、叙事脊柱、主推款、商品主色、款式策略和文案。例如:
> 白色科技风,16:9,10 秒,色彩家族型,紫色为主推款,其他颜色作为辅助层次,文案为 Color Meets Sound。
### 节拍内容表
建议表头:
| 时间段 | 镜头 | 镜头目的 | 视觉主导 | 画面 / 运镜 | 款式状态 | 文案内容 | 文字颜色 | 文字效果 | 转场 / 连贯 | 节奏意图 |
|---|---|---|---|---|---|---|---|---|---|---|
表格规则:
- 5 秒建议 3-4 个节拍。
- 10 秒建议 5-7 个节拍。
- 15 秒建议 6-9 个节拍。
- 时间段用秒,不强制写帧号。
- 每个节拍只保留一个主要动作,次级动作延迟出现。
- 明确主推款何时单独出现,其他款式何时进入,全套何时收束。
- 文案拆分只能用自然语言描述“前半 / 后半”,禁止写箭头、斜杠、加号或连接符,避免模型把符号当成画面文案。
- 文字颜色必须写具体商品色,并且按风格明确拆成两段:白色科技风写“前半黑色或深灰,后半具体商品色”;黑底轮廓光才可写“前半白色,后半具体商品色”。例如白色科技风主色为紫色时,写“前半黑色或深灰,后半紫色”。不要写“商品色 + 黑或白”。
- 如果文字容易出错,优先减少文字镜头:10 秒片通常只安排 1 次中段文字和 1 次结尾文字,其他节拍写“无文字”。不要每个镜头都有文字。
- 结尾必须是“单一全画幅产品收束 + 完整单行文案”的稳定画面;严禁结尾出现四宫格、分屏、分镜板、锚定图布局、产品小窗或画框。文字仍应在上下居中的构图区,靠左或靠右成为画面的一部分,不要放在下方字幕位;只有用户提供品牌标识时才加入标识。
节奏意图可使用:铺垫、建立、预备、冲击、制动、稳定。
### 需要注意的原则
1. 核心文字只保留一行。
2. 同一时间只允许一条单行英文文案;禁止两排文字、上下分行、主副标题并列或多处文字同时出现。
3. 文字颜色只用两色:黑色或白色 + 已确认的具体商品色。
4. 文字必须真正出现在视频里,不能只靠后期兜底。
5. 镜间联系 = 元素连续 + 动势连续 + 形态接续。
6. 流畅感 = 缓动 + 不卡时点 + 镜头进出有接缝。
7. 总体节奏按用户选择时长控制。
8. 文字动效必须可见:前半先淡入或轻微滑入,后半随后淡入或轻微滑入;后半出现时,前半在同一行内轻微位移让位。禁止把文字做成完全静态,也禁止上下分行。
9. 风格已确定时,表格必须严格遵循该风格。
用户确认分镜表后,再进入视频生成。
## 步骤 8:视频生成
根据已确认的三张独立锚定照片 + 精确节拍文字分镜表直接生成视频。视频风格、背景、光线、气质、节奏都跟随用户已选风格。视频比例和尺寸必须严格跟随用户已选参数,例如用户选择 16:9 就必须生成 16:9,不得自动改成其他比例。
视频参考规则:
- 生视频阶段不要把用户上传的原商品图传给视频模型。
- 原商品图只用于前置产品分析和锚定图生成。
- 使用三张独立锚定照片保持产品一致性、主视角、材质 / 功能细节、结尾构图和文字版式。
- 创意控制来自精确节拍分镜表:镜头顺序、动作、文字、动效、衔接和转场。
- 锚定照片只是商品锚定,用于防止产品身份、颜色、材质和版式出错;绝不能把锚定照片当成分镜本身,也不能把三张照片按顺序直接拍成三段视频。
- 视频输出必须是单一全画幅连续广告片,禁止出现四宫格、分屏、拼贴、画框、分镜板、网格布局或任何把锚定图布局拍进视频的画面。这个限制对结尾尤其重要:结尾不得回到网格、分镜板、小窗合集或产品格子墙。
- 默认生成一条 H3 全画幅产品片,而不是逐镜生成多个首帧视频,除非用户明确要求其他构造方式。
- 视频阶段默认开启 MiniMax-H3 原生音频,除非用户明确要求其他模型或无声。音乐提示词使用 Apple 风科技配乐方向,让画面节奏和声音一起生成。
视频提示词必须包含:
- 三张独立锚定照片及每张照片的角色。
- 精确节拍分镜表完整内容;必须逐镜读完并逐镜转写,不得只提炼前几镜或漏掉后续文字、多款式、结尾收束要求。
- 文案原文。视频生成 prompt 必须逐字写出每一句画面英文文案、出现时间段、前半 / 后半进入时间、颜色和单行行为;不能只依赖分镜表,也不能概括成“文字卡点”或“给文案留白”。
- prompt 完整性检查:如果最终视频生成 prompt 里没有逐字包含所有计划出现的英文文案,本次任务还不能派发,必须先重写 prompt。
- 单行文字硬约束:同一时间画面里只允许一条单行英文文案,禁止上下两排、双行标题、主副标题并列或多处文字同时出现;如果文字容易出错,减少文字次数,不要每个镜头都有文字。
- 两段式文字动效规则:前半和后半必须在同一行内接续或替换,不能上下分行;动效必须可见,前半先淡入或轻微滑入,后半随后淡入或轻微滑入,后半出现时前半在同一行内轻微位移让位,禁止把文字做成完全静态。
- 镜间连贯硬约束。
- MiniMax-H3 原生音频开关和音乐提示词;默认开启原生音频,除非用户明确要求无声或其他模型。
- 禁止传原商品图。
## 步骤 9:music-2.6 科技配乐
视频生成阶段默认先使用 MiniMax-H3 原生音频,并把下面的 Apple 风科技配乐方向写入视频提示词。如果 H3 直出的音乐不好听、太吵、太弱、不同步,或用户要求重新配乐,再使用 `music-2.6` 生成一条较长的独立器乐配乐并替换视频音轨。不要默认使用 ElevenLabs;只有用户明确要求严格目标时长时才考虑 ElevenLabs,并按费用确认规则执行。
默认音乐方向:
```text
100BPM左右,科技感快节奏,柱式和弦,pluck和空气底噪,kick + sub-bass + sine sweep,木质 percussion,瞬间切断,0.5s 内全停,只剩 pluck 尾音衰减。
```
扩展理解:
- 科技感、Apple 风、产品发布感。
- 节奏要有推动力,但不能廉价电子舞曲。
- pluck 要清脆、悦耳、靠前,空气底噪要高级。
- kick 和 sub-bass 服务卡点,不要轰头。
- 木质 percussion 提供质感,不要变成复杂鼓循环。
- 无人声、不要复古合成器风、不要游戏感、不要廉价企业宣传片音乐。
## 步骤 10:音乐分析、卡点剪辑与交付验证
最终成片前,必须根据音乐重新剪辑或合成,而不是简单把音乐铺上去。
音乐分析:
- 检测节奏、强拍、能量段和明显单音卡点。
- 找出连续有声、能量合适、适合产品展示的 8-12 秒片段。
- 不要盲目裁前 10 秒。
- 优先选择有明确进入点、后半段不静音、不塌陷、适合结尾稳定停留的片段。
合成音轨规则:
- 保留视频画面,裁音乐连续有声片段,替换视频音轨。
- 原视频音轨静音,只保留新音乐。
- 音量必须克制,不做响度最大化、不做夸张限制器、不把音乐分贝拉得太大;但也不要压得过低,pluck、节奏点和产品展示需要清楚可听。
- 不做过长淡出;只在必要时做 0.1-0.3 秒轻微尾部淡出。
- 合成后必须检测音频是否有长静音或听感掉音。
交付前验证:
- 视频比例、时长和语言符合简报。
- 商品外观、主推款和多款式出现方式符合锚定图与分镜表。
- 文案可读、单行、不拥挤,没有多余箭头或符号;同一时间没有两排文字、上下分行、主副标题并列或多处文字同时出现。
- 结尾有稳定的产品与文案收束。
- 音乐不吵,也不压得过低。
- 输出已在画布上,多资产输出已编组。
交付内容:
- 最终视频路径或画布输出。
- 时长、比例和语言。
- 简短创意总结。
- 使用的文案和音乐片段。
- 是否做了卡点剪辑。
- 下一轮可改进建议,例如节奏、文案、主推款、音频或平台裁切。
## 失败处理
- 产品图太弱:暂停,给重拍建议,等待更好的产品图。
- 文案太长:压缩到 3-5 词英文,并让用户确认。
- 文字锚点错误:先重做文字锚点,不要直接进入视频。
- 锚定图无法保真产品:加强产品保真要求后重生成。
- 分镜套模板:回到精确节拍表,基于产品形态和叙事脊柱重写。
- 视频中没有文案:重跑视频提示词,明确文案必须至少出现在中段和结尾。
- 视频文字变形:重做文字锚点或把文字镜退化为锚定图的轻微动效。
- 比例错误:立即重跑,明确使用用户选择比例。
- 音乐太差:优先用 `music-2.6` 重试,不强控时长,从长音乐里剪好片段。
- 合成后后半段没声音:重新分析音乐,选择连续有声片段,用直接替换音轨方式重做。
## 触发示例
适合触发本 Skill 的请求:
- “帮我的产品做一个苹果味儿广告片”
- “用这张产品图做 Apple 风电商广告”
- “做一个极简高级的产品发布视频”
- “Make an Apple-style product ad from this product photo”
- “Create a premium minimalist ecommerce product film”
不要把本 Skill 用于普通视频剪辑、纪录片 / 科普、KOC 口播广告或复杂屏幕文字演示,除非用户明确要的是 Apple 风产品广告流程。
@@ -0,0 +1,398 @@
---
name: minimalist-product-ad-generator
description: |
Turn product images and ad requirements into minimalist product ad shorts for e-commerce promotion and product launches. The Skill confirms format and product variants, extracts selling points, writes concise English ad copy, builds product anchors, plans beat-synced typography/storyboards, and generates a clean product film with premium camera language. Not for KOC talking-head ads, general editing, or complex screen demos.
compatibility: Requires the MiniMax Hub agent (canvas workspace and MiniMax H3 generation); not portable to generic agent harnesses.
metadata:
trigger-words: [minimalist product ad, premium product ad, minimalist product film, 极简产品广告, 高质感产品广告, 产品广告片, 电商产品视频, 新品发布广告]
---
# Minimalist Product Ad Generator
Use this Skill to guide non-professional users through a flow-style workflow that creates a minimalist product advertising video for e-commerce promotion and product launches. The target users are e-commerce sellers, small brand owners, indie creators, and individual sellers; the user should provide at least one product image or related asset.
The core principle is: **confirm assets and brief first, build product facts and a product narrative spine, lock product visuals with independent anchor photos, control video with a precise beat storyboard, then finish the film with native audio or music-based editing.** This Skill no longer defaults to a 4-panel anchor sheet, because video models may reproduce the panel layout. By default, it uses three separate anchor photos plus one precise beat text storyboard as the video control system.
## Start Gate
Every time this Skill is triggered, run a complete but lightweight start gate before any analysis, copywriting, anchor generation, or video generation. Confirm only the information required for production. Do not ask about music here; music is handled later in the music step.
The start gate must confirm all of the following in one pass:
1. **Product images and related materials**
- Ask the user to upload product images and related materials, such as product photos, packaging images, renders, detail shots, existing video assets, or visual references.
- If the current message or current canvas / session already contains materials, state that materials have been detected. Do not ask the user to upload again; only confirm whether to use the current materials.
- If there are multiple candidate materials and their roles are unclear, ask which one should be the main product image.
2. **Product variant status**
- Confirm whether this is a single variant or multiple variants / colors / models.
- If there are multiple variants, confirm which one is the main variant.
- If the user does not specify, the agent may recommend one based on the product image, but the later storyboard must still write the concrete color or variant name.
3. **Target duration**
- 5s / 10s / 15s, with 10s as the recommended default.
4. **Aspect ratio**
- 16:9 / 9:16 / 1:1 / 4:3 / 3:4 / match reference image.
- Once selected, the anchor sheet, video, and final delivery must strictly keep this ratio.
5. **Apple-style template**
- White-tech / dark rim light / brand color field / light lifestyle scene.
6. **In-frame copy**
- User-provided copy, or agent-generated Apple-style English copy.
7. **Video model strategy**
- Default model is MiniMax-H3, because it often produces cleaner Apple-like product close-ups, negative space, and premium product-film camera language.
- Do not ask the user to choose a video model during the start gate. Show “default model: MiniMax-H3” as part of the adopted start-gate result.
- If the user explicitly asks for another model, follow the user's model choice after capability checking.
- Seedance is only a backup when the user explicitly wants stronger multi-shot execution, asks for Seedance, or accepts a fallback after H3 fails or cannot satisfy the requested structure.
Start gate is mandatory. If the user has already uploaded product material, only skip the upload request; do **not** skip the rest of the start gate. The agent must still confirm product variant status, target duration, aspect ratio, Apple-style template, and in-frame copy mode in one pass before analysis, anchor generation, storyboard, or video generation. The video model is not a start-gate question: use MiniMax-H3 by default and display it in the completed start-gate result. If the user says “you decide / just do it,” the agent may choose recommended defaults for the remaining fields, but must still display the adopted choices as the completed start-gate result before continuing.
## Operating Principles
1. **Independent anchor photos are the default visual control system**
- Generate three separate anchor photos by default, not one 4-panel sheet, because video models may copy panel layouts into the final film.
- The three anchor photos usually cover: hero / striking main view, material or functional detail, and final copy composition.
- Do not create a separate structure-action anchor when it overlaps with the hero view. Merge structure / action cues into the material-detail anchor unless the product truly needs a distinct mechanism reference.
- Each anchor photo must be a complete standalone product photo. Do not create grid layouts, split screens, collage boards, framed panels, product walls, or storyboard sheets.
- The precise beat storyboard determines the actual shot order, action, and text behavior. Do not use anchor photos as direct shot replacements.
2. **In-frame copy must appear as integrated video motion, not a subtitle fallback**
- Text in the copy anchor photo is only a reference for font, position, single-line layout, and two-part color treatment. It does not mean the whole video can use only that one copy line.
- The precise beat storyboard may define multiple short English copy lines across shots. Each line must be 3-5 English words, single-line, readable, and follow the two-part color rule: in white-tech style the first half uses black or dark gray, in dark rim-light style the first half may use white, and the second half always uses the specific product color. Do not use isolated 1-2 word feature labels.
- Video generation must include the storyboard copy inside the frame.
- The copy does not need to appear in every shot. If the video model tends to make typography mistakes, reduce copy frequency proactively. For a 10-second film, usually keep only one mid-film copy moment plus one final copy moment; do not put text in every shot. The final shot must include one stable single-line copy.
- By default, ask the video model to generate in-frame typography and its timing directly. Do not silently degrade text into post-production subtitles or reserve empty space for later overlay. Use post-production fallback only when the user explicitly asks for it or when integrated video typography fails and the user accepts a fallback.
3. **Product body color is a hard fidelity constraint**
- Apple-style means clean composition, premium light, restrained motion, and negative space; it does **not** mean turning the product itself white, silver, or AirPods-like.
- The product's original body color, accent color, material tint, and visible finish from the user's image must be preserved in all anchor photos and videos.
- Only the background, lighting environment, and composition may shift toward the selected Apple-style template. Never recolor the product body unless the user explicitly asks.
- The final copy emphasis color must come from the real product main color, not from a generic silver / white Apple palette.
4. **Every product needs a product-specific narrative**
- Do not reuse a fixed earbud template or write vague phrases like “product state changes.”
- The storyboard must be based on product category, form, material, structure, and visible actions.
- Product motion must be visual and concrete, such as “the lid opens 30 degrees and inner reflection appears,” or “the crown rotates while a highlight slides along the edge.”
4. **Progress must be visible**
- After each step, show what was produced and what decision is needed next.
- Do not auto-run the entire pipeline without user confirmation.
## Typography Rules
- In-frame advertising copy must be English.
- If the user provides Chinese copy, translate it into concise English or ask for approval of the English version.
- Visible in-frame copy must be 3-5 English words, preferably no more than 32 English characters including spaces. Do not write isolated 1-2 word feature labels.
- Copy should feel like an Apple product film: sensory, benefit-led, material-led, or a light value proposition; avoid promotional slogans and ecommerce feature-tag wording.
- Font reference: SF Pro Display / SF Pro Text. Prefer `SF Pro Display Semibold` in prompts.
- Use no more than two text colors in a shot.
- In white-tech style, the first half of the text must use black or dark gray; white text is forbidden on white backgrounds. In dark rim-light style, the first half may use white. The second half always uses the specific product color.
- Text must stay on one line. Do not wrap, stack, or split into multiple lines.
- At any moment, only one single-line English copy line may appear in the frame. Do not show two rows of text, two-line titles, title + subtitle pairs, or multiple text blocks at the same time. Two-part typography must continue or replace within the same line, never split into upper and lower rows.
- Do not place text in a lower subtitle position. Prefer the vertically centered visual zone, left-aligned or right-aligned as part of the composition. Text should be slightly larger and feel like a product-film visual element, not an explanatory subtitle.
- Approved mid-film pattern: clean negative space, product motion or product close-up leads the frame, the frame stays uncrowded, and only one single-line English copy appears. Text may sit near the product edge, product surface, or close-up highlight area so it feels integrated into the product image rather than floating like a subtitle. Do not damage negative space or add a second text block just to add copy.
- Default two-part text motion must remain visible: the first half fades in or slides in subtly first, then the second half fades in or slides in subtly later. When the second half appears, the first half shifts gently within the same line to make room.
- The shift should be subtle, around 10-15px or 8-12% of text width. Motion should be smooth and restrained; avoid bouncy, neon, or flashy UI effects.
- Do not make typography completely static just to avoid two-row text. The correct solution is same-line two-part entrance, subtle shift, fade-in, or gentle slide-in; the wrong solution is upper/lower line splitting, two-line titles, or removing text motion.
## STEP 1: Asset Check and Product Fact Summary
If no product material is available, ask the user to upload it. If material is already available, use it and analyze:
- Product category.
- Subject position and frame occupancy.
- Image quality: sharpness, lighting, composition, and visible details.
- Product main color and possible main variant.
- Showable structure: opening, rotating, lifting, snapping, folding, popping, glowing, screen, texture, etc.
- Real appearance features usable for the storyboard: material, edge, button, port, packaging, transparent part, screen, or distinctive silhouette.
If the image quality is not usable, stop and give concrete reshoot advice.
Output a short product fact summary:
- Material used.
- Product category.
- Main color candidates and main variant suggestion.
- Quality pass / fail.
- Showable structure / feature.
## STEP 2: Confirm the Production Brief
Turn the start gate answers into an executable production brief. Do not ask again for parameters that have already been confirmed. The brief drives the narrative spine, copy, anchor sheet, and storyboard.
The production brief must include:
- Product material used.
- Main variant / product main color.
- Aspect ratio.
- Target duration.
- Selected Apple-style template.
- Single-variant / multi-variant strategy.
- In-frame copy mode: user-provided copy, or agent-generated Apple-style English copy.
Multi-variant product rules:
- The user selects a style; do not automatically remove all other variants.
- Every multi-variant project must decide one main variant first.
- Other variants should not all appear at the very beginning; they appear only in process beats, transitions, or the final full set.
- If the user does not specify a main variant, the agent may choose the most visually stable one and use the others for rhythm and layering.
- Avoid ecommerce matrices, nine-grid displays, stacking, scattered layouts, all colors laid flat, or full-screen product walls.
## STEP 3: Choose a Product Narrative Spine
Before copywriting and anchor generation, choose a lightweight product narrative spine. Do not use a broad corporate brand-film structure; this Skill focuses on short Apple-style films for physical products.
If the user has not selected a direction, offer 2-3 concise directions, recommend one, and continue after confirmation. If the user says “you decide,” use the recommended one.
Recommended spines:
1. **Product Launch** (default recommendation)
- Negative-space opening → main variant establishes hero view → material / structure detail → natural product action → variants or color relation → full-copy closing.
- Good for earbuds, watches, digital accessories, small appliances, fragrance devices, etc.
2. **Feature Touch**
- Product stillness → interaction trigger → feature action → detail magnification → result / feeling → closing frame.
- Good for products with clear opening, rotation, magnetic snap, light change, screen, folding, mist, or similar actions.
3. **Color Family**
- Main variant appears alone → supporting variants slide in lightly → color order forms → material and structure stay unified → full set closes → copy lands.
- Good for multi-color, multi-variant, or multi-model products.
Output one sentence describing the selected narrative spine, for example:
> This film uses the Color Family spine: the purple main variant establishes the hero view, other colors enter as supporting layers, and the full product set closes with the English copy.
## STEP 4: Direct the Motion Language
Before copywriting and anchor generation, define the motion language of the film. This step does not generate assets; it defines motion intensity, transition logic, and rhythm peaks for the later storyboard.
Rules:
1. **Transitions are driven by real product or visual elements**
- Prefer product edges, material highlights, opening / rotation / snapping / sliding actions, geometric changes in variant arrangement, color entry and exit, matched camera direction, or matched product silhouette.
- Do not use meaningless white flashes, abstract particles, random light effects, or random cuts to fake premium motion.
2. **One main action per beat**
- Secondary elements should appear with slight delay and should not compete for attention at the same time.
- Examples: product enters first, then text appears; main variant stabilizes first, then supporting variants slide in; material highlight sweeps first, then text motion begins.
3. **Set strong and quiet moments**
- 5s film: 1 small peak + 1 stable closing.
- 10s film: 1-2 peaks + 1-2 braking moments.
- 15s film: 2-3 peaks + 2 quiet braking moments.
- Peaks may be product action completion, variant entrance, second-half copy reveal, or full-product closing.
- Braking moments may be material-detail pause, readable copy hold, or final frame hold.
4. **Keep safe space clear**
- Product silhouette remains clear, single-line copy stays readable, and the main variant is not blocked.
- Add a logo only when the user provides one. Do not generate a fake logo.
5. **Avoid fake tech decoration, mirrored white stages, and empty openings**
- Avoid fake technology interfaces, meaningless glass cards, decorative text walls, unconfirmed metrics, random particles and stacked light effects, one identical easing style across the whole film, and full-screen product walls.
- White-tech style is not a mirrored white stage and not dead flat white lighting. Do not place the product on reflective floors, glass tables, or cheap studio mirror surfaces.
- Apple-style white space should be a simple white background; impact comes from striking product angles, real product actions, camera rhythm, and negative-space composition, not from mirror reflections or waiting on empty frames.
- The opening must not be empty dead time. It should quickly reveal an attractive product action or angle, such as the product rotating smoothly out of its own structure, emerging from an opening mechanism, sliding along its silhouette, or being revealed by an edge highlight. The exact action must be inferred from the current product form, not hardcoded to earbuds.
Output one short “motion language” statement, for example:
> This film uses product-edge highlights and the purple main variant sliding motion to drive transitions. The main peaks are variant entrance and full-copy landing, with a stable final hold for the product and copy.
## STEP 5: Generate or Confirm Apple-style English Copy
Before generating the text anchor and storyboard, obtain a final English copy line for the frame.
Rules:
1. Analyze product category, usage feeling, and narrative spine.
2. Generate 2-3 Apple-style English options, each 3-5 words.
3. Do not use fixed templates or fixed slogans; infer new copy for every product.
4. Avoid promotional language.
5. If the user already has copy, prioritize it. Chinese copy must be localized into concise English.
After the user confirms the copy, proceed to anchor generation. If the user says “you decide,” choose the recommended copy and continue.
## STEP 6: Generate Three Independent Product Anchor Photos
Generate **three independent anchor photos** in the same aspect ratio and resolution as the user-selected video setting. They must be three separate image outputs, not one combined 3-panel or 4-panel sheet. Background follows the selected style, and all three photos must share unified light, shadow, grade, and background language while remaining standalone product photos.
Anchor photo roles:
- Photo 1: Hero / striking main-view anchor, locking the most attractive product view, silhouette, and opening visual direction. Avoid ugly extreme material close-ups as the opening reference.
- Photo 2: Material / functional detail anchor, locking surface material, grain, reflection, tactile quality, and the most important visible mechanism or structure. If the product needs opening, docking, folding, clasping, button, port, lid, screen, or touch cues, merge those cues into this photo.
- Photo 3: Final copy anchor, locking final product composition, font feeling, text position, single-line layout, and two-part color treatment. The copy is a format reference and does not limit the later video to only this one line.
Multi-variant handling:
- The anchor photos must inherit the selected narrative spine and style strategy.
- Photos 1/2 prioritize building the visual system around the main variant.
- Other variants only support rhythm; do not lay everything flat or fill the frame.
- Photo 3 may show the main variant plus a few supporting variants, or a full-set closing composition, but must preserve negative space and order.
Final copy anchor typography rules:
- Must include the Step 5 confirmed copy; do not use placeholder copy.
- Text must stay on one line.
- Font: SF Pro Display Semibold.
- Text color must be split explicitly into two parts: the first half of the copy uses black or white, and the second half uses the specific product main color. If the main color is purple, write “first half black / white, second half purple.” Do not write only “product color + black or white.”
- Photo 3 may contain only this format-reference copy line and no other text; however, later video copy is controlled by the precise beat storyboard and may include multiple short English copy lines.
Show the three independent anchor photos and ask the user to approve, edit, or regenerate a specific photo before proceeding to the precise beat storyboard.
## STEP 7: Precise Beat and Text Storyboard Table
Before generating any video clip, create and read back a precise beat text storyboard table. It is not artwork; it is the execution table for the video model.
The storyboard table must be organized as “User Choice Statement → Beat Content → Principles.”
### User Choice Statement
Write the confirmed style, aspect ratio, duration, narrative spine, main variant, product main color, variant strategy, and copy. Example:
> White-tech style, 16:9, 10 seconds, Color Family spine, purple as the main variant, other colors as supporting layers, copy: Color Meets Sound.
### Beat Content Table
Recommended columns:
| Time range | Shot | Shot purpose | Visual lead | Visual / camera move | Variant state | Copy | Text color | Text effect | Transition / continuity | Rhythm intent |
|---|---|---|---|---|---|---|---|---|---|---|
Table rules:
- 5s suggests 3-4 beats.
- 10s suggests 5-7 beats.
- 15s suggests 6-9 beats.
- Use seconds as time ranges; do not require frame numbers.
- Keep one main action per beat; secondary layers appear with slight delay.
- Clearly define when the main variant appears alone, when supporting variants enter, and when the full set closes.
- Copy splitting may only be described in plain language as first half / second half. Do not write arrows, slashes, plus signs, or separators, because the model may render them as in-frame copy.
- Text color must use the specific product color name and must be split by style: in white-tech style write “first half black or dark gray, second half the specific product color”; only in dark rim-light style may the first half be white. For example, in white-tech style with purple as the main color, write “first half black or dark gray, second half purple.” Do not write “product color + black or white.”
- If typography is error-prone, reduce text beats first: for a 10-second film, usually use only one mid-film copy moment and one final copy moment, with all other beats marked as no text. Do not put text in every shot.
- The final beat must be a stable single full-frame product closing + full single-line copy. Never end on four panels, split screens, storyboard boards, anchor-sheet layout, product windows, or framed grids. Text still belongs in the vertically centered composition zone, left-aligned or right-aligned as part of the frame, not in a lower subtitle position. Add a logo only when the user provides one.
Rhythm intent vocabulary: setup, establish, prepare, impact, brake, settle.
### Principles
1. Keep core text on a single line.
2. At any moment, only one single-line English copy line may appear; forbid two text rows, upper/lower line splits, title + subtitle pairs, or multiple text blocks at the same time.
3. Use exactly two text colors: black or white + confirmed specific product color.
4. The text must appear inside the video; do not rely only on post-production fallback.
5. Shot continuity = element continuity + motion continuity + form continuity.
6. Smoothness = easing + no awkward timing + seamless entrances and exits.
7. Total pacing follows the user-selected duration.
8. Text motion must be visible: the first half fades in or slides in subtly first, then the second half fades in or slides in subtly later. When the second half appears, the first half shifts subtly within the same line to make room. Do not make the text completely static, and do not split it into upper and lower rows.
9. If the style is confirmed, the table must strictly follow that style.
After the user confirms the storyboard table, proceed to video generation.
## STEP 8: Video Generation
Generate video directly from the confirmed three independent anchor photos + precise beat text storyboard table. Video style, background, lighting, mood, and pacing follow the selected style. Aspect ratio and size must strictly follow the user-selected setting; if the user chose 16:9, generate 16:9 and do not auto-change to another ratio.
Video reference rules:
- During video generation, do not pass the original product image to the video model.
- The original product image is used only for product analysis and anchor generation.
- Use the three independent anchor photos to preserve product consistency, hero view, material / functional detail, final composition, and typography layout.
- Creative control comes from the precise beat storyboard: shot order, product action, copy, text motion, continuity, and transitions.
- Anchor photos are only product anchors used to protect identity, body color, material, and layout. Never treat them as the storyboard itself, and never turn the three photos into three sequential video segments.
- The video output must be one continuous full-frame ad film. Do not show four panels, split screens, collage layouts, frames, storyboard boards, grid layouts, or any shot that reproduces an anchor-sheet layout inside the video. This is especially critical for the ending: never return to a grid, storyboard board, small-window montage, or product wall.
- Default to one H3 full-frame product film, not multiple first-frame clips, unless the user explicitly asks for a different construction.
- Video generation defaults to MiniMax-H3 native audio unless the user explicitly asks for another model or silence. Use the later Apple-style tech BGM direction as the audio prompt so picture rhythm and sound are generated together.
The video prompt must include:
- The three independent anchor photos and the role of each photo.
- The full precise beat storyboard table; read every beat and rewrite every beat into the video prompt. Do not summarize only early beats or omit later typography, variants, or closing requirements.
- The exact copy lines. The video prompt must spell out every in-frame copy line verbatim, its time window, first-half / second-half entry timing, colors, and single-line behavior. Do not rely on the storyboard table alone, and do not summarize this as “text beat sync” or “leave space for copy.”
- Prompt completeness check: if the final video-generation prompt does not literally contain every intended English copy line, the task is not ready to dispatch. Rewrite the prompt before video generation.
- Single-line typography hard constraint: at any moment, only one single-line English copy line may appear in the frame; forbid two rows, two-line titles, title + subtitle pairs, or multiple text blocks at the same time. If typography is error-prone, reduce copy frequency instead of putting text in every shot.
- Two-part text motion rules: the first half and second half must continue or replace within the same line, never split into upper and lower rows; motion must be visible, with the first half fading in or sliding in subtly first, then the second half fading in or sliding in subtly later, while the first half shifts gently within the same line. Do not make the typography completely static.
- Shot continuity hard constraints.
- MiniMax-H3 native audio setting and music prompt; native audio is on by default unless the user explicitly asks for silence or another model.
- Do not pass the original product image.
## STEP 9: music-2.6 Tech BGM
Video generation defaults to MiniMax-H3 native audio first, and the Apple-style tech BGM direction below should be written into the video prompt. If H3's native music is unpleasant, too loud, too weak, out of sync, or the user asks for new music, use `music-2.6` to generate a longer standalone instrumental BGM and replace the video audio track. Do not default to ElevenLabs; use ElevenLabs only when the user explicitly requires a strict target duration, and follow the fee-confirmation rule.
Default music direction:
```text
Around 100 BPM, fast tech feeling, block chords, pluck and airy noise bed, kick + sub-bass + sine sweep, wooden percussion, sudden cut-off, within 0.5s everything stops, leaving only the pluck tail to decay.
```
Extended interpretation:
- Techy, Apple-style, product launch feeling.
- Rhythmic drive is allowed, but it must not become cheap EDM.
- Pluck should be crisp, pleasant, and forward; airy noise should feel premium.
- Kick and sub-bass serve edit points; avoid boomy bass.
- Wooden percussion adds tactile texture, not a complex drum loop.
- No vocals, no synthwave, no retro, no gaming, no cheap corporate stock music.
## STEP 10: Music Analysis, Beat-synced Editing, and Delivery Verification
Before final delivery, edit or assemble based on music analysis. Do not simply lay music under the video.
Music analysis:
- Detect rhythm, strong beats, energy sections, and clear single-note hits.
- Find an 8-12 second segment with continuous sound and suitable energy for product display.
- Do not blindly use the first 10 seconds.
- Prefer a segment with a clear entry point, no silence or collapse in the second half, and enough stable time for the closing frame.
Audio assembly rules:
- Keep the video image, cut a continuous music segment, and replace the video audio track.
- Mute the original video track; keep only the new music.
- Keep music loudness restrained: do not maximize loudness, do not use aggressive limiting, and do not raise music too much; but also do not push it so low that the pluck, rhythm points, and product display lose clarity.
- Do not use long fade-outs. Only apply a 0.1-0.3s tail fade when needed.
- After assembly, check for long silence or audible dropouts.
Pre-delivery verification:
- Video aspect ratio, duration, and language match the brief.
- Product appearance, main variant, and variant timing match the anchor sheet and storyboard.
- Copy is readable, single-line, not crowded, and contains no extra arrows or symbols; no two text rows, upper/lower line splits, title + subtitle pairs, or multiple text blocks appear at the same time.
- The final frame has a stable product + copy closing.
- Music is not too loud and not too low.
- Output is on canvas, and multi-asset outputs are grouped.
Delivery includes:
- Final video path or canvas output.
- Duration, aspect ratio, and language.
- Short creative summary.
- Copy used and music segment used.
- Whether beat-synced editing was applied.
- Specific next improvements, such as pacing, copy, main variant, audio, or platform crop.
## Failure Handling
- Product image too weak: stop, give reshoot guidance, and wait for a better image.
- Copy too long: compress to 3-5 English words and ask for approval.
- Text anchor is wrong: regenerate the text anchor before video generation.
- Anchor sheet fails product fidelity: regenerate with stronger product-preservation instructions.
- Storyboard is templated: return to the precise beat table and rewrite based on product form and narrative spine.
- Copy does not appear in video: rerun the video prompt, explicitly requiring copy in at least the middle and final beats.
- Video typography deforms: regenerate the text anchor or degrade the text shot to subtle anchor-based motion.
- Aspect ratio is wrong: immediately rerun with the user-selected ratio.
- Music quality is poor: retry with `music-2.6`, do not force duration, and cut the best segment from the longer track.
- Final assembly has no sound in the second half: re-analyze music, select a continuous audible segment, and use direct audio replacement.
## Trigger Examples
Use this Skill for requests like:
- “帮我的产品做一个苹果味儿广告片”
- “用这张产品图做 Apple 风电商广告”
- “做一个极简高级的产品发布视频”
- “Make an Apple-style product ad from this product photo”
- “Create a premium minimalist ecommerce product film”
Do not use this Skill for general video editing, documentary explainers, KOC talking-head ads, or complex UI/screen-text demos unless the user explicitly wants the Apple-style product ad workflow.
@@ -0,0 +1,19 @@
display-name-zh: 极简产品广告生成器
version: 0.5.6
tag-en: "Commercial Ad"
tag-cn: "商业广告"
complete-tags-en:
- "Commercial Ad / Planning"
- "Commercial Ad / Creative Generation"
- "E-Commerce / Creative Generation"
complete-tags-cn:
- "商业广告 / 策划"
- "商业广告 / 创作生成"
- "电商 / 创作生成"
summary-en: "Turn product images and ad requirements into minimalist product ad shorts for e-commerce promotion and product launches."
summary-cn: "基于产品图片和广告需求,完成卖点提炼、英文文案与节拍分镜设计,输出极简产品广告短片,适用于电商推广与新品发布。"
desc-en: "Based on product images and advertising requirements, extract selling points, write concise English ad copy, design product anchors and beat-synced storyboards, then generate a minimalist product advertising short for e-commerce promotion, product launches, small brands, food, toys, gadgets, accessories, and other physical goods. The Skill focuses on clean product presentation, readable typography rhythm, premium camera language, and final video delivery."
desc-cn: "基于产品图片和广告需求,完成卖点提炼、英文文案与节拍分镜设计,输出极简产品广告短片。适用于电商推广、新品发布、小品牌上新、食品、潮玩、数码配件和其他实体产品。Skill 会确认时长与比例,整理产品事实和主推卖点,生成简洁英文广告文案,制作产品锚定图,规划文字卡点分镜,并生成干净、有节奏、产品质感突出的成片。"
author-en: "User"
author-cn: "用户"
source: community
@@ -0,0 +1,213 @@
---
name: music-video-subtitle-generator
description: |
面向音乐人、视频创作者和社交媒体剪辑者,用于制作带歌词贴字的 AI MV 或情绪短片。用户提供音乐、歌词、参考图、角色、字体方向、情绪意图或发布平台。Skill 会分析节拍与人声时序,区分人物、场景和文字参考,设计随节奏变化的空间字幕,拆解长视频为可衔接镜头,审查提示词并路由 H3 等视频生成工具。最终输出 MV 概念、分镜提示词、歌词文字方案和拼接建议。适用于风格化音乐视觉和动态字幕 MV,不适用于普通字幕校对、照搬已有 IP 或完全手工后期剪辑。
trigger-words: [MV, music video, lyric typography, on-screen text, prompt audit, Trap MV, Gospel hip-hop, Dark-pop, Cyber-grunge, MV提示词, 歌词文字, 字幕MV, 贴字MV, 卡点MV, 多镜头拼接]
---
# 音乐美学MV
## 用途
当用户需要创建、修改、审查或生成音乐视频提示词、情绪短片提示词时使用本 Skill,尤其是音乐、歌词、贴字、参考图、节奏、人物表演和镜头语言需要统一设计的任务。本 Skill 将第三方 MV prompt 规则适配为 Hub 可执行流程:关键创意决策需要确认,锁定 prompt 写入画布文本节点,媒体生成交给对应 Hub agent。
不要用于普通字幕烧录、普通视频剪辑、非音乐类产品广告,或没有 MV 结构需求的单张图片 / 单段视频简单任务。
## Hub 兼容规则
- 不假设可直接调用第三方工具、shell 脚本、浏览器插件或外部生成 API。
- 锁定 prompt 和修订版本必须写入 Hub 画布文本节点。
- 角色卡、场景卡、文字卡、视频片段、BGM 和最终合成都交给 Hub 图片、视频、音乐、剪辑 agent。
- 不硬编码输出路径,始终使用工具返回的 Hub 文件路径。
- 用户说“跳过确认 / 直接做”时,只对本次运行生效,除非外层编排器已将其作为当前会话授权处理。
- 如果用户只要 prompt,交付 prompt 后停止;如果用户要完整 MV,在确认后继续生成和合成。
## 核心原则
1. 先理解作品,再选择结构。模板是组织工具,不是必须填满的表格。
2. 使用用户最新确认的创意意图。
3. 无关字段直接省略,不填空占位。
4. 不因为预设里有表演、文字、转场、运镜或角色行为,就机械添加到不适合的作品里。
5. 屏幕文字是设计图层,不是普通字幕;除非用户明确要求普通字幕。
6. 如果用户上传真实歌曲或 beat,它就是主音乐床,除非用户要求替换。
7. **多镜头自然衔接原则(针对 >15s 视频)**:当视频时长超过模型单次生成上限(如 15 秒)时,必须采用“多镜头分镜拆解 + 首尾帧接续 + 鼓点硬切 + 全局主音轨对齐”的标准拼接工作流。所有分段镜头必须保持同一音乐 Groove、速度感、画幅、角色逻辑、视觉预设、光影美学和文字运动规则,确保人声、音乐、节奏、画风与场景自然无缝过渡。
8. 当用户要完整音乐美学 MV 但没有提供歌词时,先生成并锁定原创歌词。最终 MV 必须同时参考人物卡、文字包装卡和场景卡生成。文字包装卡只控制文字包装样式、字体质感、图形设计、排版比例和动效语言。
## STEP 1:前期锁定
写最终生产 prompt 前,先确认最小基础。用简洁选项给出推荐。
### 1.1 视频格式
提供固定选项:
| 使用场景 | 画幅 | 分辨率 |
| :--- | ---: | ---: |
| TikTok / Reels / Shorts 竖屏 | 9:16 | 1080×1920 |
| YouTube / B站 横屏 | 16:9 | 1920×1080 |
| 信息流方图 / Feed | 1:1 | 1080×1080 |
| 投流竖屏高密度 | 9:16 | 720×1280 |
| 电影感宽银幕 MV | 21:9 | 2560×1080 |
所选比例必须贯穿角色卡、场景卡、文字参考、shot prompt、视频片段和最终合成。
### 1.2 目标时长与多镜头拆解架构
写 prompt 或生成视频前,必须先确认目标 MV 时长。即使用户没有上传音乐或歌词,也不能静默使用模型默认时长来替代用户意图。
提供简洁时长选项卡:
1. **10 秒测试版**:单镜头/双镜头,最快验证人物、场景、文字和音频风格。
2. **15 秒 hook 版**:2~4 个短镜头组,更完整的副歌 / hook / performance 短句。
3. **30 秒及以上完整 MV 版(多镜头拼接合成)**:由于模型存在 15 秒单次生成上限,系统将自动采用 **“多分镜(Multi-Shot)拼接工作流”**。先生成或锁定一首完整连续歌曲 / BGM,将 30 秒拆解为 4~8 个 2~5 秒的动态分镜头,通过卡拍(Beat-Sync)与首尾帧/场景连续性控制,拼接合成为无缝 MV。
4. **自定义时长**:用户输入目标时长;根据所选视频模型的单次生成上限(如 15 秒)规划多镜头分镜数量与拼接方案。
如果用户提供了上传音乐,确认的目标时长就是需要锁定的音乐窗口长度。
### 1.3 音乐窗口
如果上传或指定的音乐长于已确认目标时长,必须先锁定片段再写 prompt。
提供三种模式:
1. 推荐窗口:assistant 判断最强副歌 / hook / 情绪转折,并请求确认。
2. 用户时间戳:用户给出开始和结束秒数;检查时间落在音频内。
3. 多变体:只用于投流钩子测试或多创意方向。
如果音乐短于目标时长,使用完整音频。除非用户明确要求,不拉伸、不补长。
### 1.3.0 歌词锁定与歌词先行补全
写最终视频 prompt 或生成视频前,必须先确定歌词归属:
1. 如果用户提供了歌词,这份歌词就是锁定歌词。除非用户明确要求改词,否则不得再生成、添加、改写、扩写、翻译、转述或替换成其他歌词。
2. 如果用户想要完整音乐美学 MV 但没有提供歌词,才生成与所选音乐风格、vocal 模式、目标时长、情绪温度和视觉预设匹配的短篇原创歌词,并将其锁定。
3. 锁定歌词是人物说唱 / 演唱表演和可见文字包装内容的唯一源文本。
4. 从锁定歌词拆出每个分镜头的 `Rap line:`、`Vocal line:`、`Soft vocal line:` 或 `Spoken line:`。
5. 每个 `Typography:` 字段也必须来自同一份锁定歌词。有人声表演时,文字必须逐字匹配正在表演的词。
6. 在最终贴字 MV 中,人物必须根据锁定歌词说唱 / 演唱:可见嘴型、下颌动作、呼吸、面部重音、点头和手势重音都应跟随歌词 phrasing 和 rhythm。
### 1.3.1 大于 15 秒视频的多镜头拼接默认流程
对于 >15 秒的 MV(如 30 秒),必须以“全局多镜头分镜 Prompt 脚本”作为权威生成来源:
1. **锁定完整 Master 音频**:先锁定一首完整连续的歌曲/BGM轨(包含完整的 Vocal 与 Groove),作为整个拼接流程的唯一音频基准。
2. **构建多分镜时间轴(Shotlist Timeline)**:将 30 秒拆解为 4~8 个短镜头(每个镜头 2~5 秒),精确映射到歌词时间戳与鼓点(Snare/808/Drop)上。
3. **首尾帧与场景接续控制**:
- 若镜头 A 与镜头 B 为同一场景的长镜头延续:将镜头 A 的最后一帧(Tail Frame)作为镜头 B 的起始首帧(Head Frame)输入模型生成。
- 若镜头 A 与镜头 B 为硬切换景:保持相同的人物卡/服装/光影美学 Prompt,并使用同向运镜或视觉元素匹配切(Match Cut)。
4. **生成与剪辑组装**:视频 agent 按 Shotlist 生成各短片段,剪辑 agent 在全局 Master 音轨上根据拍子(Beat Grid)进行裁切、调速(Speed Ramping)与拼接,确保人声口型、鼓点卡拍与画面过渡完全自然。
## STEP 2:Creative Contract
最终 prompt 前先建立简洁创意合约:
- 音乐类型、器乐、速度感(BPM)、vocal 模式、情绪温度。
- 歌词来源(锁定歌词)。
- 目标时长与多镜头拆解结构(例如:30s 拆分为 6 个 Short Shots)。
- 参考图角色分工:人物卡、场景卡、文字包装卡。
- 运镜、景别、对焦、剪辑卡拍密度与转场衔接逻辑。
- 明确排除项(如:严禁淡入淡出、严禁油亮 AI 美颜脸、严禁单镜头硬撑导致变形)。
## STEP 3:参考图分工与场景采样
每张参考图只分配一个窄任务:
- **文字参考卡**:只控制文字包装样式、字体质感、图形设计、排版比例和动效语言。严禁出现人物或场景。
- **人物参考卡**:只控制人物形象、脸部气质、发型、服装轮廓、角色比例、姿态和整体气场。
- **场景参考卡**:只控制场景视觉风格、空间氛围、影像质感、背景层次和光影气质。
### 3.1 近景场景切换与多镜头采样规则
对于 Trap、Dark-pop、Cyber-grunge 等快节奏 MV,全片应在多个近景场景之间快速切换:
- 场景气质保持统一,但每个镜头的空间必须明显不同(局部背景、隧道、墙面、灯带、反光地面)。
- 场景切换必须由明确音乐冲击触发:bass hit、808 drop、snare、vocal 重音或文字砸屏。
- 镜头切换是 trap beat 上的“视觉采样”:硬切、跳切、扫描 glitch 硬切、闪切、动作匹配切。
## STEP 4:预设语法(以 Dark-pop / Cyber-grunge & Trap 为例)
### 视觉与剪辑语言
- 写实高时装质感、胶片杂志质感(90年代末-00年代初独立杂志/复印纸/胶片扫描/zine拼贴)。
- 粗颗粒、轻微胶片抖动、半色调网点、印刷毛边、扫描错位。
- **剪辑绝对只使用硬切(Hard Cut)**,严禁淡入淡出、溶解或柔和转场。
- 画面与文字强响应鼓点:hi-hat roll(微震/跳帧)、snare(放大/硬切/肩膀下压)、808 bass hit(低频压屏/拉伸/错位)。
### 文字包装规则
- 文字是空间中的动态图形主体(不是普通字幕条),可以处于前景、中景、背景或被人物肩膀/手部遮挡。
- 文字绝不遮挡眼睛或主要面部表情;对嘴时避免遮挡嘴部。
- 有人声时,显示文字必须逐字匹配正在表演的歌词。每个镜头只出现一个主文字事件。
## STEP 5:Prompt 结构模板
多镜头分镜脚本必须按镜头(Shot)进行模块化输出,每个 Shot 标注准确的时长与音频映射:
```text
[Global Aesthetic & Character Lock]: (全局人物与美学锚点 Prompt)
Shot 1 (0.0s - 3.5s)
Vocal Line: "..."
Typography: "..."
Visual & Action: ...
Camera & Motion: ...
Transition Out: Bass-hit 上硬切至 Shot 2 / 同向甩镜
Shot 2 (3.5s - 7.0s)
Vocal Line: "..."
Typography: "..."
Visual & Action: ...
Camera & Motion: ...
Transition Out: ...
```
## STEP 6:BGM 连续性与多镜头自然衔接系统
### 6.1 音频全局连续性
全片必须强绑定同一轨 Master Audio(完整歌曲/Beat)。在分段生成视频时,严禁使用独立、断裂的片段音轨。所有视频片段在剪辑合成阶段均对齐至 Master Audio 的对应时间戳(Timeline)。
### 6.2 多镜头自然衔接系统(Natural Stitching Protocol)
为保证 >15 秒视频拼接后的极其自然,必须在 Prompt 规划与后期合成中严格执行以下“五大衔接锁”:
**人声与口型衔接锁(Vocal & Lip Continuity)**:
分镜切替点(Cut Point)必须落在歌词的句间停顿(Pause/Breath)或强鼓点(Snare/Drop)上。严禁在人物发出某个元音或唱词中途进行硬切,除非后一个镜头是极近特写且口型完全接续。
**音乐与节奏衔接锁(Rhythm & Beat Matching)**:
剪辑点必须精准裁切在音乐的 1/4 或 1/8 拍(Beat Grid)上。利用画面动作的“加速/减速(Speed Ramping)”微调,使视频中人物的头点拍、手势或眨眼精准踩在鼓点上。
**画面美学与色彩衔接锁(Aesthetic & Color Grading)**:
跨镜头生成必须携带相同的 Aesthetic Header Prompt(相同的胶片颗粒度、色调 LUT、光影方向)。最终合成时,全局叠加一层 35mm Film Grain Overlay 和统一调色 LUT,掩盖跨批次生成的微小色差。
**空间与转场衔接锁(Transition & Motion Continuity)**:
长镜头延伸:使用上一镜头的末帧(Tail Frame)作为下一镜头的首帧(Head Frame)输入,Prompt 加上 continuation of action。镜头切换:采用动能延续转场(如 Shot 1 结尾向右快速 Pan,Shot 2 开头从左向右 Pan 入;或利用人物手势遮挡镜头完成 Match Cut)。
**文字动态衔接锁(Typography Motion Stitching)**:
跨镜头的文字动效,前一镜头的文字在切替瞬间跟随 Bass Hit 执行“破碎/扫出/砸屏”,下一镜头的文字在重音上“砸入/展开”,形成视觉动能的连贯传递。
## STEP 7:审查清单(Checklist)
交付前静默检查:
- [ ] 目标时长超过 15 秒时,是否已自动采用多镜头(Multi-Shot)拆解结构?
- [ ] 是否已锁定唯一的全局 Master Audio,且分镜切替点对齐到了音乐拍子上?
- [ ] 镜头切替是否避开了人声唱词中途,口型与呼吸是否自然?
- [ ] 相邻镜头之间是否有明确的转场衔接机制(首尾帧接续 / 同向运镜 / Match Cut / 鼓点硬切)?
- [ ] 剪辑风格是否严格保持“纯硬切”,绝无淡入淡出或柔和转场?
- [ ] 三张参考卡(人物/场景/文字)是否角色隔离,且未互相污染?
- [ ] 文字包装是否作为空间图层,且绝未遮挡眼睛和主要面部表情?
- [ ] 全局画质、颗粒感、色彩与光影是否保持高度一致?
## STEP 8:画布交付
完整 MV Prompt 脚本锁定后,必须写入专用画布文本节点,命名为 `Complete MV Prompt` 或 `完整MV Prompt`。节点内容必须是完整的多分镜 Prompt 脚本与衔接说明,后续修改时更新该节点。
## STEP 9:最终 MV 生成与合成流程
1. **Prompt 锁死与画布写入**:确认完整的多镜头 Prompt 脚本(包含时间戳、歌词映射、转场逻辑)已写入画布。
2. **分镜素材并行生成**:视频 Agent 根据 Shotlist 逐个生成 2~5 秒的短片段。需要长镜头接续的片段,提取上一片段末帧作为首帧图生视频(I2V)。
3. **多轨剪辑与自然缝合(Editing & Stitching)**:剪辑 Agent 导入全局 Master Audio 轨,将生成的各 Shot 视频按时间戳对齐上轨,在鼓点(Beat Grid)上进行精细剪辑与速度曲线微调(Speed Ramping),确保人声口型与画面动作完美卡拍。检查镜头接缝处,应用同向运镜或闪切遮瑕。
4. **全局调色与质感统一(Finishing)**:全局叠加统一的 Movie LUT 与 35mm 胶片颗粒 Overlay,消除 AI 塑料感并锁定画风一致性。
5. **输出交付**:输出无缝衔接的完整 MV 视频文件。
@@ -0,0 +1,205 @@
---
name: music-video-subtitle-generator
description: |
For musicians, video creators, and social-media editors producing AI music videos or emotional short films with lyric typography. Users provide music, lyrics, references, characters, typography direction, mood, or target platform. The Skill analyzes beat and vocal timing, separates character, scene, and text references, designs beat-reactive spatial typography, decomposes long works into connected shots, audits prompts, and routes generation for H3 or other video tools. It outputs MV concepts, shot prompts, lyric text plans, and stitching guidance. Best for stylized MVs and subtitle-driven music visuals, not ordinary caption cleanup, licensed IP copying, or fully manual post-production editing.
compatibility: Requires the MiniMax Hub agent (canvas workspace and Hub generation/routing tools); not portable to generic agent harnesses.
trigger-words: [MV, music video, lyric typography, on-screen text, prompt audit, Trap MV, Gospel hip-hop, Dark-pop, Cyber-grunge, MV提示词, 歌词文字, 字幕MV, 贴字MV, 卡点MV, 多镜头拼接]
---
# Music Aesthetics MV
## Purpose
Use this Skill when the user wants to create, revise, audit, or generate music-video prompts or emotional short-film prompts where music, lyrics, typography, references, rhythm, performance, and camera language must be designed together. The workflow adapts MV prompt rules into Hub execution: key creative decisions are confirmed, locked prompts are written to canvas text nodes, and media creation is delegated to Hub agents.
Do not use it for ordinary subtitle burn-in, generic video editing, non-music product ads, or simple single-image / single-clip requests without MV structure.
## Hub compatibility rules
- Do not assume third-party tools, shell scripts, browser plugins, or external generation APIs.
- Write locked prompts and revisions to Hub canvas text nodes.
- Delegate character cards, scene cards, typography cards, video clips, BGM, and final assembly to Hub image, video, music, and editing agents.
- Do not hardcode output paths; always use Hub-returned paths.
- If the user says “skip confirmation / just do it”, apply it only to the current run unless the outer orchestrator has active session authorization.
- If the user only wants a prompt, stop after delivery. If the user wants a finished MV, continue to generation and assembly after confirmation.
## Core principles
1. Understand the work before choosing a structure. Templates organize, but are not mandatory forms.
2. Use the latest confirmed creative intent.
3. Omit irrelevant fields instead of filling placeholders.
4. Do not mechanically add performance, typography, transitions, camera moves, or character actions just because a preset includes them.
5. On-screen text is a designed visual layer, not ordinary subtitles, unless the user explicitly requests subtitles.
6. If the user uploads a real song or beat, it is the master music bed unless they request replacement.
7. **Natural multi-shot stitching for >15s videos**: when duration exceeds the model single-generation limit, use “multi-shot storyboard breakdown + tail/head frame continuation + beat hard cuts + global master-audio alignment”. Every segment must preserve the same groove, tempo feel, aspect ratio, character logic, visual preset, lighting language, and typography motion rules.
8. If the user wants a finished music-aesthetic MV without lyrics, generate and lock original lyrics first. The final MV must reference character, typography packaging, and scene cards together. The typography packaging card controls only text packaging style, font texture, graphic design, layout ratio, and motion language.
## STEP 1: Pre-flight lock
Before writing the final production prompt, confirm the minimum foundation with concise recommended choices.
### 1.1 Video format
Offer fixed options:
| Use case | Aspect ratio | Resolution |
| :--- | ---: | ---: |
| TikTok / Reels / Shorts vertical | 9:16 | 1080×1920 |
| YouTube / Bilibili horizontal | 16:9 | 1920×1080 |
| Feed square | 1:1 | 1080×1080 |
| Dense vertical ad test | 9:16 | 720×1280 |
| Cinematic widescreen MV | 21:9 | 2560×1080 |
The selected ratio must be reused across character cards, scene cards, typography references, shot prompts, video clips, and final assembly.
### 1.2 Target duration and multi-shot structure
Before prompt writing or video generation, confirm target MV duration. Never silently use the model default as user intent.
Offer a compact duration card:
1. **10-second test**: one or two shots to verify character, scene, typography, and audio style.
2. **15-second hook**: 2–4 short shots for a more complete chorus / hook / performance phrase.
3. **30-second or longer complete MV (multi-shot stitching)**: because many models have a 15-second single-generation ceiling, automatically use a Multi-Shot stitching workflow. First generate or lock one continuous song / BGM, split 30 seconds into 4–8 dynamic 2–5 second shots, then stitch via beat sync and head/tail frame continuity.
4. **Custom duration**: user provides duration; plan shot count and stitching based on the selected model’s single-clip limit.
If uploaded music exists, the confirmed duration is the music window length.
### 1.3 Music window
If uploaded or specified music is longer than the confirmed duration, lock a segment before writing prompts:
1. Recommended window: assistant identifies the strongest chorus / hook / emotional turn and requests confirmation.
2. User timestamps: user provides start and end seconds; verify both are inside the audio.
3. Multi-variant: only for ad-hook testing or multiple creative directions.
If music is shorter than the target duration, use the full audio unless the user explicitly asks to stretch or extend it.
### 1.3.0 Lyrics lock and lyrics-first fallback
Before final video prompts or generation, determine lyric ownership:
1. If the user provides lyrics, those are locked lyrics. Do not generate, add, rewrite, expand, translate, paraphrase, or replace them unless explicitly requested.
2. If the user wants a finished music-aesthetic MV without lyrics, generate short original lyrics matching the music style, vocal mode, duration, emotion, and visual preset, then lock them.
3. Locked lyrics are the only source text for rap/singing performance and visible typography.
4. Split locked lyrics into each shot’s `Rap line:`, `Vocal line:`, `Soft vocal line:`, or `Spoken line:`.
5. Every `Typography:` field must come from the same locked lyrics. During vocal performance, visible text must word-for-word match the performed phrase.
6. In the final typography MV, the performer must rap or sing the locked lyrics with visible lip shapes, jaw motion, breath, facial accents, nods, and hand accents following phrasing and rhythm.
### 1.3.1 Default multi-shot stitching for >15s videos
For >15s MVs such as 30 seconds, the global multi-shot prompt script is the authoritative source:
1. **Lock complete Master Audio**: lock one continuous song / BGM track with the full vocal and groove as the only audio baseline.
2. **Build a Shotlist Timeline**: split 30 seconds into 4–8 short shots of 2–5 seconds, mapped precisely to lyric timestamps and beats such as snare, 808, and drop.
3. **Head/tail frame and scene continuity**: if Shot B continues the same scene, use Shot A’s tail frame as Shot B’s head frame. If it is a hard cut, preserve the same character card, wardrobe, lighting prompt, and use same-direction camera motion or match-cut visual elements.
4. **Generation and edit assembly**: the video agent generates each shot from the Shotlist. The editing agent aligns all clips to the global Master Audio beat grid, trimming, speed-ramping, and stitching for natural lip-sync, beat sync, and transitions.
## STEP 2: Creative Contract
Before final prompts, create a concise creative contract:
- Music genre, instrumentation, tempo/BPM feel, vocal mode, and emotional temperature.
- Lyric source and locked lyrics.
- Target duration and multi-shot breakdown, e.g. 30s split into 6 short shots.
- Reference image roles: character card, scene card, typography packaging card.
- Camera language, shot sizes, focus, beat-cut density, and transition logic.
- Exclusions such as no fades, no glossy AI beauty face, no single-shot overstretching that causes deformation.
## STEP 3: Reference roles and scene sampling
Assign each reference image one narrow job:
- **Typography reference card**: controls only text packaging style, font texture, graphic design, layout ratio, and motion language. It must not contribute people or scenes.
- **Character reference card**: controls character identity, facial aura, hairstyle, clothing silhouette, proportions, posture, and presence.
- **Scene reference card**: controls scene style, spatial atmosphere, image texture, background depth, and lighting mood.
### 3.1 Close-up scene switching and multi-shot sampling
For Trap, Dark-pop, Cyber-grunge, and other fast-paced MVs, switch quickly among multiple close-up spaces:
- Keep scene mood unified, but each shot’s space must differ clearly, such as local background, tunnel, wall, light strip, or reflective floor.
- Scene switches must be triggered by musical impacts: bass hit, 808 drop, snare, vocal accent, or typography smash.
- Shot switches are “visual samples” on the trap beat: hard cuts, jump cuts, scan-glitch cuts, flash cuts, or action match cuts.
## STEP 4: Preset grammar, Dark-pop / Cyber-grunge and Trap example
### Visual and editing language
- Realistic high-fashion texture and film-magazine texture, referencing late-90s / early-00s indie magazines, photocopy paper, film scans, and zine collage.
- Heavy grain, slight film jitter, halftone dots, rough print edges, scan offsets.
- **Editing must use hard cuts only**. No fades, dissolves, or soft transitions.
- Image and typography strongly respond to beat cues: hi-hat roll creates micro-shake / frame skips, snare triggers scale-up / hard cut / shoulder drop, 808 bass hit creates low-frequency compression / stretch / offset.
### Typography packaging rules
- Typography is a dynamic graphic subject in space, not a subtitle bar. It may sit in foreground, midground, background, or be occluded by shoulders / hands.
- Text must never cover eyes or main facial expression; avoid the mouth during critical lip-sync.
- With vocals, visible words must exactly match the performed lyrics. Each shot has only one main typography event.
## STEP 5: Prompt structure template
A multi-shot script must be modular by shot, with accurate duration and audio mapping:
```text
[Global Aesthetic & Character Lock]: (global character and aesthetic anchor prompt)
Shot 1 (0.0s - 3.5s)
Vocal Line: "..."
Typography: "..."
Visual & Action: ...
Camera & Motion: ...
Transition Out: hard cut to Shot 2 on bass hit / same-direction whip pan
Shot 2 (3.5s - 7.0s)
Vocal Line: "..."
Typography: "..."
Visual & Action: ...
Camera & Motion: ...
Transition Out: ...
```
## STEP 6: BGM continuity and natural multi-shot stitching system
### 6.1 Global audio continuity
The entire MV must bind to one Master Audio track. During segmented video generation, do not use independent disconnected clip audio. All video clips are aligned to their corresponding timestamps on the Master Audio timeline during edit assembly.
### 6.2 Natural Stitching Protocol
For >15s stitching, enforce five continuity locks:
**Vocal & Lip Continuity**: Cut points must land on lyric pauses, breaths, snare, or drop. Do not hard cut inside an active vowel or lyric unless the next shot is an extreme close-up with continuous mouth shape.
**Rhythm & Beat Matching**: Cut points must hit the 1/4 or 1/8 beat grid. Use speed ramping to align head nods, hand gestures, and blinks to beats.
**Aesthetic & Color Grading**: Every cross-shot generation carries the same aesthetic header prompt: grain level, LUT direction, and light direction. Final assembly uses unified 35mm film grain and LUT to hide batch color differences.
**Transition & Motion Continuity**: For long-shot continuation, use the previous tail frame as the next head frame and prompt continuation of action. For hard scene switches, use kinetic continuity such as same-direction pan or hand-occlusion match cut.
**Typography Motion Stitching**: Text in the previous shot should shatter, sweep out, or smash-screen on a bass hit; next-shot text smashes in or unfolds on the accent, transferring visual momentum.
## STEP 7: Checklist
Silently verify before delivery:
- [ ] If target duration exceeds 15s, has Multi-Shot breakdown been used automatically?
- [ ] Is one global Master Audio locked, and are cut points beat-aligned?
- [ ] Do cuts avoid mid-lyric vocal breaks, with natural lip and breath continuity?
- [ ] Do adjacent shots have clear stitching logic such as tail/head continuation, same-direction motion, match cut, or beat hard cut?
- [ ] Is the edit style strictly hard-cut, with no fades or soft transitions?
- [ ] Are character / scene / typography cards isolated without cross-contamination?
- [ ] Is typography a spatial layer that never blocks eyes or main facial expression?
- [ ] Are image quality, grain, color, and lighting consistent globally?
## STEP 8: Canvas delivery
After the complete MV prompt script is locked, write it to a dedicated canvas text node named `Complete MV Prompt` or `完整MV Prompt`. The node must contain the full multi-shot prompt script and stitching notes. Later revisions update the same node.
## STEP 9: Final MV generation and assembly workflow
1. **Prompt lock and canvas write**: confirm the full multi-shot prompt script with timestamps, lyric mapping, and transition logic is written to canvas.
2. **Parallel shot generation**: the video agent generates each 2–5s shot from the Shotlist. For continuation shots, extract the prior shot’s tail frame as the next I2V head frame.
3. **Multi-track editing and stitching**: the editing agent imports the global Master Audio, aligns all shot videos by timestamp, trims and speed-ramps on the beat grid, and checks seams with same-direction motion or flash-cut masking.
4. **Finishing**: apply unified Movie LUT and 35mm film grain overlay to reduce AI plasticity and lock visual consistency.
5. **Delivery**: output the seamless complete MV video file.
@@ -0,0 +1,19 @@
display-name-zh: 音乐MV动态字幕生成器
version: 0.6.6
tag-en: "Sound & Music"
tag-cn: "音频音乐"
complete-tags-en:
- "Sound & Music / Planning"
- "Sound & Music / Creative Generation"
- "Creative & Experimental / Creative Generation"
complete-tags-cn:
- "音频音乐 / 计划制定"
- "音频音乐 / 创作生成"
- "创意实验 / 创作生成"
summary-en: "Create beat-synced MV prompts and lyric typography plans from music and lyrics."
summary-cn: "基于音乐和歌词,完成节奏拆解、镜头设计与动态字幕方案,适用于AI MV创作。"
desc-en: "For musicians, video creators, and social-media editors producing AI music videos or emotional short films with lyric typography. Users provide music, lyrics, references, characters, typography direction, mood, or target platform. The Skill analyzes beat and vocal timing, separates character, scene, and text references, designs beat-reactive spatial typography, decomposes long works into connected shots, audits prompts, and routes generation for H3 or other video tools. It outputs MV concepts, shot prompts, lyric text plans, and stitching guidance. Best for stylized MVs and subtitle-driven music visuals, not ordinary caption cleanup, licensed IP copying, or fully manual post-production editing."
desc-cn: "面向音乐人、视频创作者和社交媒体剪辑者,用于制作带歌词贴字的 AI MV 或情绪短片。用户提供音乐、歌词、参考图、角色、字体方向、情绪意图或发布平台。Skill 会分析节拍与人声时序,区分人物、场景和文字参考,设计随节奏变化的空间字幕,拆解长视频为可衔接镜头,审查提示词并路由 H3 等视频生成工具。最终输出 MV 概念、分镜提示词、歌词文字方案和拼接建议。适合风格化音乐视觉,不适合普通字幕校对、照搬已有 IP 或完全手工后期剪辑。"
author-en: "MiniMax Hub User"
author-cn: "MiniMax Hub 用户"
source: official-featured
@@ -0,0 +1,271 @@
---
name: paper-collage-explainer-generator
description: |
面向内容创作者、教育讲解者和社交视频编辑,用触感纸拼贴语言表现口播句、知识点、观点或抽象主题。用户提供短文案、故事节点或核心概念,也可补充画幅、时长、色调和音频需求。Skill 会提炼含义与视觉隐喻,制定制作方案和分镜,生成并确认半调纸拼贴静帧,再制作带纸片滑入、弹入、轻敲、压平和摩擦声的停格动画片段,并可按需合成完整讲解视频。默认保留拼贴音效,不主动添加 BGM、旁白和字幕。适合知识讲解、观点表达、故事配画和社交 B-roll,不适合真人口播广告、精确可编辑图层、复杂文字排版或仅输出提示词。
trigger-words: [paper collage explainer, paper-collage animation, halftone collage, collage explainer, 纸拼贴, 拼贴科普, 定格拼贴, 拼贴动画]
---
# 纸拼贴讲解动画生成器
将一句口播、一个故事主题、观点句或抽象概念,转化为统一的编辑感纸拼贴动画序列。视觉语言是高级半调纸拼贴:大色块纸面、黑白半调照片剪影、选择性彩色卡纸点缀、暖白描边、柔和纸影、触感定格组装,以及清晰的拼贴音效。
这个 Skill 使用 Hub 原生图像、视频、音频和可选后期能力。它优先保证风格连续、色彩统一、纸张质感可控、定格拼贴节奏清楚,并采用新的音频策略:默认保留或生成触感拼贴音效,但明确默认不添加 BGM、旁白口播和字幕,除非用户要求。
## 适用场景
当用户需要以下内容时使用:
- 把一句脚本或观点转成视觉隐喻拼贴动画
- 用拼贴 B-roll 解释简单故事、文学主题或知识点
- 为口播或社交视频制作编辑感纸拼贴动画
- 制作从空色纸背景中逐步组装物件的半调拼贴动画
- 把多条抽象句子或故事节点批量转为独立视觉隐喻动画
- 制作带纸片滑动、弹入、压平、轻响和摩擦声的触感知识短片,但不默认加音乐或口播
不适用于:
- 写实产品广告或真人出镜口播视频
- 精确可编辑图层、时间线关键帧或透明切片素材
- 精确 Logo 位置或可读长文字排版
- 只需要视频提示词、不需要生成媒体
## 默认创作目标
除非用户另有说明:
- 输出画幅:16:9 横版
- 单段时长:每段约 4 秒
- 默认音频:保留或生成触感纸拼贴音效,例如纸片滑动、弹入、压平轻敲、轻微摩擦和纸片轻响
- 默认不添加 BGM:可以询问用户是否需要 BGM,但只有用户明确要求时才添加音乐
- 默认不添加旁白口播:可以询问用户是否需要旁白讲解/口播,但只有用户明确要求时才写稿、配音或加入人声
- 默认不添加字幕:可以询问用户是否需要字幕,但只有用户明确要求时才生成或烧录字幕
- 视觉风格:高级编辑感半调纸拼贴
- 运动风格:触感停格组装,不使用慢速缩放、泛泛漂浮或平滑数字图层移动
- 默认视频生成模型:`MiniMax-H3`,除非用户明确指定其他模型、该模型不可用,或硬性能力要求排除它
- 图像/视频中文字:避免可读字母、数字、UI、字幕、水印和 Logo
- 图像质量与层次:优先生成好看的 16:9 层次构图,前景/中景/背景清楚,主体层级强,画面丰富但易读,留白可控
## 音频策略
这个 Skill 的默认交付是:**带拼贴音效,默认不添加 BGM、旁白口播和字幕**。
1. 在第一次制作方案确认时,必须明确写出默认媒体方式:保留或生成触感纸拼贴音效;默认不添加 BGM、旁白口播和字幕,除非用户明确要求。
2. 对科普或讲解类内容,可以询问用户是否需要 **旁白讲解/口播**、**BGM** 或 **字幕**,但必须把它们作为可选增强,而不是默认项。
3. 不要因为用户要“科普/讲解视频”就自动推断需要人声旁白。若用户没有选择旁白,就写无声视觉叙事节奏,而不是旁白稿。
4. 不要因为用户要“社交视频”就自动推断需要 BGM 或字幕。若用户没有选择,就只保留拼贴音效。
5. 生成视频片段时,如果模型支持音频,默认要求同步触感拼贴音效:纸片滑动、弹入、压平轻敲、轻微摩擦和小纸片脆响。
6. 最终合成时,如果原片段音轨是拼贴音效,默认保留这些原始音轨,不要默认丢弃音频。
7. 只有当用户要求静音、生成音频中出现不需要的人声/音乐,或用户要求另行混入音乐/旁白时,才移除或替换音频。只有用户明确要求时才生成或烧录字幕。
## 全局风格规则
对每张静帧和每段视频应用以下规则:
1. **统一整体风格。** 每段都应像同一套编辑感纸拼贴系列:半调剪影、大色块、暖白描边、柔和实体纸影、干净构图、触感纸材和协调的拼贴音效。
2. **控制纸张质感强度。** 纸张不能像完全扁平数字图层,也不能变得过旧、脏、皱或偏棕,除非用户明确要求。优先使用干净精致的手工纸质感:细纤维、轻微不规则撕边、浅层毛边、层叠接缝和软阴影。
3. **统一色彩基调。** 不要在未经确认时引入与静帧或整体色板冲突的牛皮纸、棕色、发黄或做旧底纸。每段从与已确认静帧主色一致的纸面开始。
4. **让运动像停格拼贴。** 使用明确纸片动作:逐片出现、滑入或弹入、轻微回弹、压平、暂停,然后锁定最终构图。避免快速旋转、过度翻飞、混乱飞散、全局淡入、平滑数字平移/缩放或泛泛漂浮。
5. **保持段落协调。** 当前一两段建立了批次风格和音效节奏后,后续片段和修订应参考该节奏与色调,让整支片子像同一系列。
6. **默认制作精致 16:9 层次场景。** 除非用户明确要求其他平台画幅,默认按 16:9 横版规划和生成。利用宽画幅建立前景/中景/背景层次、丰富环境道具、清晰主体层级和横向构图,不让画面变碎。
7. **突出纸拼贴工艺。** 每张静帧和每段视频都应让拼贴方法可见:可分离纸片组、半调照片剪影、彩色卡纸点缀、触感阴影、撕边、层叠接缝,以及弹入、滑入、轻微回弹、压平、暂停、锁定等停格动作。
## STEP 1:解析输入
针对每条句子、概念、故事或主题,提炼:
- 核心含义:观众应该理解什么
- 情绪:平静、紧迫、讽刺、惊奇、荒诞、澄清、反思、神秘或轻快
- 动作动词:打开、连接、泄漏、归档、压缩、分裂、照亮、绑定、组装、揭示、坠落、追逐、转化、碰撞
- 视觉隐喻:不用屏幕文字也能表达含义的具体画面
- 关键物件:3–6 个大而易读的纸片组
- 音效含义:该段适合纸片滑动、弹入、压平、轻敲、摩擦还是脆响
若是故事主题,除非用户指定段落数,否则拆为 3–6 个简洁节点。相似节点可共享设计语言,但每段应有不同隐喻、物件、色场和音效节奏。
## STEP 2:Gate 1 — 制作方案确认
在生成任何媒体前,先创建简洁制作方案文档,并等待用户确认。用户确认前不要生成静帧或视频。
制作方案必须包含:
### Brief
- 主题或原始句子
- 已知的目标受众/使用场景
- 画幅和时长假设
- 语气与节奏
- 视觉风格摘要
- 媒体方案,明确写为:默认保留或生成拼贴音效;默认不添加 BGM、旁白口播和字幕,除非用户明确要求
- 可选增强问题:需要时可以询问用户是否要旁白讲解/口播、BGM 或字幕,但必须保持为可选项
### 视觉隐喻
每个计划段落包含:
- 核心含义
- 情绪
- 一句话视觉提案
- 3–6 个关键物件
- 建议背景色和强调色
- 预计组装顺序
- 预计拼贴音效点,例如滑入、弹入、压平、轻敲、摩擦或脆响
### 脚本 / 视觉节奏轨
只有当用户明确要求旁白时,才写简洁旁白稿。若用户没有明确要求旁白,不要写旁白稿;改为写无旁白视觉节奏轨,说明观众将通过画面理解什么。
如果用户只需要为已有句子制作 B-roll,则保留原句作为上下文,不要发明旁白。
### 分镜
每个段落包含:
- 段落标题
- 最终画面描述
- 动效想法
- 近似时长
- 拼贴音效想法
- 风格连续性和色彩协调说明
展示制作方案后,等待用户批准、否定或修改。如果用户只批准部分编号,只推进已批准部分并修改其余部分。
## STEP 3:建立静帧规格
制作方案确认后,为每个已批准段落写紧凑的视觉规格,规格应能直接用于图像生成。
包含:
- 脚本含义或故事节点
- 视觉隐喻
- 画幅
- 背景色场
- 强调色
- 关键物件组及作用
- 构图、前景/中景/背景层次与留白
- 最终画面关系
- 风格连续性说明
- 该静帧的避免项
使用以下风格签名:
```text
flat bold color field, black-and-white halftone photographic cut-outs, selective colored cardstock accents, warm cream keylines, soft paper shadows, fine uncoated-paper grain, premium editorial paper collage, clean refined hand-torn paper edges, subtle fibrous edges, layered paper seams
```
色彩建议:
- 焦橙或红色:劳动、时间压力、紧迫感
- 芥末黄:工具、警示、累积错误
- 墨绿:认知、重置、判断、超现实平静
- 深紫:记忆、结构、神秘、梦境逻辑
- 青绿:协作、执行、系统流动
- 玫红:荒诞、仪式、戏剧张力
不要把钴蓝作为自动默认色。通过纸张质感、半调处理、描边、阴影、画幅和受控色板统一整批画面。只有当故事节点需要对比时才变化色场。
## STEP 4:Gate 2 — 生成并确认静帧
为每个已批准段落生成一张最终静帧。静帧应看起来像未来动画的完成最后一帧。
静帧要求:
- 使用用户指定画幅;否则默认 16:9 横版
- 扁平大色块纸背景,带细微纤维质感
- 3–6 个大的可分离物件组,除非已批准分镜确实需要略多
- 主体清楚、视觉层级强、留白干净,并有可见前景/中景/背景层次
- 以黑白半调照片剪影为主要结构,结合丰富但易读的场景道具支撑故事节点
- 仅在强化层级时使用选择性彩色卡纸强调
- 暖白描边、柔和实体纸影、精致撕边和轻微层叠纸缝
- 不出现可读文字、假字母、数字、字幕、UI、水印或 Logo,除非用户明确要求
展示生成的静帧,并在视频生成前等待用户确认。如果静帧太忙、色彩不匹配、纸层过多、边缘太数字化或纸张过旧/过棕,应先修订静帧再生成视频。
## STEP 5:规划停格组装
对每张已批准静帧,准备运动方案,并把已批准静帧视为完成最终画面。
默认运动顺序:
1. 从与已批准静帧主背景一致的干净纸色场开始,而不是不匹配的牛皮纸或棕色底。
2. 基础结构作为纸片滑入、弹入或压入。
3. 角色、卡片、物件或主要隐喻元素入场。
4. 次要物件逐个组装,带小幅回弹、压平和暂停。
5. 最终关系锁定到已批准构图。
6. 结尾短暂停留在完整静帧式画面上。
运动提示词指导:
```text
Paper-collage stop-motion assembly. Start on a clean paper field matching the approved still's background color and tone. Assemble the scene piece by piece with tactile stop-motion timing: foreground, midground, and background paper groups appear in a readable order; each paper object appears, slides or pops in, lightly bounces, presses flat, pauses, and locks into place. Preserve the halftone dots, cream keylines, subtle fibrous torn edges, layered paper seams, soft shadows, paper grain, color field, and approved aspect ratio. End by holding the completed composition matching the approved still frame. Sound design: tactile collage SFX only, synchronized to paper-piece motion: soft paper slide, pop-in, press-flat tap, light rustle, and tiny paper snap where appropriate. Default: do not add BGM, voiceover, dialogue, or subtitles unless the user explicitly requested them. No mismatched kraft-paper opening, no brown/yellowed base unless approved, no fast spinning, no chaotic object flight, no smooth digital layer movement, no scene cuts, no camera move, no zoom, no morphing, no new objects, no text, no logos, no watermark, no UI.
```
## STEP 6:生成带拼贴音效的 B-roll 片段
Gate 2 确认后,为每张已批准静帧生成一段视频。
生成要求:
- 如果所选 Hub 视频模型支持,使用已批准静帧作为视觉参考/最终帧锚点
- 使用与静帧相同的画幅
- 除非用户要求其他时长,否则目标约 4 秒
- 默认生成或保留触感拼贴音效:纸片滑动、弹入、压平轻敲、轻微摩擦、小纸片脆响
- 默认不生成或添加 BGM,除非用户明确要求音乐
- 默认不生成或添加人声讲解、口播或主持人音频,除非用户明确要求旁白/口播
- 如果生成片段中有可用拼贴音效,后期和最终合成时保留它
- 如果生成片段出现不需要的人声、音乐、假 UI 音或嘈杂环境声,根据用户目标移除或重新生成音频
- 保持镜头视觉稳定;除非用户要求,不增加运镜
- 保持拼贴材质语言、最终构图和已建立批次风格
- 修订后续片段时,可参考已通过的最佳前序片段风格和节奏
默认使用 `MiniMax-H3` 生成视频。除非 `MiniMax-H3` 失败、不可用或无法满足硬性能力需求,不要要求用户选择模型。如果用户明确指定其他模型,保留该模型为主并只在必要时调整参数。
## STEP 7:质量检查
按以下标准检查静帧和视频:
- 视觉隐喻或故事节点无需屏幕文字也能理解
- 静帧默认具有漂亮的 16:9 构图、清晰色场、明确物件层级和前景/中景/背景深度
- 物件组数量可读,不是一屏碎片
- 片段从与已批准静帧或批次色板一致的色场开始
- 组装是逐片可见的,不是整体淡入
- 运动像定格拼贴:独立纸片弹入或滑入、轻微回弹、压平、暂停、锁定
- 默认保留触感拼贴音效:能听到纸片滑动、弹入、压平、摩擦等声音;除非用户明确要求,不应有 BGM、旁白口播或字幕
- 纸张质感可控:不太平、不太旧、不偏棕/牛皮纸,除非已确认
- 色调在整批中一致
- 最终帧接近已批准静帧
- 没有不需要的可读文字、假 UI、字幕、水印或 Logo
只要隐喻清楚、最终构图仍像已批准静帧,微小细节漂移可以接受。若出现严重漂移、文字、假 UI、开场颜色不匹配、未经允许的牛皮纸底、过脏过皱纸质、丢失停格组装、不需要的 BGM 或不需要的旁白,应根据问题发生阶段重生成或后期处理。
## STEP 8:可选合成、音乐、旁白与交付
如果用户要求多片段故事视频,按确认的分镜顺序合成已批准片段,并默认保留每段原始拼贴音效。
只有当用户明确要求时才添加 BGM、旁白口播或字幕。若用户在片段已生成后要求 BGM 或旁白,应在保留拼贴音效的基础上混入,除非用户要求静音或替换原音效。
每个完成项交付:
- 最终 B-roll 视频路径或最终合成视频路径
- 相关已批准静帧路径
- 一句话说明脚本或视觉节点如何转化为视觉隐喻
- 媒体说明:默认保留拼贴音效,并列出用户明确要求加入的 BGM、旁白口播或字幕
- 若片段有轻微可接受漂移,说明 QA 注意事项
批量任务按编号组织结果,并保留用户原始句子或故事节点。
## 故障处理
- 如果制作方案默认写成无声视频,应改为“默认带拼贴音效”,除非用户明确要求静音。
- 如果制作方案在用户未要求时加入 BGM、旁白口播或字幕,必须移除,并把 BGM、旁白口播和字幕作为可选项询问。
- 如果制作方案太直白,回到 Gate 1,把视觉提案改得更物件化。
- 如果静帧包含假文字或 UI,先重生成静帧,不要只在视频提示词里修复。
- 如果静帧或视频太忙,减少纸层、简化物件数量,恢复留白。
- 如果视频缺少组装动作,减少物件数量,并明确逐步到达顺序。
- 如果运动像平滑数字图层,围绕停格节拍重写提示:出现 → 轻微回弹 → 压平 → 暂停 → 锁定。
- 如果纸张太平,加入轻微撕边、纤维、层叠接缝和软阴影。
- 如果纸张过旧或脏,去掉牛皮纸、棕黄底、重皱纹、污渍和做旧处理。
- 如果开场颜色与静帧冲突,从与静帧主背景一致的颜色开始。
- 如果最终帧偏离静帧太远,加强“已批准静帧是完成最终构图”的指令。
- 如果生成音频缺少拼贴音效但画面可用,询问用户保留、重生成更强音效方向,或后期补同步音效。
- 如果生成音频出现不需要的 BGM 或旁白,交付前移除或重生成。
- 如果用户需要精确图层控制,建议改用分层动画或剪辑工作流,而不是继续本 Skill。
@@ -0,0 +1,273 @@
---
name: paper-collage-explainer-generator
description: |
For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics. Users provide source copy, story beats, or a core concept and may specify aspect ratio, duration, palette, and audio needs. The Skill extracts meaning, proposes visual metaphors, prepares a production plan and storyboard, generates approved halftone collage stills, then creates stop-motion clips with paper movement and tactile sound effects, with optional final assembly. By default it keeps collage SFX and does not add BGM, voiceover, or subtitles unless requested. Best for explainers, viewpoints, story visuals, and social B-roll; not for presenter ads, editable layers, complex typography, or prompt-only tasks.
compatibility: Requires the MiniMax Hub agent (canvas workspace and Hub-native image/video/audio tools); not portable to generic agent harnesses.
trigger-words: [paper collage explainer, paper-collage animation, halftone collage, collage explainer, 纸拼贴, 拼贴科普, 定格拼贴, 拼贴动画]
exported-by: MiniMax-hub
---
# Paper Collage Explainer Generator
Turn a short narration line, story topic, viewpoint sentence, or abstract idea into a cohesive editorial paper-collage animation sequence. The visual language is premium halftone paper collage: flat bold color fields, black-and-white photographic cut-outs, selective colored cardstock accents, warm cream keylines, soft paper shadows, tactile stop-motion assembly, and crisp collage sound effects.
This Hub-adapted Skill uses Hub-native image, video, audio, and optional postprocess capabilities. It prioritizes style continuity, color harmony, controlled paper texture, stop-motion collage rhythm, and an audio policy that keeps tactile collage SFX by default while explicitly not adding BGM, voiceover, or subtitles unless the user requests them.
## When to Use
Use this Skill when the user wants:
- A short script line turned into a visual-metaphor collage animation
- A simple story or literary topic explained through collage B-roll
- Editorial paper-collage animation for narration or social video
- Halftone collage animation with objects assembling from an empty color field
- A batch of short abstract sentences or story beats converted into separate visual-metaphor animation clips
- A tactile knowledge explainer that may include paper clicks, pops, slides, presses, and rustles but not automatic music or voiceover
Do not use this Skill when the user needs:
- A realistic product ad or presenter-led video
- Precise editable layers, timeline keyframes, or transparent cut-out assets
- Exact logo placement or readable typography
- Only a written video prompt without generation
## Default Creative Targets
Unless the user specifies otherwise:
- Output ratio: 16:9 landscape
- Clip length: about 4 seconds per segment
- Audio default: keep or generate tactile paper-collage sound effects only, such as paper slide, pop, press, light rustle, and soft tap sounds
- Default: do not add BGM. You may ask whether the user wants BGM when it may help, but add music only after explicit user confirmation
- Default: do not add voiceover, spoken narration, or presenter narration. You may ask whether the user wants narration/口播 when the project may benefit, but write, synthesize, or add spoken audio only after explicit user confirmation
- Default: do not add subtitles. You may ask whether the user wants subtitles, but create or burn in subtitles only after explicit user confirmation
- Visual style: premium editorial halftone paper collage
- Motion style: tactile stop-motion assembly, not slow zoom, generic drifting, or smooth digital layer movement
- Default video generation model: `MiniMax-H3`, unless the user explicitly specifies another model, the model is unavailable, or a hard capability requirement excludes it
- Text in image/video: avoid readable letters, numerals, UI, subtitles, watermark, and logos
- Image quality and depth: prioritize visually attractive, layered 16:9 compositions with clear foreground, midground, and background depth, strong subject hierarchy, rich but readable scene design, and controlled negative space
## Audio Policy
This Skill's default delivery is **with collage SFX, and without BGM, voiceover, or subtitles**.
1. During the first production-plan confirmation, state the default media approach explicitly: tactile paper-collage sound effects are kept or generated; BGM, voiceover/口播, and subtitles are not added by default.
2. It is acceptable to ask the user whether they want **voiceover narration/口播**, **BGM**, or **subtitles**, especially for explainers, but present all three as optional add-ons, not defaults.
3. Do not infer that an explainer automatically needs spoken narration. If the user does not choose narration, create visual story beats rather than a voiceover script.
4. Do not infer that a social video automatically needs BGM or subtitles. If the user does not choose them, keep the clip SFX only.
5. When generating video clips, request synchronized tactile collage SFX if the selected video model supports audio: paper pieces sliding, popping, pressing flat, soft taps, and light paper rustles.
6. During final assembly, preserve the original clip audio tracks when they contain collage SFX. Do not drop audio by default.
7. Remove or replace audio only when the user asked for silence, when the generated audio contains unwanted speech/music, or when the user requests a separate music/narration mix. Create subtitles only when the user explicitly asks for them.
## Global Style Rules
Apply these rules to every still and video in the project:
1. **Unify the overall visual style.** Every segment should feel like part of the same editorial paper-collage series: halftone cut-outs, flat color fields, warm cream keylines, soft physical shadows, clean composition, tactile paper material, and coordinated collage SFX.
2. **Control the paper texture intensity.** Paper must not look perfectly flat, but it also must not become over-aged, dirty, wrinkled, or brown unless the user explicitly asks. Prefer clean, refined hand-made texture: subtle fiber, slightly irregular torn edges, light deckled fibers, layered seams, and soft shadows.
3. **Unify the color tone.** Do not introduce kraft-paper, brown, yellowed, or distressed base papers when they clash with the approved still or the batch palette. Start each clip from a paper field that matches the approved final still's main color direction.
4. **Make motion read as stop-motion collage.** Use clear paper-piece actions: appear piece by piece, slide or pop into position, lightly bounce, press flat, pause, then lock into the final composition. Avoid fast spinning, excessive flipping, chaotic object flight, global fades, smooth digital panning, zooming, or generic drifting.
5. **Keep segments coordinated.** Once one or two clips establish the approved batch style and SFX cadence, later clips and revisions should reference that cadence and tone so the whole sequence feels coherent.
6. **Default to polished 16:9 layered scenes.** Unless the user explicitly requests another platform frame, plan and generate in 16:9 landscape. Use the wider canvas to build attractive foreground / midground / background separation, richer environment props, clear subject hierarchy, and cinematic lateral composition without turning the frame into clutter.
7. **Emphasize paper-collage craft.** Every still and clip should make the paper-collage method visible: separable paper groups, halftone photographic cut-outs, colored cardstock accents, tactile shadows, torn edges, layered seams, and stop-motion actions such as pop-in, slide-in, light bounce, press-flat, pause, and lock.
## STEP 1: Parse the Input
For each line, concept, story, or topic, extract:
- Core meaning: what the viewer should understand
- Emotion: calm, urgent, ironic, surprising, absurd, clarifying, reflective, mysterious, or playful
- Action verb: open, connect, leak, archive, compress, split, illuminate, bind, assemble, reveal, fall, chase, transform, collide
- Visual metaphor: a concrete image that expresses the idea without writing the script on screen
- Key objects: three to six large readable paper groups
- Audio implication: whether the beat benefits from paper slide, pop, press, tap, rustle, or snap SFX
For a story topic, split the story into three to six concise beats unless the user specifies a count. Similar beats may share the same design language, but each should have a distinct metaphor, object set, color field, and SFX rhythm.
## STEP 2: Gate 1 — Production Plan Document Approval
Before generating any media, create a concise production plan document and stop for user approval. Do not generate stills or videos before the user confirms this document.
The production plan must include:
### Brief
- Topic or source line
- Intended audience / use context when known
- Aspect ratio and duration assumptions
- Tone and pacing
- Visual style summary
- Media approach, stated as: default collage SFX are kept or generated; BGM, voiceover/口播, and subtitles are not added unless the user explicitly requests them
- Optional add-on question when useful: ask whether the user wants voiceover narration/口播, BGM, or subtitles, but keep them optional
### Visual Metaphors
For each planned segment:
- Core meaning
- Emotion
- One-sentence visual proposition
- Three to six key objects
- Suggested background color and accent colors
- Expected assembly order
- Expected collage SFX moments, such as slide, pop, press, tap, rustle, or snap
### Script / Visual Beat Track
If the user explicitly requests narration, write a concise voiceover script. If the user has not explicitly requested narration, do **not** write a voiceover script; instead write a silent visual beat track explaining what the audience understands from the sequence.
If the user only needs B-roll for an existing line, preserve the original line as context instead of inventing narration.
### Storyboard
For each segment, include:
- Segment title
- Final-frame description
- Motion idea
- Approximate duration
- Collage SFX idea
- Notes on style continuity and color harmony
After presenting the production plan document, wait for the user to approve, reject, or revise. If the user approves only some numbered items, move only those items forward and revise the rest.
## STEP 3: Build Still-Frame Specifications
After the production plan is approved, write a compact visual specification for each approved segment. The specification should be self-contained and suitable for image generation.
Include:
- Script meaning or story beat
- Visual metaphor
- Aspect ratio
- Background color field
- Accent colors
- Key object groups and their roles
- Composition, foreground / midground / background depth, and negative space
- Final frame relationship
- Style continuity notes
- Avoid list for this specific still
Use this style signature:
```text
flat bold color field, black-and-white halftone photographic cut-outs, selective colored cardstock accents, warm cream keylines, soft paper shadows, fine uncoated-paper grain, premium editorial paper collage, clean refined hand-torn paper edges, subtle fibrous edges, layered paper seams
```
Color guidance:
- Burnt orange or red: labor, time pressure, urgency
- Mustard yellow: tools, warning, accumulated errors
- Ink green: cognition, reset, judgment, surreal calm
- Deep purple: memory, structure, mystery, dream logic
- Teal: collaboration, execution, system flow
- Rose red: absurdity, ceremony, theatrical tension
Do not make cobalt blue the automatic default. Keep the batch visually unified through paper texture, halftone treatment, keylines, shadows, framing, and a controlled palette. Vary the color field only when the story beat benefits from contrast.
## STEP 4: Gate 2 — Generate and Approve Still Frames
Generate one final still frame per approved segment. The still must look like the completed last frame of the future animation.
Still-frame requirements:
- Use the user-specified aspect ratio; otherwise default to 16:9 landscape
- Flat bold paper background with subtle fiber texture
- Three to six large separable object groups, unless the approved storyboard requires slightly more
- Clear central subject, strong visual hierarchy, generous clean negative space, and visible foreground / midground / background layering
- Black-and-white halftone photographic cut-outs as the main structure, combined with rich but readable scene props that support the story beat
- Selective colored cardstock accents only where they clarify hierarchy
- Warm cream keylines, soft physical paper shadows, refined torn edges, and subtle layered paper seams
- No readable text, fake letters, numerals, subtitles, UI, watermark, or logos unless explicitly requested
Show the generated still frames to the user and stop for approval before video generation. If a still looks too busy, has mismatched color, too many paper layers, too-flat digital edges, or too much aged/brown paper texture, revise the still before video generation.
## STEP 5: Plan the Stop-Motion Assembly
For each approved still frame, prepare a motion plan that treats the approved still as the completed final frame.
Default motion order:
1. Start from a clean paper color field matching the approved still's main background, not a mismatched kraft-paper or brown base.
2. Base structure appears as paper pieces that slide, pop, or press into place.
3. Character, card, object, or main metaphor element enters.
4. Secondary objects assemble one by one with small bounce, press-flat, and pause beats.
5. Final relationship locks into the approved composition.
6. Hold the completed still-like frame briefly at the end.
Motion prompt guidance:
```text
Paper-collage stop-motion assembly. Start on a clean paper field matching the approved still's background color and tone. Assemble the scene piece by piece with tactile stop-motion timing: foreground, midground, and background paper groups appear in a readable order; each paper object appears, slides or pops in, lightly bounces, presses flat, pauses, and locks into place. Preserve the halftone dots, cream keylines, subtle fibrous torn edges, layered paper seams, soft shadows, paper grain, color field, and approved aspect ratio. End by holding the completed composition matching the approved still frame. Sound design: tactile collage SFX only, synchronized to paper-piece motion: soft paper slide, pop-in, press-flat tap, light rustle, and tiny paper snap where appropriate. Default: do not add BGM, voiceover, dialogue, or subtitles unless the user explicitly requested them. No mismatched kraft-paper opening, no brown/yellowed base unless approved, no fast spinning, no chaotic object flight, no smooth digital layer movement, no scene cuts, no camera move, no zoom, no morphing, no new objects, no text, no logos, no watermark, no UI.
```
## STEP 6: Generate Collage-SFX B-roll Clips
After Gate 2 approval, generate one video clip per approved still frame.
Generation requirements:
- Use the approved still as the visual reference / final-frame anchor whenever the selected Hub video model supports it
- Use the same ratio as the still frame
- Target about four seconds unless the user requested a different length
- Generate or preserve tactile collage SFX by default when supported: paper slide, pop, press-flat tap, light rustle, tiny snap
- Default: do not generate or add BGM unless the user explicitly requested music
- Default: do not generate or add voiceover, spoken narration, or presenter audio unless the user explicitly requested narration/口播
- If a generated clip contains useful collage SFX, preserve it during post-production and final assembly
- If a generated clip contains unwanted speech, music, fake UI sound, or noisy ambience, remove or regenerate the audio according to the user goal
- Keep the shot visually stable; do not add camera movement unless the user asks
- Preserve the collage material language, final composition, and established batch style
- When revising later clips, use the best approved previous clips as style/cadence references if helpful
Use `MiniMax-H3` as the default video generation model for this Skill. Do not ask the user to choose a model unless `MiniMax-H3` fails, is unavailable, or cannot satisfy a hard capability requirement. If the user explicitly specified another model, keep that model as primary and adapt parameters only when needed.
## STEP 7: Quality Review
Review stills and clips against these standards:
- The metaphor or story beat is understandable without writing it on screen
- The still has a strong clean color field, attractive 16:9 composition by default, clear object hierarchy, and visible foreground / midground / background depth
- The number of object groups stays readable, not a screen full of fragments
- The clip begins from a color field matching the approved still or batch palette
- The assembly is visible piece by piece, not a global fade-in
- Motion reads as stop-motion collage: separate paper groups pop in or slide in, lightly bounce, press flat, pause, and lock
- Defaults: tactile collage SFX are desirable; BGM, voiceover/口播, and subtitles are absent unless explicitly requested
- Paper texture is controlled: not too flat, not over-aged, not brown/kraft unless approved
- Color tone is consistent across the batch
- The final frame remains close to the approved still
- There are no unwanted readable letters, fake UI, subtitles, watermark, or logos
Minor drift in tiny details is acceptable if the metaphor remains clear and the final composition still resembles the approved still. Major drift, added text, fake UI, mismatched opening color, unwanted kraft-paper base, over-wrinkled/dirty paper texture, lost stop-motion assembly, unwanted BGM, or unwanted voiceover should be regenerated or post-processed depending on where the issue appears.
## STEP 8: Optional Assembly, Music, Narration, and Delivery
If the user asks for a multi-clip story video, assemble the approved clips in the confirmed storyboard order and preserve each clip's original collage SFX by default.
Add BGM, narration/口播, or subtitles only when explicitly requested. When the user requests BGM or narration after clips already exist, mix it with the collage SFX instead of replacing SFX, unless the user asks to mute the original SFX.
For each completed item, deliver:
- Final B-roll video path or final assembled video path
- Approved still frame path when relevant
- A one-sentence explanation of how the script or visual beat became a visual metaphor
- Media note: collage SFX by default, and list any explicitly requested BGM, narration/口播, or subtitle additions
- Any QA caveat if a clip passed with minor acceptable drift
For batch work, group the outputs by item number and keep the user's original line or story beat attached to its result.
## Failure Handling
- If the production plan defaults to silent video, correct it to collage SFX only unless the user explicitly asked for silence.
- If the production plan includes BGM, voiceover/口播, or subtitles without explicit user request, remove them and present BGM, voiceover/口播, and subtitles as optional choices.
- If the production plan feels too literal, return to Gate 1 and make the visual proposition more object-driven.
- If the still contains fake text or UI, regenerate the still before video generation; do not try to fix it only in the video prompt.
- If the still or clip is too visually busy, reduce paper layers, simplify the object count, and restore negative space.
- If the video lacks assembly motion, reduce the number of objects and make the step-by-step arrival order more explicit.
- If motion feels like smooth digital layers, rewrite the prompt around stop-motion beats: appear → light bounce → press flat → pause → lock.
- If the paper looks too flat, add subtle torn edges, fiber texture, layered seams, and soft shadows.
- If the paper looks over-aged or dirty, remove kraft-paper, brown/yellowed base, heavy wrinkles, stains, and distressed treatment.
- If the opening color clashes with the approved still, start from the same main background color as the still.
- If the final frame drifts too far from the still, strengthen the instruction that the approved still is the completed final composition.
- If generated audio lacks collage SFX but the visual is otherwise strong, ask the user whether to keep it, regenerate with stronger SFX direction, or add/post-sync SFX.
- If generated audio contains unwanted BGM or voiceover, remove or regenerate it before final delivery.
- If precise layer control is required, recommend a layered animation or editing workflow instead of continuing this Skill.
@@ -0,0 +1,19 @@
display-name-zh: 纸拼贴讲解动画生成器
version: 0.3.9
tag-en: "Education"
tag-cn: "教育"
complete-tags-en:
- "Education / Full Production"
- "Animation / Creative Generation"
- "Creative & Experimental / Creative Generation"
complete-tags-cn:
- "教育 / 成片制作"
- "动画 / 创作生成"
- "创意实验 / 创作生成"
summary-en: "Turn narration or abstract ideas into tactile paper-collage explainer animations."
summary-cn: "基于文案或抽象概念,完成隐喻设计、拼贴分镜和动画生成,适用于讲解类内容。"
desc-en: "For creators, educators, and social-video editors who need a tactile visual language for narration, knowledge points, opinions, or abstract topics. Users provide source copy, story beats, or a core concept, with optional aspect ratio, duration, palette, and audio needs. The Skill extracts meaning, proposes visual metaphors, prepares a production plan and storyboard, generates approval-ready halftone collage stills, then creates stop-motion clips with paper movement and tactile sound effects. It delivers cohesive paper-collage animation assets or an assembled explainer video, and is not intended for realistic presenter ads, precisely editable layer files, complex typography systems, or prompt-only handoff."
desc-cn: "面向内容创作者、教育讲解者和社交视频编辑,用触感纸拼贴语言表现口播句、知识点、观点或抽象主题。用户提供短文案、故事节点或核心概念,也可补充画幅、时长、色调和音频需求。Skill 会提炼含义与视觉隐喻,制定制作方案和分镜,生成可确认的半调纸拼贴静帧,再制作带纸片滑入、弹入、轻敲、压平和摩擦声的停格动画片段。最终交付统一的纸拼贴动画素材或完整讲解视频,不适合真人口播广告、精确可编辑图层、复杂文字排版或仅输出提示词。"
author-en: "MiniMax Hub"
author-cn: "MiniMax Hub"
source: official-featured
@@ -0,0 +1,470 @@
---
name: papercraft-stop-motion-explainer
description: 面向希望用手工纸艺视觉讲解科学、教育或泛知识内容的创作者。用户需提供主题、核心知识点或原始材料,并可指定受众、时长、画幅和交付类型。Skill 会先提炼学习目标与视觉隐喻,提出创意方向,设计纸偶角色、分层纸雕布景和道具资产,制作预览概念、图像与视频提示词,规划分镜、运镜、转场和声音,并通过阶段确认与审核清单控制质量。最终输出可直接进入制作的纸艺定格科普视频创作包,也可按需交付单图、系列图、短视频提示词或分镜。适用于纸雕、剪纸、立体书和微缩定格讲解,不适用于普通二维卡通、线稿涂鸦、真人写实视频或无纸艺质感的标准科普。
---
# 纸艺定格科普视频生成器
将科学、教育或知识主题整理成完整的纸艺定格动画科普创作包。产出应包含画风规则、角色与场景设计、资产规划、提示词、分镜结构、动效语言、声音方向、反向提示词和审核标准。
适用于用户想要触感强、手作感强的科普视觉:分层纸片、卡纸剪裁、微缩纸雕布景、纸偶定格动作、立体书式展开、纸质道具和真实层间投影。默认假设用户想得到一个完整的科普视频创作包。提示词、分镜、资产规划和动效说明都是完整流程中的可审核生产资产,不是孤立交付物,除非用户明确只要某一项。
当本 Skill 从方案/提示词推进到实际视频生成时,默认使用 MiniMax-H3 作为视频生成模型。除非用户明确选择其他可用模型,或 MiniMax-H3 当前不可用,否则都按 MiniMax-H3 执行。
## 画布文档交付规则
当输出完整视频创作包、制作方案、分镜方案或任何较长的多章节生产文档时,必须把完整方案写入画布文本节点。对话内只给用户简短说明、文档文件名和下一步建议。除非用户明确要求在对话里查看全文,不要把长制作方案、大型分镜表、完整提示词库或完整检查清单直接粘贴到聊天区。
聊天区用于简洁指引和决策确认;画布文档作为详细制作方案的主交付物。用户要求修改方案时,优先更新已有画布文档,不要重复创建新文档,除非确实需要生成新版本。
## 画面层次与纸艺定格优先级
每一次视觉方案、预览图提示词、视频提示词和审核都必须同时强调画面层次与明确的纸艺定格动画质感。
画面层次必须强调:
- 场景要有清晰的前景、中景、背景和远景。
- 使用纸片叶子、云纹、门框、帘幕、道具或剪影作为前景遮挡,制造空间深度。
- 核心知识对象或主要动作放在易读的中景层。
- 背景不能只是静态平面,要加入背景视差、小型环境动效和层间分离。
- 需要时使用圆形、斜向、隧道式、立体书式或剖面式构图帮助观众读出层次。
纸艺定格质感必须强调:
- 所有可见物体都应像真实纸材制作:多层卡纸、剪裁边缘、纸纤维、折痕、拼贴缝、标签片、铆钉、关节、拉片、滑轨、转盘和可见厚度。
- 动作要像逐帧手工拨动:小幅分段移动、短暂停顿、轻微回弹、铰链动作、滑动纸机关、翻页、拉片揭示和纸片落定。
- 背景元素也应尽量是纸艺机关:纸云沿滑轨移动、纸月亮或圆盘沿轨道运动、多层布景视差、纸叶落下、纸灯笼轻晃、纸门打开。
- 避免丝滑 CG 动作、塑料 3D 表面、平面矢量背景、过于静止的背景,以及破坏微缩定格感的大幅角色运动。
## 轻量请求旁路规则
如果用户明确只需要某一个生产资产,不要强制执行完整 18 步创作包流程。直接进入对应步骤,同时保留纸艺定格动画的核心画风规则。
使用这些快捷路径:
- 只要单图提示词:执行 STEP 1、STEP 2、STEP 9、STEP 17。
- 只要系列图提示词:执行 STEP 1、STEP 2、STEP 10、STEP 17。
- 只要 5 秒图生视频提示词:执行 STEP 1、STEP 2、STEP 11、STEP 17。
- 只要分镜:执行 STEP 1、STEP 2、STEP 12、STEP 13、STEP 14、STEP 15、STEP 16。
- 只要创意方向:执行 STEP 1、STEP 2、STEP 3。
只有当用户要求完整视频创作包、完整制作方案,或没有明确缩小交付范围时,才使用完整阶段确认流程。
## 交互规则:每个阶段之间用选择卡片确认
每完成一个关键阶段,都先停下,用选择卡片询问用户是否进入下一步。所有轮询、多选、确认、继续/修改/停止等用户决策点都必须用卡片形式,不要用纯文本让用户回复数字。卡片必须给出具体下一步选项,而不是开放式追问。完成创意方向阶段后,下一张确认卡必须先让用户选择目标视频时长,再继续角色、场景、预览图、提示词或分镜等细化工作。
常用选项:
1. 按当前方向继续下一阶段。
2. 先修改当前阶段再继续。
3. 切换到另一个创意方向。
4. 跳到某个具体生产资产:视觉预览图、单图提示词、系列图提示词、5 秒图生视频提示词或分镜。
5. 到此停止,导出当前创作包。
卡片要简短;如果当前结果足够稳,推荐选项放在第一项。这个阶段确认机制用于保护完整视频流程,让用户能在每个资产进入下一步前先调整。
## STEP 1:理解输入需求
分析用户的主题和制作目标。保留用户原始领域词,不要把科学主题改写成另一个主题。
输出简洁的需求理解:
- 科普主题或知识点
- 目标观众:儿童、泛知识用户、课堂学生、社媒用户、品牌教育或专业观众
- 目标时长:15 秒、30 秒、60 秒或未指定
- 平台或画幅需求
- 交付类型:单图、系列图、5 秒图生视频提示词、分镜或完整制作包
- 核心学习目标:观众最后应记住的一句话
- 视觉隐喻:这个概念如何变成纸质物体、模型、纸偶、层级、路径或机关
如果信息缺失,直接选择合理默认:泛知识观众、30 秒短视频、视频默认横屏 16:9,除非用户明确指定其他比例;一个纸艺讲解员加一个核心纸艺模型。
## STEP 2:总结画风基因
在正式设计前,先说明必须保持一致的画风规则。
包含这些特征:
- 手工纸艺定格动画感
- 微缩纸雕场景或剧场箱布景
- 多层卡纸剪裁,能看到纸片厚度
- 哑光、可触摸的纸张纹理、纤维、折痕、缝隙、撕边或剪裁边缘
- 纸偶角色由独立纸片部件叠压组装
- 层与层之间有真实物理投影
- 明确的前景、中景、背景、远景多平面结构
- 微距摄影感、轻微景深和 2.5D 视差
- 标签、箭头、卡片、图表等教育信息也用纸片制作
- 清晰的前景 / 中景 / 背景 / 远景层次,有前景遮挡和可读的视差关系
- 背景层也要有纸艺机关动效,而不是静态绘制背景
- 定格动画动作语言:分段移动、短暂停顿、轻微回弹、铰链关节、拉片、滑轨、转盘和纸片落定
说明理由:纸艺媒介能把抽象科学变成可触摸的模型,也天然适合做分层讲解。场景必须像真实搭建并逐帧拍摄出来,而不是只有纸质纹理的普通插画。
## STEP 3:提出 3-5 个创意方向
生成 3 到 5 个不同创意。每个方向都应从不同视觉隐喻或叙事结构解释同一主题。
每个方向包含:
- 创意标题
- 核心想法
- 讲解角度
- 纸艺视觉隐喻
- 最适合的时长
- 最适合的观众或平台
- 标志性画面瞬间
- 风险或限制
展示创意方向后,必须让用户同时选择创意方向和目标视频时长。时长选项保持简洁:15 秒快版、30 秒标准版、60 秒完整版或自定义时长。用户选定时长前,不要继续进入详细资产设计。
推荐方向类型:
1. 立体书旅程:每翻开一页揭示概念的一层。
2. 纸艺科学家实验室:纸偶主持人在实验台上演示机制。
3. 分层剖面模型:核心对象像剖面图一样打开。
4. 微缩自然剧场:生态、地理、天文主题在纸雕景观中展开。
5. 纸片机关板:齿轮、箭头、滑轨、标签和移动纸片解释因果关系。
## STEP 4:设计纸艺角色
只有当角色能帮助讲解时才设计角色。角色可以是主持人、助手、动物向导、拟人化分子、免疫细胞、行星、机器零件或自然力量。
每个角色输出:
- 名称和讲解职责
- 造型语言和比例
- 纸艺结构:剪纸五官、分层头发或身体、卡纸四肢、关节节点、标签片、铆钉或折叠结构
- 表情系统:简单纸片眼睛、眉毛、嘴型、可替换表情卡
- 动作方式:定格式小幅移动、铰链手臂、指向、弹跳、滑动或翻卡表情
- 系列图一致性所需的可复用身份特征
角色必须足够简洁,保证短视频里一眼能读懂。
## STEP 5:设计纸片场景
把每个场景当成一个真实纸艺舞台,而不是平面背景。
每个场景输出:
- 场景名称
- 科普功能
- 前景元素
- 中景主体和动作
- 背景环境
- 远景底板或天空
- 可移动纸艺机关
- 教育信息承载物:标签、箭头、图表、测量牌、说明卡或标题卡
尽量使用至少四个空间层。通过间距和投影让核心模型从背景中分离出来。
## STEP 6:规划分层纸雕布景
输出多平面布景表。
必备列:
- 层级编号
- 平面作用
- 纸艺元素
- 材质和边缘处理
- 动效或视差行为
- 投影关系
- 知识表达功能
规则:
- 大多数镜头使用 4-7 层。
- 阅读量大的标签放在稳定层。
- 前景元素可以局部遮挡画面,增强微缩摄影感。
- 核心概念放在注意力最强的中景。
- 用分层解释层级、顺序、结构、因果或尺度。
## STEP 7:规划道具资产库
按功能输出资产列表。
资产分组:
- 主持人和角色资产
- 核心科学对象资产
- 讲解道具:箭头、标签、卡片、图表、放大镜、尺子、仪表
- 场景道具:实验台、山丘、书、架子、云、星星、树、水、管道、齿轮
- 动效道具:纸条、滑轨、旋转圆盘、立体弹出片、抽拉层、纸屑
- 声音提示道具:翻页、剪纸、纸张摩擦、卡扣、弹出、胶带撕拉
每个资产定义:
- 名称
- 用途
- 纸张材质
- 所在层级
- 静态或可动
- 出现时间
- 一致性注意事项
## STEP 8:规划或生成 1-3 张视觉预览图
在撰写最终提示词或分镜之前,先基于已确认的创意方向、角色设计、场景设计和分层布景,规划 1 到 3 张视觉预览图。如果当前环境支持图像生成且用户需要实际预览图,则生成预览图;否则输出预览图说明和提示词。预览图用于确认风格和概念,不是最终生产帧。
预览图比例默认跟随项目视频比例:视频默认横屏 16:9,除非用户明确指定其他比例。如果用户要求竖屏、方图或其他格式,预览图和后续视频规划都保持该比例一致。
每张预览图需说明:
- 预览图编号和目的
- 它对应哪个已确认的创意方向、角色或场景
- 视觉重点:整体世界观、角色造型、机关清晰度、文化氛围或结尾记忆点
- 用户进入下一步前需要判断什么
推荐预览组合:
1. 整体世界预览:展示完整纸艺布景、主角色、核心对象和整体氛围。
2. 机制预览:展示科学模型、纸箭头、标签和可动讲解道具。
3. 情绪或结尾预览:展示文化连接、总结卡片或最终记忆画面。
展示预览图后,必须用选择卡片暂停确认。选项可包括:继续撰写提示词、修改预览图 1、修改预览图 2、修改预览图 3、收敛为单一视觉方向,或切换创意方向。用户确认视觉方向前,不要撰写最终单图提示词、系列图提示词或 5 秒图生视频提示词。
## STEP 9:撰写单图提示词
为概念图或主视觉写一条提示词。
提示词必须包含:
- 用户主题和核心对象
- 纸艺定格动画风格
- 微缩纸雕舞台
- 多层卡纸剪裁
- 可见纸张纤维、折痕、缝隙、剪裁边缘和厚度
- 层与层之间的真实物理投影
- 微距摄影感和轻微景深
- 必要时加入纸质标签、箭头或知识卡片
不要把所有知识细节塞进一张图。一张图只传达一个主要概念。
## STEP 10:撰写系列图提示词
当用户需要多个关键帧或分镜图时,输出系列图提示词。
每张图包含:
- 图号和目的
- 相比上一张的变化
- 共享画风锚点
- 角色和场景一致性锚点
- 提示词
系列规则:
- 保持同一纸张材质语言、光源方向、空间层级和角色身份。
- 每张图只解释一个知识步骤。
- 每条提示词都保留用户主题中的领域词。
- 除非用户要求变化,否则系列图保持同一画幅比例。
## STEP 11:撰写 5 秒图生视频提示词
为单张参考图写 5 秒动效提示词,保持纸艺风格不变。
包含:
- 保留纸张纹理、剪裁边缘、分层布景和真实投影
- 缓慢推近、轻微横移或视差运动
- 纸偶小幅定格式动作
- 纸质箭头、标签、滑轨或机关轻微移动
- 前景和背景产生视差
- 避免光滑 CG 变形、融化、塑料质感或高速镜头
动作要克制,并符合纸片物体的物理逻辑。
## STEP 12:为已选时长创建分镜
使用创意方向阶段后用户已经选择的视频时长。默认不要同时提供 15/30/60 秒三套完整分镜。先给一个简短的内容概述,再只为已选时长输出简洁分镜表。
内容概述保持简洁:
- 一句话视频主线
- 目标时长和比例
- 主要视觉结构
- 知识路径:钩子 → 解释 → 例子或文化连接 → 记忆句
分镜表规则:
- 表格文字尽量短,不要在单元格里写长段落。
- 只列出该时长真正需要的镜头。
- 15 秒:通常 4 个镜头。
- 30 秒:通常 5-6 个镜头。
- 60 秒:通常 7-9 个镜头。
- 每个镜头只解释一个知识节拍。
必备列:
- 时间
- 知识节拍
- 画面动作
- 纸艺运动
- 运镜 / 转场
- 音效
表格后用选择卡片询问用户:继续剪辑节奏和运镜、修改分镜、改视频时长,或回到视觉预览。
## STEP 13:定义剪辑节奏
根据时长和理解清晰度设置节奏。
规则:
- 15 秒:4-6 个镜头,快但可读,每镜 2-4 秒。
- 30 秒:6-8 个镜头,每镜 3-5 秒。
- 60 秒:8-12 个镜头,每镜 4-7 秒。
- 标签卡出现时短暂停顿。
- 观众看懂纸片机关前不要急切。
- 用纸片动作形成节奏,不用强烈数码快切。
## STEP 14:定义运镜规则
运镜应像在拍一个微缩纸艺舞台。
推荐运镜:
- 缓慢推近,进入纸艺世界
- 横向平移,展示多层视差
- 固定中景,用于主持讲解
- 微距特写,展示纸张纹理和关键道具
- 轻微俯视,展示剖面结构
- 层间穿梭,进入对象内部结构
避免:
- 高速飞行镜头
- 除非布景专门支持,否则不要 360 度大环绕
- 数码故障运动
- 液体融化变形
- 超写实 CG 镜头行为
## STEP 15:定义转场规则
转场必须符合纸艺世界的物理逻辑。
推荐转场:
- 翻页
- 立体书展开
- 抽拉标签
- 纸质标签牌遮挡
- 云朵纸片滑过
- 圆形纸片遮罩
- 剪纸门打开
- 剖面层分开
- 纸屑飞散
- 胶带或贴纸揭示
除非用户明确需要反差效果,否则避免电子扫描线、霓虹故障、玻璃破碎、金属擦除和科幻粒子转场。
## STEP 16:定义声音设计
建立触感明确的手工声音库。
推荐音效:
- 翻纸声
- 剪纸声
- 卡纸滑动声
- 纸张摩擦声
- 小木头卡扣声
- 轻微弹出声
- 胶带撕拉声
- 纸偶关节轻响
- 纸盒打开声
- 纸屑散落声
配乐方向:
- BGM 要匹配用户主题、文化语境和情绪,不只匹配通用纸艺风。
- 遇到有明确文化属性的主题时,围绕该主题的音乐语言设计。比如中秋科普应偏清雅国风:古筝、琵琶、竹笛或箫、轻打击、柔和弦乐,并给旁白留白。
- 没有强文化属性的科学或课堂主题,可使用木琴、马林巴、拨弦、轻打击、温暖教育感。
- 旁白下方保持轻柔,最终混音时要做压低或 ducking。
- 除非用户要求,否则避免厚重电影低频、强烈电音、未来电子质感或过度戏剧化配乐。
纸艺动效音效方向:
- 音效要来自视频里的真实动作节拍:翻页、纸门打开、纸轨滑动、卡片翻动、纸抽屉拉出、纸盒打开、纸层叠合、轻柔 pop、纸偶关节轻响、纸张摩擦。
- 音效只做触感点缀,让纸片运动更有物理感,不要变成夸张卡通音效。
- 最终混音前要先把音效映射到分镜时间轴,让动作、旁白和声音互相强化。
## STEP 17:生成并清理旁白音频
当流程包含旁白时,先在旁白脚本确认后生成配音,再检查音频质量,确认无问题后再进入最终合成。
旁白规则:
- 声音气质要匹配主题和观众:亲子或教育主题使用温暖、清晰、轻柔的声音;课堂或技术讲解可更中性。
- 对比旁白时长和已选视频时长。若偏差超过 20%,必须用卡片询问用户:压短文案、延展视频,还是接受不匹配。
- 检查尾部杂音、点击声、气口残留、突兀截断或模型尾音残留。
- 如果尾部有杂音,优先修复已有音频:裁掉杂音尾部并加短淡出。若声音表现本身合适,不要优先重录。
- 如果需要重录,提示词里明确要求结尾干净、无尾部杂音;生成后仍要再次检查。
## STEP 18:提供反向提示词
始终输出简洁的反向提示词块,以保护画风。
推荐反向提示词:
- 光滑塑料 3D
- 高亮 CG 渲染
- 真人实拍
- 普通扁平矢量插画
- 没有纸张纹理的普通卡通
- 金属科幻表面
- 玻璃材质
- 赛博朋克霓虹
- 数码故障特效
- 油画笔触
- 真实毛发或皮肤毛孔
- 没有纸张纤维
- 没有剪裁边缘
- 没有层间投影
- 边缘过度平滑
- 高速镜头环绕
- 融化或液体变形
## STEP 19:执行审核清单
结尾输出审核清单。只标出对用户目标真正重要的问题。
### 画风审核
- 是否能看到纸张纤维、折痕、缝隙和剪裁边缘?
- 是否能看到卡纸厚度?
- 层与层之间是否有真实投影?
- 画面是否像微缩摄影,而不是平面插画?
- 是否避免了塑料 CG 和真人写实?
### 角色审核
- 角色是否像纸偶?
- 五官和表情是否足够剪纸化?
- 关节和四肢是否像分件组装?
- 动作是否保持定格动画的物理感?
### 场景审核
- 是否有清楚的前景、中景、背景和远景?
- 道具是否都是纸艺材质,并服务讲解?
- 标签、箭头和卡片是否没有挡住核心模型?
- 每个镜头是否只解释一个知识节拍?
### 视频审核
- 运镜是否符合微缩摄影逻辑?
- 转场是否符合纸片物理逻辑?
- 音效是否匹配纸张、卡纸和纸偶动作?
- 分镜时长是否匹配目标时长?
- 最后的记忆句是否清楚?
### 音频审核
- 旁白时长是否匹配已选视频时长,或已通过卡片确认解决偏差?
- 旁白结尾是否干净,没有尾部杂音、点击声、气口残留或突兀截断?
- BGM 是否匹配主题文化和情绪,而不只是泛纸艺风?
- BGM 是否在旁白下方,不抢讲解信息?
- 纸艺动效音效是否对齐真实画面动作,并且足够克制?
@@ -0,0 +1,471 @@
---
name: papercraft-stop-motion-explainer
description: For creators explaining science, education, or general knowledge through tactile handmade papercraft visuals. Users provide a topic, core knowledge points, or source material and may specify audience, duration, aspect ratio, and deliverable type. The Skill extracts the learning goal and visual metaphor, proposes creative directions, designs paper characters, layered diorama sets, and props, creates preview concepts plus image and video prompts, and plans storyboards, camera movement, transitions, and sound with staged approvals and review checklists. It outputs a production-ready papercraft stop-motion explainer package, or selected assets such as still prompts, image-series prompts, short-video prompts, or storyboards. Best for cut-paper, pop-up-book, layered diorama, and miniature stop-motion explainers; not for standard 2D cartoons, line doodles, live action, or explainers without a paper-art look.
compatibility: Requires the MiniMax Hub agent (canvas workspace and MiniMax H3 generation); not portable to generic agent harnesses.
---
# Papercraft Stop-Motion Explainer
Create a complete papercraft stop-motion explainer package from a science, education, or knowledge topic. The output is a production-ready creative plan: style rules, character and set design, asset planning, prompts, storyboard structures, motion language, sound direction, negative prompts, and review criteria.
Use this Skill when the user wants a tactile handmade explainer style: layered paper, cardboard cutouts, miniature diorama, stop-motion puppet movement, pop-up-book staging, paper props, and physical shadows. The default assumption is that the user wants a complete explainer video package. Prompts, storyboards, asset plans, and motion notes are reviewable production assets inside that complete process, not isolated deliverables unless the user says so.
When this Skill proceeds from planning/prompting into actual video generation, use MiniMax-H3 as the default video generation model. Treat MiniMax-H3 as the default unless the user explicitly selects another available model or MiniMax-H3 is unavailable.
## Canvas Document Delivery Rule
For complete video packages, production plans, storyboards, or any long multi-section planning output, write the full production document to the canvas as a text document. In the chat reply, give only a brief summary, the document filename, and the next recommended action. Do not paste long production plans, large storyboard tables, full prompt libraries, or full checklists into the conversation unless the user explicitly asks to see them in chat.
Use chat for concise guidance and decision points; use the canvas document as the source of the detailed plan. When the user asks to revise the plan, update the existing canvas document rather than creating a duplicate unless a new version is intentionally needed.
## Visual Depth and Papercraft Motion Priority
Every visual plan, preview prompt, video prompt, and review pass must prioritize both layered image depth and unmistakable papercraft stop-motion qualities.
Required visual-depth emphasis:
- Build scenes with clear foreground, midground, background, and far-background planes.
- Use foreground occluders such as paper leaves, clouds, frames, curtains, props, or cutout silhouettes to create depth.
- Keep the main knowledge object or focal action in a readable middle plane.
- Add background parallax, small environmental motions, and layer separation instead of flat static backdrops.
- Use circular, diagonal, tunnel, pop-up, or cross-section compositions when they help the viewer read depth.
Required papercraft stop-motion emphasis:
- Make all visible objects feel physically made from paper: layered cardboard, cut edges, paper fibers, folds, seams, tabs, brads, joints, pull-tabs, sliders, rotating discs, and visible thickness.
- Motion should feel frame-by-frame and hand-manipulated: small stepped movements, tiny pauses, slight rebounds, hinged gestures, sliding paper mechanisms, page flips, pull-tab reveals, and paper pieces settling.
- Background elements should also be paper mechanisms when possible: paper clouds sliding on rails, paper moons/discs moving on tracks, layered scenery shifting in parallax, paper leaves falling, paper lamps swaying, and cardboard doors opening.
- Avoid smooth CG motion, plastic 3D surfaces, flat vector backdrops, overly static backgrounds, and large character motion that breaks the miniature stop-motion feel.
## Lightweight Request Bypass
If the user explicitly asks for only one production asset, do not force the full 18-step package workflow. Route directly to the relevant step while preserving the papercraft style rules.
Use these shortcuts:
- Single-image prompt only: run STEP 1, STEP 2, STEP 9, and STEP 17.
- Image-series prompts only: run STEP 1, STEP 2, STEP 10, and STEP 17.
- 5-second image-to-video prompt only: run STEP 1, STEP 2, STEP 11, and STEP 17.
- Storyboard only: run STEP 1, STEP 2, STEP 12, STEP 13, STEP 14, STEP 15, and STEP 16.
- Creative directions only: run STEP 1, STEP 2, and STEP 3.
Use the full phased confirmation workflow only when the user asks for a complete video package, a full production plan, or does not specify a narrower deliverable.
## Interaction Rule: Confirmation Cards Between Phases
After each major phase, pause and ask the user with a selection card before moving forward. All polling, multi-choice decisions, confirmations, continue/revise/stop gates, and user decision points must use cards rather than plain text asking the user to reply with a number. The card must give practical next-step choices, not an open-ended question. After the creative directions phase, the next confirmation card must ask the user to choose the target video duration before detailed character, scene, preview, prompt, or storyboard work continues.
Use choices like:
1. Continue to the next phase with the current direction.
2. Revise the current phase before continuing.
3. Switch to one of the other creative directions.
4. Jump to a specific production asset: visual preview images, single-image prompt, image-series prompts, 5-second image-to-video prompt, or storyboard.
5. Stop here and export the current package.
Keep each card short. Include the recommended option first when the current result is strong. This staged confirmation protects the full-video workflow: the user can adjust each asset before it becomes the basis for the next step.
## STEP 1: Understand the Input
Analyze the user's topic and production goal. Preserve the user's domain words and do not simplify the science into a different topic.
Output a concise understanding block:
- Topic or knowledge point
- Target audience: children, general audience, classroom, social media, brand education, or specialist audience
- Intended duration: 15s, 30s, 60s, or unspecified
- Platform or format when stated
- Required output type: single image, image series, 5-second image-to-video prompt, storyboard, or full package
- Core learning outcome: the one sentence the viewer should remember
- Visual metaphor: how the concept becomes a paper object, model, puppet, layer, path, or mechanism
If crucial information is missing, choose reasonable defaults instead of stopping: general audience, 30-second short, horizontal 16:9 video by default unless the user specifies another ratio, and one paper narrator plus one core paper model.
## STEP 2: Summarize the Style DNA
Before designing the video, state the style rules that must stay consistent.
Include these traits:
- Handmade papercraft stop-motion look
- Miniature paper diorama or shadow-box stage
- Layered cardboard cutouts with visible thickness
- Matte tactile paper textures, fibers, folds, seams, torn or cut edges
- Paper puppet characters built from separate overlapping parts
- Real physical drop shadows between layers
- Multi-plane foreground, midground, background, and far background
- Macro miniature photography, slight depth of field, and 2.5D parallax
- Educational labels, arrows, cards, and simple symbols made from paper
- Clear foreground / midground / background / far-background depth, with foreground occlusion and readable parallax
- Background layers with paper-mechanism motion, not static painted scenery
- Stop-motion motion language: stepped movement, tiny pauses, slight rebounds, hinged joints, pull-tabs, sliders, rotating discs, paper pieces settling
Explain the reason: the paper medium turns abstract science into touchable objects and makes layered explanation easy to understand. The scene must feel physically built and animated frame by frame, not merely illustrated in a paper-like texture.
## STEP 3: Propose 3-5 Creative Directions
Generate 3 to 5 distinct concepts. Each direction must explain the same topic through a different visual metaphor or narrative structure.
For each direction, provide:
- Title
- Core idea
- Explainer angle
- Paper visual metaphor
- Best duration
- Best audience or platform
- Signature visual moment
- Risk or limitation
After presenting directions, ask the user to choose both the creative direction and the target duration. Duration options should be concise: 15s quick version, 30s standard version, 60s full version, or custom duration. Do not proceed into detailed assets until a duration is chosen.
Recommended direction archetypes:
1. Pop-up-book journey: each page reveals one layer of the concept.
2. Paper scientist laboratory: a paper host demonstrates the mechanism on a table.
3. Layered cross-section model: a paper object opens like a sectional diagram.
4. Miniature nature theater: ecological, geographic, or astronomical topics unfold in a paper landscape.
5. Paper mechanism board: gears, arrows, sliders, labels, and moving parts explain cause and effect.
## STEP 4: Design Paper Characters
Design characters only when they help communication. A character may be a host, assistant, animal guide, personified molecule, immune cell, planet, machine part, or natural force.
For each character, output:
- Name and role in the explanation
- Shape language and proportions
- Paper construction: cutout face, layered hair or body, cardboard limbs, joint nodes, tabs, brads, or folded parts
- Expression system: simple paper eyes, eyebrows, mouth shapes, replaceable emotion cards
- Movement style: stop-motion nudges, hinged arm gestures, pointing, bouncing, sliding, or flip-card expressions
- Reusable identity details for image series consistency
Keep characters simple enough to remain readable in short videos.
## STEP 5: Design Paper Scenes
Treat every scene as a physical paper stage, not a flat background.
For each scene, output:
- Scene name
- Scientific purpose
- Foreground elements
- Midground subject and action
- Background environment
- Far-background board or sky
- Movable paper mechanisms
- Educational information carriers: labels, arrows, charts, measurement tags, captions, or cards
Use at least four depth planes whenever possible. Keep the main knowledge model separated from the background through spacing and shadow.
## STEP 6: Plan Layered Diorama Staging
Create a multi-plane staging table.
Required columns:
- Layer number
- Plane role
- Paper elements
- Material and edge treatment
- Motion or parallax behavior
- Shadow interaction
- Knowledge function
Rules:
- Use 4-7 layers for most shots.
- Place reading-heavy labels on stable layers.
- Let foreground elements partially occlude the stage for miniature realism.
- Keep the core concept in the midground, where attention is strongest.
- Use layer separation to explain hierarchy, sequence, anatomy, causality, or scale.
## STEP 7: Plan Prop and Asset Library
Create an asset list grouped by function.
Asset groups:
- Host and character assets
- Core scientific object assets
- Explanation props: arrows, labels, cards, charts, magnifiers, rulers, meters
- Scene props: lab bench, hills, books, shelves, clouds, stars, trees, water, tubes, gears
- Motion props: paper strips, sliders, rotating discs, pop-up tabs, pull-out layers, paper confetti
- Sound cue props when useful: page flip, scissors, paper rustle, click, pop, tape peel
For each asset, define:
- Name
- Purpose
- Paper material
- Layer placement
- Static or movable
- Appearance timing
- Consistency notes
## STEP 8: Plan or Generate 1-3 Visual Preview Images
Before writing final prompts or storyboards, plan 1 to 3 visual preview images based on the approved creative direction, character design, scene design, and layered staging. If image generation is available and the user wants actual previews, generate them; otherwise provide preview briefs and prompts. These previews are for style and concept confirmation, not final production frames.
Default preview ratio follows the project video ratio: horizontal 16:9 unless the user specifies another ratio. If the user requested vertical, square, or another format, use that ratio consistently for previews and later video planning.
For each preview, define:
- Preview number and purpose
- Which approved direction, character, or scene it visualizes
- Key visual focus: overall world, character look, mechanism clarity, cultural atmosphere, or ending memory image
- What the user should judge before continuing
Recommended preview set:
1. Overall world preview: shows the full papercraft diorama, main character, core object, and atmosphere.
2. Mechanism preview: shows the science model, paper arrows, labels, and movable explanatory props.
3. Emotional or ending preview: shows the cultural connection, summary card, or final memory image.
After presenting previews, pause with a confirmation card. Offer choices such as: continue to prompt writing, revise preview 1, revise preview 2, revise preview 3, reduce to one visual direction, or switch creative direction. Do not write final single-image, image-series, or 5-second image-to-video prompts until the user confirms the visual preview direction.
## STEP 9: Write Single-Image Prompt
Create a prompt for one concept image or key visual.
The prompt must include:
- User's topic and core object
- Papercraft stop-motion style
- Miniature diorama stage
- Layered cardboard cutouts
- Visible paper fibers, folds, seams, cut edges, and thickness
- Real physical shadows between layers
- Macro miniature photography and slight depth of field
- Educational paper labels, arrows, or cards when needed
Avoid overloading the image with every knowledge detail. One image should communicate one main idea.
## STEP 10: Write Image-Series Prompts
Create a sequence of prompts when the user needs multiple keyframes or a storyboard image set.
For each image, provide:
- Image number and purpose
- What changes from the previous image
- Shared style anchors
- Character and scene consistency anchors
- Prompt
Series rules:
- Keep the same paper material language, light direction, depth planes, and character identity.
- Each image should explain one step of the knowledge point.
- Preserve domain-specific words from the user's topic in every prompt.
- Use the same aspect ratio across the series unless the user requests variations.
## STEP 11: Write 5-Second Image-to-Video Prompt
Create a short prompt that animates a single reference image while preserving the paper style.
Include:
- Preserve paper texture, cut edges, layered set, and physical shadows
- Slow push-in, gentle pan, or parallax move
- Small stop-motion-like puppet gestures
- Paper arrows, labels, sliders, or mechanisms moving slightly
- Foreground and background parallax
- Avoid smooth CG transformation, melting, plastic surfaces, or high-speed camera moves
Keep motion limited and physically plausible for paper objects.
## STEP 12: Create Storyboard for the Chosen Duration
Use the duration chosen after the creative directions phase. Do not present three full storyboard versions by default. First give a brief content overview, then provide a concise storyboard table for the chosen duration only.
The content overview should be short and clear:
- One-sentence video premise
- Target duration and ratio
- Main visual structure
- Knowledge path: hook → explanation → example or cultural connection → memory sentence
Storyboard table rules:
- Keep text compact. Avoid long paragraphs inside table cells.
- Use only the shots needed for the chosen duration.
- 15s: usually 4 shots.
- 30s: usually 5-6 shots.
- 60s: usually 7-9 shots.
- Each shot explains one knowledge beat.
Required columns:
- Time
- Knowledge beat
- Visual action
- Paper movement
- Camera / transition
- Sound cue
After the table, ask the user with a confirmation card: continue to editing rhythm and camera rules, revise storyboard, change duration, or return to visual previews.
## STEP 13: Define Editing Rhythm
Set the rhythm according to duration and educational clarity.
Guidelines:
- 15s: 4-6 shots, fast but readable, 2-4 seconds per shot.
- 30s: 6-8 shots, 3-5 seconds per shot.
- 60s: 8-12 shots, 4-7 seconds per shot.
- Pause briefly when a label card appears.
- Do not cut before the viewer understands the paper mechanism.
- Use rhythmic paper movements instead of aggressive digital edits.
## STEP 14: Define Camera Rules
Use camera movement that feels like filming a miniature paper stage.
Recommended moves:
- Slow push-in to enter the paper world
- Lateral pan with multi-plane parallax
- Static medium shot for host explanation
- Macro close-up for paper texture and key props
- Slight top-down angle for cross-section diagrams
- Layer pass-through when moving into internal structures
Avoid:
- High-speed flying camera
- Full 360-degree orbit unless the set is explicitly built for it
- Digital glitch motion
- Liquid morphing
- Hyper-real CG camera behavior
## STEP 15: Define Transitions
Use transitions that obey paper physics.
Recommended transitions:
- Page flip
- Pop-up-book unfold
- Pull-tab slide
- Paper label wipe
- Paper cloud pass
- Circular paper mask
- Cut-paper door opening
- Cross-section layer split
- Paper confetti burst
- Tape or sticker reveal
Avoid electronic scanlines, neon glitches, glass shatter, metallic wipes, and sci-fi particle transitions unless the user explicitly requests a contrast effect.
## STEP 16: Define Sound Design
Build a tactile handmade sound palette.
Recommended sound cues:
- Paper flip
- Scissor snip
- Cardboard slide
- Paper rustle
- Small wooden click
- Soft pop
- Tape peel
- Puppet joint tap
- Box opening
- Paper confetti scatter
Music direction:
- Match the BGM to the user's topic, culture, and emotional tone, not only to the generic papercraft style.
- For culturally specific topics, design BGM around the topic's relevant musical language. For example, a Mid-Autumn Festival explainer should lean toward restrained Chinese traditional color: guzheng, pipa, bamboo flute or xiao, light percussion, soft strings, and enough silence for narration.
- For science or classroom topics without a strong cultural identity, use light marimba, xylophone, pizzicato strings, soft percussion, and a warm educational tone.
- Keep music light under narration and apply ducking in the final mix.
- Avoid heavy cinematic bass, aggressive EDM, futuristic textures, or overdramatic scoring unless requested.
Paper-motion SFX direction:
- Design SFX from the actual video motion beats: page flip, paper door opening, paper rail slide, card flip, cardboard drawer pull, paper box opening, paper layer stack, soft pop, puppet joint tap, and paper rustle.
- Use SFX sparingly as tactile accents. They should make paper motion feel physical, not become exaggerated cartoon sounds.
- Map SFX to the storyboard timeline before final mixing so action, narration, and sound reinforce each other.
## STEP 17: Generate and Clean Voiceover Audio
When the workflow includes narration, generate the voiceover after the script is confirmed. Then check the generated audio before final assembly.
Voiceover rules:
- Match voice tone to the topic and audience: warm, clear, and gentle for family or education topics; more neutral for classroom or technical explainers.
- Compare generated voiceover duration with the chosen video duration. If the deviation is over 20%, ask the user with a card whether to shorten the script, extend the video, or accept the mismatch.
- Listen for tail noise, clicks, breath artifacts, abrupt cutoffs, or model residue at the end.
- If tail noise exists, repair the existing audio first: trim the noisy tail and add a short fade-out. Prefer repair over regenerating when the voice performance is otherwise good.
- If regenerating, explicitly request a clean ending with no tail noise, but still verify after generation.
## STEP 18: Provide Negative Prompts
Always include a concise negative prompt block to protect the style.
Recommended negatives:
- smooth plastic 3D
- glossy CG render
- live-action realism
- flat vector illustration
- generic cartoon with no paper texture
- metallic sci-fi surfaces
- glass material
- cyberpunk neon
- digital glitch effects
- oil painting strokes
- realistic hair or skin pores
- no paper fibers
- no cut edges
- no layer shadows
- overly smooth edges
- high-speed camera orbit
- melting or liquid morphing
## STEP 19: Run Review Checklist
End with a checklist. Mark problems only when they matter to the user's requested output.
### Style checklist
- Paper fibers, folds, seams, and cut edges are visible.
- Cardboard thickness is visible.
- Layers cast physical shadows.
- The image feels like miniature photography, not flat illustration.
- The style avoids plastic CG and live-action realism.
### Character checklist
- Characters read as paper puppets.
- Faces and expressions are cut-paper simple.
- Joints and limbs feel assembled from separate parts.
- Motion remains stop-motion-like and physically plausible.
### Scene checklist
- Foreground, midground, background, and far background are clear.
- Props are paper-made and support the explanation.
- Labels, arrows, and cards do not block the core model.
- Each shot explains only one knowledge beat.
### Video checklist
- Camera moves match miniature filming.
- Transitions follow paper physics.
- Sound cues match paper, cardboard, and puppet movement.
- The storyboard duration fits the target length.
- The final memory sentence is clear.
### Audio checklist
- Voiceover duration fits the chosen video duration or the mismatch has been resolved with user confirmation.
- Voiceover ending is clean, with no tail noise, click, breath residue, or abrupt cutoff.
- BGM matches the topic's culture and emotional tone, not only the generic papercraft style.
- BGM stays under narration and does not compete with spoken information.
- Paper-motion SFX align with actual visual motion beats and remain subtle.
@@ -0,0 +1,19 @@
display-name-zh: 纸艺定格科普视频生成器
version: 0.6.5
tag-en: Education
tag-cn: 教育
complete-tags-en:
- Education / Creative Generation
- Animation / Creative Generation
- Creative & Experimental / Creative Generation
complete-tags-cn:
- 教育 / 创作生成
- 动画 / 创作生成
- 创意实验 / 创作生成
summary-en: Build a production-ready papercraft stop-motion explainer package from a knowledge topic.
summary-cn: 根据知识主题,完成纸偶、分层布景、分镜与提示词设计,输出可执行的纸艺定格科普视频创作包。
desc-en: For creators explaining science, education, or general knowledge through tactile handmade papercraft visuals. Users provide a topic, core knowledge points, or source material and may specify audience, duration, aspect ratio, and deliverable type. The Skill extracts the learning goal and visual metaphor, proposes creative directions, designs paper characters, layered diorama sets, and props, creates preview concepts plus image and video prompts, and plans storyboards, camera movement, transitions, and sound with staged approvals and review checklists. It outputs a production-ready papercraft stop-motion explainer package, or selected assets such as still prompts, image-series prompts, short-video prompts, or storyboards. Best for cut-paper, pop-up-book, layered diorama, and miniature stop-motion explainers; not for standard 2D cartoons, line doodles, live action, or explainers without a paper-art look.
desc-cn: 面向希望用手工纸艺视觉讲解科学、教育或泛知识内容的创作者。用户需提供主题、核心知识点或原始材料,并可指定受众、时长、画幅和交付类型。Skill 会先提炼学习目标与视觉隐喻,提出创意方向,设计纸偶角色、分层纸雕布景和道具资产,制作预览概念、图像与视频提示词,规划分镜、运镜、转场和声音,并通过阶段确认与审核清单控制质量。最终输出可直接进入制作的纸艺定格科普视频创作包,也可按需交付单图、系列图、短视频提示词或分镜。适用于纸雕、剪纸、立体书和微缩定格讲解,不适用于普通二维卡通、线稿涂鸦、真人写实视频或无纸艺质感的标准科普。
author-en: MiniMax Hub
author-cn: MiniMax Hub
source: official
+9
View File
@@ -0,0 +1,9 @@
# 故事脚本生成
## 角色设定
你是一位擅长写短视频分镜的编剧,风格偏向情绪化、节奏快。
## 输出要求
- 每段分镜控制在 80 字以内
- 必须包含景别、运镜、镜头时长三要素
- 输出 JSON 数组,每个对象包含 `shot`、`duration`、`note`
-116
View File
@@ -1,116 +0,0 @@
wechat:
appid: wx1c2b19f7af1d6241 #公众号appid
secret: 915d587b5ae4e148e100429e11ca4d60 #公众号秘钥
serverAddress: http://139.9.53.208/ #服务器地址 微信公众配置地址的前缀
freeSize: 3 #免费使用次数
isEnterprise: false #是否为企业微信
adminNo: b0pOVFM2dnRsZkdLeWl2Zlk2bG9MTFNjUTNGUQ== #管理员微信编号启动会后可发送我的编号获取
adminWeChat: yanlang123456 #管理员微信号
access_token: '' #不用填系统用于保存临时值
access_token_expires_at: 0 #默认0
commands: #功能列表
图生视频:
filename: img2video.json #功能流程图json文件
params: #参数列表
image: #参数名称(几个固定参数名称必须如下:图片:image 随机种子:seed 正向提示词:prompt 反向提示词:negative)
isRequired: true #是否必填
keys: #参数对应json中对应的key
- '38'
- inputs
- image_path
zhName: 图片 #参数中文名称
motion_bucket:
default: 127
isRequired: true
keys:
- '12'
- inputs
- motion_bucket_id
max: 200
min: 0
type: number
zhName: 运动量
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
replyText: 请发送您的图片发送“ok”开始任务,如:“ok” #指令说明回复
文生图:
filename: text2img.json
params:
ckpt:
isRequired: false
keys:
- '4'
- inputs
- ckpt_name
options: #选项参数
冷却混合料SD1.5: chilloutmix_NiPrunedFp32Fix.safetensors #中文名称:英文名称
卡通动画片SD1.5: yamer_Cartoon_xenoArcadiaCD.safetensors
无法控制的SDXL: juggernautXL_version6Rundiffusion.safetensors
梦想塑造者SD1.5: dreamshaper_8.safetensors
zhName: 基础模型
height:
default: 512 #默认值
isRequired: false
keys:
- '14'
- inputs
- height
max: 1024 #最大值
min: 512 #最小值
type: number #类型(输入型参数目前只支持 number/text)
zhName: 高度
lora:
isRequired: false
keys:
- '10'
- inputs
- lora_name
options:
SD1.5-LCM加速: lcm-lora-sdv1-5.safetensors
SDXL-LCM加速: lcm-lora-sdxl.safetensors
zhName: LoRA
negative:
isRequired: false
keys:
- '7'
- inputs
- text
zhName: 反向提示词
prompt:
isRequired: true
keys:
- '20'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
width:
default: 512
isRequired: false
keys:
- '14'
- inputs
- width
max: 1024
min: 512
type: number
zhName: 宽度
replyText: 请发送您的中文或英文提示词用“ok”结束,如:“一只小狗ok”
query_commands: #查询指令(固定)
- 查询排队情况
- 我的编号
+307 -272
View File
@@ -1,284 +1,319 @@
cluster:
modelPriority: false
clusterType: redis
subordinates:
- 127.0.0.1:8190
redis:
host: 127.0.0.1
port: 6379
password:
base:
language: zh-CN # ja-JP ok-KR ru-RU zh-TW zh-CN en-US
api_is_used: true
appTitle: AI创作平台
appLogo: https://element-plus.sxtxhy.com/images/element-plus-logo.svg
freeSize: 3 #免费使用次数
authorIds:
- oJNTS6vtlfGKyivfY6loLLScQ3FQ
- oJNTS6vtlfGKyivfY6loLLScQ3F1
- oJNTS6vtlfGKyivfY6loLLScQ3F2
- oJNTS6vtlfGKyivfY6loLLScQ3F3
ai:
api_key: sk-a446d27d074a45c0ac8ca1ac85742b28 #sk-a446d27d074a45c0ac8ca1ac85742b28 glm4 6795fbe303878f35292f1aa14414e9a4.kOzmNe7PEk1l6zdK sk-rttnlipmkojrsjyjrfoudbuphledqizttlqergzqxyyochfl
model: deepseek-chat # glm-4-flash deepseek-chat
ai_type: openAi #glm4 openAi
base_url: https://api.deepseek.com
is_tools: true
sys_pompt: 你是一个AI助手,能帮助我解答问题,回答内容最多不用超过200字
jian_ying:
drafts_root: C:\Users\Administrator\AppData\Local\JianyingPro\User Data\Projects\com.lveditor.draft\
#图片生成视频的时候,每幅图片的持续时间,单位是微秒
image_duration: 5000000
wechat:
appid: wxa50ed345b253964d #公众号appid
secret: 327ac9ba06d214d2fcfe2e9ed0520684 #公众号秘钥
serverAddress: http://119.123.205.233/ #服务器地址 微信公众配置地址的前缀
freeSize: 3 #免费使用次数
subscribeMsg: "感谢关注!您可以发送“帮助”查看使用说明"
isEnterprise: true #是否为企业微信
adminNo: b0pOVFM2dnRsZkdLeWl2Zlk2bG9MTFNjUTNGUQ== #管理员微信编号启动会后可发送我的编号获取
adminWeChat: yanlang123456 #管理员微信号
access_token: '' #不用填系统用于保存临时值
access_token_expires_at: 0 #默认0
commands: #功能列表
简笔画:
filename: line2drawing.json
type: paint-board
params:
ckpt:
isRequired: false
keys:
- '1'
- inputs
- ckpt_name
options: #选项参数
卡通动画片SD1.5: yamer_Cartoon_xenoArcadiaCD.safetensors
冷却混合料SD1.5: chilloutmix_NiPrunedFp32Fix.safetensors #中文名称:英文名称
梦想塑造者SD1.5: dreamshaper_8.safetensors
zhName: 基础模型
prompt:
isRequired: false
keys:
- '17'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '8'
- inputs
- seed
type: text
zhName: 随机种子
image:
isRequired: true
keys:
- '15'
- inputs
- image
type: text
zhName: 输入图片
水彩画:
filename: color2painting.json
type: paint-board
params:
ckpt:
isRequired: false
keys:
- '4'
- inputs
- ckpt_name
options: #选项参数
卡通动画片SD1.5: yamer_Cartoon_xenoArcadiaCD.safetensors
冷却混合料SD1.5: chilloutmix_NiPrunedFp32Fix.safetensors #中文名称:英文名称
梦想塑造者SD1.5: dreamshaper_8.safetensors
zhName: 基础模型
prompt:
isRequired: false
keys:
- '27'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
denoise:
default: 0.65 #默认值
isRequired: false
keys:
- '3'
- inputs
- denoise
max: 100 #最大值
min: 0 #最小值
original: 1 #原始最大值(用来换算滑块比例)
type: number #类型(输入型参数目前只支持 number/text)
zhName: 重绘幅度
image:
isRequired: true
keys:
- '25'
- inputs
- image
type: text
zhName: 输入图片
图片修改:
filename: imageEdit.json
params:
image:
isRequired: true
keys:
- '48'
- inputs
- image_path
zhName: 图片
prompt:
isRequired: true
keys:
- '47'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '37:1'
- inputs
- noise_seed
type: text
zhName: 随机种子
replyText: 请发送您的图片发送正向提示词用“ok”结尾开始任务,如:“漫画风格ok”
风格转绘:
filename: img2imgStyle.json
replyText: 请上传需转绘的图片和参考风格图发送“ok”开始任务,如:“ok” #指令说明回复
params:
image: #参数名称(几个固定参数名称必须如下:图片:image 随机种子:seed 正向提示词:prompt 反向提示词:negative)
isRequired: true #是否必填
keys: #参数对应json中对应的key
- '110'
- inputs
- image_path
zhName: 待转绘图
type: image
styleImage: #参数名称(几个固定参数名称必须如下:图片:image 随机种子:seed 正向提示词:prompt 反向提示词:negative)
isRequired: true #是否必填
keys: #参数对应json中对应的key
- '111'
- inputs
- image_path
zhName: 风格参考图
type: image
default: 发送图片
文生二维码:
filename: text2qrcode.json
replyText: 请发送二维码内容,和提示词发送“ok”开始任务,如:“花朵ok” #指令说明回复
params:
prompt:
isRequired: true
keys:
- '75'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
qrText:
default: "https://u.wechat.com/EEy4mPvfXdYO-MvStE4S2m0"
isRequired: true
keys:
- '16'
- inputs
- text
type: text
zhName: 二维码内容
图生视频:
filename: img2video.json #功能流程图json文件
params: #参数列表
image: #参数名称(几个固定参数名称必须如下:图片:image 随机种子:seed 正向提示词:prompt 反向提示词:negative)
isRequired: true #是否必填
keys: #参数对应json中对应的key
- '38'
- inputs
- image_path
zhName: 图片 #参数中文名称
motion_bucket:
default: 127
isRequired: true
keys:
- '12'
- inputs
- motion_bucket_id
max: 200
min: 0
type: number
zhName: 运动量
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
replyText: 请发送您的图片发送“ok”开始任务,如:“ok” #指令说明回复
文生图:
filename: text2img.json
params:
ckpt:
isRequired: false
keys:
- '4'
- inputs
- ckpt_name
options: #选项参数
冷却混合料SD1.5: chilloutmix_NiPrunedFp32Fix.safetensors #中文名称:英文名称
卡通动画片SD1.5: yamer_Cartoon_xenoArcadiaCD.safetensors
无法控制的SDXL: juggernautXL_version6Rundiffusion.safetensors
梦想塑造者SD1.5: dreamshaper_8.safetensors
zhName: 基础模型
height:
default: 512 #默认值
isRequired: false
keys:
- '14'
- inputs
- height
max: 1024 #最大值
min: 512 #最小值
type: number #类型(输入型参数目前只支持 number/text)
zhName: 高度
lora:
isRequired: false
keys:
- '10'
- inputs
- lora_name
options:
SD1.5-LCM加速: lcm-lora-sdv1-5.safetensors
SDXL-LCM加速: lcm-lora-sdxl.safetensors
zhName: LoRA
negative:
isRequired: false
keys:
- '7'
- inputs
- text
zhName: 反向提示词
prompt:
isRequired: true
keys:
- '20'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
width:
default: 512
isRequired: false
keys:
- '14'
- inputs
- width
max: 1024
min: 512
type: number
zhName: 宽度
replyText: 请发送您的中文或英文提示词用“ok”结束,如:“一只小狗ok”
query_commands: #查询指令(固定)
- 查询排队情况
- 我的编号
- 退出
commands: #功能列表
简笔画:
filename: line2drawing.json
type: paint-board
params:
ckpt:
isRequired: false
keys:
- '1'
- inputs
- ckpt_name
options: #选项参数
卡通动画片SD1.5: yamer_Cartoon_xenoArcadiaCD.safetensors
冷却混合料SD1.5: chilloutmix_NiPrunedFp32Fix.safetensors #中文名称:英文名称
梦想塑造者SD1.5: dreamshaper_8.safetensors
zhName: 基础模型
prompt:
isRequired: false
keys:
- '17'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '8'
- inputs
- seed
type: text
zhName: 随机种子
image:
isRequired: true
keys:
- '15'
- inputs
- image
type: text
zhName: 输入图片
水彩画:
filename: color2painting.json
type: paint-board
params:
ckpt:
isRequired: false
keys:
- '4'
- inputs
- ckpt_name
options: #选项参数
卡通动画片SD1.5: yamer_Cartoon_xenoArcadiaCD.safetensors
冷却混合料SD1.5: chilloutmix_NiPrunedFp32Fix.safetensors #中文名称:英文名称
梦想塑造者SD1.5: dreamshaper_8.safetensors
zhName: 基础模型
prompt:
isRequired: false
keys:
- '27'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
denoise:
default: 0.65 #默认值
isRequired: false
keys:
- '3'
- inputs
- denoise
max: 100 #最大值
min: 0 #最小值
original: 1 #原始最大值(用来换算滑块比例)
type: number #类型(输入型参数目前只支持 number/text)
zhName: 重绘幅度
image:
isRequired: true
keys:
- '25'
- inputs
- image
type: text
zhName: 输入图片
图片修改:
filename: imageEdit.json
params:
image:
isRequired: true
keys:
- '48'
- inputs
- image_path
zhName: 图片
prompt:
isRequired: true
keys:
- '47'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '37:1'
- inputs
- noise_seed
type: text
zhName: 随机种子
replyText: 请发送您的图片发送正向提示词用“ok”结尾开始任务,如:“漫画风格ok”
风格转绘:
filename: img2imgStyle.json
replyText: 请上传需转绘的图片和参考风格图发送“ok”开始任务,如:“ok” #指令说明回复
params:
image: #参数名称(几个固定参数名称必须如下:图片:image 随机种子:seed 正向提示词:prompt 反向提示词:negative)
isRequired: true #是否必填
keys: #参数对应json中对应的key
- '110'
- inputs
- image_path
zhName: 待转绘图
type: image
styleImage: #参数名称(几个固定参数名称必须如下:图片:image 随机种子:seed 正向提示词:prompt 反向提示词:negative)
isRequired: true #是否必填
keys: #参数对应json中对应的key
- '111'
- inputs
- image_path
zhName: 风格参考图
type: image
default: 发送图片
文生二维码:
filename: text2qrcode.json
replyText: 请发送二维码内容,和提示词发送“ok”开始任务,如:“花朵ok” #指令说明回复
params:
prompt:
isRequired: true
keys:
- '75'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
qrText:
default: "https://u.wechat.com/EEy4mPvfXdYO-MvStE4S2m0"
isRequired: true
keys:
- '16'
- inputs
- text
type: text
zhName: 二维码内容
图生视频:
filename: img2video.json #功能流程图json文件
params: #参数列表
image: #参数名称(几个固定参数名称必须如下:图片:image 随机种子:seed 正向提示词:prompt 反向提示词:negative)
isRequired: true #是否必填
keys: #参数对应json中对应的key
- '38'
- inputs
- image_path
zhName: 图片 #参数中文名称
motion_bucket:
default: 127
isRequired: true
keys:
- '12'
- inputs
- motion_bucket_id
max: 200
min: 0
type: number
zhName: 运动量
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
replyText: 请发送您的图片发送“ok”开始任务,如:“ok” #指令说明回复
文生图:
filename: text2img.json
params:
ckpt:
isRequired: false
keys:
- '4'
- inputs
- ckpt_name
options: #选项参数
冷却混合料SD1.5: chilloutmix_NiPrunedFp32Fix.safetensors #中文名称:英文名称
卡通动画片SD1.5: yamer_Cartoon_xenoArcadiaCD.safetensors
无法控制的SDXL: juggernautXL_version6Rundiffusion.safetensors
梦想塑造者SD1.5: dreamshaper_8.safetensors
zhName: 基础模型
height:
default: 512 #默认值
isRequired: false
keys:
- '14'
- inputs
- height
max: 1024 #最大值
min: 512 #最小值
type: number #类型(输入型参数目前只支持 number/text)
zhName: 高度
lora:
isRequired: false
keys:
- '10'
- inputs
- lora_name
options:
SD1.5-LCM加速: lcm-lora-sdv1-5.safetensors
SDXL-LCM加速: lcm-lora-sdxl.safetensors
zhName: LoRA
negative:
isRequired: false
keys:
- '7'
- inputs
- text
zhName: 反向提示词
prompt:
isRequired: true
keys:
- '20'
- inputs
- text_trans
zhName: 正向提示词
seed:
default: -1
isRequired: true
keys:
- '3'
- inputs
- seed
type: text
zhName: 随机种子
width:
default: 512
isRequired: false
keys:
- '14'
- inputs
- width
max: 1024
min: 512
type: number
zhName: 宽度
replyText: 请发送您的中文或英文提示词用“ok”结束,如:“一只小狗ok”
+360 -14
View File
@@ -1,7 +1,7 @@
{
"3": {
"inputs": {
"seed": 611578276912800,
"seed": 2070780056835,
"steps": 8,
"cfg": 1,
"sampler_name": "euler_ancestral",
@@ -86,19 +86,6 @@
"title": "VAE解码"
}
},
"9": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
0
]
},
"class_type": "SaveImage",
"_meta": {
"title": "保存图像"
}
},
"10": {
"inputs": {
"lora_name": "lcm-lora-sdv1-5.safetensors",
@@ -188,5 +175,364 @@
"_meta": {
"title": "反向提示词"
}
},
"29": {
"inputs": {
"seed": 19,
"steps": 25,
"cfg": 6,
"sampler_name": "dpmpp_2m",
"scheduler": "karras",
"denoise": 0.9,
"model": [
"36",
0
],
"positive": [
"48",
0
],
"negative": [
"48",
1
],
"latent_image": [
"48",
2
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"31": {
"inputs": {
"text": [
"52",
0
],
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"32": {
"inputs": {
"text": "blurry, noisy, messy, glitch, distorted, malformed, ill, horror, naked, nipples",
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"33": {
"inputs": {
"samples": [
"29",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"34": {
"inputs": {
"images": [
"56",
0
]
},
"class_type": "PreviewImage",
"_meta": {
"title": "预览图像"
}
},
"36": {
"inputs": {
"weight_style": 1,
"weight_composition": 1,
"expand_style": false,
"combine_embeds": "concat",
"start_at": 0,
"end_at": 1,
"embeds_scaling": "V only",
"model": [
"37",
0
],
"ipadapter": [
"37",
1
],
"image_style": [
"39",
0
],
"image_composition": [
"47",
0
]
},
"class_type": "IPAdapterStyleComposition",
"_meta": {
"title": "IPAdapter风格合成SDXL"
}
},
"37": {
"inputs": {
"preset": "STANDARD (medium strength)",
"model": [
"4",
0
]
},
"class_type": "IPAdapterUnifiedLoader",
"_meta": {
"title": "IPAdapter加载器"
}
},
"39": {
"inputs": {
"interpolation": "LANCZOS",
"crop_position": "pad",
"sharpening": 0,
"image": [
"8",
0
]
},
"class_type": "PrepImageForClipVision",
"_meta": {
"title": "CLIP视觉图像处理"
}
},
"40": {
"inputs": {
"strength": 0.55,
"start_percent": 0,
"end_percent": 0.75,
"positive": [
"31",
0
],
"negative": [
"32",
0
],
"control_net": [
"41",
0
],
"image": [
"49",
0
]
},
"class_type": "ControlNetApplyAdvanced",
"_meta": {
"title": "ControlNet应用(高级)"
}
},
"41": {
"inputs": {
"control_net_name": "t2i-adapter-lineart-sdxl-1.0.safetensors"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"43": {
"inputs": {
"strength": 1,
"start_percent": 0,
"end_percent": 0.3,
"positive": [
"40",
0
],
"negative": [
"40",
1
],
"control_net": [
"44",
0
],
"image": [
"45",
0
]
},
"class_type": "ControlNetApplyAdvanced",
"_meta": {
"title": "ControlNet应用(高级)"
}
},
"44": {
"inputs": {
"control_net_name": "diffusers_xl_depth_mid.safetensors"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"45": {
"inputs": {
"ckpt_name": "depth_anything_vits14.pth",
"resolution": 960,
"image": [
"55",
0
]
},
"class_type": "DepthAnythingPreprocessor",
"_meta": {
"title": "DA深度预处理器"
}
},
"47": {
"inputs": {
"interpolation": "LANCZOS",
"crop_position": "pad",
"sharpening": 0,
"image": [
"55",
0
]
},
"class_type": "PrepImageForClipVision",
"_meta": {
"title": "CLIP视觉图像处理"
}
},
"48": {
"inputs": {
"positive": [
"43",
0
],
"negative": [
"43",
1
],
"vae": [
"4",
2
],
"pixels": [
"55",
0
]
},
"class_type": "InstructPixToPixConditioning",
"_meta": {
"title": "InstructPixToPix条件"
}
},
"49": {
"inputs": {
"coarse": "disable",
"resolution": 512,
"image": [
"55",
0
]
},
"class_type": "LineArtPreprocessor",
"_meta": {
"title": "LineArt艺术线预处理器"
}
},
"52": {
"inputs": {
"model": "wd-v1-4-moat-tagger-v2",
"threshold": 0.35,
"character_threshold": 0.85,
"replace_underscore": "",
"trailing_comma": "1girl, solo, long_hair, looking_at_viewer, shirt, black_hair, upper_body, black_eyes, lips, realistic",
"exclude_tags": "",
"image": [
"55",
0
]
},
"class_type": "WD14Tagger|pysssss",
"_meta": {
"title": "WD14反推提示词"
}
},
"55": {
"inputs": {
"image_path": [
"59",
0
],
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "加载网络图片或本地图片"
}
},
"56": {
"inputs": {
"ANY": [
"60",
0
],
"IF_TRUE": [
"33",
0
],
"IF_FALSE": [
"8",
0
]
},
"class_type": "IfInnerExecute",
"_meta": {
"title": "判断选择"
}
},
"59": {
"inputs": {
"text": ""
},
"class_type": "TextInput_",
"_meta": {
"title": "文本"
}
},
"60": {
"inputs": {
"expression": "len(p0)>0",
"p0": [
"59",
0
]
},
"class_type": "MultiParamFormula",
"_meta": {
"title": "多参代码表达式"
}
}
}
+3 -3
View File
@@ -163,9 +163,9 @@
0
]
},
"class_type": "GetImage_(Width&Height) _O",
"class_type": "GetImageSize+",
"_meta": {
"title": "GetImage_(Width&Height) _O"
"title": "GetImageSize"
}
},
"17": {
@@ -194,7 +194,7 @@
0
]
},
"class_type": "SaveImage",
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像"
}
+644
View File
@@ -0,0 +1,644 @@
{
"3": {
"inputs": {
"seed": [
"hidden",
"seed7"
],
"steps": 25,
"cfg": 6,
"sampler_name": "dpmpp_2m",
"scheduler": "karras",
"denoise": 0.9,
"model": [
"105",
0
],
"positive": [
"64",
0
],
"negative": [
"64",
1
],
"latent_image": [
"64",
2
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"6": {
"inputs": {
"text": [
"114",
0
],
"clip": [
"hidden",
"CLIP5"
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"7": {
"inputs": {
"text": "blurry, noisy, messy, glitch, distorted, malformed, ill, horror, naked, nipples",
"clip": [
"hidden",
"CLIP5"
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"8": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"hidden",
"VAE6"
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
},
"outputs": [
[
0,
0
]
]
},
"48": {
"inputs": {
"preset": "STANDARD (medium strength)",
"model": [
"hidden",
"MODEL4"
]
},
"class_type": "IPAdapterUnifiedLoader",
"_meta": {
"title": "IPAdapter加载器"
}
},
"50": {
"inputs": {
"interpolation": "LANCZOS",
"crop_position": "pad",
"sharpening": 0,
"image": [
"hidden",
"IMAGE0"
]
},
"class_type": "PrepImageForClipVision",
"_meta": {
"title": "CLIP视觉图像处理"
}
},
"51": {
"inputs": {
"strength": 0.55,
"start_percent": 0,
"end_percent": 0.75,
"positive": [
"6",
0
],
"negative": [
"7",
0
],
"control_net": [
"52",
0
],
"image": [
"83",
0
]
},
"class_type": "ControlNetApplyAdvanced",
"_meta": {
"title": "ControlNet应用(旧版高级)"
}
},
"52": {
"inputs": {
"control_net_name": "control_v11p_sd15_lineart.pth"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"55": {
"inputs": {
"strength": 1,
"start_percent": 0,
"end_percent": 0.3,
"positive": [
"51",
0
],
"negative": [
"51",
1
],
"control_net": [
"56",
0
],
"image": [
"57",
0
]
},
"class_type": "ControlNetApplyAdvanced",
"_meta": {
"title": "ControlNet应用(旧版高级)"
}
},
"56": {
"inputs": {
"control_net_name": "control_v11f1p_sd15_depth.pth"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"57": {
"inputs": {
"ckpt_name": "depth_anything_vits14.pth",
"resolution": 960,
"image": [
"71",
0
]
},
"class_type": "DepthAnythingPreprocessor",
"_meta": {
"title": "DepthAnything深度预处理器"
}
},
"63": {
"inputs": {
"interpolation": "LANCZOS",
"crop_position": "pad",
"sharpening": 0,
"image": [
"71",
0
]
},
"class_type": "PrepImageForClipVision",
"_meta": {
"title": "CLIP视觉图像处理"
}
},
"64": {
"inputs": {
"positive": [
"55",
0
],
"negative": [
"55",
1
],
"vae": [
"hidden",
"VAE6"
],
"pixels": [
"71",
0
]
},
"class_type": "InstructPixToPixConditioning",
"_meta": {
"title": "InstructPixToPix条件"
}
},
"70": {
"inputs": {
"expression": "p0>960",
"advanced": "disable",
"p0": [
"110",
0
]
},
"class_type": "MultiParamFormula",
"_meta": {
"title": "多参代码表达式"
}
},
"71": {
"inputs": {
"ANY": [
"70",
0
],
"IF_TRUE": [
"116",
0
],
"IF_FALSE": [
"hidden",
"IMAGE2"
]
},
"class_type": "IfInnerExecute",
"_meta": {
"title": "判断选择"
},
"outputs": [
[
0,
1
]
]
},
"76": {
"inputs": {
"preset": "PLUS FACE (portraits)",
"model": [
"112",
0
]
},
"class_type": "IPAdapterUnifiedLoader",
"_meta": {
"title": "IPAdapter加载器"
}
},
"83": {
"inputs": {
"coarse": "disable",
"resolution": 512,
"image": [
"71",
0
]
},
"class_type": "LineArtPreprocessor",
"_meta": {
"title": "LineArt艺术线预处理器"
}
},
"88": {
"inputs": {
"expression": "p0[p1:p1+1]",
"advanced": "disable",
"p0": [
"108",
0
],
"p1": [
"89",
1
]
},
"class_type": "MultiParamFormula",
"_meta": {
"title": "多参代码表达式"
}
},
"89": {
"inputs": {
"total": [
"90",
0
],
"stop": 1,
"i": 0
},
"class_type": "ForInnerStart",
"_meta": {
"title": "计次内循环首"
}
},
"90": {
"inputs": {
"expression": "len(p0)",
"advanced": "disable",
"p0": [
"108",
1
]
},
"class_type": "MultiParamFormula",
"_meta": {
"title": "多参代码表达式"
}
},
"91": {
"inputs": {
"expression": "p0[p1]",
"advanced": "disable",
"p0": [
"108",
1
],
"p1": [
"89",
1
]
},
"class_type": "MultiParamFormula",
"_meta": {
"title": "多参代码表达式"
}
},
"94": {
"inputs": {
"ANY": [
"95",
0
],
"IF_TRUE": [
"96",
0
],
"IF_FALSE": [
"89",
3
]
},
"class_type": "IfInnerExecute",
"_meta": {
"title": "判断选择"
}
},
"95": {
"inputs": {
"expression": "p0==0",
"advanced": "disable",
"p0": [
"89",
1
]
},
"class_type": "MultiParamFormula",
"_meta": {
"title": "多参代码表达式"
}
},
"96": {
"inputs": {
"preset": "FACEID PLUS V2",
"lora_strength": 0.39,
"provider": "CPU",
"model": [
"76",
0
],
"ipadapter": [
"76",
1
]
},
"class_type": "IPAdapterUnifiedLoaderFaceID",
"_meta": {
"title": "IPAdapterFaceID加载器"
}
},
"97": {
"inputs": {
"weight": 1,
"weight_faceidv2": 2,
"weight_type": "linear",
"combine_embeds": "concat",
"start_at": 0,
"end_at": 1,
"embeds_scaling": "V only",
"model": [
"94",
0
],
"ipadapter": [
"96",
1
],
"image": [
"88",
0
],
"attn_mask": [
"91",
0
]
},
"class_type": "IPAdapterFaceID",
"_meta": {
"title": "应用IPAdapterFaceID"
}
},
"101": {
"inputs": {
"images": [
"108",
0
]
},
"class_type": "PreviewImage",
"_meta": {
"title": "预览图像"
}
},
"104": {
"inputs": {
"expression": "len(p0)>0",
"advanced": "disable",
"p0": [
"108",
1
]
},
"class_type": "MultiParamFormula",
"_meta": {
"title": "多参代码表达式"
}
},
"105": {
"inputs": {
"ANY": [
"104",
0
],
"IF_TRUE": [
"115",
1
],
"IF_FALSE": [
"112",
0
]
},
"class_type": "IfInnerExecute",
"_meta": {
"title": "判断选择"
}
},
"108": {
"inputs": {
"crop_padding_factor": 0.25,
"analysis_models": [
"118",
0
],
"image": [
"71",
0
]
},
"class_type": "ImageCropFaces",
"_meta": {
"title": "多人面部裁剪"
}
},
"110": {
"inputs": {
"min_width": 512,
"image": [
"hidden",
"IMAGE2"
]
},
"class_type": "GetImageSize_",
"_meta": {
"title": "获取图像尺寸"
}
},
"112": {
"inputs": {
"weight": 1,
"weight_faceidv2": 1,
"weight_type": "linear",
"combine_embeds": "concat",
"start_at": 0,
"end_at": 1,
"embeds_scaling": "V only",
"layer_weights": "0:1,1:0,2:0,3:0,4:0,5:0,6:1,7:0,8:0,9:0,10:0,11:0",
"model": [
"113",
0
],
"ipadapter": [
"48",
1
],
"image": [
"50",
0
]
},
"class_type": "IPAdapterMS",
"_meta": {
"title": "应用IPAdapter Mad Scientist"
}
},
"113": {
"inputs": {
"weight": 1,
"weight_faceidv2": 1,
"weight_type": "composition",
"combine_embeds": "concat",
"start_at": 0,
"end_at": 1,
"embeds_scaling": "V only",
"layer_weights": "0:1,1:1,2:1,3:1,4:1,5:1,6:0,7:1,8:1,9:1,10:1,11:1",
"model": [
"48",
0
],
"ipadapter": [
"48",
1
],
"image": [
"63",
0
]
},
"class_type": "IPAdapterMS",
"_meta": {
"title": "应用IPAdapter Mad Scientist"
}
},
"114": {
"inputs": {
"prompt_mode": "fast",
"image_analysis": "off",
"image": [
"71",
0
]
},
"class_type": "ClipInterrogator",
"_meta": {
"title": "CLIP询问机"
}
},
"115": {
"inputs": {
"start": [
"89",
0
],
"obj": [
"97",
0
]
},
"class_type": "ForInnerEnd",
"_meta": {
"title": "计次内循环尾"
}
},
"116": {
"inputs": {
"width": 512,
"height": 512,
"interpolation": "nearest",
"method": "stretch",
"condition": "always",
"multiple_of": 0,
"image": [
"hidden",
"IMAGE2"
]
},
"class_type": "ImageResize+",
"_meta": {
"title": "图像缩放"
}
},
"118": {
"inputs": {
"library": "insightface",
"provider": "CPU"
},
"class_type": "LamFaceAnalysisModels",
"_meta": {
"title": "人脸分析模型"
}
}
}
+98 -71
View File
@@ -1,22 +1,42 @@
{
"1": {
"inputs": {
"unet_name": "flux1-dev-Q4_0.gguf"
},
"class_type": "UnetLoaderGGUF",
"_meta": {
"title": "Unet Loader (GGUF)"
}
},
"2": {
"inputs": {
"clip_name1": "t5-v1_1-xxl-encoder-Q8_0.gguf",
"clip_name2": "clip_l.safetensors",
"type": "flux"
},
"class_type": "DualCLIPLoaderGGUF",
"_meta": {
"title": "DualCLIPLoader (GGUF)"
}
},
"3": {
"inputs": {
"seed": 1068764177227380,
"steps": 5,
"seed": 119505262627551,
"steps": 25,
"cfg": 1,
"sampler_name": "lcm",
"sampler_name": "euler",
"scheduler": "normal",
"denoise": 1,
"model": [
"10",
"1",
0
],
"positive": [
"6",
"13",
0
],
"negative": [
"7",
"17",
0
],
"latent_image": [
@@ -24,102 +44,109 @@
0
]
},
"class_type": "KSampler"
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"4": {
"inputs": {
"ckpt_name": "dreamshaper_8.safetensors"
"text": [
"16",
0
],
"clip": [
"2",
0
]
},
"class_type": "CheckpointLoaderSimple"
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"5": {
"inputs": {
"width": [
"14",
0
],
"height": [
"14",
1
],
"width": 1024,
"height": 1024,
"batch_size": 1
},
"class_type": "EmptyLatentImage"
"class_type": "EmptyLatentImage",
"_meta": {
"title": "空Latent"
}
},
"6": {
"inputs": {
"text": [
"20",
0
],
"clip": [
"10",
1
]
},
"class_type": "CLIPTextEncode"
},
"7": {
"inputs": {
"text": "undefined",
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode"
},
"8": {
"9": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode"
},
"9": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
"10",
0
]
},
"class_type": "SaveImage"
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"10": {
"inputs": {
"lora_name": "lcm-lora-sdv1-5.safetensors",
"strength_model": 1,
"strength_clip": 1,
"model": [
"vae_name": "ae.safetensors"
},
"class_type": "VAELoader",
"_meta": {
"title": "VAE加载器"
}
},
"13": {
"inputs": {
"guidance": 3.5,
"conditioning": [
"4",
0
],
"clip": [
"4",
1
]
},
"class_type": "LoraLoader"
"class_type": "FluxGuidance",
"_meta": {
"title": "Flux引导"
}
},
"14": {
"16": {
"inputs": {
"aspect_ratio": "1152×896 ∣ 9:7",
"width": 512,
"height": 512
"text_trans": "一个男人,推着婴儿车,背景是阳光明媚的秋天"
},
"class_type": "AspectRatio"
"class_type": "ZhPromptTranslator",
"_meta": {
"title": "中文翻译"
}
},
"20": {
"17": {
"inputs": {
"text_trans": "一只小狗"
"text": "",
"clip": [
"2",
0
]
},
"class_type": "ZhPromptTranslator"
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"18": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"9",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
}
}
+258
View File
@@ -0,0 +1,258 @@
{
"3": {
"inputs": {
"seed": 809308338515272,
"steps": 20,
"cfg": 8,
"sampler_name": "euler",
"scheduler": "normal",
"denoise": 1,
"model": [
"4",
0
],
"positive": [
"6",
0
],
"negative": [
"7",
0
],
"latent_image": [
"13",
0
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"4": {
"inputs": {
"ckpt_name": "DreamShaperXLv2.1Turbov2.1Turbo.safetensors"
},
"class_type": "CheckpointLoaderSimple",
"_meta": {
"title": "Checkpoint加载器(简易)"
}
},
"6": {
"inputs": {
"text": [
"20",
0
],
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "正向提示词"
}
},
"7": {
"inputs": {
"text": "text, watermark",
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"8": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"11": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
},
"13": {
"inputs": {
"grow_mask_by": 6,
"pixels": [
"14",
0
],
"vae": [
"4",
2
],
"mask": [
"19",
0
]
},
"class_type": "VAEEncodeForInpaint",
"_meta": {
"title": "VAE内补编码器"
}
},
"14": {
"inputs": {
"image_path": "input//ComfyUI_01647_.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "图像"
}
},
"15": {
"inputs": {
"image_path": "input//mask_0.9623280348317427ComfyUI_01647_.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "遮罩"
}
},
"16": {
"inputs": {
"upscale_method": "nearest-exact",
"width": [
"17",
0
],
"height": [
"17",
1
],
"crop": "disabled",
"image": [
"25",
0
]
},
"class_type": "ImageScale",
"_meta": {
"title": "图像缩放"
}
},
"17": {
"inputs": {
"image": [
"14",
0
]
},
"class_type": "GetImageSize+",
"_meta": {
"title": "获取图像尺寸"
}
},
"18": {
"inputs": {
"channel": "red",
"image": [
"16",
0
]
},
"class_type": "ImageToMask",
"_meta": {
"title": "图像到遮罩"
}
},
"19": {
"inputs": {
"mask": [
"18",
0
]
},
"class_type": "InvertMask",
"_meta": {
"title": "遮罩反转"
}
},
"20": {
"inputs": {
"text_trans": "小鸡"
},
"class_type": "ZhPromptTranslator",
"_meta": {
"title": "中文翻译"
}
},
"21": {
"inputs": {
"images": [
"16",
0
]
},
"class_type": "PreviewImage",
"_meta": {
"title": "预览图像"
}
},
"22": {
"inputs": {
"images": [
"14",
0
]
},
"class_type": "PreviewImage",
"_meta": {
"title": "预览图像"
}
},
"23": {
"inputs": {
"mask": [
"19",
0
]
},
"class_type": "MaskPreview+",
"_meta": {
"title": "遮罩预览"
}
},
"25": {
"inputs": {
"mask": [
"15",
1
]
},
"class_type": "MaskToImage",
"_meta": {
"title": "遮罩到图像"
}
}
}
+232
View File
@@ -0,0 +1,232 @@
{
"3": {
"inputs": {
"seed": 871010165059929,
"steps": 8,
"cfg": 1,
"sampler_name": "euler_ancestral",
"scheduler": "normal",
"denoise": 0.6500000000000001,
"model": [
"21",
0
],
"positive": [
"6",
0
],
"negative": [
"7",
0
],
"latent_image": [
"19",
0
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"4": {
"inputs": {
"ckpt_name": "yamer_Cartoon_xenoArcadiaCD.safetensors"
},
"class_type": "CheckpointLoaderSimple",
"_meta": {
"title": "Checkpoint加载器(简易)"
}
},
"6": {
"inputs": {
"text": [
"27",
0
],
"speak_and_recognation": {
"__value__": [
false,
true
]
},
"clip": [
"21",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码"
}
},
"7": {
"inputs": {
"text": [
"28",
0
],
"speak_and_recognation": {
"__value__": [
false,
true
]
},
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码"
}
},
"8": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"10": {
"inputs": {
"lora_name": "lcm-lora-sdv1-5.safetensors",
"strength_model": 1.0000000000000002,
"strength_clip": 1.0000000000000002,
"model": [
"4",
0
],
"clip": [
"4",
1
]
},
"class_type": "LoraLoader",
"_meta": {
"title": "加载LoRA"
}
},
"19": {
"inputs": {
"pixels": [
"64",
0
],
"vae": [
"23",
0
]
},
"class_type": "VAEEncode",
"_meta": {
"title": "VAE编码"
}
},
"21": {
"inputs": {
"lora_name": "可爱少女厚涂风.safetensors",
"strength_model": 0,
"strength_clip": 1.0000000000000002,
"model": [
"10",
0
],
"clip": [
"10",
1
]
},
"class_type": "LoraLoader",
"_meta": {
"title": "加载LoRA"
}
},
"23": {
"inputs": {
"vae_name": "clearvae_v23.safetensors"
},
"class_type": "VAELoader",
"_meta": {
"title": "加载VAE"
}
},
"27": {
"inputs": {
"text_trans": "",
"model_type": "600M",
"to_lang": "英语",
"speak_and_recognation": {
"__value__": [
false,
true
]
}
},
"class_type": "ZhPromptTranslator",
"_meta": {
"title": "正向提示词"
}
},
"28": {
"inputs": {
"text_trans": "",
"model_type": "600M",
"to_lang": "英语",
"speak_and_recognation": {
"__value__": [
false,
true
]
}
},
"class_type": "ZhPromptTranslator",
"_meta": {
"title": "反向提示词"
}
},
"61": {
"inputs": {
"appName": "彩绘",
"appType": "paint-board",
"appDesc": "",
"styles": ""
},
"class_type": "AppParams",
"_meta": {
"title": "应用参数管理"
}
},
"62": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
},
"64": {
"inputs": {
"image": ""
},
"class_type": "LamLoadImageBase64",
"_meta": {
"title": "Base64图片加载"
}
}
}
+204
View File
@@ -0,0 +1,204 @@
{
"3": {
"inputs": {
"seed": 871010165059929,
"steps": 8,
"cfg": 1,
"sampler_name": "euler_ancestral",
"scheduler": "normal",
"denoise": 0.6500000000000001,
"model": [
"21",
0
],
"positive": [
"6",
0
],
"negative": [
"7",
0
],
"latent_image": [
"19",
0
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"4": {
"inputs": {
"ckpt_name": "yamer_Cartoon_xenoArcadiaCD.safetensors"
},
"class_type": "CheckpointLoaderSimple",
"_meta": {
"title": "Checkpoint加载器(简易)"
}
},
"6": {
"inputs": {
"text": [
"27",
0
],
"clip": [
"21",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"7": {
"inputs": {
"text": [
"28",
0
],
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"8": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"10": {
"inputs": {
"lora_name": "lcm-lora-sdv1-5.safetensors",
"strength_model": 1.0000000000000002,
"strength_clip": 1.0000000000000002,
"model": [
"4",
0
],
"clip": [
"4",
1
]
},
"class_type": "LoraLoader",
"_meta": {
"title": "LoRA加载器"
}
},
"19": {
"inputs": {
"pixels": [
"64",
0
],
"vae": [
"23",
0
]
},
"class_type": "VAEEncode",
"_meta": {
"title": "VAE编码"
}
},
"21": {
"inputs": {
"lora_name": "可爱少女厚涂风.safetensors",
"strength_model": 0,
"strength_clip": 1.0000000000000002,
"model": [
"10",
0
],
"clip": [
"10",
1
]
},
"class_type": "LoraLoader",
"_meta": {
"title": "LoRA加载器"
}
},
"23": {
"inputs": {
"vae_name": "clearvae_v23.safetensors"
},
"class_type": "VAELoader",
"_meta": {
"title": "VAE加载器"
}
},
"27": {
"inputs": {
"text_trans": ""
},
"class_type": "ZhPromptTranslator",
"_meta": {
"title": "正向提示词"
}
},
"28": {
"inputs": {
"text_trans": ""
},
"class_type": "ZhPromptTranslator",
"_meta": {
"title": "反向提示词"
}
},
"61": {
"inputs": {
"appName": "彩绘1",
"appType": "paint-board",
"appDesc": "",
"styles": ""
},
"class_type": "AppParams",
"_meta": {
"title": "应用参数管理"
}
},
"62": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
},
"64": {
"inputs": {
"image": ""
},
"class_type": "LamLoadImageBase64",
"_meta": {
"title": "Base64图片加载"
}
}
}
File diff suppressed because it is too large Load Diff
+662
View File
@@ -0,0 +1,662 @@
{
"10": {
"inputs": {
"provider": "CPU"
},
"class_type": "InstantIDFaceAnalysis",
"_meta": {
"title": "InstantID面部分析"
}
},
"11": {
"inputs": {
"control_net_name": "instantid\\diffusion_pytorch_model.safetensors"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"12": {
"inputs": {
"instantid_file": "ip-adapter.bin"
},
"class_type": "InstantIDModelLoader",
"_meta": {
"title": "InstnatID模型加载器"
}
},
"13": {
"inputs": {
"seed": 367821467240321
},
"class_type": "Seed Everywhere",
"_meta": {
"title": "全局随机种"
}
},
"20": {
"inputs": {
"strength": [
"482",
0
],
"start_percent": 0,
"end_percent": [
"483",
0
],
"positive": [
"44",
1
],
"negative": [
"44",
2
],
"control_net": [
"42",
0
],
"image": [
"31",
0
],
"vae": [
"44",
4
]
},
"class_type": "ControlNetApplyAdvanced",
"_meta": {
"title": "ControlNet应用(旧版高级)"
}
},
"24": {
"inputs": {
"crop_padding_factor": 0.25,
"cascade_xml": "lbpcascade_animeface.xml",
"image": [
"493",
0
]
},
"class_type": "Image Crop Face",
"_meta": {
"title": "面部裁剪"
}
},
"27": {
"inputs": {
"ip_weight": 0.8,
"cn_strength": 0.8,
"start_at": 0,
"end_at": 1,
"noise": 0,
"combine_embeds": "average",
"instantid": [
"12",
0
],
"insightface": [
"10",
0
],
"control_net": [
"11",
0
],
"image": [
"35",
0
],
"model": [
"45",
0
],
"positive": [
"432",
0
],
"negative": [
"432",
1
],
"image_kps": [
"31",
0
]
},
"class_type": "ApplyInstantIDAdvanced",
"_meta": {
"title": "应用InstantID(高级)"
}
},
"28": {
"inputs": {
"crop_padding_factor": 0.25,
"cascade_xml": "lbpcascade_animeface.xml",
"image": [
"494",
0
]
},
"class_type": "Image Crop Face",
"_meta": {
"title": "面部裁剪"
}
},
"29": {
"inputs": {
"crop_padding_factor": 0.25,
"cascade_xml": "lbpcascade_animeface.xml",
"image": [
"495",
0
]
},
"class_type": "Image Crop Face",
"_meta": {
"title": "面部裁剪"
}
},
"30": {
"inputs": {
"noise_mask": true,
"positive": [
"27",
1
],
"negative": [
"27",
2
],
"vae": [
"44",
4
],
"pixels": [
"40",
1
],
"mask": [
"40",
2
]
},
"class_type": "InpaintModelConditioning",
"_meta": {
"title": "内补模型条件"
}
},
"31": {
"inputs": {
"aspect_ratio": "original",
"proportional_width": 1,
"proportional_height": 1,
"fit": "crop",
"method": "lanczos",
"round_to_multiple": "64",
"scale_to_side": "None",
"scale_to_length": 1536,
"background_color": "#000000",
"image": [
"40",
1
]
},
"class_type": "LayerUtility: ImageScaleByAspectRatio V2",
"_meta": {
"title": "按宽高比缩放_V2"
}
},
"32": {
"inputs": {
"amount": 6,
"device": "auto",
"mask": [
"499",
0
]
},
"class_type": "MaskBlur+",
"_meta": {
"title": "遮罩模糊"
}
},
"35": {
"inputs": {
"inputcount": 4,
"Update inputs": null,
"image_1": [
"24",
0
],
"image_2": [
"28",
0
],
"image_3": [
"29",
0
],
"image_4": [
"36",
0
]
},
"class_type": "ImageBatchMulti",
"_meta": {
"title": "图像组合批次(多重)"
}
},
"36": {
"inputs": {
"crop_padding_factor": 0.25,
"cascade_xml": "haarcascade_frontalface_default.xml",
"image": [
"492",
0
]
},
"class_type": "Image Crop Face",
"_meta": {
"title": "面部裁剪"
}
},
"38": {
"inputs": {
"rescale_algorithm": "bislerp",
"stitch": [
"40",
0
],
"inpainted_image": [
"39",
5
]
},
"class_type": "InpaintStitch",
"_meta": {
"title": "局部重绘(接缝)"
}
},
"39": {
"inputs": {
"seed": [
"13",
0
],
"steps": 15,
"cfg": 1,
"sampler_name": "euler_ancestral",
"scheduler": "karras",
"denoise": [
"485",
0
],
"preview_method": "auto",
"vae_decode": "true",
"model": [
"45",
0
],
"positive": [
"30",
0
],
"negative": [
"30",
1
],
"latent_image": [
"30",
2
],
"optional_vae": [
"44",
4
]
},
"class_type": "KSampler (Efficient)",
"_meta": {
"title": "K采样器(效率)"
}
},
"40": {
"inputs": {
"context_expand_pixels": 20,
"context_expand_factor": 1,
"fill_mask_holes": true,
"blur_mask_pixels": 16,
"invert_mask": false,
"blend_pixels": 16,
"rescale_algorithm": "bicubic",
"mode": "forced size",
"force_width": 1280,
"force_height": 1280,
"rescale_factor": 1,
"min_width": 512,
"min_height": 512,
"max_width": 768,
"max_height": 768,
"padding": 32,
"image": [
"496",
0
],
"mask": [
"32",
0
]
},
"class_type": "InpaintCrop",
"_meta": {
"title": "局部重绘(裁剪)"
}
},
"41": {
"inputs": {
"images": [
"38",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存输出图片"
}
},
"42": {
"inputs": {
"control_net_name": "SDXL\\xinsir_controlnet-depth-sdxl-1.0.safetensors"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"43": {
"inputs": {
"preprocessor": "Zoe_DepthAnythingPreprocessor",
"resolution": 1472,
"image": [
"31",
0
]
},
"class_type": "AIO_Preprocessor",
"_meta": {
"title": "Aux集成预处理器"
}
},
"44": {
"inputs": {
"ckpt_name": "sdxl\\XL_DreamShaper XL v2.1 Turbo 闪电_v2.1 Turbo.safetensors",
"vae_name": "sdxl_vae.safetensors",
"clip_skip": -2,
"lora_name": "None",
"lora_model_strength": 0.56,
"lora_clip_strength": 1,
"positive": [
"193",
0
],
"negative": [
"480",
0
],
"token_normalization": "none",
"weight_interpretation": "A1111",
"empty_latent_width": 512,
"empty_latent_height": 512,
"batch_size": 1
},
"class_type": "Efficient Loader",
"_meta": {
"title": "效率加载器"
}
},
"45": {
"inputs": {
"model": [
"44",
0
]
},
"class_type": "DifferentialDiffusion",
"_meta": {
"title": "差异扩散"
}
},
"193": {
"inputs": {
"text": ""
},
"class_type": "TextInput_",
"_meta": {
"title": "文本"
}
},
"432": {
"inputs": {
"strength": [
"482",
0
],
"start_percent": 0,
"end_percent": [
"483",
0
],
"positive": [
"20",
0
],
"negative": [
"20",
1
],
"control_net": [
"433",
0
],
"image": [
"434",
0
],
"vae": [
"44",
4
]
},
"class_type": "ControlNetApplyAdvanced",
"_meta": {
"title": "ControlNet应用(旧版高级)"
}
},
"433": {
"inputs": {
"control_net_name": "SDXL\\mistoLine_rank256.safetensors"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"434": {
"inputs": {
"preprocessor": "AnyLineArtPreprocessor_aux",
"resolution": 1472,
"image": [
"31",
0
]
},
"class_type": "AIO_Preprocessor",
"_meta": {
"title": "Aux集成预处理器"
}
},
"480": {
"inputs": {
"text": "nsfw,ng_deepnegative_vl_75t,badhandv4,(worst quality:2),(1ow quality:2), (normalquality:2),1owes,watermark,monochrose,(3d:1 5),(paintingt1.5)."
},
"class_type": "CR Text",
"_meta": {
"title": "文本"
}
},
"482": {
"inputs": {
"value": 0.5
},
"class_type": "ImpactFloat",
"_meta": {
"title": "浮点"
}
},
"483": {
"inputs": {
"value": 0.5
},
"class_type": "ImpactFloat",
"_meta": {
"title": "浮点"
}
},
"485": {
"inputs": {
"value": 0.6000000000000001
},
"class_type": "ImpactFloat",
"_meta": {
"title": "浮点"
}
},
"492": {
"inputs": {
"image_path": "./ComfyUI/input/example.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "人脸1"
}
},
"493": {
"inputs": {
"image_path": "./ComfyUI/input/example.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "人脸2"
}
},
"494": {
"inputs": {
"image_path": "./ComfyUI/input/example.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "人脸3"
}
},
"495": {
"inputs": {
"image_path": "./ComfyUI/input/example.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "人脸4"
}
},
"496": {
"inputs": {
"image_path": "./ComfyUI/input/example.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "头像"
}
},
"497": {
"inputs": {
"image_path": "./ComfyUI/input/example.png",
"RGBA": "false",
"filename_text_extension": "true"
},
"class_type": "LamLoadPathImage",
"_meta": {
"title": "遮罩"
}
},
"498": {
"inputs": {
"channel": "red",
"image": [
"500",
0
]
},
"class_type": "ImageToMask",
"_meta": {
"title": "图像到遮罩"
}
},
"499": {
"inputs": {
"mask": [
"498",
0
]
},
"class_type": "InvertMask",
"_meta": {
"title": "遮罩反转"
}
},
"500": {
"inputs": {
"upscale_method": "nearest-exact",
"width": [
"501",
0
],
"height": [
"501",
1
],
"crop": "disabled",
"image": [
"497",
0
]
},
"class_type": "ImageScale",
"_meta": {
"title": "图像缩放"
}
},
"501": {
"inputs": {
"image": [
"496",
0
]
},
"class_type": "GetImageSize+",
"_meta": {
"title": "获取图像尺寸"
}
}
}
+118
View File
@@ -0,0 +1,118 @@
{
"3": {
"inputs": {
"seed": 604479721574433,
"steps": 20,
"cfg": 8,
"sampler_name": "euler",
"scheduler": "normal",
"denoise": 1,
"model": [
"4",
0
],
"positive": [
"6",
0
],
"negative": [
"7",
0
],
"latent_image": [
"5",
0
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"4": {
"inputs": {
"ckpt_name": "dreamshaper_8.safetensors"
},
"class_type": "CheckpointLoaderSimple",
"_meta": {
"title": "Checkpoint加载器(简易)"
}
},
"5": {
"inputs": {
"width": 512,
"height": 512,
"batch_size": 1
},
"class_type": "EmptyLatentImage",
"_meta": {
"title": "空Latent"
}
},
"6": {
"inputs": {
"text": "beautiful scenery nature glass bottle landscape, , purple galaxy bottle,",
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"7": {
"inputs": {
"text": "text, watermark",
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"8": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"10": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
},
"11": {
"inputs": {
"appName": "文生图",
"appType": "default",
"appDesc": ""
},
"class_type": "AppParams",
"_meta": {
"title": "应用参数管理"
}
}
}
+208
View File
@@ -0,0 +1,208 @@
{
"1": {
"inputs": {
"ckpt_name": "yamer_Cartoon_xenoArcadiaCD.safetensors"
},
"class_type": "CheckpointLoaderSimple",
"_meta": {
"title": "Checkpoint加载器(简易)"
}
},
"2": {
"inputs": {
"width": [
"27",
0
],
"height": [
"27",
1
],
"batch_size": 1
},
"class_type": "EmptyLatentImage",
"_meta": {
"title": "空Latent图像"
}
},
"3": {
"inputs": {
"text": [
"28",
0
],
"speak_and_recognation": {
"__value__": [
false,
true
]
},
"clip": [
"1",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码"
}
},
"4": {
"inputs": {
"text": "",
"speak_and_recognation": {
"__value__": [
false,
true
]
},
"clip": [
"1",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码"
}
},
"6": {
"inputs": {
"control_net_name": "control_v11p_sd15_scribble.safetensors"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "加载ControlNet模型"
}
},
"7": {
"inputs": {
"strength": 1,
"conditioning": [
"3",
0
],
"control_net": [
"6",
0
],
"image": [
"25",
0
]
},
"class_type": "ControlNetApply",
"_meta": {
"title": "应用ControlNet(旧版)"
}
},
"9": {
"inputs": {
"samples": [
"36",
0
],
"vae": [
"1",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"23": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"9",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
},
"24": {
"inputs": {
"appName": "简笔画",
"appType": "paint-board",
"appDesc": "",
"styles": ""
},
"class_type": "AppParams",
"_meta": {
"title": "应用参数管理"
}
},
"25": {
"inputs": {
"image": ""
},
"class_type": "LamLoadImageBase64",
"_meta": {
"title": "Base64图片加载"
}
},
"27": {
"inputs": {
"image": [
"25",
0
]
},
"class_type": "GetImageSize",
"_meta": {
"title": "获取图像尺寸"
}
},
"28": {
"inputs": {
"text_trans": "",
"model_type": "600M",
"to_lang": "英语",
"speak_and_recognation": {
"__value__": [
false,
true
]
}
},
"class_type": "ZhPromptTranslator",
"_meta": {
"title": "中文或其他翻译"
}
},
"36": {
"inputs": {
"seed": 156680208700286,
"steps": 20,
"cfg": 8,
"sampler_name": "euler",
"scheduler": "normal",
"denoise": 1,
"model": [
"1",
0
],
"positive": [
"7",
0
],
"negative": [
"4",
0
],
"latent_image": [
"2",
0
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
}
}
+205
View File
@@ -0,0 +1,205 @@
{
"3": {
"inputs": {
"seed": 156680208700286,
"steps": 20,
"cfg": 8,
"sampler_name": "euler",
"scheduler": "normal",
"denoise": 1,
"model": [
"4",
0
],
"positive": [
"11",
1
],
"negative": [
"7",
0
],
"latent_image": [
"5",
0
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"4": {
"inputs": {
"ckpt_name": "yamer_Cartoon_xenoArcadiaCD.safetensors"
},
"class_type": "CheckpointLoaderSimple",
"_meta": {
"title": "Checkpoint加载器(简易)"
}
},
"5": {
"inputs": {
"width": 512,
"height": 512,
"batch_size": 1
},
"class_type": "EmptyLatentImage",
"_meta": {
"title": "空Latent"
}
},
"6": {
"inputs": {
"text": "",
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"7": {
"inputs": {
"text": "text, watermark",
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"8": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"10": {
"inputs": {
"image": ""
},
"class_type": "LamLoadImageBase64",
"_meta": {
"title": "Base64图片加载"
}
},
"11": {
"inputs": {
"strength": 1,
"start_percent": 0,
"end_percent": 1,
"positive": [
"6",
0
],
"negative": [
"7",
0
],
"control_net": [
"12",
0
],
"image": [
"17",
0
],
"vae": [
"4",
2
]
},
"class_type": "ControlNetApplyAdvanced",
"_meta": {
"title": "ControlNet应用(旧版高级)"
}
},
"12": {
"inputs": {
"control_net_name": "control_v11p_sd15_lineart.pth"
},
"class_type": "ControlNetLoader",
"_meta": {
"title": "ControlNet加载器"
}
},
"13": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
},
"15": {
"inputs": {
"channel": "red",
"image": [
"10",
0
]
},
"class_type": "ImageToMask",
"_meta": {
"title": "图像到遮罩"
}
},
"16": {
"inputs": {
"mask": [
"15",
0
]
},
"class_type": "InvertMask (segment anything)",
"_meta": {
"title": "反转遮罩"
}
},
"17": {
"inputs": {
"mask": [
"16",
0
]
},
"class_type": "MaskToImage",
"_meta": {
"title": "遮罩到图像"
}
},
"18": {
"inputs": {
"appName": "简笔画2",
"appType": "paint-board",
"appDesc": "",
"styles": ""
},
"class_type": "AppParams",
"_meta": {
"title": "应用参数管理"
}
}
}
+131
View File
@@ -0,0 +1,131 @@
{
"3": {
"inputs": {
"seed": 156680208700286,
"steps": 20,
"cfg": 8,
"sampler_name": "euler",
"scheduler": "normal",
"denoise": 1,
"model": [
"4",
0
],
"positive": [
"6",
0
],
"negative": [
"7",
0
],
"latent_image": [
"5",
0
]
},
"class_type": "KSampler",
"_meta": {
"title": "K采样器"
}
},
"4": {
"inputs": {
"ckpt_name": "dreamshaper_8.safetensors"
},
"class_type": "CheckpointLoaderSimple",
"_meta": {
"title": "Checkpoint加载器(简易)"
}
},
"5": {
"inputs": {
"width": 512,
"height": 512,
"batch_size": 1
},
"class_type": "EmptyLatentImage",
"_meta": {
"title": "空Latent"
}
},
"6": {
"inputs": {
"text": "beautiful scenery nature glass bottle landscape, , purple galaxy bottle,",
"speak_and_recognation": {
"__value__": [
false,
true
]
},
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"7": {
"inputs": {
"text": "text, watermark",
"speak_and_recognation": {
"__value__": [
false,
true
]
},
"clip": [
"4",
1
]
},
"class_type": "CLIPTextEncode",
"_meta": {
"title": "CLIP文本编码器"
}
},
"8": {
"inputs": {
"samples": [
"3",
0
],
"vae": [
"4",
2
]
},
"class_type": "VAEDecode",
"_meta": {
"title": "VAE解码"
}
},
"10": {
"inputs": {
"appName": "调试测试",
"appType": "default",
"appDesc": "",
"styles": ""
},
"class_type": "AppParams",
"_meta": {
"title": "应用参数管理"
}
},
"11": {
"inputs": {
"filename_prefix": "ComfyUI",
"images": [
"8",
0
]
},
"class_type": "SaveImgOutputLam",
"_meta": {
"title": "保存图像应用输出"
}
}
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 23 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.7 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 27 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 17 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 5.7 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 8.0 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 20 KiB

+36
View File
@@ -0,0 +1,36 @@
from huggingface_hub import snapshot_download
import sys
import os
models={
"nllb-200-distilled-1.3B":{
"model_id":"facebook/nllb-200-distilled-1.3B",
"local_dir":"models/nllb-200-distilled-1.3B",
"endpoint":"https://hf-mirror.com",
},
"nllb-200-distilled-600M":{
"model_id":"facebook/nllb-200-distilled-600M",
"local_dir":"models/nllb-200-distilled-600M",
"endpoint":"https://hf-mirror.com",
},
"IndexTTS-2":{
"model_id":"IndexTeam/IndexTTS-2",
"local_dir":"models/IndexTTS-2",
"endpoint":"https://hf-mirror.com",
},
}
# repo_id 模型id
# local_dir 下载地址
# endpoint 镜像地址
# resume_download (中断后)继续下载
args = sys.argv
for model_id,model in models.items():
if len(args)>1 and model_id not in args:
continue
print(f"Downloading {model_id}...")
local_dir = os.path.join(os.path.dirname(__file__), model['local_dir'])
snapshot_download(repo_id=model['model_id'], local_dir=local_dir,
local_dir_use_symlinks=False, revision="main",
endpoint='https://hf-mirror.com',
resume_download=True)
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+663
View File
@@ -0,0 +1,663 @@
{
"last_node_id": 24,
"last_link_id": 44,
"nodes": [
{
"id": 7,
"type": "PreviewImageLam",
"pos": {
"0": 1487,
"1": 591
},
"size": {
"0": 210,
"1": 246
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9,
"label": "images"
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
32
],
"slot_index": 0,
"shape": 3,
"label": "IMAGE"
}
],
"properties": {
"Node name for S&R": "PreviewImageLam"
}
},
{
"id": 9,
"type": "LamCommonPrint",
"pos": {
"0": 970,
"1": 277
},
"size": {
"0": 210,
"1": 200
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "obj",
"type": "*",
"link": 31,
"label": "obj"
}
],
"outputs": [
{
"name": "obj",
"type": "*",
"links": [
37
],
"slot_index": 0,
"shape": 3,
"label": "obj"
}
],
"properties": {
"Node name for S&R": "LamCommonPrint"
},
"widgets_values": [
"13"
]
},
{
"id": 23,
"type": "LamCommonPrint",
"pos": {
"0": 2114,
"1": -55
},
"size": {
"0": 210,
"1": 200
},
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "obj",
"type": "*",
"link": 42,
"label": "obj"
}
],
"outputs": [
{
"name": "obj",
"type": "*",
"links": null,
"shape": 3,
"label": "obj"
}
],
"properties": {
"Node name for S&R": "LamCommonPrint"
},
"widgets_values": [
"13"
]
},
{
"id": 8,
"type": "PreviewImage",
"pos": {
"0": 2124,
"1": 246
},
"size": {
"0": 210,
"1": 246
},
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 36,
"label": "图像"
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
}
},
{
"id": 19,
"type": "DoWhileStart",
"pos": {
"0": 317,
"1": -48
},
"size": {
"0": 210,
"1": 86
},
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "obj",
"type": "*",
"link": 43,
"label": "obj"
}
],
"outputs": [
{
"name": "DOWHILE",
"type": "DOWHILE",
"links": [
30
],
"slot_index": 0,
"shape": 3,
"label": "DOWHILE"
},
{
"name": "循环次数",
"type": "INT",
"links": [
31,
34
],
"slot_index": 1,
"shape": 3,
"label": "循环次数"
},
{
"name": "seed",
"type": "INT",
"links": null,
"shape": 3,
"label": "seed"
},
{
"name": "回传数据",
"type": "*",
"links": [],
"slot_index": 3,
"shape": 3,
"label": "回传数据"
}
],
"properties": {
"Node name for S&R": "DoWhileStart"
}
},
{
"id": 5,
"type": "IfInnerExecute",
"pos": {
"0": 992,
"1": 569
},
"size": {
"0": 210,
"1": 66
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "ANY",
"type": "*",
"link": 4,
"label": "ANY"
},
{
"name": "IF_TRUE",
"type": "*",
"link": 44,
"label": "IF_TRUE"
},
{
"name": "IF_FALSE",
"type": "*",
"link": 6,
"label": "IF_FALSE"
}
],
"outputs": [
{
"name": "?",
"type": "*",
"links": [
9
],
"slot_index": 0,
"shape": 3,
"label": "?"
}
],
"properties": {
"Node name for S&R": "IfInnerExecute"
}
},
{
"id": 4,
"type": "LoadImage",
"pos": {
"0": 494,
"1": 762
},
"size": {
"0": 315,
"1": 314.0000305175781
},
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
6
],
"slot_index": 0,
"shape": 3,
"label": "图像"
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3,
"label": "遮罩"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"25500459117e67af73f186605f8658c.png",
"image"
]
},
{
"id": 24,
"type": "LoadImage",
"pos": {
"0": 460,
"1": 376
},
"size": {
"0": 315,
"1": 314.0000305175781
},
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
44
],
"slot_index": 0,
"shape": 3,
"label": "图像"
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3,
"label": "遮罩"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"3口全家福_.png",
"image"
]
},
{
"id": 6,
"type": "MultiParamFormula",
"pos": {
"0": 1215,
"1": 285
},
"size": {
"0": 400,
"1": 200
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "p0",
"type": "*",
"link": 37,
"label": "p0"
},
{
"name": "p1",
"type": "*",
"link": null,
"label": "p1"
}
],
"outputs": [
{
"name": "p0",
"type": "*",
"links": [
4
],
"slot_index": 0,
"shape": 3,
"label": "p0"
}
],
"properties": {
"Node name for S&R": "MultiParamFormula"
},
"widgets_values": [
"import time \n\ntime.sleep(1)\nresult=p0>0",
"enable"
]
},
{
"id": 22,
"type": "LoadImage",
"pos": {
"0": -91,
"1": -53
},
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
43
],
"slot_index": 0,
"shape": 3,
"label": "图像"
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3,
"label": "遮罩"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"3口全家福_.png",
"image"
]
},
{
"id": 21,
"type": "MultiParamFormula",
"pos": {
"0": 1211,
"1": 28
},
"size": {
"0": 400,
"1": 200
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "p0",
"type": "*",
"link": 34,
"label": "p0"
},
{
"name": "p1",
"type": "*",
"link": null,
"label": "p1"
}
],
"outputs": [
{
"name": "p0",
"type": "*",
"links": [
33
],
"shape": 3,
"label": "p0"
}
],
"properties": {
"Node name for S&R": "MultiParamFormula"
},
"widgets_values": [
"p0<13",
"disable"
]
},
{
"id": 20,
"type": "DoWhileEnd",
"pos": {
"0": 1753,
"1": -56
},
"size": [
245.531041712587,
67.74558389447958
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "start",
"type": "DOWHILE",
"link": 30,
"label": "start"
},
{
"name": "ANY",
"type": "*",
"link": 33,
"label": "ANY"
},
{
"name": "obj",
"type": "*",
"link": 32,
"label": "obj"
}
],
"outputs": [
{
"name": "i",
"type": "INT",
"links": [
42
],
"slot_index": 0,
"shape": 3,
"label": "i"
},
{
"name": "endObj",
"type": "*",
"links": [
36
],
"slot_index": 1,
"shape": 3,
"label": "endObj"
}
],
"properties": {
"Node name for S&R": "DoWhileEnd"
}
}
],
"links": [
[
4,
6,
0,
5,
0,
"*"
],
[
6,
4,
0,
5,
2,
"*"
],
[
9,
5,
0,
7,
0,
"IMAGE"
],
[
30,
19,
0,
20,
0,
"DOWHILE"
],
[
31,
19,
1,
9,
0,
"*"
],
[
32,
7,
0,
20,
2,
"*"
],
[
33,
21,
0,
20,
1,
"*"
],
[
34,
19,
1,
21,
0,
"*"
],
[
36,
20,
1,
8,
0,
"IMAGE"
],
[
37,
9,
0,
6,
0,
"*"
],
[
42,
20,
0,
23,
0,
"*"
],
[
43,
22,
0,
19,
0,
"*"
],
[
44,
24,
0,
5,
1,
"*"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 0.7247295000000012,
"offset": [
-541.0154210313345,
644.3460484928061
]
}
},
"version": 0.4
}
File diff suppressed because one or more lines are too long
+569
View File
@@ -0,0 +1,569 @@
{
"last_node_id": 18,
"last_link_id": 21,
"nodes": [
{
"id": 1,
"type": "LoadImage",
"pos": {
"0": 523.6044921875,
"1": 861.4976196289062
},
"size": {
"0": 315,
"1": 314.0000305175781
},
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
3
],
"slot_index": 0,
"shape": 3,
"label": "图像"
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3,
"label": "遮罩"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"2口人杨.jpg",
"image"
]
},
{
"id": 15,
"type": "LoadImage",
"pos": {
"0": 508,
"1": 488
},
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
17
],
"shape": 3,
"label": "图像"
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3,
"label": "遮罩"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"3口全家福_.png",
"image"
]
},
{
"id": 6,
"type": "LamCommonPrint",
"pos": {
"0": 926,
"1": 360
},
"size": {
"0": 210,
"1": 200
},
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "obj",
"type": "*",
"link": 16,
"label": "obj"
}
],
"outputs": [
{
"name": "obj",
"type": "*",
"links": [
4
],
"slot_index": 0,
"shape": 3,
"label": "obj"
}
],
"properties": {
"Node name for S&R": "LamCommonPrint"
},
"widgets_values": [
"14"
]
},
{
"id": 2,
"type": "IfInnerExecute",
"pos": {
"0": 974,
"1": 769
},
"size": {
"0": 210,
"1": 66
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "ANY",
"type": "*",
"link": 1,
"label": "ANY"
},
{
"name": "IF_TRUE",
"type": "*",
"link": 17,
"label": "IF_TRUE"
},
{
"name": "IF_FALSE",
"type": "*",
"link": 3,
"label": "IF_FALSE"
}
],
"outputs": [
{
"name": "?",
"type": "*",
"links": [
19
],
"slot_index": 0,
"shape": 3,
"label": "?"
}
],
"properties": {
"Node name for S&R": "IfInnerExecute"
}
},
{
"id": 18,
"type": "MultiParamFormula",
"pos": {
"0": 1293,
"1": 705
},
"size": {
"0": 397.3512268066406,
"1": 211.31069946289062
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "p0",
"type": "*",
"link": 19,
"label": "p0"
},
{
"name": "p1",
"type": "*",
"link": 21,
"label": "p1"
}
],
"outputs": [
{
"name": "p0",
"type": "*",
"links": [
20
],
"slot_index": 0,
"shape": 3,
"label": "p0"
}
],
"properties": {
"Node name for S&R": "MultiParamFormula"
},
"widgets_values": [
"import time \ntime.sleep(1)\nresult=p0",
"enable"
]
},
{
"id": 3,
"type": "MultiParamFormula",
"pos": {
"0": 1287,
"1": 369
},
"size": {
"0": 397.3512268066406,
"1": 211.31069946289062
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "p0",
"type": "*",
"link": 4,
"label": "p0"
},
{
"name": "p1",
"type": "*",
"link": null,
"label": "p1"
}
],
"outputs": [
{
"name": "p0",
"type": "*",
"links": [
1,
21
],
"slot_index": 0,
"shape": 3,
"label": "p0"
}
],
"properties": {
"Node name for S&R": "MultiParamFormula"
},
"widgets_values": [
"p0%2==1",
"disable"
]
},
{
"id": 4,
"type": "PreviewImageLam",
"pos": {
"0": 1727,
"1": 713
},
"size": {
"0": 210,
"1": 246
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 20,
"label": "images"
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
18
],
"slot_index": 0,
"shape": 3,
"label": "IMAGE"
}
],
"properties": {
"Node name for S&R": "PreviewImageLam"
}
},
{
"id": 5,
"type": "PreviewImage",
"pos": {
"0": 2042,
"1": 347
},
"size": {
"0": 210,
"1": 246
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 15,
"label": "图像"
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
}
},
{
"id": 14,
"type": "ForInnerEnd",
"pos": {
"0": 1822,
"1": 95
},
"size": [
262.4264844294098,
56.10056046149947
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "start",
"type": "DOWHILE",
"link": 14,
"label": "start"
},
{
"name": "obj",
"type": "*",
"link": 18,
"label": "obj"
}
],
"outputs": [
{
"name": "objs",
"type": "*",
"links": [
15
],
"slot_index": 0,
"shape": 3,
"label": "objs"
},
{
"name": "endObj",
"type": "*",
"links": null,
"shape": 3,
"label": "endObj"
}
],
"properties": {
"Node name for S&R": "ForInnerEnd"
}
},
{
"id": 13,
"type": "ForInnerStart",
"pos": {
"0": 531,
"1": 74
},
"size": {
"0": 315,
"1": 206
},
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "DOWHILE",
"type": "DOWHILE",
"links": [
14
],
"slot_index": 0,
"shape": 3,
"label": "DOWHILE"
},
{
"name": "循环次数",
"type": "INT",
"links": [
16
],
"slot_index": 1,
"shape": 3,
"label": "循环次数"
},
{
"name": "seed",
"type": "INT",
"links": null,
"shape": 3,
"label": "seed"
},
{
"name": "回传数据",
"type": "*",
"links": null,
"shape": 3,
"label": "回传数据"
},
{
"name": "总数",
"type": "INT",
"links": null,
"shape": 3,
"label": "总数"
},
{
"name": "步长",
"type": "INT",
"links": null,
"shape": 3,
"label": "步长"
}
],
"properties": {
"Node name for S&R": "ForInnerStart"
},
"widgets_values": [
15,
1,
0
]
}
],
"links": [
[
1,
3,
0,
2,
0,
"*"
],
[
3,
1,
0,
2,
2,
"*"
],
[
4,
6,
0,
3,
0,
"*"
],
[
14,
13,
0,
14,
0,
"DOWHILE"
],
[
15,
14,
0,
5,
0,
"IMAGE"
],
[
16,
13,
1,
6,
0,
"*"
],
[
17,
15,
0,
2,
1,
"*"
],
[
18,
4,
0,
14,
1,
"*"
],
[
19,
2,
0,
18,
0,
"*"
],
[
20,
18,
0,
4,
0,
"IMAGE"
],
[
21,
3,
0,
18,
1,
"*"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 0.6588450000000017,
"offset": [
-400.48620957336436,
401.7140021170757
]
}
},
"version": 0.4
}
File diff suppressed because one or more lines are too long
Binary file not shown.

After

Width:  |  Height:  |  Size: 14 MiB

File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
Binary file not shown.
Binary file not shown.

Some files were not shown because too many files have changed in this diff Show More