20% em todo GPT 🎉 30% no DeepSeekSaiba mais

Prompts e exemplos do MiniMax Hailuo 3

Os casos oficiais vêm diretamente da ficha de modelo da MiniMax no HuggingFace, com o mesmo pedido que os gerou.

Short Drama12

A Two-Character Scene Driven by Exact Dialogue
Comunidade@cocktailpeanut

A Two-Character Scene Driven by Exact Dialogue

14s · 带音频

cinematic sci-fi suspense video. A realistic spaceship crew in full futuristic uniforms follows a mysterious distress signal and boards a dark, abandoned spacecraft drifting in space. Flashlights cut through darkness as they move cautiously through narrow metallic corridors. Warning lights flicker, metal creaks, and their boots echo on the floor. One crew member says, "Signal's coming from deeper inside." They enter a large chamber filled with hundreds of empty cryogenic pods, all open and abandoned. Another crew member says nervously, "Where did everyone go?" They approach a final closed door at the far end. The captain slowly opens it, and instead of another ship compartment, it reveals a completely normal present-day elementary school classroom in bright daylight. Children sit at desks, backpacks on the floor, colorful posters on the walls, and a teacher at the front turns to the crew and scolds them: "You're late again. Take your seats." The astronauts stand frozen in confusion at the doorway. Realistic tone, smooth camera movement, strong suspense buildup, with the final reveal clear, sudden, and absurd. Include realistic audio: radio static, low ship hum, footsteps, pod machinery ambience, door hiss, then ordinary classroom sounds and the teacher's voice.

Starship Bridge — The Hyperspace Jump
OficialMiniMax

Starship Bridge — The Hyperspace Jump

10s · 768p · t2v · 带音频

integrated_multimodal_description: [Shot 1] Cinematic, medium wide shot, pushing in slowly. In the cavernous, dimly lit bridge of a starship, sleek metallic consoles with glowing amber displays flank a massive, curved observation window. A female captain, in her late 40s with an athletic build and short silver-streaked black hair, stands in the center midground. She wears a structured, high-collared dark navy military tunic with silver chest insignias. Her back is to the camera, silhouetted against the cool, ambient starlight pouring through the thick glass. She stands perfectly still with her hands clasped tightly behind her back. Outside the window, a massive armada of jagged, dark grey dreadnoughts hovers in tight formation against a deep purple space nebula. The fleet's massive rear thrusters begin to glow with an intense, escalating bright blue light. [Shot 2] At 00:04.500, the camera cuts to a close-up of the captain's face and shakes strongly. The brilliant blue-white light from the fleet's gathering energy reflects vividly in her dark eyes. Suddenly, a blinding white flash floods through the window, completely washing out the background as the fleet jumps to hyperspace. The sheer spatial force violently jolts the bridge, causing the captain from Shot 1 to stagger slightly forward, her shoulders tensing as she visibly braces herself against the physical tremors. As the intense white light fades abruptly, leaving only the dim, empty expanse of the purple nebula reflected on her starkly lit skin, her jaw clenches, and she slowly closes her eyes in the newly emptied space. overall_soundscape: A low, resonant hum of the ship's ambient life support systems serves as the baseline, soon drowned out by an audible, escalating, high-pitched electronic whine as the fleet outside charges its hyperdrives. A massive, deafening, bass-heavy boom and sharp crackle erupts during the blinding flash, accompanied by the loud metallic creaking, rattling, and deep thuds of the bridge's bulkheads vibrating under immense physical stress. The intense roaring impact then cuts abruptly back to a hollow, echoing room tone, leaving only the faint, steady hum of the isolated bridge. non_diegetic_music: Cinematic space-opera orchestral score, slow tempo, featuring a solitary, mournful French horn melody over deep, sustained string dissonances that build rapidly in volume and intensity, swelling to a massive orchestral peak before snapping immediately into silence right after the jump.

A Morning Begins from Bed to a Seaside Terrace
Comunidade@ayzalnooor24521

A Morning Begins from Bed to a Seaside Terrace

15s · 带音频

Create a 15-second ultra-realistic cinematic lifestyle vlog video, vertical 9:16, featuring the same young woman throughout the entire video. Preserve her facial identity, facial proportions, hairstyle, skin tone and overall appearance consistently in every shot. She wears the same outfit throughout: fitted white V-neck T-shirt with a small subtle logo, blue denim jeans, natural makeup, long softly wavy brown hair. 0:00–0:01 — Wake-up: Close-up inside a beautiful bright bedroom. The woman is lying comfortably on the bed, slowly wakes up, stretches naturally and opens her eyes. She is NOT filming a vlog yet and does not hold a phone or camera. Soft morning sunlight enters through the curtains. 0:01–0:02 — Gets up: Medium shot. She sits up on the bed, smiles softly, fixes her hair and gets ready to start her morning. Natural, effortless movement. 0:02–0:03 — Walks to window: She walks toward the large glass balcony door/window. Camera follows her naturally from behind/side. 0:03–0:04 — Seaside reveal: She opens the curtains/door and looks outside. Reveal a breathtaking blue ocean, coastal hills, flowers, balcony and beautiful morning sunlight. She smiles happily while taking in the view. 0:04–0:05 — Steps outside: She walks out onto the seaside terrace. Gentle ocean breeze moves her hair naturally. Wide cinematic shot showing the beautiful surroundings. 0:05–0:06 — VLOG START: Only now she starts filming herself in handheld selfie-vlog style. She looks into the camera with a bright natural smile and says: “Good morning!” 0:06–0:07 — Show the view: She turns the camera away from herself and slowly pans across the stunning ocean, coastal mountains, flowers and terrace. Smooth handheld vlog movement. 0:07–0:08 — Back to selfie: Selfie shot. She looks into the camera and happily says: “This place is just perfect!” 0:08–0:09 — Location reveal: Wide cinematic shot of the cozy seaside terrace with wooden table, chairs, plants and flowers overlooking the ocean. 0:09–0:10 — Walk to table: Medium tracking shot as she walks toward the table, enjoying the view. Her hair and T-shirt move gently in the sea breeze. 0:10–0:11 — Sit and relax: She sits at the seaside table, smiling peacefully and enjoying the ocean view. A refreshing orange-colored juice is placed on the table. 0:11–0:12 — Juice close-up: Cinematic close-up of her hand picking up the glass of fresh orange juice. Beautiful ocean bokeh in the background, natural sunlight reflecting through the glass. 0:12–0:13 — Vlog toast: Selfie shot. She raises the juice toward the camera with a cheerful smile and says: “Cheers to good days!” 0:13–0:14 — Happy close-up: Beautiful close-up of her smiling naturally at the camera, ocean and warm sunlight softly blurred behind her. 0:14–0:15 — Ending: Camera moves from her toward the sparkling ocean and peaceful coastal landscape. Warm sunlight, gentle waves and a relaxing cinematic ending. Overall Style Ultra-realistic, cinematic travel vlog, natural handheld camera movement, realistic human motion, smooth transitions, soft morning sunlight, realistic ocean waves, gentle wind in hair and clothes, beautiful coastal atmosphere, premium lifestyle aesthetic, natural expressions, authentic vlog feeling, shallow depth of field, cinematic composition, realistic skin texture, high detail, 4K quality.

Ad Creative18

Turning an Educational Vision into an Experience
Comunidade@umesh_ai

Turning an Educational Vision into an Experience

15s · 带音频

Create a 15-second animated educational video that teaches young children the letters A, B, C, and D. The learning pattern for every letter must be: LETTER → SOUND → OBJECT → PLAYFUL ACTION → OBJECT NAME Target audience: children ages 3 to 6. Visual style: Use adorable rounded 3D characters, soft pastel colors, gentle facial expressions, and simple recognizable objects. Combine this with a premium minimalist technology aesthetic featuring clean white space, elegant composition, soft studio lighting, subtle reflections, smooth gradients, rounded geometry, crisp typography, and extremely polished transitions. The animation should feel playful and child-friendly while remaining calm, uncluttered, and beautifully designed. Use a clean off-white background with a different soft color glow behind each letter. 0:00–0:01 | Introduction A small smiling star mascot bounces into the center of the screen. Colorful letters briefly float around it. Display the text: “Let’s learn!” The mascot taps the screen, creating a soft ripple that reveals the first letter. 0:01–0:04 | A is for Apple Show a large uppercase “A” and smaller lowercase “a” beside it. Use thick, rounded, highly readable typography. The narrator says: “A. A says ah. A is for Apple.” The uppercase A gently inflates and transforms into a shiny red apple. Its top point becomes the apple stem, and a small green leaf unfolds from the side. The apple gains a cute smiling face and performs one soft bounce. Display the word: “APPLE” Highlight the first letter A in red. Add a soft pop and a tiny crunchy sound. 0:04–0:07 | B is for Ball The apple rolls across the screen and leaves behind a curved red trail. The trail loops twice and forms a large uppercase “B,” with a lowercase “b” appearing beside it. The narrator says: “B. B says buh. B is for Ball.” The two rounded sections of the B expand and merge into a colorful striped ball. The ball bounces twice with playful squash-and-stretch animation. Display the word: “BALL” Highlight the first letter B in blue. Synchronize each bounce with a soft musical note. 0:07–0:10 | C is for Cat On its final bounce, the ball stretches into a curved shape and becomes a large uppercase “C.” A lowercase “c” slides gently into place beside it. The narrator says: “C. C says kuh. C is for Cat.” The C rotates and becomes the curled tail of a cute orange cat. The rest of the cat forms from soft rounded shapes. The cat stretches, blinks, and gives one gentle wave with its paw. Display the word: “CAT” Highlight the first letter C in orange. Add a quiet and friendly “meow.” 0:10–0:13 | D is for Duck The cat’s tail uncurls and transforms into the curved side of a large uppercase “D.” A lowercase “d” pops up beside it. The narrator says: “D. D says duh. D is for Duck.” The straight line of the D becomes the duck’s neck. The curved section becomes its round yellow body. A small orange beak and two tiny wings pop into place. The duck waddles forward, flaps its wings, and gives one cheerful quack. Display the word: “DUCK” Highlight the first letter D in yellow. Add tiny water ripples beneath its feet. 0:13–0:15 | Recap The apple, ball, cat, and duck slide into four clean rounded tiles. Place their letters above them: “A B C D” The mascot returns and points to each object as they bounce once in sequence. Narrator: “A, B, C, D. Great job!” Finish with the text: “Great job!” Use a small sparkle animation and a warm musical chime. Animation requirements: Keep each letter fully visible for a moment before it transforms. Show uppercase and lowercase versions clearly. Make every object instantly recognizable. Use smooth shape morphing so children can visually understand how the letter becomes the object. Maintain stable spelling, clean letterforms, accurate object shapes, and consistent character design. Use gentle squash-and-stretch, soft motion blur, subtle shadows, polished lighting, and precisely synchronized sound effects. Avoid fast camera movement, cluttered backgrounds, harsh colors, tiny text, warped letters, random symbols, duplicated objects, scary expressions, or overly complex transformations. The final video should feel cute, educational, memorable, calming, and exceptionally polished.

A Lacquer-Thread Vase from Clay Rolling to Display
Comunidade@derek_wall90176

A Lacquer-Thread Vase from Clay Rolling to Display

15s · 带音频

【参考图规则】 @图片1 是最终漆线雕作品的唯一权威参考。 必须严格保持图片1中的作品造型、主体比例、漆面颜色、金色漆线纹样、纹样走向、器型和整体气质。 不得随意增加龙、凤、佛像、文字或不存在的装饰。 @图片2 是包装盒、包装纸、品牌标识以及产品标签的唯一权威参考。 所有包装结构、颜色和标识均按照图片2执行。 @图片3 是厦门传统漆线雕工作室的环境与工匠服装参考。 保持真实闽南手工作坊质感,不做古装影视化处理,不制造虚假的“古代作坊”。 【目标】 制作一支15秒、9:16竖版的厦门漆线雕非遗工艺短片。 主题: “一根线,走完三百年的手艺。” 影片完整呈现: 备料 → 舂打漆线土 → 搓线 → 盘线成纹 → 安金 → 贴金 → 清理完成 → 包装 → 上架 快速剪辑, 字幕驱动, 即使完全静音观看也能理解整个制作过程。 视觉采用真实高端人文纪录广告摄影: 35毫米胶片电影质感, 细腻颗粒, 浅景深, 大量微距, 自然手持, 运动非常克制, 不做旅游宣传片, 不做华丽国潮特效, 不做博物馆宣传片。 重点永远是: 手, 线, 漆, 金箔, 纹样, 时间。 【金色就是时间】 整支影片必须存在一条不可逆的视觉变化: 画面中的“金色”从无到有,并越来越多。 影片开始: 只有黑、 深褐、 砖灰、 暗红, 几乎看不到金色。 漆线开始盘绕以后, 画面出现第一点暖金。 贴金以后, 金色迅速扩大。 作品完成时, 精细金色纹样覆盖主体。 最后上架时, 整个作品被温暖自然光照亮。 每一个镜头都必须比前一个镜头拥有更多的金色。 这个变化不能倒退。 金色就是这支影片的时间。 【固定字幕牌】 整个影片始终只有一个字幕牌。 位置: 画面下方三分之一, 水平居中, 所有镜头完全相同的位置。 造型: 非常克制的小型圆角矩形, 深朱砂红底, 暖米白文字。 两行文字: 第一行: 较细字体, 显示“工序”。 第二行: 粗体, 显示一句极短的动作描述。 字幕牌: 不移动, 不放大, 不缩小, 不淡入淡出, 不跳动。 只在指定剪辑点瞬间更换文字。 字体必须清楚、正确。 【镜头序列】 0—1.7秒 主镜头。 深夜般昏暗的传统工作台。 一盏暖色工作灯只照亮双手。 桌面上可以看到: 陈年砖粉质感的细粉、 深色大漆材料、 正在被反复捶打揉合的深褐色漆线土。 工匠双手将材料反复捶、压、揉, 逐渐形成柔软、富有韧性的泥团。 周围环境全部沉入黑暗。 没有金色。 字幕: “第一道” “捶土成线” 1.7—3.3秒 极端微距。 一小块漆线土放在传统搓线板之间。 工匠双手稳定向前推动搓板。 原本粗厚的泥条逐渐被搓成长而均匀的细线。 摄影机贴得非常近。 能够清楚看到: 漆线轻微湿润的表面, 细小纹理, 手指压力, 线条被不断拉细。 背景完全虚化。 字幕保持: “第一道” “捶土成线” 3.3—5.0秒 主镜头。 朱红漆面的器物坯体第一次出现。 工匠用极细工具, 将刚刚搓好的柔软漆线轻轻落在器物表面。 第一根漆线贴上去。 只有一根。 然后第二根。 线条开始形成第一个小小的卷云纹。 画面第一次出现非常微弱的暖色高光。 字幕: “第二道” “一线起纹” 5.0—6.8秒 微距插入镜头。 镜头几乎贴着器物表面横向观察。 工匠用竹制细工具推动漆线: 盘, 绕, 结, 堆。 几根不足毫米级视觉尺度的柔软漆线, 逐渐形成具有明显高度差的浮凸纹样。 线条之间非常紧密。 卷云、缠枝、如意形态逐渐出现。 重点展示: 线并不是画上去的, 而是真正一根一根盘出来的。 字幕保持: “第二道” “一线起纹” 6.8—8.6秒 主镜头。 器物已经完成大部分漆线纹样。 摄影机非常缓慢地向前推近。 工匠旋转器物。 光线擦过密集漆线表面。 能够看到: 高低起伏, 盘绕结构, 层层叠叠的线性浮雕。 这时依然没有真正的金箔。 只有漆线自身的暖棕色。 字幕: “第三道” “盘线成雕” 8.6—10.2秒 全片最重要的声音与视觉高潮。 极端微距。 一张极薄金箔悬在空气里轻微颤动。 几乎没有重量。 工匠屏住动作, 用传统工具将金箔缓缓落向已经完成的漆线纹样。 金箔接触表面的瞬间, 贴住。 镜头保持。 不要快速切走。 细小金箔自然贴合漆线的高低起伏, 金色第一次大面积出现。 字幕: “第四道” “一片金落下” 10.2—11.7秒 微距。 柔软毛刷轻轻扫过作品表面。 多余金箔碎屑被一点点扫开。 随着刷毛经过, 清晰、 锐利、 细密的金色漆线纹样从杂乱金箔中显现。 这是第二个视觉高潮。 镜头必须清楚表现: 杂乱金箔 变成 精确金色纹样。 字幕: “第五道” “金纹醒来” 11.7—12.8秒 主镜头。 完成后的作品放在深色木质工作台上。 工匠用双手缓慢旋转检查。 没有任何戏剧表演。 只检查: 纹样, 漆面, 金箔, 边缘, 细节。 自然光已经明显进入工作室。 金色纹样被晨光照亮。 字幕: “第六道” “最后一眼” 12.8—13.8秒 插入镜头。 作品被轻轻放入定制包装盒。 暖米色软布包裹器物。 盒盖缓慢合上。 手将包装盒推向画面前方。 纸张、 布料、 木盒均保持真实触感。 不要奢侈品浮夸包装。 字幕: “第七道” “入盒” 13.8—15秒 最终镜头。 切到极简现代非遗精品店。 温暖自然晨光。 深色木制陈列架。 一只手将完成的漆线雕作品从画面侧面轻轻放到展示台中央。 手退出画面。 作品保持完全静止。 金色漆线在自然光中形成极精细的浮雕反光。 背景虚化。 最后0.6秒不再运动。 字幕: “第八道” “上架” 画面中央仅允许出现一句极小的中文: “厦门漆线雕” 下面更小: “一线成雕” 最终画面必须像高端工艺品牌的静态产品摄影。 绝对不要继续运动。 【色彩】 整支影片限制在以下颜色: 陈年砖粉灰, 生漆深褐, 木质黑褐, 朱砂红, 暖米白, 金箔金。 影片前半段: 黑褐色占主导。 中段: 朱红逐渐出现。 后半段: 金色逐渐占据视觉中心。 禁止: 青蓝科技色, 霓虹, 紫色, 赛博朋克, 高饱和国潮配色。 金色必须来自真实金箔和真实光线。 不能使用发光特效制造金色。 【摄影】 全片使用真实人文纪录摄影语言。 大量: 100毫米微距镜头, 近距离手部特写, 极浅景深, 35毫米胶片颗粒。 摄影机运动非常少。 允许: 轻微呼吸感手持, 一次非常缓慢的轴向推进, 极细微的跟随动作。 禁止: 无人机, 环绕运镜, 快速推拉, 旋转镜头, 甩镜, 变焦冲击, 慢动作, 速度渐变, 镜头炫技。 镜头永远尊重手艺本身。 【声音】 全片没有任何音乐。 没有旁白。 没有人物对白。 没有广告配音。 只使用现场真实声音。 0—3.3秒: 凌晨安静工作室。 远处非常轻的环境底噪。 漆线土落在木板上的闷响。 手掌揉压材料的湿润摩擦声。 搓板来回运动产生细密摩擦声。 3.3—8.6秒: 整体声音进一步降低。 细工具轻碰器物表面的声音。 手指移动。 漆线被放下时极轻微的黏连声音。 器物缓慢旋转时木架产生细小摩擦声。 8.6—10.2秒: 这是整支影片最重要的声音时刻。 其他环境声音全部压低。 只留下: 金箔极薄的纸张颤动声, 轻微呼吸, 工具触碰金箔的细小声音。 金箔落到漆线表面的瞬间, 接近安静。 让观众感觉自己离这张金箔只有几厘米。 10.2—12.8秒: 软毛刷扫过金箔。 细碎金箔摩擦。 器物在木架上轻轻旋转。 12.8—15秒: 包装纸折叠。 布料摩擦。 盒盖轻轻合上。 随后进入店铺: 作品底座接触木质展台, “嗒”。 最后接近完全安静。 只保留非常微弱的室内环境声。 【核心视觉原则】 不要把它拍成: “传统文化宣传片”。 要把它拍成: “世界顶级手工艺品牌新品诞生纪录”。 传统来自材料和工艺, 高级感来自摄影、节奏和克制。 【禁止项】 不得出现任何与产品无关的文字。 不得出现旅游宣传标语。 不得出现: “匠心传承” “非遗之美” “千年文化” “东方美学” 等泛化宣传口号。 不得出现不存在的历史年代。 不得出现虚假的皇宫或寺庙场景。 不得增加和尚、古装人物、舞狮、灯笼等所谓中国元素。 不得把漆线雕表现成木雕、石雕或刀刻。 重点必须明确表现: 搓线, 放线, 盘线, 绕线, 堆线, 贴金。 不得用三维特效让纹样自动生长。 不得让漆线自行移动。 所有纹样必须由人的手完成。 不得使用魔法粒子。 不得让金箔发光。 不得使用人工光晕。 不得使用慢动作。 不得加入速度渐变。 不得加入转场音效。 不得使用呼啸声。 不得出现手机界面。 不得出现电商界面。 不得出现价格。 不得出现二维码。 不得出现网站。 不得出现社交媒体账号。 不得出现水印。 不得增加第十个镜头。 最终画面完全静止。

A Woven Handbag on the Riviera
Comunidade@Dustfinger2077

A Woven Handbag on the Riviera

15s · ref2va · 带音频

Create a 15 second, 16:9 luxury fashion commercial using @Image1@Image2 as exact visual references. @Image1 is the sole product reference. Preserve the bag’s exact proportions, pale grey white woven textile, irregular rectangular stitching, soft structure, curved flap, silver clasp, four black eyelets, side construction, zipper details, and hardware. Do not reinterpret it as leather. The bag has two separate short straps made from silver chain woven with matching textile. Preserve their exact construction, attachment points, resting positions, and maximum lengths. The front strap connects only the two front eyelets. The rear strap connects only the two rear eyelets. They must never cross, merge, disappear, duplicate, lengthen, or be rerouted. @Image2 is the exact character reference. Preserve the woman’s face, green eyes, straight blonde bob, natural proportions, and understated makeup. Dress her consistently in a sophisticated beige silk summer mini dress with a square neckline, narrow shoulder straps, fitted waist, and softly moving skirt. Add ivory slingback heels and small pearl earrings. The bag is the visual anchor throughout. It must always be the largest, brightest, sharpest, or most prominently framed element. The woman, hotel, landscape, and car are supporting elements only. 0 to 1.75 seconds: Tight product shot of the bag alone on a sunlit limestone console beside a softly shimmering Riviera pool. Directional sunlight reveals the textile weave, stitching, two straps, four eyelets, side details, and clasp. 1.75 to 3.25 seconds: Keep the bag large and tack sharp in the foreground while the woman approaches through the bright hotel suite in soft focus. She has not touched it yet. 3.25 to 5 seconds: Extreme close up of her fingers lifting both straps together. Keep the front and rear straps visibly separate, correctly routed, and uncrossed. Show realistic chain tension, textile threaded through the links, and accurate hands. 5 to 7 seconds: She places both short straps together over one shoulder and wears the bag as a side shoulder bag. It hangs close beneath her arm, high against the side of her upper waist. It must never reach her hip or thigh. Both straps remain visibly separate, uncrossed, and at their true maximum length. 7 to 9 seconds: Close tracking shot through a sunlit hotel colonnade, composed mainly around the bag moving naturally against the beige silk dress. The straps respond realistically to her steps, shoulder movement, tension, and gravity. 9 to 11 seconds: The bag stands upright and tack sharp on a pale stone terrace table. Its two straps settle naturally and remain identifiable. The woman sits behind it in soft focus, looking toward the sea. 11 to 13 seconds: She carries the bag on one shoulder toward a cream vintage convertible. In a tight product focused montage, she opens the driver door, removes both straps from her shoulder, and carefully places the bag upright on the front passenger seat cushion. The bag’s entire base must rest directly on the horizontal cushion where a passenger sits. It is not on the backrest, headrest, seat edge, center console, or top of the seat. The passenger backrest rises behind it. Both straps settle naturally beside and partly behind the bag without disappearing. She enters the driver seat. Keep the car cropped and secondary. 13 to 15 seconds: The car moves along the Riviera coast. Keep the bag upright, large, and tack sharp on the passenger seat cushion while the woman holds steering wheel and drives in the soft background. The sea and road pass through the windows in gentle motion blur. Moving sunlight reveals the woven textile, stitching, eyelets, two distinct straps, and hardware. End with one controlled flash across the silver clasp. Use bright, fashion forward European luxury cinematography inspired by a restored 1990s campaign film. Directional Mediterranean sunlight, dimensional shadows, layered reflections, warm 35 mm grain, gentle haliation, creamy highlights, subtle film weave, natural motion blur, and tactile textile detail. Use elegant editorial cuts, realistic movement, accurate hands, and consistent spatial logic. Add a sophisticated original fashion score with crisp percussion, warm bass, airy electronic textures, subtle chain sounds, clasp click, silk movement, footsteps, coastal ambience, car door movement, and quiet engine sound. No dialogue. No product redesign, leather texture, colour change, altered stitching, missing or merged straps, crossed straps, extra straps, extended straps, changed eyelet routing, crossbody carry, low hip placement, warped chains, changed hardware, duplicated bag, distorted hands, outfit changes, blue dress, car hero shot, wide vehicle profile, empty road shot, incorrect seat placement, text, or subtitles.

Keyframe Animation9

Sword Formations Above a Moonlit Immortal City
Comunidade@Stellakjbk

Sword Formations Above a Moonlit Immortal City

15s · fl2va · 带音频

5秒,4:3,顶级东方仙侠3A游戏CG,电影级高端宣传片,严格参考图1的月下仙城场景与美术气质:巨大明月、云海天宫、环形仙城、悬浮仙山、瀑布、白玉仙桥、金顶楼阁、冷蓝月光。整体画面必须唯美、恢弘、神圣、危险、酷炫,像大厂正式游戏开场CG,不像普通AI视频。 双主角设定:男剑仙,黑色长发,白蓝战袍,冰蓝雷光长剑,英俊冷峻,强大沉稳;女剑仙,黑色长发,白青仙裙,月白灵剑与月华法术,清冷绝美,高贵飘逸。反派为黑金重甲魔修,手持黑色巨剑,周身暗红黑色魔气,压迫感极强。人物脸和服装全程稳定统一,保持高级游戏角色建模质感。 开场镜头从巨大明月前高速穿越云海与浮空仙山,俯冲进入宏伟仙城,来到白玉高台,男女主背对镜头立于高台边缘,衣袍长发飞扬,英雄感极强。远处天空突然裂开暗红色空间裂缝,魔修高速坠落,轰然落地,引发碎石、冲击波和魔气爆发。 男剑仙瞬间拔剑冲出,与魔修在仙桥中央高速对撞,剑刃碰撞爆发蓝红双色能量冲击、空间波纹、火星与剑气。近战动作凌厉、清晰、有重量感,包含高速斩击、格挡、反击、短暂慢动作碰撞特写,镜头低机位跟拍与高速环绕,极具电影感。 随后女剑仙御剑升空,展开巨大蓝白仙阵,发动“万剑齐发”,数百柄发光飞剑如流星暴雨般射向魔修,夜空被剑光点亮。魔修爆发黑红魔气,凝聚巨大的魔龙反击,空中连续爆炸。男剑仙御剑穿梭爆炸火光与云海,一剑斩裂魔龙,空战场面华丽震撼。 最后魔修释放百米高的黑红魔神法相,遮蔽天空。男女主并肩悬空,男剑仙开启覆盖半个天空的金蓝剑阵,女剑仙召唤巨大的白蓝月华莲华法相,二人联手凝聚数百米长的金蓝神剑,与魔神巨剑正面碰撞。碰撞瞬间进入超级慢动作,蓝色雷电、金色符文、白色月华、红色魔焰与空间裂痕同时凝固,随后爆发毁天灭地的大爆炸,形成巨大的环形云洞与能量冲击波。 结尾爆炸光芒散去,巨大明月重现,男女主并肩立于悬空巨剑之上俯瞰整座仙城,周围灵剑环绕,远处天宫亮起金蓝光柱,镜头缓慢拉远,形成极强的高级游戏主视觉结尾,并预留后期添加Logo空间。 镜头要求:超广角航拍、俯冲推进、低机位英雄视角、高速跟拍、环绕镜头、御剑追踪、短暂升格慢动作、终极碰撞震动、史诗级拉远收尾。镜头必须流畅稳定,不乱晃,不跳轴。 视觉要求:特效炫丽但高级干净,不能廉价光污染。强调冰蓝雷电剑气、金色仙纹、月华粒子、飞剑阵、空间裂缝、黑红魔焰、魔龙、巨大法相、神剑、冲击波、体积云、瀑布水雾、发光碎片。人物主体始终清晰,画面精致通透,PBR材质,真实布料与长发动效,电影级光影与空间层次。 音效:原生立体声。包含高空风声、远处钟鸣、低沉鼓点、剑刃碰撞、破风声、雷电炸裂、万剑掠空、魔龙低吼、终极爆炸低频冲击。无对白,无字幕,无水印,无乱码。

Ramen Dinner — Rack Focus to Family
OficialMiniMax

Ramen Dinner — Rack Focus to Family

8s · 768p · fl2va(仅首帧,未提供尾帧) · 带音频

Primeiro fotogramaRamen Dinner — Rack Focus to Family · Primeiro fotograma · picture1 ramen family dinner keyframe

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] This is a live-action, cinematic shot with a shallow depth of field. The camera holds a perfectly static shot throughout the entire eight-second duration, capturing a cozy family gathering in a traditional Japanese dining room. The scene opens with a large, intricately patterned blue and white ceramic bowl of ramen in the immediate foreground, rendered in crisp, sharp focus. The bowl sits on a smooth, polished long wooden table. Inside the bowl, a rich, oily golden-brown broth surrounds yellow wavy noodles, topped with two thick, round slices of chashu pork featuring visible fat marbling and a distinct spiral meat pattern. A generous mound of freshly chopped, bright green scallions rests in the center, and a crisp, dark green rectangular sheet of nori seaweed is tucked into the right edge. To the left of the bowl, a pair of light brown wooden chopsticks rests horizontally on a small, dark rectangular chopstick rest, near a small cylindrical ceramic teacup with blue painted patterns. On the right side of the table, a spherical paper lantern with a ribbed bamboo frame sits on a black wooden base. In the background, a large family of seven is gathered around the table, initially appearing as a soft, blurred presence. Behind them, traditional Japanese sliding shoji screens with wooden lattice frames are open, revealing a bright outdoor scene with lush green trees. Early in the clip, the thick, white steam rising from the hot ramen broth immediately intensifies, billowing upwards in thick, swirling clouds that dance continuously above the bowl. As the clip progresses into the middle seconds, the camera maintains its static position while the focus begins a deliberate, smooth shift deeper into the room. The foreground ramen bowl, its vibrant ingredients, and the rising steam gradually soften into a hazy, out-of-focus blur. Simultaneously, the family members in the background come into sharp, detailed clarity. The heavy steam continues to rise from the foreground, creating a dynamic, translucent veil between the camera and the family. With the focus now firmly locked on the background, the vibrant family dinner comes alive. The man in the dark navy blue long-sleeved shirt on the left leans forward, his mouth moving animatedly in a silent exchange. The young girl in the crisp white short-sleeved t-shirt beside him smiles brightly, looking toward the center of the table. The woman on the far left, wearing a soft light blue long-sleeved blouse, turns her head slightly, smiling gently. Across the table, the woman in the light grey button-down shirt smiles broadly, her eyes crinkling, as she rests her hands near her plate. The woman in the dark grey top further back uses her wooden chopsticks to pick up a small piece of food from a central ceramic dish filled with bright red pickled vegetables. The woman in the center back in the light grey sweater smiles gently, her hands clasped softly in front of her, observing the interaction. Throughout the remainder of the clip, the family continues their lively physical interaction, their mouths moving in continuous, silent cadences of conversation, while the thick, white steam from the blurred ramen bowl in the foreground never stops rising, adding a comforting atmosphere to the warm gathering. overall_soundscape: The soundscape begins with a quiet room tone mixed with the faint, airy rustle of the thick steam billowing from the hot ramen bowl in the foreground, accompanied by the subtle, continuous hissing and bubbling of the rich broth. As the visual focus shifts deeper into the room, the physical sounds of the bustling family dinner become dominant in the foreground. The clear, sharp clinking of ceramic bowls and wooden chopsticks touching plates is clearly heard as the family members reach for food. This is followed by the faint, muffled thud of a cup being set down on the smooth wooden table, and the subtle, rhythmic rustle of cotton and wool clothing as the family members lean forward and gesture, perfectly capturing the lively, physical atmosphere of the shared meal. non_diegetic_music: A gentle, heartwarming acoustic guitar melody plays softly in the background, accompanied by the subtle, resonant notes of a traditional Japanese koto. The music maintains a slow, comforting tempo that enhances the cozy, nostalgic, and joyful atmosphere of the family gathering.

Neon Mountain Hover-Bike Race
Comunidade@yourPlugAI

Neon Mountain Hover-Bike Race

15s · fl2va · 带音频

Cinematic Anime Video Scene Generate a 15-second horizontal 16:9 original high-speed hover-bike racing anime video from the provided first frame. CRITICAL ENTITY LOCK: There must be exactly 2 racers and 2 bikes in the entire video: RENJI on VALKYRIE-01 (cyan/black drift bike) and ELENA on AERO-X (crimson/white draft bike). Do not add extra racers, drone support vehicles, spectators, or traffic. Maintain total visual consistency for both bikes, helmet visors, suit patterns, repulsor spark colors, and bike liveries throughout the sequence. Entity identity: VALKYRIE-01: Matte-black and cyan angular hover-bike, exposed repulsor pads, lateral drift brakes, blue plasma exhaust trails, ridden by Renji (cyan trim suit). AERO-X: Pearl-white and neon-crimson aerodynamic hover-bike, enclosed canopy, crimson energy draft aura, white-hot central booster, ridden by Elena (crimson/gold visor suit). Video style: High-budget modern sports anime, sakuga-level velocity animation, crisp line art, vibrant neon lighting contrast, high-speed camera tracking, hyper-realistic friction and energy particle effects. Set on a wet downhill mountain pass at dawn. Camera and pacing: Continuous forward velocity, zero slow-motion interruptions: 0.0s - 3.0s: High-speed rear-tracking shot diving into the first downhill hairpin curve; instant drift initiation. 3.0s - 7.5s: Tight side-parallel tracking shot as bikes navigate rock debris and trade positions through S-curves. 7.5s - 11.5s: Close camera lock on the draft-slingshot maneuver; high-energy particle displacement as booster ignition occurs. 11.5s - 15.0s: Low-angle front-facing camera lock on the final sprint to the finish line bridge, ending on a hyper-speed photo-finish freeze. Action timing: 0.0s - 1.5s: Sequence begins at speed. VALKYRIE-01 leads downhill; AERO-X locks onto its rear bumper. Anti-gravity repulsors spray road water and blue sparks into the frame. 1.5s - 4.0s: First sharp hairpin. VALKYRIE-01 deploys lateral drift airbrakes with a burst of blue thruster fire, sliding sideways at 300 km/h. AERO-X stays glued inside its slipstream aura. 4.0s - 7.0s: Mountain debris hazard. VALKYRIE-01 hops over a boulder using a repulsor burst. AERO-X ducks under it, scraping the neon magenta guardrail in a cloud of friction sparks. 7.0s - 10.0s: S-Curve exchange. Bikes lean side-by-side; their repulsor fields collide, creating a bright electrical shockwave. ELENA pulls the overdrive lever; AERO-X's rear fins extend. 10.0s - 13.0s: Slingshot maneuver. AERO-X bursts out of VALKYRIE-01's draft, igniting its central white plasma booster. Both bikes roar down the final straightaway side-by-side. 13.0s - 15.0s: Final sprint toward the finish light gate. Water sprays violently behind them. Both nose cones cross the finish line simultaneously in a flash of light. Final freeze frame. Motion quality: Fluid 2D animation, extreme speed-line integration, stable bike geometry, flawless vehicle reflection rendering, zero limb or body clipping, high-frame-rate kinetic realism. Environment: Wet mountain pass asphalt, sheer cliff walls, neon cyan and magenta guardrail lights, early dawn sky with pink/purple clouds, water spray, floating spark particles. Final output: 15 seconds, horizontal 16:9, original high-budget sports racing anime, exactly 2 racers, relentless kinetic pacing, dynamic cinematography, no subtitles, no watermarks, no logos.

Video Extend & Edit8

A Five-Second Greeting from an Anime Idol
Comunidade@riddi0908

A Five-Second Greeting from an Anime Idol

5s · ref2va · 带音频

プロンプト置いときます✋ Japanese anime style, high-end TV anime production quality, soft pastel lighting, clean line art, subtle cel shading, bright and cheerful idol stage atmosphere, glowing lens flare, gentle bokeh in background, smooth vibrant 60fps animation feel. Scene overview: An adorable anime idol girl stands center stage, looking directly at the camera with a brilliant smile. She holds a colorful microphone, energetic and joyful, giving a cute greeting to her fans. A short, charming 5-second anime idol greeting scene. Storyboard: [0s-1.5s] Shot 1: Medium close-up: The idol girl steps forward with a big smile, bringing the microphone to her mouth, blinking playfully as stage lights twinkle behind her. [1.5s-3.5s] Shot 2: Close-up on her face: She speaks clearly into the microphone with precise lip-sync: "みんな、こんにちは!今日も一緒に楽しもうね!" (Minna, konnichiwa! Kyou mo issho ni tanoshimou ne!). She winks on "tanoshimou ne!". [3.5s-5s] Shot 3: Slightly pulled-back medium shot: She strikes a cute victory-pose/heart-pose with her free hand, tilting her head with a beaming smile as sparkle particles drift around her. Camera: Smooth pan-in at the start, staying centered and static during the lip-sync for maximum clarity, subtle handheld bounce on the final pose at 3.5s. Audio: Crisp female anime idol voice clearly pronouncing "みんな、こんにちは!今日も一緒に楽しもうね!" in Japanese. Underneath, an upbeat synth-pop idol BGM playing, accompanied by a soft microphone "pop" sound at 0.5s and a sparkling magic SFX hit at 3.5s during the pose. No real-life live action rendering, no realistic skin texture, no 3D CG look, no English subtitles or on-screen text clutter, keep the pure 2D anime aesthetic.

Lip-Synced Redub — Editing an Existing Clip
OficialMiniMax

Lip-Synced Redub — Editing an Existing Clip

5s · 768p · ref2va(视频编辑 + 语音音色参考) · 带音频

Vídeo de entradaLip-Synced Redub — Editing an Existing Clip · Vídeo de entrada · video1 source clip pink suit lamb

subject_definitions: <Subject 1> is the young man with short wavy blonde hair, wearing a bright pink suit jacket, matching pink trousers, an unbuttoned white shirt, and silver rings, holding a small black lamb in his arms in <Video 1>. <Video 1> is the source video for the editing task. <Audio 1> is the synchronized audio track of <Video 1>, providing the background music. <Audio 2> is the voice timbre reference for <Subject 1>'s voice, containing a spoken male voiceover. summary: [video editing + audio reference + audio reuse] The target video is an edited version of <Video 1>. <Subject 1>, wearing a bright pink suit and holding a black lamb, stands in a grassy field with other white lambs in the background. The edit animates <Subject 1>'s face to speak the user-provided dialogue. <Audio 1> is partially reused as the continuous background music, while the target references the calm male voice timbre of <Audio 2> for <Subject 1>'s spoken lines. retention_analysis: <Subject 1> (appears in [Shot 1]): fully_preserved - the man retains his identity, wavy blonde hair, pink suit, white shirt, accessories, and the black lamb he holds, with his mouth newly animated to speak. <Video 1> (source video editing): fully_preserved - the original camera framing, warm golden hour lighting, grassy hill setting, and background white lambs are maintained while the central character is edited. <Audio 1>: partially_copy - the atmospheric background music from <Audio 1> is reused in the target video, mixed beneath the newly added spoken dialogue. <Audio 2>: reference - the target audio references the male voice timbre from <Audio 2> to generate <Subject 1>'s spoken dialogue. detailed_description: The target video is in realistic photographic style. [Shot 1] The shot begins from the source <Video 1>, showing <Subject 1>, a young man with short wavy blonde hair, wearing a bright pink suit jacket, matching pink trousers, and a casually unbuttoned white shirt. He stands confidently in a sunlit green pasture, gently holding a small black lamb securely in his arms. The warm, golden hour lighting casts soft shadows across his face and the bright pink fabric of his suit. Behind him, several white lambs stand and graze on the rolling grassy hill against a clear, pale blue sky. The atmospheric background music from <Audio 1> plays continuously throughout the scene. <Subject 1> physically speaks, his mouth movements naturally syncing to the new dialogue, with his voice timbre referencing the calm male delivery from <Audio 2>. Looking thoughtfully forward, <Subject 1> (S1) speaks softly, <d>[English] Follow the wind, live free.</d> As he delivers the line, he subtly shifts his weight, cradling the resting black lamb while the camera slowly pushes in. <Subject 1> (S1) continues his thought, <d>[English] Leave worries behind, enjoy the moment.</d> Exactly as his voice stops, his lips meet in a relaxed, peaceful smile, and his jaw ceases speaking motion. He then turns his gaze slightly away toward the horizon, gently stroking the black lamb's fleece with his fingers as the camera holds on this tranquil, sunlit state through the end of the video. overall_soundscape: The soundscape consists of the continuous, atmospheric background music from <Audio 1>, overlaid with the clear, calm male dialogue spoken by the main character, referencing the voice timbre of <Audio 2>. non_diegetic_music: The atmospheric, sustained background music from <Audio 1> is reused as the continuous score, playing quietly beneath the spoken dialogue.

Adding Japanese Type to a Coastal Cycling Video
Comunidade@MireilleDartois

Adding Japanese Type to a Coastal Cycling Video

ref2va · 带音频

Video editing, add texts into the <Video 1> while keep the layout, expressions, movement, color in <Video 1>. The bold white sans-serif word "この世界は" and "アシンメトリック" snaps upward individually behind the girl driving bicycle and in front of background in sync with the timing of the <Audio 1> audio track titled <d>[Japanese] この世界はアシンメトリック</d>.

Music Video14

Orange-and-Black Street Dance Commercial
Comunidade@tanabe_fragm

Orange-and-Black Street Dance Commercial

15s · t2va · 带音频

Multiplication Tables with Fish Pairs
Comunidade@tanabe_fragm

Multiplication Tables with Fish Pairs

15s · t2va · 带音频

Silver Earbuds in the City at Night
Comunidade@ayzalnooor24521

Silver Earbuds in the City at Night

15s · t2va · 带音频

Create a premium 15-second cinematic ad for Audionic Trance Airbud 850. Start with a close-up of the silver earbuds and case, then show a stylish Korean girl wearing them and walking through a luxurious Korean shopping street while enjoying a modern Korean K-pop song. Use beautiful city lights, soft bokeh, elegant fashion, natural expressions and smooth cinematic camera movement. End with one clean hero shot of the silver earbuds and case only. No phone screen, Bluetooth connection, discount text, logo end card or extra ending scene.

Dialogue & Lip-Sync12

Tasting a Burger from a First-Person View
Comunidade@oggii_0

Tasting a Burger from a First-Person View

15s · 带音频

Duration: 15 seconds | Aspect Ratio: 16:9 | Style: Authentic UGC / iPhone selfie-vlog, handheld, natural light, TikTok/Reels aesthetic. Product Reference: Use the uploaded gourmet burger image as the only product reference. Preserve the bun shape, patty thickness, cheese melt, lettuce, tomato, sauces, and proportions exactly in every shot. Character Description Name: Hana A young Japanese woman in her early 20s with natural beauty, long dark hair in a loose ponytail, oversized cream sweatshirt, minimal makeup, bright smile, friendly lifestyle-vlogger personality. Shot Breakdown SHOT 1 (0–2s) — Selfie showing the burger box. Dialogue: "Burger night!" SHOT 2 (2–4s) — Opens the box. SHOT 3 (4–6s) — Quick zoom on the burger. SHOT 4 (6–8s) — Hands lifting the burger with cheese stretching naturally. SHOT 5 (8–10s) — Bite reaction. Dialogue: "Okay... that's incredible." SHOT 6 (10–12s) — Casual close-up b-roll while reaching for fries. SHOT 7 (12–14s) — Toasting the burger toward the camera. Dialogue: "You need this." SHOT 8 (14–15s) — Freeze frame with overlay: "burger cravings = solved 🍔" Look & Feel Warm apartment lighting, genuine phone footage, slight grain, natural autofocus breathing, handheld imperfections, fast jump cuts. Negative Prompt cinematic grading, commercial production, CGI burger, fake cheese, distorted hands, warped food, perfect stabilization, studio lighting, text glitches, logo distortion.

Punchline Timing in an AI Stand-up Special
Comunidade@azed_ai

Punchline Timing in an AI Stand-up Special

15s · t2va · 带音频

Style: Live stand-up comedy special, intimate comedy club, professional multi-camera production, warm stage lighting, packed audience around small tables, sharp HD broadcast look, natural facial expressions, authentic comedic timing, clean microphone audio, realistic audience reactions, subtle handheld audience camera, 15-second video, 5 cinematic cuts 0–3s: [Wide Establishing → Medium Push-In] A comedian stands center stage holding a microphone as the audience settles. The comedian smiles and says: Comedian: “I asked AI to organize my life yesterday” Brief beat 3–6s: [Medium Close-Up] The comedian maintains a completely serious expression Comedian: “It looked at my schedule and said, ‘Actually... I’m just a language model.’” Audience immediately laughs 6–9s: [Side Angle + Audience Reaction] The comedian waits for the laughter, then slowly nods Comedian: “Even artificial intelligence has boundaries” Quick cut to the front row laughing and clapping 9–12s: [Tight Close-Up] The comedian leans toward the microphone Comedian: “My mom thinks AI is listening to everything we say” Beat Comedian: “Mom... nobody wants that podcast” Audience erupts into louder laughter 12–15s: [Medium → Wide Ending] The comedian waits for silence, then delivers the final line: Comedian: “AI isn't replacing us. It saw our search history and declined the position” Big audience laugh. The comedian smiles and lowers the microphone as the camera pulls out to reveal the applauding club Production details: Keep the same comedian, outfit, microphone, stage and audience throughout. Prioritize precise comedic timing, short pauses before punchlines, believable facial expressions and natural audience reactions. Never cut during a punchline. Audience laughter starts only after the punchline lands. No canned laughter, no overlapping dialogue, coherent eyelines, realistic club acoustics, professional live comedy special editing

A High-Stakes Acquisition in the Boardroom
Comunidade@adithatipalli

A High-Stakes Acquisition in the Boardroom

15s · t2va · 带音频

cinematic 15-second ultra-realistic sequence inside a luxurious modern boardroom during daytime. Soft natural light enters through large glass windows. Four powerful industrialists in expensive tailored suits sit around a long polished dark wooden table. The atmosphere is tense and high-stakes. The senior industrialist at the head of the table leans forward slightly and says firmly: “This acquisition will change the entire market. Either we move now, or someone else will.” Another industrialist across the table adjusts his cufflinks and replies calmly: “Moving now is risky. The numbers still don’t support a full takeover.”A third industrialist places both hands on the table and speaks with intensity: “Risk is the price of control. The investors are already waiting for our decision.”The camera slowly pushes in as the most powerful industrialist leans back, looks at everyone, and delivers the final line with quiet authority: “Then we take it. No more delays.” Cinematic lighting, sharp details on suits and the wooden table, subtle tension in their faces, realistic skin textures, high-end corporate atmosphere, shallow depth of field, filmic look.

Cinematic10

A Vintage 1947 Town Under Fire
Comunidade@RuzainaMeer

A Vintage 1947 Town Under Fire

15s · t2va · 带音频

Create a 15-second ultra-photorealistic live-action war sequence set in the United States in 1947, designed to look like authentic historical footage captured on a 1940s film camera. The entire scene must feel grounded, documentary-like, raw, and physically realistic. Environment: A rural American town in 1947 with wooden houses, old brick buildings, telephone poles, dirt roads, vintage American cars from the 1940s, wooden fences, farmland, and period-accurate street details. Overcast afternoon light, light fog, drifting smoke, dust in the air, damaged buildings, scattered debris, and a tense wartime atmosphere. Characters: American soldiers wearing historically accurate late-1940s military uniforms, helmets, boots, and equipment. Civilians wear authentic 1940s American clothing. Natural faces, realistic skin texture, sweat, dirt, fatigue, and believable body movements. 0–3s — Establishing Shot: Wide handheld shot of a quiet rural American street suddenly filled with smoke and confusion. Vintage 1940s vehicles are parked along the road while soldiers move quickly between wooden buildings. Civilians rush toward safer areas. 3–6s — Tension: Camera moves through the street at shoulder height, following several soldiers as distant gunfire is heard. They immediately react and take cover behind a vintage vehicle and a brick wall. Their movements are cautious and realistic. 6–10s — Combat: Fast handheld tracking shot as the soldiers move between cover while distant gunfire impacts the environment. Small pieces of wood, dust, and debris fall naturally from nearby impacts. Weapon recoil, movement, and body weight must be physically accurate. Keep the violence realistic and restrained. 10–13s — Human Moment: Camera briefly focuses on a soldier helping an injured civilian move behind cover. Their breathing, facial expressions, body language, and movement should feel natural and unscripted. 13–15s — Final Shot: Camera pulls back into a wide shot of the American town as smoke slowly moves through the street. Soldiers remain behind cover while vintage vehicles and damaged buildings fill the background. The scene ends with an authentic, tense 1940s documentary feeling. Visual Style: Ultra-photorealistic live-action, authentic 1940s American environment, vintage 35mm film texture, subtle film grain, natural imperfections, realistic exposure, handheld documentary cinematography, muted historical color palette, realistic smoke and dust, natural shadows, accurate depth of field. Physics: Strictly obey real-world gravity, momentum, inertia, friction, recoil, weight, collision physics, and human biomechanics. No exaggerated explosions, impossible movements, superhero behavior, or choreographed-looking combat. Negative Prompt: modern buildings, modern cars, smartphones, modern clothing, modern weapons, futuristic technology, CGI appearance, video-game graphics, fantasy, superhero action, excessive explosions, excessive blood, gore, impossible physics, unrealistic recoil, slow-motion physics, distorted faces, extra limbs, floating objects, plastic skin, artificial-looking environments.

Hand-Painted Blizzard Survival
Comunidade@sada_ai

Hand-Painted Blizzard Survival

15s · t2va · 带音频

Visual Medium & Craft: Traditional frame-by-frame hand-painted oil-on-paper animation, physical 12 fps stop-motion cadence with deliberate frame holds and natural line jitter between frames. Raw brushstroke textures, strictly zero CGI, no 3D rendering, no digital interpolation, no game engine aesthetics. Cinematography & Lighting: Masterful cinematography in the style of Emmanuel Lubezki and Roger Deakins, shot with physical cine lenses and a classic 180° shutter blur. Lit exclusively by a single overhead light source: pale, freezing moonlight cutting through dense atmospheric haze. Strong contre-jour backlighting with the camera positioned on the shadow side. Uniform cold blue moonlight washes over skin, clothing, snow, and tree trunks, casting sharp cold rim lights on shoulders and hoods while keeping faces in soft shadow. Atmosphere & Environment: A violent, blinding blizzard with near-zero visibility (3 meters maximum), dissolving the background into a solid white fog wall. Gale-force horizontal winds whip thick snow and freezing fog through the frame. Deep snow quickly accumulates on garments. Exhaled breath instantly vaporizes and tears away in the wind. Characters & Emotion: A desperate mother and her frightened young child battling for survival, lost in the storm at the absolute limit of their endurance. The mother conceals her panic to protect her child, while the child cries openly in fear. Both lean heavily into the gale, squinting against the biting wind, shielding their faces with forearms and ducking into their collars. Technical Details: Color palette dominated by 60% snow-blue, 30% deep trunk-black, and 10% cold highlight accents. Strict physical realism with grounded weight, true inertia, and accurate contact shadows on snow. Dynamic compositional balance using the rule of thirds and golden ratio. Uncompromising continuity across all character designs, garments, and environmental assets. 8K IMAX aesthetic.

INK AFTER DARK in Pulp-Noir Motion
Comunidade@opener_ai

INK AFTER DARK in Pulp-Noir Motion

15s · t2va · 带音频

[FORMAT] Create a 15-second, 16:9 retro graphic opening-title sequence titled "INK AFTER DARK". A stylish 1960s pulp-noir animation about a mysterious night writer. Bold, minimal, elegant, slightly dangerous, with dry visual wit. [IDENTITY] Keep one female writer visually consistent throughout: sharp angular bob haircut, narrow almond-shaped eyes, long black turtleneck, high-waisted trousers, slim silhouette, black leather gloves, calm expression, upright self-assured posture, and hard graphic highlights along one side of her face. Use flat hand-inked 2D illustration with rough screen-print texture, imperfect paper grain, and slightly uneven ink edges. [BEATS] [0–3 seconds] Extreme close-up of the writer's face emerging from near-total black. A narrow cream-colored strip of light slides across her eyes. The frame abruptly opens sideways like torn paper, revealing her full silhouette walking across a burnt-red background while loose sheets of paper trail behind her like a long ribbon. [3–6 seconds] The paper ribbon sweeps across frame and becomes a graphic wipe. Reveal her in profile at a desk. She strikes one typewriter key. On impact, the screen fractures into three bold rectangular panels: her gloved fingers, the metal typebar snapping forward, and a giant black ink letter striking paper. [6–9 seconds] Rapid rhythmic close-ups: spinning typewriter ribbon spool, carriage return lever snapping sideways, black ink spreading through rough paper fibers. Thin cream lines cut diagonally across the screen and reorganize the images into an asymmetric editorial collage. [9–12 seconds] Pull wide. The writer stands alone beside an enormous abstract typewriter rendered as a black geometric silhouette. She pulls one endless page upward. The rising page becomes a full-frame cream vertical wipe while scattered black letters tumble downward like physical debris. [12–15 seconds] The letters rapidly assemble into the exact title "INK AFTER DARK" centered large on a burnt-red paper field. The writer's small black silhouette crosses beneath the title and exits frame. The title appears through sharp letter-by-letter mechanical impacts over 0.5 seconds, holds completely still through the ending. No bouncing, spinning, stretching, or fly-in typography. [CAMERA] Use aggressive graphic changes in scale: extreme facial close-up → full-body silhouette → macro mechanical inserts → monumental wide composition. Camera movement should feel designed rather than realistic: fast lateral pushes, sudden graphic crops, one rapid pull-out, and precise locked compositions. Avoid conventional cinematic orbit shots. [LIGHT] Limited palette only: burnt red, aged cream, deep black, with tiny muted silver highlights on typewriter metal. Hard noir side-lighting translated into flat graphic shapes. Rough vintage print stock, subtle paper scratches, coarse ink grain, slight registration offsets, and occasional frame jitter. [EDIT] Fast editorial rhythm with hard cuts, paper wipes, diagonal panel slices, oversized object masks, and split-screen recompositions. Transitions must be motivated by paper, ink, typewriter mechanisms, or moving silhouettes. No soft dissolves. Make every shot feel newly composed rather than simply zooming into the previous image. [AUDIO] Audio: dry typewriter key strikes, paper slides, ribbon-spool clicks, carriage-return snaps, faint room hum, and one heavy mechanical impact when the final title locks. BGM: an original 15-second cue, 65% noir suspense and 35% cool jazz. Upright bass, brushed snare, muted vibraphone, sparse low piano, and short clipped brass accents. Begin almost empty, introduce bass at 3 seconds, rhythmic percussion at 6 seconds, a brief brass accent at 10 seconds, then freeze the final 2 seconds on one bass note and the mechanical title hit. Do not imitate an existing melody. [NEGATIVE] No subtitles, extra on-screen text, watermarks, platform logos, or stickers. Do not introduce Chinese text, garbled characters, misspellings, or additional title variations. Render "INK AFTER DARK" once only. No 3D CGI, photorealism, anime styling, glossy modern motion graphics, neon cyberpunk aesthetics, or smooth vector-clean surfaces. Never a slideshow. Maintain active graphic motion, physical visual transitions, and continuously evolving compositions.

Action & VFX12

The Sky Descends Like a Giant Wall over a Suburban Street
Comunidade@cocktailpeanut

The Sky Descends Like a Giant Wall over a Suburban Street

14s · 带音频

A hyper-realistic handheld phone video of a quiet suburban backyard on a bright normal afternoon. A person casually waters plants near a fence while birds chirp, distant traffic hums, and everything feels completely ordinary and mundane. The camera drifts naturally like a real home video, capturing the lawn, patio furniture, and blue sky with soft white clouds. After several seconds of normal peaceful activity, a deep cracking sound suddenly comes from above. Without warning, the entire sky begins dropping straight downward like a gigantic solid ceiling, with the blue sky and clouds moving as one physical surface descending toward the yard. The person looks up in shock just as the sky rapidly fills the frame, swallowing the scene in a terrifying instant. The camera jerks and falls to the ground at the last moment. No cinematic buildup, no surreal visual style, it should feel like a totally normal real-life recording interrupted by one sudden impossible and visceral event.

The Diner That Rewinds
Comunidade@icreat_ai

The Diner That Rewinds

15s · t2va · 带音频

Photorealistic cinematic 1950s American diner, chrome stools, red vinyl, neon glow and checkerboard floor, shot with modern lived-in realism and soft natural window light. Subtle handheld texture, warm practicals, rich period detail, heavy film grain. 0-4s: [Medium Wide] A striking young woman in her early 20s sits alone at the counter, calm and slightly amused, slowly sipping a tall thick milkshake through a straw. Behind her a young waitress in classic uniform approaches with a tray of eggs and bacon in one hand and a full glass coffee pot in the other. An older lady starts rising from a nearby booth. 4-8s: [Dynamic Tracking] The older lady collides hard into the waitress. Tray, plate, eggs, bacon and coffee pot explode upward in chaotic slow motion. Coffee erupts into long liquid ribbons and perfect suspended droplets. Camera immediately begins a smooth continuous orbit around the impact. Time locks completely at the peak of the spill. Every face freezes in pure shock. Only the girl at the counter keeps moving, completely unfazed. 8-17s: [Slow 360° Orbital] Camera glides in a full elegant orbit through the frozen diner. Coffee hangs in mid-air as glassy ribbons and spheres with perfect volume and surface tension. Bacon strips, eggs and the spinning tray float weightlessly. Patrons and waitress remain locked in startled expressions. The girl at the counter takes one slow, deliberate sip, eyes half-lidded, almost bored, while the entire frozen world (except her) begins an elegant reverse: every droplet, every piece of food and every person rewinds smoothly back to the exact starting positions. 17-24s: [Medium Shot] Rewind lands perfectly. Waitress stands balanced again with tray and coffee pot. The girl lifts her eyes, raises two fingers in a small casual gesture and softly calls the waitress by name. The waitress turns toward her just before the older lady begins to stand, completely avoiding the collision. A tiny private smile crosses the girl’s face. 24-30s: [Extreme Close-Up] Hard cut to her face as she takes one last slow sip. Soft knowing smile, eyes almost closed in quiet satisfaction, like she has done this a hundred times. Shallow depth of field, creamy bokeh of the neon diner behind her. Photorealistic, ultra-detailed fluid physics, perfect motion blur only on moving elements, stable characters, cinematic lighting, heavy natural film grain, no artifacts, movie-level temporal coherence, high rewatch value.

A Green Locomotive Mechanically Transforms into a Robot
Comunidade@EndFolding79421

A Green Locomotive Mechanically Transforms into a Robot

15s · fl2va · 带音频

Use <Image 0> the exact FIRST FRAME and <Image 1>as the exact LAST FRAME. Create one continuous 15-second photorealistic cinematic transformation connecting these two states. Maintain the EXACT same railroad location throughout: same curving track, green grass, surrounding trees, distant rolling forested hills and mountains, sky, morning atmosphere, perspective, and overall natural lighting. The environment does NOT transform, shift, regenerate, or change location. No additional people, vehicles, trains, or robots appear. The locomotive itself physically transforms into the robot shown in <Image 1>. This must be a true complex mechanical transformation, NOT a dissolve, magical morph, liquid-metal effect, crossfade, replacement, or sudden character swap. Every major robot component originates visibly from physical locomotive components. Preserve the locomotive's recognizable green-and-yellow painted metal, weathering, markings, wheels, panels, vents, grilles, lights, and mechanical construction as they reorganize into the final robot. [00:00–00:02.5] Begin exactly on <Image 0>. Hold the intact locomotive briefly. The engine idles subtly. Then small mechanical movements begin: suspension compresses, exterior panels unlock and separate along real seams, vents shift, internal gears engage, and heavy metal sections begin sliding apart. The camera makes a subtle cinematic adjustment to accommodate the increasing vertical scale. [00:02.5–00:06] The transformation becomes larger and clearly mechanical. The locomotive's wheel assemblies separate and rotate downward, reorganizing into two enormous articulated legs while remaining visibly composed of train machinery. Axles rotate, wheel trucks pivot, steel panels telescope outward, internal pistons extend, and mechanical joints lock together. The locomotive remains physically planted over the same section of track. The main locomotive body begins rising upward as the newly formed legs extend beneath it. No pieces float independently and nothing materializes from nowhere. [00:06–00:10] The locomotive body folds and compacts into the robot torso. Large green exterior panels rotate inward and interlock across the chest while recognizable yellow locomotive markings remain incorporated into the armor. Side machinery unfolds outward into shoulders. Long mechanical assemblies telescope and rotate downward into articulated arms. Hands assemble from smaller interlocking mechanical components at the ends of the forearms. Visible gears turn, pistons compress, cables tighten, panels slide across one another, and heavy locking mechanisms engage. The camera smoothly pulls back and tilts upward only as necessary to keep the entire transformation readable while preserving the original environment and spatial orientation. [00:10–00:13] The upper locomotive housing separates and folds inward. Smaller panels rapidly rotate, collapse and interlock around an emerging mechanical neck assembly. The robot head rises mechanically from inside the reorganized machinery; face plates slide into position, the final head structure locks together, and the eyes illuminate. Remaining locomotive panels snap precisely into their final locations across the shoulders, chest, back, arms and legs. Final pistons compress and locking mechanisms engage with tremendous physical weight. [00:13–00:15] The transformation completes. The enormous robot settles its weight onto the railway tracks with a heavy suspension-like compression through its legs. Tiny secondary panels finish closing and locking. The camera settles naturally into the exact composition required by <Image 1>. The robot becomes the precise completed design, proportions, orientation and pose shown in <Image 1>, standing in the same landscape. End EXACTLY on <Image 1> and hold the completed robot briefly. No further transformation. overall_soundscape: Sound ON. Highly detailed realistic mechanical transformation audio synchronized precisely to visible movement: deep diesel engine rumble, enormous steel panels sliding and rotating, gears grinding smoothly, hydraulic pistons releasing and compressing, servomotors, heavy hinges, wheel assemblies rotating, suspension groans, metal components telescoping, powerful mechanical locking CLUNKS and deep resonant impacts. Transformation sounds increase in scale as larger components move. During the final assembly, several massive sequential mechanical locks culminate in one deep heavy final THUNK as the robot reaches full form. Afterward, quieter servo settling, cooling metal ticks and a low mechanical idle remain. Maintain subtle natural countryside ambience underneath: light morning wind through grass and trees and distant birds. No dialogue. non_diegetic_music: None. No score, no musical tones, no trailer music. Let the mechanical transformation and natural environment provide the entire soundtrack.

Animation & Stylized16

Learning the Days of the Week
Comunidade@Strength04_X

Learning the Days of the Week

15s · 带音频

Create a 15-second animated educational video that introduces young children to the first four days of the week: Monday, Tuesday, Wednesday, Thursday. Learning pattern: DAY NAME → SOUND → FUN ACTIVITY ICON → PLAYFUL ACTION → DAY NAME REPEAT Target audience: children ages 3 to 6. Visual style: Same premium-cute aesthetic — pastel tones, rounded 3D icons, soft lighting, off-white background, unique glow per day, clean bold typography, gentle calendar-page-flip style transitions. 0:00–0:01 Intro: Mascot bounces onto a soft floating calendar page, text "Let's learn the days!", tap flips to reveal Monday. ▪️0:01–0:04 | Monday — School Day: Calendar page shows "Monday" in big rounded text. Narrator: "Monday. Time for school!" A tiny backpack icon bounces onto the page and wiggles happily. Word "MONDAY" highlighted in soft blue. ▪️0:04–0:07 | Tuesday — Park Day: Page flips gently to Tuesday. Narrator: "Tuesday. Time to play!" A small swing icon appears and swings back and forth once. Word "TUESDAY" highlighted in soft orange. ▪️0:07–0:10 | Wednesday — Art Day: Page flips to Wednesday. Narrator: "Wednesday. Time to paint!" A tiny paintbrush icon appears and draws one curved rainbow line. Word "WEDNESDAY" highlighted in soft purple. ▪️0:10–0:13 | Thursday — Story Day: Page flips to Thursday. Narrator: "Thursday. Time for stories!" A small open book icon appears, pages gently flutter. Word "THURSDAY" highlighted in soft green. ▪️0:13–0:15 Recap: Backpack, swing, paintbrush, book appear in four tiles with day names above. Mascot points to each; Narrator: "Monday, Tuesday, Wednesday, Thursday. Great job!" Sparkle + chime.

An 1890s Bird Plate Draws Itself
Comunidade@Dheepanratnam

An 1890s Bird Plate Draws Itself

15s · fl2va · 带音频

@ Image1 is the sole visual authority for the folio spread, the wren, the hazel twig, the egg study, the foot sketch, the handwriting, the printed type, the paper, the foxing, the lighting, the palette, the layout and the horizontal composition. Animate this exact page without redesigning it. Preserve every existing mark, letter and colour in its final state. [Goal] Create a 15-second sequence in which this plate makes itself, and then the bird inside it comes quietly alive. One continuous, locked, straight-on overhead composition. No camera movement, no cuts, no cropping, no perspective change, no page turning. [Sequence] 0-2.4 seconds: Begin on blank aged ivory paper — the empty spread, correct in paper texture, chain lines, foxing and lighting, but carrying no image, no writing and no type. A faint graphite underdrawing appears across the right page in light searching strokes, finding the wren's posture and the line of the twig. A few construction marks overshoot and are left uncorrected. 2.4-6 seconds: Engraved ink linework draws itself over the pencil, following the exact contours of the finished plate — outline first, then hatching building the shadow under the wing, the barring of the tail, the texture of the twig. The rectangular plate mark presses into the paper around the image with a soft physical impression, lifting the paper very slightly at its edge. 6-10.5 seconds: Watercolour floods in. Washes enter the printed outlines and settle into them, warm russet across the back, buff along the eyebrow stripe, muted olive on the twig, each wash landing exactly within the shapes of the finished plate and overrunning the line in the same few places the reference does. The colour blooms wet and then visibly dries, darkening slightly at the edges. On the left page, the sepia handwriting writes itself line by line from the top, in a period cursive hand that is elegant, cramped and only partly legible. The small egg study fills with colour and the pencil foot sketch appears faintly beside it. 10.5-13 seconds: The printed type impresses into the right page — TROGLODYTES TROGLODYTES, then The Common Wren beneath it, then PLATE XVII in the top right corner. Foxing spots deepen. The last details resolve: the highlight in the bird's eye, the fine barring on the flank. The page settles. Everything on the spread is now exactly as in @[ref img]. 13-15 seconds: The page stays completely still and completely flat — it is paper and it behaves like paper. But the printed bird begins to breathe. Its flank rises and falls very slightly. It blinks once. It turns its head a few degrees and holds, alert, looking past the viewer at something outside the page. The cocked tail gives one small twitch. It does not fly, it does not leave the plate, it does not become three-dimensional, and it does not look at the viewer. Settle into a living idle: the bird breathing, one more slow blink, the paper still. Render as photoreal archival photography of a real object under warm north-window light. Keep the printed page flat and matte with no glow, no emission, no glint and no digital sheen. Keep the two motion layers distinct: the page is a made thing being made, the bird is alive. Keep all type, handwriting, marks and geometry fixed once they appear. Do not introduce any new text, letters, numerals, marks, birds, objects, hands, tools, logos, captions, borders, watermarks or page numbers. Do not turn or lift the page. Do not move, rotate or scale the camera. Do not zoom. Do not add glow, neon, coloured light, lens flare, bokeh or shallow-focus racking. Do not make the bird photoreal or three-dimensional — it must remain a hand-coloured engraving that happens to be alive. Do not have it fly, hop, open its beak wide, or leave the plate. Audio: fine graphite scratching on rough paper; the dry precise scrape of a steel nib; the deep soft compression of a plate press; a sable brush loaded with water moving across paper; paper fibres settling; the quiet room tone of a study with a slow long-case clock ticking somewhere behind. At 13 seconds, one Eurasian wren song — sudden, intricate and far louder than anything that small should produce — then the room tone returns and the clock keeps going. No music, no dialogue, no narration, no electronic tones.

Circle, Square, Triangle, and Star
Comunidade@Strength04_X

Circle, Square, Triangle, and Star

15s · 带音频

Create a 15-second animated educational video that teaches young children the shapes Circle, Square, Triangle, and Star. Learning pattern for every shape: SHAPE → NAME SOUND → OBJECT → PLAYFUL ACTION → OBJECT NAME Target audience: children ages 3 to 6 Visual style: Same premium-cute aesthetic pastel colors, rounded 3D forms, soft lighting, minimalist off-white background, unique glow color per shape, crisp bold typography, polished smooth morphing transitions. 0:00–0:01 Intro: Mascot bounces in, shapes float briefly, text "Let's learn shapes!", tap reveals first shape with a ripple. 0:01–0:04 | Circle is for Orange: A perfect circle appears. Narrator: "Circle. Round and round. Circle is for Orange." Circle gains texture and a tiny stem, becomes a smiling orange, bounces once. Word "ORANGE," highlight shape outline in orange. 0:04–0:07 | Square is for Gift Box: The orange rolls and reshapes into a square. Narrator: "Square. Four equal sides. Square is for Gift Box." Square gains a ribbon and bow, wiggles playfully. Word "GIFT BOX," highlight in purple. 0:07–0:10 | Triangle is for Mountain: The ribbon folds into a triangle. Narrator: "Triangle. Three sides. Triangle is for Mountain." Triangle grows a snowy peak and a small smiling sun beside it. Word "MOUNTAIN," highlight in brown/teal. 0:10–0:13 | Star is for Sky: The mountain peak stretches into a star. Narrator: "Star. Star is for Sky." Star twinkles softly with tiny sparkles floating around it, gentle glow pulse. Word "STAR," highlight in golden yellow. 0:13–0:15 Recap: Orange, gift box, mountain, star line up in four tiles with shape outlines above. Mascot points to each in sequence. Narrator: "Circle, Square, Triangle, Star. Great job!" Sparkle + warm chime. Requirements: Each shape fully visible before morph, uppercase-clarity level precision on shape edges, no warped or duplicated forms, gentle motion only, synced sound cues, calm and beautifully polished finish.

UI & Typography9

Beat-Synced Spaghetti-Western Title Sequence
Comunidade@itsphotogptai

Beat-Synced Spaghetti-Western Title Sequence

10s · ref2va · 带音频

Use the attached character reference (Image1) and audio track as the timing guide. Create a spaghetti-western pulp anime title sequence where every cut lands precisely on the beat. Style the character as a flat black silhouette against solid desert-inspired colors (burnt orange, ochre, rust, adobe tan), using bold pop-art compositions, vintage bounty-poster framing, split panels, extreme close-ups, and wide horizon shots. Alternate fast action bursts with full-beat freeze frames (draw, stride, turn, match strike, etc.), using hard cuts only. Add worn woodblock/stencil title typography that slams into frame on the beat. Finish with the character frozen in silhouette on a burnt-orange background as the title stamps in on the final hit. Apply heavy 35mm film grain, scratches, dust, gate weave, cigarette burns, and sun-faded print wear throughout for an authentic 1960s spaghetti-western look.

Typography Flows Through the Scene
Comunidade@opener_ai

Typography Flows Through the Scene

15s · 带音频

[1] FORMAT 15-second 16:9 stylized anime action title sequence — martial-arts choreography fused with 3D kinetic typography. Painterly 2D anime illustration look with cel-shaded volumes and subtle paper grain, matching the source artwork's rendering. [2] ASSETS Use Image 1 as a strict identity reference for the female duelist KAELAN — read only the character artwork itself for her face, hair, wardrobe, weapon, and color palette. Ignore the sheet's layout, panels, printed labels, handwriting, silhouette thumbnails, and background parchment; none of them may appear in the video. [3] IDENTITY Preserve exactly: deep brown skin with warm highlights; luminous crimson-red eyes with a sharp black liner shape; jet-black hair pulled into a tight scalp-wrapped part that splits into two very long thick braids falling past the hips, moving like weighted whips; large white crescent hoop earrings on both ears; black high-neck sleeveless crop top exposing a toned midriff; a sheer white long-sleeved kimono robe worn open, translucent and lightly billowing; cream wide-leg hakama trousers with a single black diamond patch on the left thigh; black wrapped fingerless gloves with lattice cord across the back of the hand; white cloth-wrapped platform sandals; a well-worn katana with a dark blue-grey blade and a black diamond-wrapped hilt. Her expression stays calm and unblinking, never comic. [4] BEATS Shot 1 — [0–3 seconds] Kaelan stands centered in a void of warm off-white light, katana point-down, braids settling. She exhales; a slow iai draw begins. → Hard cut. Shot 2 — [3–6 seconds] Mid-draw. The blade clears the saya and a clean arc of light follows the edge. The 3D extruded word "HAILUO" flies in from screen left along the blade's arc, its letter faces catching the same light, passing BEHIND her shoulder and IN FRONT of her trailing braid before exiting right. → Whip-pan cut. Shot 3 — [6–9 seconds] Full-body kata combination: descending cut, pivot, rising horizontal sweep. Her sheer robe and both braids trail one frame behind her body. The word "MINIMAX" rotates up from the floor plane in thick extruded 3D, standing taller than her, and she cuts through the gap between its letters — the letterforms hold solid, do not shatter. → Hard cut. Shot 4 — [9–12.5 seconds] Low-angle crouch landing, sword arm extended back, dust ring pushing outward. "H3" slams in from depth as a massive chromed 3D glyph directly behind her, scaling down to lock, with her silhouette reading cleanly against it. → Flash cut on the impact frame. Shot 5 — [12.5–15 seconds] She rises into the neutral standing pose, katana returning to point-down. The full lockup "HAILUO MiniMax H3" assembles from the three pieces on a single horizontal baseline at lower third, settles, and holds perfectly still through the final frame. [5] CAMERA Anamorphic-feel 35mm for wides, 50mm for the mid shots.

Two Glaciologists Enter a Blue-Ice Cave
Comunidade@Strength04_X

Two Glaciologists Enter a Blue-Ice Cave

15s · 带音频

EXPEDITION: Two glaciologists enter a newly formed ice cave to document its internal structure before seasonal melting changes it. EXPLORERS: Two researchers, both wearing crampons, helmets and headlamps, carrying compact measurement equipment. SHOOT WINDOW: 7:05 AM – 7:20 AM Early morning outside. Overcast sky. Cold blue daylight enters through the cave entrance. LOCATION: Glacial valley, massive blue ice formation, narrow cave entrance, translucent walls, frozen water flowing beneath the surface. RECORDING STYLE: First-person expedition footage mixed with a second camera operator. Headlamps create moving pools of light across the ice. Natural breathing and footsteps remain prominent. 15-SECOND EXPLORATION: 00:00–00:03 → Researchers squeeze through the narrow entrance. 00:03–00:06 → Camera reveals deep blue translucent ice walls surrounding them. 00:06–00:09 → One researcher notices water suddenly moving beneath a transparent ice floor. 00:09–00:12 → Both carefully step backward and examine the changing surface. 00:12–00:15 → They mark the location and begin retreating toward daylight. AUDIO: Boots scraping ice, dripping water, distant cracking, breathing, muffled voices. REALISM DIRECTIVE: Ice must behave like real compressed glacier ice, not glass. Light scatters through the walls naturally. No fantasy formations, supernatural elements or impossible cave geometry.

Como escrever um prompt para o MiniMax Hailuo 3

Compilado a partir do Video Prompt Writing Guide oficial, publicado na ficha de modelo MiniMax-H3 no HuggingFace — não é um resumo nosso.

A fórmula base

integrated_multimodal_description: [Shot 1] ... overall_soundscape: ... non_diegetic_music: ...

  • integrated_multimodal_descriptionobrigatórioDescribes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.
  • overall_soundscapeobrigatórioSummarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video.
  • non_diegetic_musicobrigatórioDescribes background music that the characters cannot hear and only the audience can hear.

Modelo de frase oficial

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

Quatro regras

  1. 1. Shots and CutsDo not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration.
  2. 2. Camera MotionA complete camera-motion expression has three dimensions: the motion type defines how the camera moves, amplitude defines the range of compositional change, and speed defines the pacing of that change. Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted.
  3. 3. Speakers, Dialogue, and SingingSubjects who speak, sing, or produce an off-screen human voice use stable IDs such as (S1) and (S2). A speaker keeps the same ID across shots. Preserve every original word and punctuation mark verbatim inside <d> — do not translate or rewrite them.
  4. 4. On-Screen TextPlace any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation.

Parâmetros

overall_soundscape accepts N/A only when the user explicitly requests complete silence throughout the video; non_diegetic_music accepts N/A whenever there is no background score.

É uma recomendação do fornecedor, não uma regra rígida.

Veja o caso n.º 1 · Ponte da nave estelar, o exemplo mais completo desta estrutura em três partes.

Pronto para Criar seu Próximo AI Agent ?

3 minutos para integrar. Comece agora.

Obter API Key