SOL-H3 IS HERE.

Big imagination. One Spark. A cinematic promotional film for Sol-H3 by Sana Team / Sol-Engine.

Sol-H3 Spark · Two separate films

完整 Prompt · 展开查看与复制
打开 TXT
SOL-H3 SPARK — 高级宣传片完整制作资料 / 34 SECONDS

说明:这支成片由已有 Spark showcase 素材剪辑、品牌动效和独立 TTS 旁白组成,没有一条直接生成整支 34 秒成片的原始模型 prompt。以下保留实际使用的全部素材 prompt 和完整旁白;整片制作说明依据 v5 渲染脚本整理。

=== 1. 整片制作说明 / EDITORIAL BRIEF ===
Create a 34-second, 1344 × 756, 24 fps Sol-H3 Spark promotional film using the existing generated showcase clips listed below. Use a consistent two-line Sol-H3 / Spark lockup: upright Sol-H3, italic Spark, and a four-point spark. Palette: warm paper (#EFEFE0), charcoal (#0C1314), lime (#B4F158), deep green (#315C2B). Typeset exact text and composite official logos in post-production.

00:00–00:04 — Astronaut reentry cockpit. Sol-H3 is fully visible from the first frame. A brief lime ignition introduces Spark as a complete word underneath; the cinematic background brightens. Show SOL-ENGINE, Now on NVIDIA DGX Spark., and Big imagination. One Spark.
00:04–00:10 — Snow leopard ridge, source played at approximately 0.8×. Show CREATE IN / 56 seconds, 768p video + audio, 5-second clip · one DGX Spark, and ~56 s complete hot end-to-end latency.
00:10–00:17 — Macro fountain pen, source played at approximately 0.69×. Show Small draft. / Big detail., 384p → 768p, MINIMAX-H3 DRAFT, LTX REFINEMENT, and Two stages. One Sol-Engine pipeline.
00:17–00:20 — Wildflower portrait. Label: Cinematic / portraits.
00:20–00:23 — Dragon cloud archipelago. Label: Impossible / worlds.
00:23–00:26 — Sunset skatepark. Label: Natural / motion.
00:26–00:29 — Seaside coffee. Label: Human / stories.
For the last four showcases, use three seconds starting at source time 0.45 s. Keep full-screen moving images, consistent corner branding, a darkened lower area for readable type, and the footer GENERATED ON ONE SPARK. Animate the two-line label upward into place.
00:29–00:34 — Warm-paper brand ending with Sol-H3 / Spark, a quiet enlarged spark motif, SOL-ENGINE, IMAGINATION, ACCELERATED., and Video + audio. One NVIDIA DGX Spark. Align official NVIDIA, MiniMax and SANA logos on the dark lower strip. Fade to charcoal during the final 0.3 s.

Keep the existing v5 visual edit intact when adding narration. Retain its original score, smoothly reduced by 13 dB under speech. Normalize voice to roughly -20 dBFS RMS and limit the final mixed audio. The ~56-second result is a warm complete end-to-end request for a five-second 768p video with audio on one DGX Spark: 56.17 s mean across three warm resident requests.

=== 2. 完整旁白 / EXACT NARRATION ===
Voice: Microsoft Edge TTS, en-GB-SoniaNeural, rate -5%, pitch -2 Hz. The narration is an independent post-production audio track.

0.25–3.341 s
Display text: Sol-H3. Now on Spark.
Exact TTS input: Soul H three. Now on Spark.

4.18–8.568 s
Display text: Five seconds of video and sound, in about 56 seconds.
Exact TTS input: Five seconds of video and sound, in about fifty-six seconds.

10.15–15.670 s
Display text: MiniMax-H3 creates the draft. LTX refines the detail.
Exact TTS input: MiniMax H three creates the draft. L T X refines the detail.

17.30–20.923 s
Display text: Cinematic portraits. Impossible worlds.
Exact TTS input: Cinematic portraits. Impossible worlds.

23.30–26.278 s
Display text: Natural motion. Human stories.
Exact TTS input: Natural motion. Human stories.

29.35–33.569 s
Display text: Sol-H3 Spark. Imagination, accelerated.
Exact TTS input: Soul H three Spark. Imagination, accelerated.

=== 3. 七段素材的完整原始模型 PROMPTS / VERBATIM ===

--- 00:00–00:04 / astronaut-reentry-cockpit.mp4 ---
Source video: https://huggingface.co/datasets/Efficient-Large-Model/Sana-assets/resolve/efed7c62d78352b703adaf1d5a1d0a8de445c8fd/Sol-Engine/Sol-H3-Spark/20260910/showcase/astronaut-reentry-cockpit.mp4

integrated_multimodal_description: [Shot 1] Cinematic live-action inside a cramped spacecraft cockpit during atmospheric reentry, a close view holds on an astronaut's face behind a clear glass visor that reflects blinking red and green controls. In one continuous five-second shot, the camera shakes slightly with the vibrating cabin while beads of sweat travel down the astronaut's temple and their wide determined eyes scan the instrument panel. Fierce orange friction light flickers through a small porthole and moves dynamically across the detailed flight suit, helmet, and visor. Preserve the astronaut's appearance, helmet geometry, reflections, and cockpit layout; no cuts, captions, or logos. overall_soundscape: Deep mechanical vibration and rapid hull rattles fill the cockpit beneath intermittent electronic beeps. The astronaut's controlled breathing remains close inside the helmet as a low reentry roar rises outside. non_diegetic_music: N/A

--- 00:04–00:10 / snow-leopard-ridge.mp4 ---
Source video: https://huggingface.co/datasets/Efficient-Large-Model/Sana-assets/resolve/efed7c62d78352b703adaf1d5a1d0a8de445c8fd/Sol-Engine/Sol-H3-Spark/20260910/showcase/snow-leopard-ridge.mp4

integrated_multimodal_description: [Shot 1] Photoreal wildlife documentary in crisp mountain daylight, a medium-wide view frames one majestic snow leopard with thick spotted fur moving across a jagged, snow-covered ridge beneath towering Himalayan peaks and a clear blue sky. During one continuous five-second shot, the camera pans slowly to follow its low stalking gait as each paw sinks slightly into fresh powder. The leopard pauses near the end, its long bushy tail twitching for balance while its green eyes scan the valley; frost-dusted whiskers and individual hairs remain sharply visible. Preserve the animal's anatomy, markings, scale, paw contact, and ridge geometry throughout, with no cuts, text, or logos. overall_soundscape: Soft mountain wind passes over the ridge while paws compress fresh snow with muted crunches. A faint tail brush and the leopard's quiet breathing are audible in the cold open air. non_diegetic_music: N/A

--- 00:10–00:17 / fountain-pen-blue-ink.mp4 ---
Source video: https://huggingface.co/datasets/Efficient-Large-Model/Sana-assets/resolve/efed7c62d78352b703adaf1d5a1d0a8de445c8fd/Sol-Engine/Sol-H3-Spark/20260910/showcase/fountain-pen-blue-ink.mp4

integrated_multimodal_description: [Shot 1] Photoreal macro product film in soft window daylight, an ornate gold nib on a handcrafted fountain pen moves across textured cream parchment while its polished mahogany barrel shows fine wood grain. In one continuous five-second shot, the camera tracks tightly with the nib as dark blue ink flows smoothly from the tip, follows a short curved stroke, and soaks into individual paper fibers. The metal flexes subtly under pressure and the fresh ink retains a wet sheen before beginning to dry. Keep the same pen, nib, ink line, and parchment surface coherent throughout, with no cuts, captions, or logos. overall_soundscape: The gold nib makes a delicate textured scratch across paper while the pen barrel shifts softly in the writer's grip. Quiet room tone and a faint rustle of parchment remain underneath. non_diegetic_music: N/A

--- 00:17–00:20 / wildflower-firefly-portrait.mp4 ---
Source video: https://huggingface.co/datasets/Efficient-Large-Model/Sana-assets/resolve/efed7c62d78352b703adaf1d5a1d0a8de445c8fd/Sol-Engine/Sol-H3-Spark/20260910/showcase/wildflower-firefly-portrait.mp4

integrated_multimodal_description: [Shot 1] Cinematic live-action portrait at twilight, a young person wearing a delicate crown of wildflowers sits on a deep green mossy forest floor under cool blue moonlight. In one continuous five-second shot, the camera makes a slow small-amplitude arc around them, revealing soft petals, damp moss, and fireflies blinking warmly in the background. Tiny points of bioluminescent light play across their face as they slowly turn their head to watch one firefly settle on their shoulder, ending with the same quiet expression of wonder. Preserve the person's appearance, flower crown, seated pose, and forest layout; no cuts, captions, or logos. overall_soundscape: Quiet evening forest ambience continues beneath a light breeze through leaves and a soft rustle of clothing against moss. Sparse insects chirp in the distance as the person takes one calm breath. non_diegetic_music: N/A

--- 00:20–00:23 / dragon-cloud-archipelago.mp4 ---
Source video: https://huggingface.co/datasets/Efficient-Large-Model/Sana-assets/resolve/684fda076add94955cd52f202f60825e471d3eee/Sol-Engine/Sol-H3-Spark/20260907/media/dragon-cloud-archipelago.mp4

integrated_multimodal_description: A cinematic fantasy shot follows a single pearl-white dragon with deep blue wing membranes above an archipelago of immense green cliffs rising through clouds. In one continuous five-second shot it gives two broad powerful wingbeats and gently banks past a sunlit stone arch. The camera glides beside the dragon at a constant distance, revealing fine scales, a stable face and flowing trailing whiskers. Morning light scatters through the mist, loose cloud wisps curl in the wake, and distant islands remain geometrically stable. Elegant monumental scale, rich natural colors, coherent anatomy and motion, no riders, no text or logos. overall_soundscape: Deep leathery wingbeats, rushing high-altitude wind and a distant resonant call. non_diegetic_music: N/A

--- 00:23–00:26 / sunset-skatepark.mp4 ---
Source video: https://huggingface.co/datasets/Efficient-Large-Model/Sana-assets/resolve/684fda076add94955cd52f202f60825e471d3eee/Sol-Engine/Sol-H3-Spark/20260907/media/sunset-skatepark.mp4

integrated_multimodal_description: A cinematic five-second tracking shot of an adult woman skateboarding through an open concrete skatepark at golden hour. She wears a red windbreaker, a dark helmet and worn white sneakers. Starting with both feet on a single wooden skateboard, she rolls forward, crouches, performs one small clean ollie over a painted line, and lands smoothly on the same board. The camera tracks beside her at knee height. Sunlight catches loose strands of hair and dust; long shadows glide across the concrete. Natural athletic motion, coherent board and feet, realistic wheels and ground contact, one continuous shot, no text or logos. overall_soundscape: Wheels humming over concrete, a crisp wooden pop, a short landing clack and light wind. non_diegetic_music: N/A

--- 00:26–00:29 / seaside-coffee.mp4 ---
Source video: https://huggingface.co/datasets/Efficient-Large-Model/Sana-assets/resolve/684fda076add94955cd52f202f60825e471d3eee/Sol-Engine/Sol-H3-Spark/20260907/media/seaside-coffee.mp4

integrated_multimodal_description: A naturalistic close medium shot of an elderly couple sitting together at a small seaside cafe table in warm morning sunlight. The woman has silver curls and a blue linen shirt; the man has a neat white beard and a cream cardigan. Two small ceramic coffee cups rest on the table. Over one continuous five-second take, she says something quietly, he turns toward her, and they share an unguarded laugh. Their faces and clothing remain consistent. A sea breeze moves the loose edge of a striped awning, with a softly focused turquoise harbor behind them. Subtle handheld camera, lifelike skin texture, warm intimate documentary feeling, no captions or logos. overall_soundscape: Gentle shared laughter, a few quiet conversational syllables, distant gulls, waves against the harbor and light cafe room tone. non_diegetic_music: N/A

Original v5 archive: https://huggingface.co/spaces/Lawrence-cj/sol-h3-promo/blob/main/spark-reveal-v5/showcase-prompts.json

Premium film · with narration 34 SEC

Latest reveal-v5 visuals, now with a clear English voiceover.

Watch the two-film page and read prompts

完整 Prompt · 展开查看与复制
打开 TXT
SOL-H3 SPARK — 卡皮巴拉完整 PROMPTS / 39 SECONDS

说明:4 秒后期封面 + 四段各 8 秒的 Ref2VA 原生音画 + 3 秒品牌片尾。四段模型 prompt 为实际使用的原文;封面、精确黑板文字和片尾由后期合成完成。

=== 1. 封面与整片剪辑说明 / EDITORIAL BRIEF ===
00:00–00:04 — Show the result before the explanation. Warm-paper background; two-line Sol-H3 / Spark at left, deep-green italic Spark and a four-point spark. At right, animate in GENERATE IN / ~56 s / 5-second video + audio / 768p output. Charcoal lower strip: ONE NVIDIA DGX SPARK, Warm end-to-end run, Kapi Lai explains how, and a small arrow. Fade the original score in, then down before the speaker enters. No new voiceover is used for this four-second cover.
00:04–00:12 — Clip 01: MiniMax-H3 384p draft, motion and audio together.
00:12–00:20 — Clip 02: VAE adapter, direct H3-to-LTX latent transfer.
00:20–00:28 — Clip 03: LTX + Sol-Attn, 768p refinement, cached refiner prompt and quantization.
00:28–00:36 — Clip 04: About 56 seconds for a five-second video with audio on one DGX Spark, warm end-to-end run.
00:36–00:39 — Reuse seconds 31–34 of the original spark-reveal-v5 brand ending. Keep its Sol-H3 Spark identity and official logo row. Mute the source ending audio; let the continuous score rise after the last dialogue.

Preserve Kapi Lai's identity, classroom, native generated voice and gestures. Typeset the exact labels on the right side of the chalkboard; keep the character and pointing paw visible. Normalize spoken audio, keep music quiet underneath, and preserve the complete existing 35-second body when adding the cover.

Exact board text:
01 — 384p / MINIMAX-H3 DRAFT / A compact first pass. / Motion + audio, together.
02 — VAE ADAPTER / H3 latents → LTX latents / Direct latent transfer. / No pixel decode / re-encode.
03 — 768p / LTX + SOL-ATTN / Cached refiner prompt / Quantized. Kept resident.
04 — ~56 s / WARM END-TO-END / 5-second video + audio / One NVIDIA DGX Spark.
Shared header: SOL-H3 / ON SPARK.
Shared footer: 384p draft > latent transfer > 768p detail.

=== 2. 实际推理配置 / GENERATION SETTINGS ===
MiniMax-H3 Ref2VA + LightX2V adapter: minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors.
Exactly 8 actual DiT evaluations. In the recovered Diffusers runner, num_inference_steps=9 creates 9 scheduler grid points and 8 evaluations.
LoRA rank 128, alpha 8, scale 1.0, effective alpha/rank 0.0625; fused before inference.
Video shift 6.0; audio shift 3.0.
Each clip: 1344 × 768, 192 frames, 24 fps, 8.000 s. Final edit: 1344 × 756, 24 fps.
Two image references and one audio reference, in the order below. Reference resize policy: match. Audio is a timbre reference for new dialogue.
The promotional footage was rendered on B200. The ~56-second product result refers to the separate Spark pipeline on one DGX Spark.
Official prompt-writing skill: https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/SKILL.md

=== 3. 参考素材 / REFERENCE INPUTS ===
Public Kapi Lai reference video:
https://huggingface.co/spaces/Lawrence-cj/sol-h3-promo/resolve/main/kapi-lai-ref2va-50step-20s.mp4
<Picture 1>: character crop from the reference at 1.0 s (576 × 1089 before resize).
<Picture 2>: full classroom frame from the reference at 1.0 s (768 × 1344).
<Audio 1>: voice-timbre excerpt from 0.15–3.65 s, mono 32 kHz WAV. Do not reuse its words or soundtrack in place of new dialogue.
The old 50-step source provides reference appearance and voice only; all four new clips use the 8-step configuration above.
Brand ending source:
https://huggingface.co/spaces/Lawrence-cj/sol-h3-promo/resolve/main/spark-reveal-v5/sol-h3-spark-reveal-34s.mp4

=== 4. 四段完整原始 MODEL PROMPTS / VERBATIM ===

--- 01-draft / seed 202609110 / 8.000 s ---

subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.

summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is MiniMax-H3. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.

detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, the capybara notices the camera and raises his right forepaw into a small open welcoming gesture, keeping it below shoulder height. The gesture is slow and anatomically continuous; the same paw returns toward his side as he turns his gaze briefly toward the board. A faint chalk-like line develops into a simple landscape thumbnail on the board, showing only broad silhouettes and a consistent horizon. This is an illustrative draft, not an actual recorded intermediate output. The rest of the board stays quiet. At 00:00.600, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] One Spark. Here's how. MiniMax H three creates a small draft, with motion and sound.</d> He looks toward the viewer while naming the small draft, then glances toward the thumbnail during the final words about motion and sound. His muzzle opens only as far as the syllables require, with the same two incisors remaining attached and proportionate. Finish speaking by 00:07.200. During the final eight tenths of a second, he closes his mouth, gives a small acknowledging nod, and settles into a neutral stance. The board holds the completed simple thumbnail rather than switching to a new scene.

overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.

non_diegetic_music:
N/A

--- 02-transfer / seed 202609111 / 8.000 s ---

subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.

summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is VAE Adapter. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.

detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, the capybara turns his head slightly toward the board without moving his feet. The board shows two simple separated image panels connected by one fine line. The left panel retains a rough landscape contour, and the right panel begins with the same contour rather than a different scene. At 00:00.500, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] An adapter passes the latents straight to the refiner. No decoding and re-encoding pictures in between.</d> His right forepaw makes one compact pointing gesture toward the board's left edge while remaining entirely in the left portion of the frame. A small organized packet of lime light travels directly along the connection from left panel to right panel. The packet arrives once, accompanied by a soft dry click. Nothing explodes, dissolves into a cloud, or travels through a third image panel. The abstract visual represents direct latent transfer and makes no claim to show actual internal tensors. The capybara returns his gaze to the camera while explaining the omitted decode and re-encode. He finishes by 00:07.100, closes his mouth, and lowers the pointing paw smoothly. Hold the two connected panels and his steady attentive expression through the last frame.

overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.

non_diegetic_music:
N/A

--- 03-refine / seed 202609112 / 8.000 s ---

subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.

summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is LTX + Sol-Attn. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.

detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, a narrow green highlight passes slowly over the existing landscape thumbnail on the board. The thumbnail keeps its original horizon and shapes while a few fine edges become more distinct; the camera never enters the illustration. This is a conceptual refinement demonstration rather than a measured before-and-after reconstruction. At 00:00.500, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] Sol Attention accelerates refinement. A cached refiner prompt and quantization help keep both stages in memory.</d> Pronounce the visible method name Sol-Attn as Sol Attention, clearly and without spelling each letter. He makes one gentle open-palm presentation gesture with his right forepaw and then rests both paws near his belly, leaving his mouth and face unobstructed. While he mentions the cached refiner prompt, two small neutral rectangles illuminate beneath the main image. One represents the reused refinement context, and the other represents keeping both stages resident. Their contents remain simple visual shapes; exact explanatory labels will be added in the edit. His delivery stays matter-of-fact rather than triumphant. Finish speaking by 00:07.200, then close his mouth and let the board remain still. There are no numeric speedup counters or extra claims about skipping Stage 1 text encoding.

overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.

non_diegetic_music:
N/A

--- 04-result / seed 202609113 / 8.000 s ---

subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.

summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is 5-second video + audio / one DGX Spark. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.

detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, the capybara shifts his gaze from the now-settled board to the camera. His feet, body scale, and position remain consistent with the preceding scene. A single soft lime outline frames the central board area, preparing a simple result card; this visual will receive the exact performance text during editing. At 00:00.500, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] On a warm run, five seconds of video and sound in about fifty-six seconds. One Spark.</d> He places a clear, unhurried emphasis on the words warm run. His right forepaw opens toward the board once while the other stays near his side. He does not count with additional fingers, clap, shout, or turn his back on the viewer. As he mentions the result, his eyes briefly follow the board and return to the camera. Keep his mouth naturally synchronized to all words, including the complete number fifty-six. He finishes the short phrase One Spark by 00:07.200, closes his mouth, and gives a small satisfied nod. Hold the relaxed capybara and quiet result board through 00:08.000, providing a clean edit point for the separately composited brand end card. No laughter track, roar, scream, confetti, or bright flash interrupts the final pause.

overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.

non_diegetic_music:
N/A

Kapi Lai · How Sol-H3 Spark works 8 STEP · 39 SEC

A results-first opening: one DGX Spark, a 5-second video with audio in about 56 seconds. Then Kapi Lai explains how. Ref2VA + LightX2V 8-step.

Sol-H3 Spark · Imagination, accelerated.

Big imagination. One Spark. 34 SEC · PROMO

768p video + audio · 5-second clip in about 56 seconds · 384p draft → 768p refinement · one NVIDIA DGX Spark

Download film · Film notes and source prompts

Niu Lai · LightX2V Ref2VA Turbo v1.0 8 STEP · 768P · HQ

33.875 seconds · native 1344 x 768 · 813 frames · CRF 10 high-quality master · exact same prompts, four references, seeds, and music mix as the 4-step edition · 8 effective DiT evaluations · official 768p Ref2VA LoRA · video/audio shifts 6/3 · experimental comparison: the generated FastH3 subtitle visibly omits the leading 4- while SOL-ATTN, benchmark numbers, AND MORE, and the logo finale remain clear

Niu Lai · Official MiniMax-H3 Ref2VA base OFFICIAL 50 STEP · 16:9 · HQ

33.875 seconds · native 1344 × 768 · 813 frames · freshly regenerated · CRF 10 high-quality master · background score +2 dB versus the previous master · no long silence gaps · on-screen SOL-ATTN, spoken as “Sol Attention” · clean opening-audio prompt contract · official base model with no LoRA · 50 scheduler grid points / 49 DiT evaluations

Niu Lai · portrait SOL-ENGINE story 4 STEP

Two new LightX2V Ref2AV Turbo segments · three references · three two-card capability pages · slower narration · benchmark, NVIDIA, and SANA / SOL-ENGINE · 729 frames

Niu Lai · AI-native SOL-Attn × FastH3 HD edition

Niu Lai · SOL-Attn × FastH3 4 STEP · 16:9 · HQ

34.625 seconds · native 1344 × 768 · 831 frames · 10.25 Mbps video · FastH3 four-step source model · AI-native SOL-Attn core · B200 timings: 5s video in 2.625s, 10s in 6.839s, 15s in 12.884s · NVIDIA and SANA / SOL-ENGINE branding inside the video

Niu Lai · full 20-second pipeline comparison

LightX2V Ref2VA 4 STEP

Three still-image references · cinematic glass plaques · 486 frames · shifts 12 / 3

Fast H3 Dense I2V 4 STEP

dense-datafree · non-white frame-30 I2V · circuit-powered plaques · no VSA/VIC

Animal cast · full 20-second Ref2VA at 50 steps

Ma Lai 50 STEP

Horse + NVIDIA + SANA / SOL-ENGINE references · unobstructed cinematic reveal

Kapi Lai 50 STEP

Capybara + NVIDIA + SANA / SOL-ENGINE references · same cinematic plaque reveal