私的AI研究会 > ComfyUI9f

画像生成AI「ComfyUI」9(動画編6) == 編集中 ==

 「ComfyUI」を使ってローカル環境でのAI画像生成を検証する

▲ 目 次
※ 最終更新:2026/08/15 

MinMax H3 による音声付き動画生成

 2026年7月発表されたテキスト、画像、動画、音声をひとつの文脈として同時に処理できる最新の汎用マルチモーダル動画生成AIモデルの検証

概要

プロジェクトで作成するワークフロー

 このプロジェクトで作成するワークフローと関連データは下記にアップロードしている(更新されている場合は再度ダウンロードのこと)

動画生成のための環境構築

  1. 必要モデルのダウンロード と配置
    モデル名ファイル名(.safetensors)配置先(/StabilityMatrix/Data/)ダウンロード URL利用先
    diffusion_modelsminimax_h3_fl2va_pruned_int8_convrotModels/DiffusionModels/minimax_h3_fl2va_pruned_int8_convrot.safetensors
    minimax_h3_ref2va_pruned_int8_convrotminimax_h3_ref2va_pruned_int8_convrot.safetensors
    text_encodersqwen3vl_32b_minimax_h3_nvfp4_awqTextEncodersqwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
    vaeminimax_h3_video_vae_fp16VAE/minimax_h3_video_vae_fp16.safetensors
    minimax_h3_audio_vae_fp32minimax_h3_audio_vae_fp32.safetensors

Step 1:オフィシャルサイトの標準テンプレートからワークフローを作成

 ComfyUI オフィシャルサイトで公開されているテンプレートのシード値を固定して再現性を確保し動かしながら整理する
  1. ワークフローを選ぶ
    comfyui_995_m.jpg① 左端のメニューから「Template」を選択
    ②「MinMax H3」を選択する

    ・表示された一覧からワークフローを選ぶ
    ③「MinMax H3: Text to Video」
    ④「MinMax H3: Image to Video」
    ⑤「MinMax H3: Reference to Video」

    ・ワークフローでエラーが発生する場合はモデルの配置を確認する
    生成動画 864x480

  2. 動作確認を行ってから保存する
    テンプレート名モデルワークフロー保存ワークフロー名
    MinMax H3: Text to VideoINT8video_minimax_h3_t2v.jsonvideo_minimax_h3_t2v.json
    MinMax H3: Image to VideoINT8video_minimax_h3_i2v.jsonvideo_minimax_h3_i2v.json
    MinMax H3: Reference to VideoINT8video_minimax_h3_r2v.jsonvideo_minimax_h3_r2v.json
    ・オリジナルのワークフロー
     テンプレートのオリジナル・ワークフローに対して「① モデルの変更, ② シード値を固定(同じ動画を再現できるように)」を変更する
    プロンプト
    realistic live-action cinematic look, action movie trailer: practical film photography style, a post-rain dusk metropolis, anamorphic lens, shallow depth of field, film grain, city volumetric fog, flying-car traffic between the towers, restrained grading for a premium feel, powerful natural movement.

    Scene overview: at dusk on a cluster of skyscrapers, the protagonist is being chased, sprinting and leaping across rooftops, jumping from one building's roof to the next with pursuers closing in behind. This is the escape sequence of an action movie trailer: every leap is life-or-death, thrilling and fluid.

    Storyboard (each shot a separate scene, rapid cuts, all landing on the musical beats):
    [0s-1.5s] Shot 1: high side angle: the protagonist sprinting at the roof edge, pursuers appearing in the rooftop doorway behind him, wind catching his coat.
    [1s-2.5s] Shot 2: the protagonist leaps across the gap between buildings, body stretching mid-air, towers and flying-car light trails behind him, a slight slow-motion feel.
    [2.5s-4s] Shot 3: he lands, rolls and rises, low-angle shot, tower shadows and fog behind him, he keeps running.
    [4s-5s] Shot 4: freeze: the instant he hits the edge of the next roof and launches into the jump, silhouette, holding.

    Camera: each shot its own angle, cuts clean and hard, no dissolves, a slight frame jitter on the jumps.

    Audio: wind, rapid footsteps, city ambience, low score underneath, an accent hit on each leap, the score bursting at 4s, closing the last 1s.

    No text, subtitles, logos or watermarks of any kind, no animation or cartoon rendering, no overly-CG look, keep the live-action texture.
    リアルな実写映画のようなルック、アクション映画の予告編風:フィルム撮影の質感を活かしたスタイル、雨上がりの夕暮れの都市、アナモルフィックレンズ、浅い被写界深度、フィルムグレイン、都市に漂うボリュメトリックフォグ、高層ビル間を行き交う空飛ぶ車、高級感を演出する抑制の効いたカラーグレーディング、力強く自然な動き。

    シーンの概要:夕暮れ時の高層ビル群。追っ手が迫る中、主人公が屋上を全力疾走し、あるビルの屋上から別のビルへと飛び移りながら逃走するアクション映画の予告編の一場面。どのジャンプも命がけであり、スリリングかつ流麗な動きが展開される。

    絵コンテ(各ショットは独立したシーン、素早いカット割り、すべて音楽のビートに合わせる):
    [0秒-1.5秒] ショット1:高い位置からのサイド・アングル。屋上の縁を全力疾走する主人公と、背後の屋上入り口に現れる追っ手。風にたなびくコート。
    [1秒-2.5秒] ショット2:ビル間の隙間を飛び越える主人公。空中で体を伸ばす姿、背後に広がる高層ビル群と空飛ぶ車の光の軌跡。わずかにスローモーションの演出。
    [2.5秒-4秒] ショット3:着地、回転(ロール)、そして立ち上がる動作。ローアングル。背後にはビルの影と霧。そのまま走り続ける。
    [4秒-5秒] ショット4:フリーズ(静止画)。次の屋上の縁に到達し、ジャンプへと踏み切る瞬間。シルエット。静止状態で保持。

    カメラワーク:各ショットでアングルを変え、ディゾルブ(溶暗・溶明)なしの鋭いカット割り。ジャンプ時にはわずかなフレームの揺れ(ジッター)を加える。

    音声:風の音、素早い足音、都市の環境音、低音のBGM(スコア)。各ジャンプの瞬間にアクセントとなる音を入れ、4秒の時点でスコアが盛り上がり、最後の1秒で締めくくる。

    テキスト、字幕、ロゴ、透かし(ウォーターマーク)は一切なし。アニメーションやカートゥーン調のレンダリング、過度なCG感も排除し、実写の質感を維持する。
    MinMax H3: Text to Video ワークフローMinMax H3: Text to Video SubGraph
    comfyui_996_m.jpg comfyui_996a_m.jpg
    「MinMax/」video_minimax_h3_t2v.json
    ・ワークフローに含まれる注意書き
    MiniMax H3
    MiniMax H3は、MiniMaxが提供する汎用的なマルチモーダル生成モデルです。テキスト、画像、動画、音声を統合的に理解し、ステレオ音声を伴う動画を生成します。音声(セリフ、効果音、音楽)は後から重ね合わせるのではなく、単一の推論プロセス(フォワードパス)内で統合的に生成されます。出力仕様は、最大解像度2K、24fps、最大約15秒間です。
    主な入力項目
    prompt(プロンプト)ショットの内容、カメラワーク、および付随する音声(セリフ、効果音、音楽)を一つのブロックにまとめて記述します。
    width / height(幅 / 高さ)Resolution Selector(解像度選択ノード)で設定します。H3のネイティブなキャンバスサイズは短辺768pxを基準とし、上限は768x1344ピクセル、値は32の倍数に丸められます。
    duration (seconds)(長さ(秒))Math Expressionノードによって適切なフレーム数に変換されます。24fpsにおいて、モデルの仕様である「1ブロックあたり17フレーム(17k+5)」のグリッドに合わせて切り上げ処理が行われます。
    Size Settings Reference
    megapixelsAspectOutput (multiple = 32)
    0.216:9608 x 352
    0.316:9736 x 416
    0.416:9864 x 480
    0.516:9960 x 544
    0.616:91056 x 608
    0.716:91152 x 640
    0.816:91216 x 672
    0.916:91280 x 736
    0.9816:91344 x 768
    1.016:91376 x 768
    1.216:91504 x 832
    1.516:91664 x 928
    1.816:91824 x 1024
    2.016:91920 x 1088


 

更新履歴

 

参考資料