私的AI研究会 > ComfyUI9h
「ComfyUI」を使ってローカル環境でのAI画像生成を検証する
| 2026年3月発表された音声対応の動画生成モデル。 |
| このプロジェクトで作成するワークフローと関連データは下記にアップロードしている(更新されている場合は再度ダウンロードのこと) |
📂ComfyUI ├─📂input ← ワークフローに含まれる入力画像 └─📂user └─📂default └─📂workflows ← ワークフローの保存場所 : ├─📂LTX ├─📂LTX-2.5 ← この章で作成するワークフロー :・解凍してできる「ComfyUI/」フォルダを「StabilityMatrix/Data/Packages/ComfyUI」へ上書きコピーする
| 拡張ノード(検索名) | 拡張ノード URL | 主な機能 / 参照ページ | 特記事項 |
| rgthree | rgthree-comfy | ノードをスイッチで切り替える | すべてのノードに必須 |
| ComfyUI_Custom_Nodes_AlekPet | ComfyUI_Custom_Nodes_AlekPet | プロンプトを日本語入力する |
| モデル名 | ファイル名(.safetensors) | 配置先(/StabilityMatrix/Data/) | ダウンロード URL | サイズ | |
| diffusion_models | ltx-2.5-22b-distilled-transformer-comfy-int8-convrot | Models/ | DiffusionModels/ | ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors | 21.5GB |
| text_encoders | gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot | TextEncoders/ | gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors | 15.4GB | |
| gemma4_e2b_it_bf16 | gemma4_e2b_it_bf16.safetensors | 10.3GB | |||
| vae | ltx-2.5-video-vae-bf16 | VAE/ | ltx-2.5-video-vae-bf16.safetensors | 1.47GB | |
| ltx-2.5-audio-vae-bf16.safetensors | ltx-2.5-audio-vae-bf16.safetensors.safetensors | 0.37GB | |||
| latent_upscale_models | ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0 | latent_upscale_models/ | ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors | 1.00GB | |
| ComfyUI オフィシャルサイトで公開されているテンプレートの動作を確認する |
![]() | ① 左端のメニューから「Template」を選択 ②「LTX-2.5」を選択する ・表示された一覧からワークフローを選ぶ ③「LTX-2.5: Text to Video」 ④「LTX-2.5: Image to Video」 ⑤「LTX-2.5: FLF2V」 ・ワークフローでエラーが発生する場合はモデルの配置を確認する | |
| テンプレート名 | 保存ワークフロー名 | |
| ③ LTX-2.5: Text to Video | video_ltx2_5_t2v.json | |
| ④ LTX-2.5: Image to Video | video_ltx2_5_i2v.json | |
| ⑤ LTX-2.5: Reference to Video | video_ltx2_5_flf2v.json | |
| プロンプト |
| A close-up of an Arctic hunter's face, his eyes fixed straight ahead, narrowed and locked on something in the far distance directly in front of him, frost dusting his dark beard, one hand slowly reaching toward the rifle slung on his back, his breath quick and shallow, the tension in his jaw visible even through the scarf. The camera slowly pulls back, revealing him crouched at the edge of an ice floe, and ahead of him, in the exact direction of his gaze, a polar bear moves slowly along a distant ice ridge, its pale coat blending into the white haze, facing toward him. The pull-back continues, the hunter in the foreground and the bear far ahead, the two facing each other across the dark lead of water, the ice field and grey sea stretching around them, seabirds wheeling overhead. Midway through the shot, the complete title "LTX-2.5" fades in large and centered, plain white sans-serif letters with wide spacing, no glow, no shadow, no ornament, the lettering large enough to dominate the center of the frame, emerging slowly from the mist, resting still and blending into the snowscape, with the hunter and the bear still visible below on either side of the title, staying through the end of the scene. The wind carries the distant sound of shifting ice, a single low growl rolling across the water. |
| 北極の狩人の顔のクローズアップ。視線は真っ直ぐ前方に固定され、遠くの何かを捉えようと細められている。黒い髭には霜が降り、背負ったライフルへ片手がゆっくりと伸びる。呼吸は速く浅く、スカーフ越しにも顎の緊張がうかがえる。カメラがゆっくりと引くと、彼が流氷の縁にしゃがみ込んでいる姿が現れる。その視線の先、遠く離れた氷の隆起に沿ってホッキョクグマがゆっくりと動いており、その淡い色の毛皮は白い霞に溶け込みながらも、彼の方を向いているのがわかる。さらにカメラが引いていくと、手前の狩人と遥か前方の熊が、暗い海面(リード)を挟んで対峙する構図となる。周囲には氷原と灰色の海が広がり、頭上では海鳥が旋回している。ショットの途中で、タイトル「LTX-2.5」が画面中央に大きくフェードインしてくる。文字はシンプルな白のサンセリフ体で、字間は広く取られ、発光や影、装飾などの加工は一切ない。フレームの中央を占めるほどの大きさで、霧の中からゆっくりと浮かび上がり、雪景色に溶け込むように静止する。タイトルの下、左右には狩人と熊の姿が見えたままで、その状態がシーンの終わりまで続く。風が氷の動く遠くの音を運び、低く響く唸り声が水面を伝わってくる。 |
| プロンプト・エンハンサーにより内部で生成されるプロンプト |
| A close-up shot frames an Arctic hunter's face, captured from a front-facing angle as the camera slowly pulls back, revealing frost dusting his dark beard, his eyes fixed straight ahead, narrowed and locked on something in the far distance directly in front of him, one hand slowly reaching toward a rifle slung on his back, his breath quick and shallow, the tension in his jaw visible even through a thick scarf; simultaneously, the camera pulls back further to reveal the hunter crouched at the edge of a fractured ice floe, and ahead of him, in the exact direction of his gaze, a polar bear moves slowly along a distant ice ridge, its pale coat blending into the white haze while facing toward the hunter; the pull-back continues, placing the hunter in the foreground and the bear far ahead, the two facing each other across the dark lead of water, the vast ice field and grey sea stretching around them under cold, diffused light, seabirds wheeling overhead; midway through the shot, the complete title "LTX-2.5" fades in large and centered, rendered in plain white sans-serif letters with wide spacing, showing no glow, no shadow, and no ornament, the lettering large enough to dominate the center of the frame, emerging slowly from the mist and resting still, blending into the snowscape while the hunter and the bear remain visible below on either side of the title throughout the duration; the wind carries the distant sound of shifting ice, accompanied by a single low growl rolling across the water, creating a tense, frigid atmosphere with crisp high-resolution detail and richly saturated film-grade color. |
| 北極の狩人の顔を正面から捉えたクローズアップから始まり、カメラがゆっくりと引いていくと、彼の黒い髭に霜が降りている様子や、正面の遥か彼方にある何かを凝視して細められた鋭い眼差し、背負ったライフルへゆっくりと伸びる片手、荒く浅い呼吸、そして厚手のスカーフ越しにも見て取れる食いしばった顎の緊張感が明らかになります。さらにカメラが引くと、割れた流氷の縁にしゃがみ込む狩人の姿と、その視線の先、遥か彼方の氷の隆起に沿ってゆっくりと動くホッキョクグマの姿が現れます。白い霞に溶け込むような淡い毛色をしたその熊は、狩人の方を向いています。カメラはさらに引き続け、手前に狩人、遥か前方に熊を配置します。二者は暗い海面(リード)を挟んで対峙し、冷たく柔らかな光の下、広大な氷原と灰色の海が周囲に広がり、頭上では海鳥が旋回しています。映像の途中で、タイトル「LTX-2.5」が画面中央に大きくフェードインして現れます。装飾のないシンプルな白いサンセリフ体で、文字間隔は広く取られ、発光や影などの装飾は一切ありません。フレームの中央を占めるほどの大きさで、霧の中からゆっくりと浮かび上がり、雪景色に溶け込むように静止します。その間も、タイトルの下、左右に狩人と熊の姿が見え続けています。風が氷の動く音を運び、水面を伝う低く響く唸り声がそれに重なります。鮮明な高解像度のディテールと、豊かな彩度を持つ映画品質の色調が、緊張感に満ちた極寒の雰囲気を醸し出しています。 |
| LTX-2.5 | |
| LTX-2.5は、Lightricksが提供するオープンな動画生成モデルです。高速かつローカル環境での動作を優先し、実運用(プロダクション)に耐えうる設計となっています。ローカル環境での実行、特定のユースケースに合わせたファインチューニング、あるいは既存のパイプラインへの組み込みが可能です。 | |
| 主な特徴 | |
| Pixel Diffusion | キーフレームを先行して生成する手法を採用。高精細なキーフレームのグリッドを基にシーンを構築し、業界最高水準の画質(ピクセル品質)を実現します。 |
| Diffusion Video Decoder | 顔の鮮明さ向上、テキストの視認性確保、高速な動きにおけるブレ(スミア)の低減を実現します。 |
| ネイティブ・マルチショット | 1回の生成で、キャラクター、環境、照明、音声、スタイルをカット間で維持したまま、複数の連続したショットを生成します。 |
| カスタムGemma 4 12Bテキストエンコーダー | 複雑なプロンプトにおいても、複数の被写体、アクション、照明、細部、カメラワークの指定を正確に反映・維持します。 |
| 専用プロンプトエンハンサー | 軽量なモデルにより、短いプロンプトを、追加の計算コストをほぼかけずに、映画のような表現豊かな指示へと拡張します。 |
| 自動デュレーション(Auto Duration) | 拡散(ディフュージョン)処理の開始前に、記述されたアクションに基づいて適切なクリップの長さを予測します。 |
| 改良された蒸留モデル | 小型・高速化したモデルでありながら、画質、プロンプトへの忠実度、動きの質を大幅に維持・向上させています。 |
| LTX-2.3からの継承機能 | ネイティブ4K対応、音声と映像の同期。 |
| このワークフローの機能 | |
| この Text-to-Video(テキストから動画への変換) ワークフローは、テキストプロンプトから直接動画を生成します。プロンプトを入力し、必要に応じてプロンプトエンハンサーを有効にすると、モデルは映画のような画質と同期された音声を備えたクリップを生成します。 | |
| サブグラフのパラメータ | |
| Text to Video (LTX-2.5) ノードはサブグラフで構成されています。ノード下部にある Enter subgraph ボタンをクリックして内部グラフを開くと、パイプラインの各ステップ(コンディショニング、サンプラー、CFG ガイダー、VAE デコードなど)を確認・微調整できます。 | |
| prompt | シーン、アクション、照明、カメラワーク、スタイルを記述するテキストプロンプト |
| prompt_enhance | 内蔵プロンプトエンハンサーの切り替え:短いプロンプトを、追加の計算コストをほぼかけずに、映画のような表現豊かな指示へと拡張します |
| duration | 秒単位のクリップの長さ Clip_length (「自動デュレーション」機能は、アクションに基づいて長さを予測できます) |
| width / height | 出力解像度(ピクセル単位、例: 1280×720) |
| seed | 乱数シード:固定値を設定すると同じ結果を再現可能 |
| frame_rate | フレームレート(fps) |
| unet_name | メインモデル:LTX-2.5 蒸留済みTransformer (ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors) |
| video_vae | デコードに使用するビデオVAE (ltx-2.5-video-vae-bf16.safetensors) |
| audio_vae | 同期された音声を生成するオーディオVAE (ltx-2.5-audio-vae-bf16.safetensors) |
| clip_name | テキストエンコーダー:プロジェクション付きカスタムGemma 4 12B (gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors) |
| upscale_model | 高品質な出力のためにデコード前に適用されるLatent Upscaler(潜在空間アップスケーラー) |
| prompt_enhance_model | Prompt Enhancer(プロンプト拡張機能)で使用されるモデルチェックポイント |
| 使用方法 | |
| 1. モデルをダウンロードし、ComfyUIの `models/` フォルダに配置します。 2. プロンプトを入力します。短いプロンプトの場合は、必要に応じて prompt_enhance を有効にしてください。 3. duration / width / height / frame_rate を設定し(デフォルト値で問題なく動作します)、Queue をクリックします。 4. 詳細な設定を行うには、メインノード下部にある Enter subgraph をクリックします。ここから、内部の各ステップ(条件付け、サンプラー、CFG、VAE デコードなど)を編集できます。 | |
| 生成動画 1280x704 | Size Settings Reference | ||
| megapixels | Aspect | Output (multiple = 32) | |
| 0.2 | 16:9 | 608 x 352 | |
| 0.3 | 16:9 | 736 x 416 | |
| 0.4 | 16:9 | 864 x 480 | |
| 0.5 | 16:9 | 960 x 544 | |
| 0.6 | 16:9 | 1056 x 608 | |
| 0.7 | 16:9 | 1152 x 640 | |
| 0.8 | 16:9 | 1216 x 672 | |
| 0.9 | 16:9 | 1280 x 736 | |
| 0.98 | 16:9 | 1344 x 768 | |
| 1.0 | 16:9 | 1376 x 768 | |
| 1.2 | 16:9 | 1504 x 832 | |
| 1.5 | 16:9 | 1664 x 928 | |
| 1.8 | 16:9 | 1824 x 1024 | |
| 2.0 | 16:9 | 1920 x 1088 | |