Buckets:

|
download
raw
7.82 kB

Licensed under the Apache License, Version 2.0 (the "License");

you may not use this file except in compliance with the License.

You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software

distributed under the License is distributed on an "AS IS" BASIS,

WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.

See the License for the specific language governing permissions and

limitations under the License. -->

Echo

Echo is a long-video generation model. It adds an optional clean first frame, ordered image/audio memory slots, and a stochastic few-step Distribution Matching Distillation (DMD) sampler. The pipeline generates synchronized video and audio.

Echo is implemented as a Modular Pipeline so its text encoding, memory conditioning, stochastic DMD denoising, and decoding blocks can be run as a complete workflow or composed independently.

Convert the checkpoint

Convert the BF16 Echo release checkpoint before loading it. The converter reuses the Gemma text encoder and tokenizer directly from google/gemma-3-12b-it.

python scripts/convert_echo_to_diffusers.py \
  --checkpoint /path/to/echo15_full_dmd \
  --output-path /path/to/Echo-Diffusers \
  --repo-id jdopensource/JoyAI-Echo

The Gemma repository is gated, so users must accept its license and authenticate with Hugging Face before loading the pipeline. Pass a different --base-model only when the compatible Gemma model and tokenizer are stored together at that repository or path root. --repo-id records portable Hub references for the converted Echo components; without it, the index targets the local output path.

Inference

The released model uses 241 frames in its long-video example. The video RoPE coordinates remain at the training rate of 24 fps, independently of the output container rate.

import torch
import torchaudio
from PIL import Image

from diffusers import ComponentsManager, ModularPipeline
from diffusers.utils import encode_video

model_path = "/path/to/Echo-Diffusers"
manager = ComponentsManager()
pipe = ModularPipeline.from_pretrained(model_path, components_manager=manager)
pipe.load_components(dtype={"default": torch.bfloat16, "audio_vae": torch.float32})
manager.enable_auto_cpu_offload(device="cuda")
pipe.vae.enable_tiling()

first_frame = Image.open("first_frame.png").convert("RGB")
memory_images = [Image.open(path).convert("RGB") for path in ["memory_0.png", "memory_1.png"]]
memory_audio_with_rates = [torchaudio.load(path) for path in ["memory_0.wav", "memory_1.wav"]]
memory_audio = [waveform for waveform, _ in memory_audio_with_rates]
memory_audio_rates = [sample_rate for _, sample_rate in memory_audio_with_rates]

output = pipe(
    prompt="A cinematic dialogue scene in a quiet cafe.",
    image=first_frame,
    memory_images=memory_images,
    memory_audio_waveforms=memory_audio,
    memory_audio_sample_rates=memory_audio_rates,
    width=1280,
    height=736,
    num_frames=241,
    frame_rate=25.0,
    model_frame_rate=24.0,
    generator=torch.Generator(device="cuda").manual_seed(42),
    output_type="np",
    output=["videos", "audio"],
)

encode_video(
    output["videos"][0],
    fps=25,
    audio=output["audio"][0].float().cpu(),
    audio_sample_rate=pipe.vocoder.config.output_sampling_rate,
    output_path="echo.mp4",
)

The default DMD sigma schedule is the released eight-step schedule. It predicts x0 at every step and re-noises with fresh Gaussian noise at the next sigma, so a seeded torch.Generator controls both the initial noise and all intermediate re-noising.

Raw audio-memory encoding requires torchaudio. For reference parity, keep audio_vae in FP32 as shown above. Modular workflows can cache and reuse the condition encoder's packed token outputs by running that block separately.

EchoModularPipeline[[diffusers.EchoModularPipeline]]

diffusers.EchoModularPipeline[[diffusers.EchoModularPipeline]]

diffusers.EchoModularPipeline(blocks: diffusers.modular_pipelines.modular_pipeline.ModularPipelineBlocks | None = None, pretrained_model_name_or_path: str | os.PathLike | None = None, components_manager: diffusers.modular_pipelines.components_manager.ComponentsManager | None = None, collection: str | None = None, workflow: str | None = None, modular_config_dict: dict[str, typing.Any] | None = None, config_dict: dict[str, typing.Any] | None = None, **kwargs)

Source

A ModularPipeline for the Echo long-video DMD checkpoint.

EchoBlocks[[diffusers.EchoBlocks]]

diffusers.EchoBlocks[[diffusers.EchoBlocks]]

diffusers.EchoBlocks()

Source

Echo reference-to-video generation with clean first-frame conditioning, ordered image/audio memory slots, and stochastic DMD denoising.

Components: text_encoder (PreTrainedModel) tokenizer (PreTrainedTokenizerBase) connectors (LTX2TextConnectors) vae (AutoencoderKLLTX2Video) audio_vae (AutoencoderKLLTX2Audio) transformer (LTX2VideoTransformer3DModel) video_processor (VideoProcessor) vocoder (LTX2Vocoder)

Inputs: prompt (str): The prompt or prompts to guide image generation. max_sequence_length (int, optional, defaults to 1024): Maximum sequence length for prompt encoding. image (Image | Tensor, optional): Optional single first frame used as a clean reference condition. memory_images (list, optional): Ordered reference images, one per Echo memory slot. memory_audio_waveforms (list, optional): Ordered memory waveforms as (channels, samples) tensors. Inputs longer than 9.62 seconds are cropped to their highest-response window. Use None for a silent slot. memory_audio_sample_rates (int | list, optional): Sampling rate shared by all memory waveforms, or one rate per slot. height (int, optional, defaults to 512): The height in pixels of the generated image. width (int, optional, defaults to 704): The width in pixels of the generated image. model_frame_rate (float, optional, defaults to 24.0): Training-time frame rate used for Echo video RoPE coordinates. memory_position_offset (float, optional, defaults to 500.0): Temporal center assigned to the first memory slot. memory_position_slot_stride (float, optional, defaults to 50.0): Temporal distance between consecutive memory-slot centers. num_frames (int, optional, defaults to 241): Number of generated pixel frames; must be 1 + k * vae_temporal_compression_ratio. frame_rate (float, optional, defaults to 25.0): Frame rate of the generated video. latents (Tensor, optional): Pre-generated noisy latents for image generation. audio_latents (Tensor, optional): Optional packed initial audio noise latents. generator (Generator, optional): Torch generator for deterministic generation. num_videos_per_prompt (int, optional, defaults to 1): The number of images to generate per prompt. sigmas (list | tuple): DMD sigma schedule, including the terminal zero. attention_kwargs (dict, optional): Additional kwargs for attention processors. output_type (str, optional, defaults to pil): Output format: 'pil', 'np', 'pt'. decode_timestep (None, optional, defaults to 0.0): The timestep at which the VAE decodes the final latents. decode_noise_scale (None, optional): Noise interpolation factor applied to the latents at the decode timestep.

Outputs: videos (list): The generated videos. audio (Tensor): The generated audio waveform.

Xet Storage Details

Size:
7.82 kB
·
Xet hash:
c040a73af97d78c952f638b41db20e01929eeda39d2764883b7134b8fced2dcb

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.