Buckets:

|
download
raw
7.82 kB
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License. -->
# Echo
[Echo](https://github.com/jd-opensource/JoyAI-Echo) is a long-video generation model. It adds an optional clean first
frame, ordered image/audio memory slots, and a stochastic few-step Distribution Matching Distillation (DMD) sampler.
The pipeline generates synchronized video and audio.
Echo is implemented as a Modular Pipeline so its text encoding, memory conditioning, stochastic DMD denoising, and
decoding blocks can be run as a complete workflow or composed independently.
## Convert the checkpoint
Convert the BF16 Echo release checkpoint before loading it. The converter reuses the Gemma text encoder and
tokenizer directly from `google/gemma-3-12b-it`.
```bash
python scripts/convert_echo_to_diffusers.py \
--checkpoint /path/to/echo15_full_dmd \
--output-path /path/to/Echo-Diffusers \
--repo-id jdopensource/JoyAI-Echo
```
The Gemma repository is gated, so users must accept its license and authenticate with Hugging Face before loading the
pipeline. Pass a different `--base-model` only when the compatible Gemma model and tokenizer are stored together at
that repository or path root. `--repo-id` records portable Hub references for the converted Echo components; without
it, the index targets the local output path.
## Inference
The released model uses 241 frames in its long-video example. The video RoPE coordinates remain at the training rate
of 24 fps, independently of the output container rate.
```py
import torch
import torchaudio
from PIL import Image
from diffusers import ComponentsManager, ModularPipeline
from diffusers.utils import encode_video
model_path = "/path/to/Echo-Diffusers"
manager = ComponentsManager()
pipe = ModularPipeline.from_pretrained(model_path, components_manager=manager)
pipe.load_components(dtype={"default": torch.bfloat16, "audio_vae": torch.float32})
manager.enable_auto_cpu_offload(device="cuda")
pipe.vae.enable_tiling()
first_frame = Image.open("first_frame.png").convert("RGB")
memory_images = [Image.open(path).convert("RGB") for path in ["memory_0.png", "memory_1.png"]]
memory_audio_with_rates = [torchaudio.load(path) for path in ["memory_0.wav", "memory_1.wav"]]
memory_audio = [waveform for waveform, _ in memory_audio_with_rates]
memory_audio_rates = [sample_rate for _, sample_rate in memory_audio_with_rates]
output = pipe(
prompt="A cinematic dialogue scene in a quiet cafe.",
image=first_frame,
memory_images=memory_images,
memory_audio_waveforms=memory_audio,
memory_audio_sample_rates=memory_audio_rates,
width=1280,
height=736,
num_frames=241,
frame_rate=25.0,
model_frame_rate=24.0,
generator=torch.Generator(device="cuda").manual_seed(42),
output_type="np",
output=["videos", "audio"],
)
encode_video(
output["videos"][0],
fps=25,
audio=output["audio"][0].float().cpu(),
audio_sample_rate=pipe.vocoder.config.output_sampling_rate,
output_path="echo.mp4",
)
```
The default DMD sigma schedule is the released eight-step schedule. It predicts x0 at every step and re-noises with
fresh Gaussian noise at the next sigma, so a seeded `torch.Generator` controls both the initial noise and all
intermediate re-noising.
Raw audio-memory encoding requires `torchaudio`. For reference parity, keep `audio_vae` in FP32 as shown above.
Modular workflows can cache and reuse the condition encoder's packed token outputs by running that block separately.
## EchoModularPipeline[[diffusers.EchoModularPipeline]]
#### diffusers.EchoModularPipeline[[diffusers.EchoModularPipeline]]
```python
diffusers.EchoModularPipeline(blocks: diffusers.modular_pipelines.modular_pipeline.ModularPipelineBlocks | None = None, pretrained_model_name_or_path: str | os.PathLike | None = None, components_manager: diffusers.modular_pipelines.components_manager.ComponentsManager | None = None, collection: str | None = None, workflow: str | None = None, modular_config_dict: dict[str, typing.Any] | None = None, config_dict: dict[str, typing.Any] | None = None, **kwargs)
```
[Source](https://github.com/huggingface/diffusers/blob/vr_14696/src/diffusers/modular_pipelines/echo/modular_pipeline.py#L27)
A ModularPipeline for the Echo long-video DMD checkpoint.
## EchoBlocks[[diffusers.EchoBlocks]]
#### diffusers.EchoBlocks[[diffusers.EchoBlocks]]
```python
diffusers.EchoBlocks()
```
[Source](https://github.com/huggingface/diffusers/blob/vr_14696/src/diffusers/modular_pipelines/echo/modular_blocks_echo.py#L26)
Echo reference-to-video generation with clean first-frame conditioning, ordered image/audio memory slots, and
stochastic DMD denoising.
Components:
text_encoder (`PreTrainedModel`) tokenizer (`PreTrainedTokenizerBase`) connectors (`LTX2TextConnectors`) vae
(`AutoencoderKLLTX2Video`) audio_vae (`AutoencoderKLLTX2Audio`) transformer (`LTX2VideoTransformer3DModel`)
video_processor (`VideoProcessor`) vocoder (`LTX2Vocoder`)
Inputs:
prompt (`str`):
The prompt or prompts to guide image generation.
max_sequence_length (`int`, *optional*, defaults to 1024):
Maximum sequence length for prompt encoding.
image (`Image | Tensor`, *optional*):
Optional single first frame used as a clean reference condition.
memory_images (`list`, *optional*):
Ordered reference images, one per Echo memory slot.
memory_audio_waveforms (`list`, *optional*):
Ordered memory waveforms as `(channels, samples)` tensors. Inputs longer than 9.62 seconds are cropped to
their highest-response window. Use `None` for a silent slot.
memory_audio_sample_rates (`int | list`, *optional*):
Sampling rate shared by all memory waveforms, or one rate per slot.
height (`int`, *optional*, defaults to 512):
The height in pixels of the generated image.
width (`int`, *optional*, defaults to 704):
The width in pixels of the generated image.
model_frame_rate (`float`, *optional*, defaults to 24.0):
Training-time frame rate used for Echo video RoPE coordinates.
memory_position_offset (`float`, *optional*, defaults to 500.0):
Temporal center assigned to the first memory slot.
memory_position_slot_stride (`float`, *optional*, defaults to 50.0):
Temporal distance between consecutive memory-slot centers.
num_frames (`int`, *optional*, defaults to 241):
Number of generated pixel frames; must be `1 + k * vae_temporal_compression_ratio`.
frame_rate (`float`, *optional*, defaults to 25.0):
Frame rate of the generated video.
latents (`Tensor`, *optional*):
Pre-generated noisy latents for image generation.
audio_latents (`Tensor`, *optional*):
Optional packed initial audio noise latents.
generator (`Generator`, *optional*):
Torch generator for deterministic generation.
num_videos_per_prompt (`int`, *optional*, defaults to 1):
The number of images to generate per prompt.
sigmas (`list | tuple`):
DMD sigma schedule, including the terminal zero.
attention_kwargs (`dict`, *optional*):
Additional kwargs for attention processors.
output_type (`str`, *optional*, defaults to pil):
Output format: 'pil', 'np', 'pt'.
decode_timestep (`None`, *optional*, defaults to 0.0):
The timestep at which the VAE decodes the final latents.
decode_noise_scale (`None`, *optional*):
Noise interpolation factor applied to the latents at the decode timestep.
Outputs:
videos (`list`):
The generated videos.
audio (`Tensor`):
The generated audio waveform.

Xet Storage Details

Size:
7.82 kB
·
Xet hash:
c040a73af97d78c952f638b41db20e01929eeda39d2764883b7134b8fced2dcb

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.