KandinskyLab GitHub Report PyPI HF Demo Diffusers vLLM SGLang FastVideo ComfyUI

Kandinsky 6.0: A family of diffusion models for Video + Audio generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).

The complete Kandinsky 6.0 video super-resolution pipeline as a single 🧨 Diffusers bundle: the 1.4B flow-matching SR DiT (4 Euler steps per tile), the KVAE video VAE and the x2 / x4 latent-upscaler bank. It upscales a video by x2, x2.25 or x4 with tiled diffusion in KVAE latent space. The Diffusers pipeline is video-only; mux the source audio back in if you need it in the output.

This repository contains all components needed for inference with from_pretrained. For the π-Flow distilled model with 2 model evaluations per tile, see the distilled Diffusers bundle.

Usage with Diffusers

import torch
from diffusers import Kandinsky6SRPipeline
from diffusers.utils import export_to_video, load_video

torch._inductor.config.max_autotune = True  # required: lets inductor pick flex-attention tiles that fit the SR block mask

pipe = Kandinsky6SRPipeline.from_pretrained("kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()
pipe.transformer.set_attention_backend("flex")
pipe.transformer.compile_repeated_blocks(fullgraph=True)

video = load_video("lq.mp4")  # -> list[PIL.Image.Image]
output = pipe(
    video=video,
    resolution_scale=2.25,  # 2, 2.25 or 4
    num_inference_steps=4,  # 4 Euler steps, the default
    generator=torch.Generator("cuda").manual_seed(1137),
)
export_to_video(output.frames[0], "lq_sr.mp4", fps=24)

Video Super Resolution Models

Component Repository
Diffusers bundle, π-Flow distilled (2 evaluations per tile) Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
Diffusers bundle, flow matching (this repository) Kandinsky-6.0-VSR-5s-Diffusers

Components

Folder Class Parameters Contents
transformer/ Kandinsky6SRTransformer3DModel 1.41B Flow-matching SR DiT (4 Euler steps per tile)
vae/ Kandinsky6SRVAE 1.74B Video VAE (included in this bundle)
latent_upscaler/ Kandinsky6SRLatentUpscalerBank (x2 + x4) 3.65B x2 and x4 latent upscalers (included in this bundle)
scheduler/ FlowMatchEulerDiscreteScheduler — stock Diffusers scheduler, shift = 5.0

model_index.json names the pipeline class Kandinsky6SRPipeline.

Downloads last month
254
Safetensors
Model size
1B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers

Finetunes
1 model

Spaces using kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers 2

Collection including kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers

Paper for kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers