Research Notes

SparkDiffusion: Mitigating the High-Sparsity Trap in Video Generation

MELON Research Group

Video DiffusionSparse AttentionDistillationEfficient Inference

TL;DR: We achieve 265× speedup on a single RTX 5090 (18 seconds for 720P diffusion generation, excluding text encoding and VAE decoding), while discovering and mitigating the "high-sparsity trap" — a counterintuitive failure mode where training loss decreases but generation quality collapses at extreme sparsity (97%).

Affiliations: Peking University, Tsinghua University, Alibaba, UESTC, HIT

Links: Project | Code | Paper


Table of Contents

  1. The Problem: Why Extreme Sparsity Matters
  2. The Discovery: High-Sparsity Trap
  3. The Diagnosis: Where Errors Accumulate
  4. The Solution: Three-Stage Framework
  5. Results: 265× on Consumer Hardware

The Problem: Why Extreme Sparsity Matters

The Computational Bottleneck

Video diffusion transformers face two fundamental costs:

  1. Multiple diffusion steps
  2. Long-sequence attention (O(L²) complexity, L = frames × spatial tokens)

For Wan2.1-T2V-14B-720P:

  • Sequence length: 81 frames at 720×1280 create a long spatiotemporal token sequence
  • Single model evaluation: ~48 seconds (RTX 5090)
  • Dense baseline: 50 sampling steps with CFG (100 model evaluations), 4769 seconds / ~79.5 minutes

Sparse attention reduces O(L²) to O(L·topk) — the most direct acceleration path.

The 97% Target

SparsityRetained Sparse BlocksSparse-Branch Attention CostTheoretical Sparse-Branch Reduction
80%20%0.2L²5×
90%10%0.1L²10×
95%5%0.05L²20×
97%3%0.03L²33×

But when we actually push to 97%, something unexpected happened...


The Discovery: High-Sparsity Trap

The Counterintuitive Phenomenon

As sparsity is pushed to 95–97%, step-local validation loss remains low or continues to improve, while terminal generation quality stagnates or degrades. Extending the same step-local training to 10,000 steps does not substantially restore terminal quality.

The core observation: At 95% and especially 97%, generated videos exhibit broken structures, distorted backgrounds, and temporal inconsistency — even after substantially extended step-local training.

Ruling Out Common Causes

We tried every standard fix:

AttemptResultConclusion
Extend training (≈250→10K steps)Validation loss continues to decrease; terminal quality does not substantially recover✗ Not simple underfitting
Use an enlarged training set (~15× production warm-up set)Terminal-quality gap persists under the controlled recipe✗ Still not enough to escape the trap
Increase low-rank capacity (rank 64→full rank 128)Terminal error essentially unchanged (Table 1)✗ Not capacity bottleneck
Switch architecture (RoLa→VSA/SLA)Same failure pattern✗ Not architecture-specific

Definition: We term this phenomenon the "High-Sparsity Trap":

At extreme attention sparsity, sparse video DiTs fall into a failure regime where step-local validation loss converges while terminal generation quality stagnates or degrades, and extending step-local training fails to substantially restore terminal quality.


The Diagnosis: Where Errors Accumulate

Key Experiment: Oracle Intervention

We designed an oracle experiment to localize the problem:

Setup: During 97% sparse student sampling, replace its velocity predictions in a contiguous noise interval with the dense teacher's predictions.

Results:

Correction WindowOutcome
Top 5 highest-noise stepsRemoves most terminal error
Equal-budget low-noise correctionLeaves error essentially unchanged

Core Finding: Correcting the high-noise window removes most terminal error, while correcting low-noise steps (with equal correction budget) has minimal effect. Structural errors injected during high-noise structure generation get amplified by later steps, ultimately collapsing terminal quality.

Why Step-Local Training Fails

Traditional Flow Matching loss optimizes the instantaneous velocity field:

L_step = E[||predicted_velocity - true_velocity||²]

The Problem: This loss doesn't directly constrain the trajectory endpoint. Per-step errors can accumulate coherently along the sampling trajectory.

An idealized surrogate analysis in Appendix C shows how coherent step-local errors can accumulate into terminal error, and why terminal-aligned supervision can correct terminal-visible components that remain at a step-local optimum.

In controlled 2D experiments:

  • 95% sparse model trained until validation loss converges → multi-step trajectories deviate visibly from data distribution near the endpoint
  • After trajectory-level terminal-aligned distillation → the few-step student moves substantially closer to the data distribution near the endpoint

The Solution: Three-Stage Framework

Core Strategy:

  1. First adapt the sparse architecture into a usable prior under extreme sparsity
  2. Then correct the terminal distribution — naturally leads to SparkDiffusion
  3. Finally quantize for deployment to realize speedup

Stage 1: Sparse Warm-up

Goal: Quickly adapt the dense pretrained model to the 97%-sparse long-sequence configuration and establish a coarse generative prior

Method: RoLa compensated sparse attention

  • Sparse branch: In the 97%-sparse long-sequence configuration, block-sparse softmax attention retains only 3% of high-energy query-key blocks
  • Compensation branch: Low-rank linear attention with rank-truncated RoPE recovers global context discarded by sparsification
  • Gated fusion: A token-wise gate initialized near zero fuses the two branches, starting close to the sparse branch and opening the compensation path only where needed

Training Strategy: We fine-tune the full DiT backbone, including the low-rank compensation modules and gating parameters, for a short warm-up using the native flow-matching objective.

Stage 2: Trajectory-Mixed Distillation

Goal: Distill Stage-1 multi-step sparse model to 3 steps while correcting terminal distribution

Key Design: Split the sampling trajectory at a noise crosspoint — high-noise segment trained with PCM-style consistency to preserve structure, motion, and diversity; low-noise segment trained with DMD-style distribution matching to sharpen detail and correct terminal-visible errors.

3-Step Crosspoint Schedule: The trajectory-mixed schedule splits at a noise crosspoint (the CrossDistill default; 0.934 in the Wan2.1-T2V-14B training configuration):

  • High-noise segment (t=1.0 → t_cross=0.934): 1 PCM consistency step
  • Low-noise segment (t_cross=0.934 → t=0): 2 DMD distribution matching steps
SegmentStepsMethodPurpose
High noise1PCM consistencyPreserve coarse structure and seed-level diversity
Low noise2DMD distribution matchingCorrect terminal-visible errors and enhance detail fidelity

Why This Split? Based on the oracle intervention: high-noise errors dominate terminal quality, so the high-noise segment stays anchored to the teacher's coarse trajectory via PCM, while the low-noise segment supplies the terminal-aligned correction via DMD.

Ablation Results (Table 4, 97% sparsity):

Stage-2 ObjectiveVBenchVBench-2.0
Full Attention (dense, 50-step)83.6960.20
PCM-only (3-step)81.9456.41
DMD-only (3-step)82.5657.38
CrossDistill (1+2 hybrid)83.1558.05

The trajectory-mixed objective combines the best of both: PCM preserves coarse structure and seed-level diversity (Table 3), while DMD corrects the terminal distribution and sharpens fine details. Together with the warm-up ablation (Fig. 10), this supports the staging principle of Sec. 3: first adapt the sparse architecture into a coarse prior, then correct the terminal distribution.

Stage 3: FP8 Quantization

Goal: Convert reduced FLOPs into wall-clock speedup

Method: W8A8 FP8 (E4M3) + fused kernels

  • Static per-channel weight scaling
  • Dynamic per-token activation scaling at runtime
  • Selective quantization: Only linear projections; LayerNorm, gating operations, and sparse mask selection remain in BF16
  • Fused kernels: The per-token activation quantization path is fused into a single forward pass; normalization layers, gating operations, and sparse mask selection remain BF16

Timing: Applied purely at deployment time, after all training (Stage 1+2) is complete in BF16. With few-step sampling, the quantization error has little room to accumulate.

Quality Impact: FP8 quantization changes VBench and VBench-2.0 by at most 0.07 points.


Results: 265× on Consumer Hardware

Main Quantitative Comparison

FastWan and TurboDiffusion are evaluated with their official released operating points; Full Attention uses the paper's strongest dense BF16 implementation. We report SparkDiffusion both at matched 90% and at 97%. At matched 90% sparsity, SparkDiffusion improves over all baselines on every reported metric on Wan2.1-T2V-14B-720P.

Wan2.1-T2V-1.3B at 480×832, 81 frames:

MethodSparsityVBenchVBench-2.0RTX 5090 (s)H100 (s)Speedup (5090/H100)
Full Attention0%83.2156.02182921×/1×
FastWan (VSA)90%82.3754.632.81.265×/77×
TurboDiffusion90%82.5254.612.01.091×/92×
SparkDiffusion90%82.6455.851.30.6140×/153×

Wan2.1-T2V-14B at 720×1280, 81 frames:

MethodSparsityVBenchVBench-2.0RTX 5090 (s)H100 (s)Speedup (5090/H100)
Full Attention0%83.6960.20476917571×/1×
FastWan (VSA)90%82.7258.0454.120.588×/86×
TurboDiffusion90%82.8857.9825.316.0188×/110×
SparkDiffusion90%83.4259.3623.710.5201×/167×
SparkDiffusion97%83.1558.0518.08.0265×/220×

Wan2.2-T2V-A14B at 720×1280, 81 frames:

MethodSparsityVBenchVBench-2.0RTX 5090 (s)Speedup (5090)
Full Attention0%84.2160.3645451×
SparkDiffusion90%83.7559.7732.4140×
SparkDiffusion97%83.3658.4625.1181×

Key Observations:

  1. At 90% sparsity, SparkDiffusion outperforms FastWan and TurboDiffusion on all metrics across both model scales
  2. Pushing to 97% on the large long-sequence models: at 97%, SparkDiffusion attains slightly better aggregate quality than the strongest 90% baselines at clearly lower latency
  3. Resolution scaling: The isolated 90%→97% sparsity stretch contributes 1.31–1.32×; this is the regime where step-local-only recipes degrade (Sec. 3)

Why Higher Resolution Benefits More

Sparse attention becomes more valuable as sequences grow longer:

  • 720P contains roughly 2.3× as many spatial tokens per frame as 480P
  • Dense attention scales as O(L²), so this translates into more than 5× the attention computation at the same frame count
  • RoLa skips the vast majority of redundant query–key interactions, making high-resolution, long-sequence generation its natural sweet spot

That is why SparkDiffusion is especially compelling at 720P: Wan2.1-T2V-14B reaches 265× acceleration, versus 140× for the compact 480P-1.3B model. These are full-system results that also reflect different model scales and sparsity levels; within the same 14B-720P configuration, pushing sparsity from 90% to 97% adds a further 1.31–1.32× speedup.

Diversity Preservation

Seed-level diversity measured by pairwise cosine and ℓ₂ distances among N=5 videos sampled from the same prompt with 5 noise seeds (1,000 prompts total):

MethodSparsityV-JEPA 2 Cos ↑V-JEPA 2 ℓ₂ ↑VideoMAE V2 Cos ↑VideoMAE V2 ℓ₂ ↑
Full Attention0%0.12527.150.02522.83
FastWan (VSA)90%0.07521.310.01172.05
TurboDiffusion90%0.07821.670.01252.21
SparkDiffusion97%0.08723.040.01422.47

Despite operating at 97% sparsity, SparkDiffusion stays closest to the dense reference, whereas the few-step baselines lose a visible fraction of seed-level variation, consistent with the mode-seeking tendency of distribution matching alone.

This supports the trajectory-mixed design of Sec. 4.2: high-noise consistency matching preserves coarse structure and diversity, while low-noise distribution matching improves terminal fidelity.

Theoretical Validation: 2D Toy Distributions

On six toy 2D sequence distributions:

  • 95% sparse model trained until validation loss converges → trajectories deviate from data distribution near endpoints
  • After trajectory-mixed distillation → 3-step student trajectories tightly wrap the true distribution

This supports the staged remedy (Sec. 3.3): short sparse warm-up adapts the architecture into a coarse prior; trajectory-mixed distillation then corrects the terminal distribution.

Cross-Model Qualitative Validation

To examine whether SparkDiffusion generalizes beyond the primary benchmark setting, we conduct qualitative comparisons across multiple model scales, tasks, resolutions, and sparsity levels:

  • Wan2.1-T2V-1.3B at 480×832 with 90% sparsity
  • Wan2.1-T2V-14B at 480×832 with 90% sparsity
  • Wan2.1-T2V-14B at 720×1280 with 97% sparsity
  • Wan2.1-I2V-14B at 720×1280 with 97% sparsity
  • Wan2.2-T2V-A14B at 720×1280 with 97% sparsity

All accelerated models use 3-step inference; TurboDiffusion and FastWan use their officially released weights and inference scripts.

SparkDiffusion maintains consistent quality across:

  • Different model scales (1.3B / 14B / A14B)
  • Different task types (T2V / I2V)
  • Different resolutions (480P / 720P)
  • Different sparsity levels (90% / 97%)

Full settings given in Appendix B.1. These results suggest that SparkDiffusion maintains coherent structure, semantic alignment, and temporal detail even when attention sparsity is pushed to 97%.


Future Work

  1. Autoregressive video diffusion: The sparse module can be restricted to a causal window, and trajectory-mixed distillation operates on the noise axis, making the framework compatible with self-forced AR training.
  2. Omni-modal generative models and autoregressive world models: These are the future directions stated in the paper's conclusion.