SparkDiffusion: Mitigating the High-Sparsity Trap in Video Generation
MELON Research Group
TL;DR: We achieve 265× speedup on a single RTX 5090 (18 seconds for 720P diffusion generation, excluding text encoding and VAE decoding), while discovering and mitigating the "high-sparsity trap" — a counterintuitive failure mode where training loss decreases but generation quality collapses at extreme sparsity (97%).
Affiliations: Peking University, Tsinghua University, Alibaba, UESTC, HIT
Table of Contents
- The Problem: Why Extreme Sparsity Matters
- The Discovery: High-Sparsity Trap
- The Diagnosis: Where Errors Accumulate
- The Solution: Three-Stage Framework
- Results: 265× on Consumer Hardware
The Problem: Why Extreme Sparsity Matters
The Computational Bottleneck
Video diffusion transformers face two fundamental costs:
- Multiple diffusion steps
- Long-sequence attention (O(L²) complexity, L = frames × spatial tokens)
For Wan2.1-T2V-14B-720P:
- Sequence length: 81 frames at 720×1280 create a long spatiotemporal token sequence
- Single model evaluation: ~48 seconds (RTX 5090)
- Dense baseline: 50 sampling steps with CFG (100 model evaluations), 4769 seconds / ~79.5 minutes
Sparse attention reduces O(L²) to O(L·topk) — the most direct acceleration path.
The 97% Target
| Sparsity | Retained Sparse Blocks | Sparse-Branch Attention Cost | Theoretical Sparse-Branch Reduction |
|---|---|---|---|
| 80% | 20% | 0.2L² | 5× |
| 90% | 10% | 0.1L² | 10× |
| 95% | 5% | 0.05L² | 20× |
| 97% | 3% | 0.03L² | 33× |
But when we actually push to 97%, something unexpected happened...
The Discovery: High-Sparsity Trap
The Counterintuitive Phenomenon
As sparsity is pushed to 95–97%, step-local validation loss remains low or continues to improve, while terminal generation quality stagnates or degrades. Extending the same step-local training to 10,000 steps does not substantially restore terminal quality.
The core observation: At 95% and especially 97%, generated videos exhibit broken structures, distorted backgrounds, and temporal inconsistency — even after substantially extended step-local training.
Ruling Out Common Causes
We tried every standard fix:
| Attempt | Result | Conclusion |
|---|---|---|
| Extend training (≈250→10K steps) | Validation loss continues to decrease; terminal quality does not substantially recover | ✗ Not simple underfitting |
| Use an enlarged training set (~15× production warm-up set) | Terminal-quality gap persists under the controlled recipe | ✗ Still not enough to escape the trap |
| Increase low-rank capacity (rank 64→full rank 128) | Terminal error essentially unchanged (Table 1) | ✗ Not capacity bottleneck |
| Switch architecture (RoLa→VSA/SLA) | Same failure pattern | ✗ Not architecture-specific |
Definition: We term this phenomenon the "High-Sparsity Trap":
At extreme attention sparsity, sparse video DiTs fall into a failure regime where step-local validation loss converges while terminal generation quality stagnates or degrades, and extending step-local training fails to substantially restore terminal quality.
The Diagnosis: Where Errors Accumulate
Key Experiment: Oracle Intervention
We designed an oracle experiment to localize the problem:
Setup: During 97% sparse student sampling, replace its velocity predictions in a contiguous noise interval with the dense teacher's predictions.
Results:
| Correction Window | Outcome |
|---|---|
| Top 5 highest-noise steps | Removes most terminal error |
| Equal-budget low-noise correction | Leaves error essentially unchanged |
Core Finding: Correcting the high-noise window removes most terminal error, while correcting low-noise steps (with equal correction budget) has minimal effect. Structural errors injected during high-noise structure generation get amplified by later steps, ultimately collapsing terminal quality.
Why Step-Local Training Fails
Traditional Flow Matching loss optimizes the instantaneous velocity field:
L_step = E[||predicted_velocity - true_velocity||²]
The Problem: This loss doesn't directly constrain the trajectory endpoint. Per-step errors can accumulate coherently along the sampling trajectory.
An idealized surrogate analysis in Appendix C shows how coherent step-local errors can accumulate into terminal error, and why terminal-aligned supervision can correct terminal-visible components that remain at a step-local optimum.
In controlled 2D experiments:
- 95% sparse model trained until validation loss converges → multi-step trajectories deviate visibly from data distribution near the endpoint
- After trajectory-level terminal-aligned distillation → the few-step student moves substantially closer to the data distribution near the endpoint
The Solution: Three-Stage Framework
Core Strategy:
- First adapt the sparse architecture into a usable prior under extreme sparsity
- Then correct the terminal distribution — naturally leads to SparkDiffusion
- Finally quantize for deployment to realize speedup
Stage 1: Sparse Warm-up
Goal: Quickly adapt the dense pretrained model to the 97%-sparse long-sequence configuration and establish a coarse generative prior
Method: RoLa compensated sparse attention
- Sparse branch: In the 97%-sparse long-sequence configuration, block-sparse softmax attention retains only 3% of high-energy query-key blocks
- Compensation branch: Low-rank linear attention with rank-truncated RoPE recovers global context discarded by sparsification
- Gated fusion: A token-wise gate initialized near zero fuses the two branches, starting close to the sparse branch and opening the compensation path only where needed
Training Strategy: We fine-tune the full DiT backbone, including the low-rank compensation modules and gating parameters, for a short warm-up using the native flow-matching objective.
Stage 2: Trajectory-Mixed Distillation
Goal: Distill Stage-1 multi-step sparse model to 3 steps while correcting terminal distribution
Key Design: Split the sampling trajectory at a noise crosspoint — high-noise segment trained with PCM-style consistency to preserve structure, motion, and diversity; low-noise segment trained with DMD-style distribution matching to sharpen detail and correct terminal-visible errors.
3-Step Crosspoint Schedule: The trajectory-mixed schedule splits at a noise crosspoint (the CrossDistill default; 0.934 in the Wan2.1-T2V-14B training configuration):
- High-noise segment (t=1.0 → t_cross=0.934): 1 PCM consistency step
- Low-noise segment (t_cross=0.934 → t=0): 2 DMD distribution matching steps
| Segment | Steps | Method | Purpose |
|---|---|---|---|
| High noise | 1 | PCM consistency | Preserve coarse structure and seed-level diversity |
| Low noise | 2 | DMD distribution matching | Correct terminal-visible errors and enhance detail fidelity |
Why This Split? Based on the oracle intervention: high-noise errors dominate terminal quality, so the high-noise segment stays anchored to the teacher's coarse trajectory via PCM, while the low-noise segment supplies the terminal-aligned correction via DMD.
Ablation Results (Table 4, 97% sparsity):
| Stage-2 Objective | VBench | VBench-2.0 |
|---|---|---|
| Full Attention (dense, 50-step) | 83.69 | 60.20 |
| PCM-only (3-step) | 81.94 | 56.41 |
| DMD-only (3-step) | 82.56 | 57.38 |
| CrossDistill (1+2 hybrid) | 83.15 | 58.05 |
The trajectory-mixed objective combines the best of both: PCM preserves coarse structure and seed-level diversity (Table 3), while DMD corrects the terminal distribution and sharpens fine details. Together with the warm-up ablation (Fig. 10), this supports the staging principle of Sec. 3: first adapt the sparse architecture into a coarse prior, then correct the terminal distribution.
Stage 3: FP8 Quantization
Goal: Convert reduced FLOPs into wall-clock speedup
Method: W8A8 FP8 (E4M3) + fused kernels
- Static per-channel weight scaling
- Dynamic per-token activation scaling at runtime
- Selective quantization: Only linear projections; LayerNorm, gating operations, and sparse mask selection remain in BF16
- Fused kernels: The per-token activation quantization path is fused into a single forward pass; normalization layers, gating operations, and sparse mask selection remain BF16
Timing: Applied purely at deployment time, after all training (Stage 1+2) is complete in BF16. With few-step sampling, the quantization error has little room to accumulate.
Quality Impact: FP8 quantization changes VBench and VBench-2.0 by at most 0.07 points.
Results: 265× on Consumer Hardware
Main Quantitative Comparison
FastWan and TurboDiffusion are evaluated with their official released operating points; Full Attention uses the paper's strongest dense BF16 implementation. We report SparkDiffusion both at matched 90% and at 97%. At matched 90% sparsity, SparkDiffusion improves over all baselines on every reported metric on Wan2.1-T2V-14B-720P.
Wan2.1-T2V-1.3B at 480×832, 81 frames:
| Method | Sparsity | VBench | VBench-2.0 | RTX 5090 (s) | H100 (s) | Speedup (5090/H100) |
|---|---|---|---|---|---|---|
| Full Attention | 0% | 83.21 | 56.02 | 182 | 92 | 1×/1× |
| FastWan (VSA) | 90% | 82.37 | 54.63 | 2.8 | 1.2 | 65×/77× |
| TurboDiffusion | 90% | 82.52 | 54.61 | 2.0 | 1.0 | 91×/92× |
| SparkDiffusion | 90% | 82.64 | 55.85 | 1.3 | 0.6 | 140×/153× |
Wan2.1-T2V-14B at 720×1280, 81 frames:
| Method | Sparsity | VBench | VBench-2.0 | RTX 5090 (s) | H100 (s) | Speedup (5090/H100) |
|---|---|---|---|---|---|---|
| Full Attention | 0% | 83.69 | 60.20 | 4769 | 1757 | 1×/1× |
| FastWan (VSA) | 90% | 82.72 | 58.04 | 54.1 | 20.5 | 88×/86× |
| TurboDiffusion | 90% | 82.88 | 57.98 | 25.3 | 16.0 | 188×/110× |
| SparkDiffusion | 90% | 83.42 | 59.36 | 23.7 | 10.5 | 201×/167× |
| SparkDiffusion | 97% | 83.15 | 58.05 | 18.0 | 8.0 | 265×/220× |
Wan2.2-T2V-A14B at 720×1280, 81 frames:
| Method | Sparsity | VBench | VBench-2.0 | RTX 5090 (s) | Speedup (5090) |
|---|---|---|---|---|---|
| Full Attention | 0% | 84.21 | 60.36 | 4545 | 1× |
| SparkDiffusion | 90% | 83.75 | 59.77 | 32.4 | 140× |
| SparkDiffusion | 97% | 83.36 | 58.46 | 25.1 | 181× |
Key Observations:
- At 90% sparsity, SparkDiffusion outperforms FastWan and TurboDiffusion on all metrics across both model scales
- Pushing to 97% on the large long-sequence models: at 97%, SparkDiffusion attains slightly better aggregate quality than the strongest 90% baselines at clearly lower latency
- Resolution scaling: The isolated 90%→97% sparsity stretch contributes 1.31–1.32×; this is the regime where step-local-only recipes degrade (Sec. 3)
Why Higher Resolution Benefits More
Sparse attention becomes more valuable as sequences grow longer:
- 720P contains roughly 2.3× as many spatial tokens per frame as 480P
- Dense attention scales as O(L²), so this translates into more than 5× the attention computation at the same frame count
- RoLa skips the vast majority of redundant query–key interactions, making high-resolution, long-sequence generation its natural sweet spot
That is why SparkDiffusion is especially compelling at 720P: Wan2.1-T2V-14B reaches 265× acceleration, versus 140× for the compact 480P-1.3B model. These are full-system results that also reflect different model scales and sparsity levels; within the same 14B-720P configuration, pushing sparsity from 90% to 97% adds a further 1.31–1.32× speedup.
Diversity Preservation
Seed-level diversity measured by pairwise cosine and ℓ₂ distances among N=5 videos sampled from the same prompt with 5 noise seeds (1,000 prompts total):
| Method | Sparsity | V-JEPA 2 Cos ↑ | V-JEPA 2 ℓ₂ ↑ | VideoMAE V2 Cos ↑ | VideoMAE V2 ℓ₂ ↑ |
|---|---|---|---|---|---|
| Full Attention | 0% | 0.125 | 27.15 | 0.0252 | 2.83 |
| FastWan (VSA) | 90% | 0.075 | 21.31 | 0.0117 | 2.05 |
| TurboDiffusion | 90% | 0.078 | 21.67 | 0.0125 | 2.21 |
| SparkDiffusion | 97% | 0.087 | 23.04 | 0.0142 | 2.47 |
Despite operating at 97% sparsity, SparkDiffusion stays closest to the dense reference, whereas the few-step baselines lose a visible fraction of seed-level variation, consistent with the mode-seeking tendency of distribution matching alone.
This supports the trajectory-mixed design of Sec. 4.2: high-noise consistency matching preserves coarse structure and diversity, while low-noise distribution matching improves terminal fidelity.
Theoretical Validation: 2D Toy Distributions
On six toy 2D sequence distributions:
- 95% sparse model trained until validation loss converges → trajectories deviate from data distribution near endpoints
- After trajectory-mixed distillation → 3-step student trajectories tightly wrap the true distribution
This supports the staged remedy (Sec. 3.3): short sparse warm-up adapts the architecture into a coarse prior; trajectory-mixed distillation then corrects the terminal distribution.
Cross-Model Qualitative Validation
To examine whether SparkDiffusion generalizes beyond the primary benchmark setting, we conduct qualitative comparisons across multiple model scales, tasks, resolutions, and sparsity levels:
- Wan2.1-T2V-1.3B at 480×832 with 90% sparsity
- Wan2.1-T2V-14B at 480×832 with 90% sparsity
- Wan2.1-T2V-14B at 720×1280 with 97% sparsity
- Wan2.1-I2V-14B at 720×1280 with 97% sparsity
- Wan2.2-T2V-A14B at 720×1280 with 97% sparsity
All accelerated models use 3-step inference; TurboDiffusion and FastWan use their officially released weights and inference scripts.
SparkDiffusion maintains consistent quality across:
- Different model scales (1.3B / 14B / A14B)
- Different task types (T2V / I2V)
- Different resolutions (480P / 720P)
- Different sparsity levels (90% / 97%)
Full settings given in Appendix B.1. These results suggest that SparkDiffusion maintains coherent structure, semantic alignment, and temporal detail even when attention sparsity is pushed to 97%.
Future Work
- Autoregressive video diffusion: The sparse module can be restricted to a causal window, and trajectory-mixed distillation operates on the noise axis, making the framework compatible with self-forced AR training.
- Omni-modal generative models and autoregressive world models: These are the future directions stated in the paper's conclusion.