YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Cache-method comparison on 4-step block-causal video DiTs

Four training-free caching methods (TeaCache, TaylorSeer, FlowCache, MotionCache) implemented on two 4-step autoregressive video base models (Self-Forcing and Causal-Forcing), swept to matched denoise-DiT speedups.

Layout

repos/                     the six upstream clones (reference implementations + base models)
  Self-Forcing/            base model 1   (guandeh17/Self-Forcing)
  Causal-Forcing/          base model 2   (thu-ml/Causal-Forcing)
  TeaCache/                ali-vilab/TeaCache
  TaylorSeer/              Shenyi-Z/TaylorSeer
  FlowCache/               mikeallen39/FlowCache        (ICLR 2026)
  MotionCache/             MAC-AutoML/MotionCache       (ICML 2026)
cachelib/                  the ports -- one implementation, both base models
  methods.py               the four cache methods + baseline + calibration probe
  selective.py             token-subset forward for the KV-cached causal DiT
  patch.py                 routes CausalWanModel._forward_inference through a method
  runner.py                chunked denoising loop with cache hooks and denoise-only timing
harness.py                 model loading, prompt loading, paired A/B measurement
calibrate.py               fits the TeaCache rescale polynomial for a base model
sweep.py                   finds the parameter that hits a target speedup
finalize.py                re-targets on the timed prompts + held-out check (authoritative)
overhead_bench.py          per-step cost of a cached vs a full step (contention-robust)
retime.py                  plain stopwatch re-measurement of the settled points
verify.py                  re-measures operating points on held-out prompts
run_compare.py             single-configuration CLI
summarize.py               collects everything into one table
results/                   JSON output, plus SUMMARY.md / SUMMARY.csv
videos/                    generated mp4s per operating point, plus *_baseline references

Pipeline order: calibrate.py -> sweep.py -> finalize.py -> overhead_bench.py -> retime.py -> summarize.py.

Weights are symlinked from the pre-existing checkouts rather than re-downloaded:

base checkpoint source
self_forcing checkpoints/self_forcing_dmd.pt (generator_ema) ../Self-Forcing
causal_forcing checkpoints/chunkwise/causal_forcing.pt (generator) ../Causal-Forcing

Both also symlink wan_models/ (Wan2.1-T2V-1.3B backbone, VAE, UMT5 text encoder).

The setting

Both base models are the same architecture: a block-causal Wan2.1-1.3B DiT that generates a video chunk at a time. With num_frame_per_block=3 and 21 latent frames, a video is 7 chunks x 4 denoising steps = 28 DiT forwards, plus one untimed KV-cache refresh pass per chunk (context_noise) that rewrites the chunk's KV entries from the clean latent.

The experiment fixes the schedule to F ? ? F: step 0 and step 3 always run the full DiT, steps 1 and 2 are the ones a cache method may skip. This is equivalent to the upstream ret_steps=1 / cutoff_steps=num_steps-1 guards.

Speedup is measured on denoise-DiT time only β€” the sum of the 28 denoising forwards, timed with CUDA events. Text encoding, VAE decode and the per-chunk KV-refresh pass are excluded. That makes 2.0x the arithmetic ceiling (14 of 28 forwards), which is why the 2.0x target is exactly "skip both middle steps everywhere".

How each method was ported

All four hook the same place: the 30-block loop inside CausalWanModel._forward_inference. The DiT preamble and the head/unpatchify tail always run.

method decision granularity what a skipped step reuses knob
TeaCache whole forward the 30-block residual x_out - x_in from the last computed step thresh
FlowCache per frame group inside the chunk the same residual, banked per group thresh, group_size
TaylorSeer whole forward a Taylor forecast of each block's self-attn / cross-attn / FFN output interval, max_order
MotionCache per token, motion-weighted cached rows for un-selected tokens; selected tokens are recomputed thresh, weight_norm

Skipping the block stack leaves the chunk's KV-cache entries holding the previous computed step's keys/values. That is safe here: every later step of the chunk overwrites the same slots, and the per-chunk KV-refresh pass rewrites them from the clean latent before the next chunk reads them.

Selective (partial-token) forward

FlowCache and MotionCache recompute only part of the chunk. On a KV-cached causal model that means, per self-attention layer: compute q/k/v for the selected rows only, scatter their fresh k/v into the slots the chunk owns (leaving unselected slots holding the previous step's k/v), and attend the selected queries against the full cache. This mirrors MotionCache's forward_selective, moved from SkyReels' persistent buffers onto Self-Forcing's rolling kv_cache.

Validated: with all tokens selected the selective path is bit-exact against the stock full forward (max abs latent difference 0.0). The same check passes for TaylorSeer's recording path at interval=1.

Deviations from upstream, and why

  1. Indicator. TeaCache4Wan2.1 derives its indicator from e0 (the timestep modulation) alone. That is fine for bidirectional Wan, where one forward covers the whole video at a single timestep, but it degenerates here: every chunk runs the same four timesteps, so an e0 indicator is bit-identical across chunks and the threshold cannot adapt to content at all. The default is therefore TeaCache's original definition β€” the relative L1 distance of the timestep-modulated noisy input, the tensor block 0 feeds to its self-attention, which is what TeaCache uses on HunyuanVideo and FLUX. --indicator e0 restores the literal Wan port.

  2. Rescale polynomial. TeaCache's shipped 4th-degree coefficients were fitted for a different model, schedule and indicator, so calibrate.py reruns TeaCache's own fitting procedure here. The fit is degree 1, not 4: on this model the input distance only spans ~[1.19, 1.36], where a degree-4 fit is ill-conditioned (coefficients ~10^3 with alternating signs) and MotionCache, which evaluates the polynomial per token, would extrapolate it far outside the fitted domain. Degree 1 is monotone and safe. All degrees are recorded in the coefficient JSON.

  3. FlowCache granularity. FlowCache's contribution is that the units covered by one forward denoise at different rates and need independent policies. In SkyReels-V2 those units are the chunks of the diffusion-forcing window; here a forward covers exactly one chunk, so the corresponding units are the frame groups inside it (group_size=1, i.e. 3 groups). Its KV-cache-compression component is not ported: Self-Forcing already bounds its KV cache by local attention, so there is no growing cache to compress.

  4. TaylorSeer fractional interval. TaylorSeer's schedule is a uniform integer refresh interval, which on an F ? ? F schedule can only produce 1.0x, 1.33x and 2.0x. interval is accepted as a float and realised as a deterministic per-chunk alternation between the two bracketing integer intervals (Bresenham-style), giving the requested average interval. This keeps TaylorSeer's content-independent schedule rather than bolting a threshold onto it.

  5. Fused forecast. TaylorSeer's cached step is ~90 bandwidth-bound elementwise ops per forward. Upstream puts @torch.compile on the equivalent wan_attention_cache_forward; the port does the same. Without it the forecast ate most of the saving (1.72x measured at a 2.0x FLOPs setting); with it, it does not. Set CACHELIB_NO_COMPILE=1 to disable.

Scope change (Sep 4)

FlowCache is frozen: its finished operating points stay in the tables, nothing new is generated for it. The continuing study is three methods (TeaCache, TaylorSeer, MotionCache) x three base models (Self-Forcing, Causal-Forcing, HY-WorldPlay) x three points: FxxF at ~1.3x, FxxF at 1.7-1.8x, and Fxxx at 2.7-3.0x -- all targeted on measured wall-clock speedup, not on the compute fraction. The reason is MotionCache: matching compute at 2.0x under FxxF forces its per-token selection to pick zero tokens, which makes it bit-identical to TeaCache and reproduces nothing. Operating points are now chosen as the fastest setting that still keeps the token-level decision active (>=25% of cacheable steps partial, >=5% of tokens selected on those steps); the activity statistics are recorded per video in cache_diagnostics.

Schedules beyond FxxF

Every entry point takes --schedule (sweep.py, finalize.py, run_compare.py; the evaluation reads it from the final_*.json rows). F marks a step that always runs the full DiT, x one the method may serve from cache: FxxF is the study above and the default; FFxx gives TaylorSeer two computed points before its first forecast so the first-order term is actually exercised; Fxxx lifts the cool-down guard and raises the arithmetic ceiling to 4x. Operating points for FFxx (TaylorSeer, 1.3/1.6/2.0x) and Fxxx (the other three, ~3x) were searched on 3 prompts (results/sweep_*_{FFxx,Fxxx}.json, collected in results/final_<base>_newsched.json) and evaluated on Extended-251.

All results, speed and quality, are collected in Experiment.md (regenerate with python experiment_md.py). The headline of the second round: caching the last denoising step -- which every Fxxx point and every FFxx point at >= 1.6x does -- costs 15-25 VBench points on both base models, while FxxF at the same compute costs 1-2. On a 4-step distilled model the final step is not optional.

Measurement protocol

This is a shared node and it was busy throughout. Other tenants' jobs move absolute timings by 30%+ and drift them within a single sweep, so a baseline measured once at the start silently inflates every later configuration. Four things follow, and they are the difference between numbers worth reading and numbers that are noise:

  1. Search on compute, not on time. The parameter search runs on the compute fraction (how many of the 28 DiT forwards were actually computed), which is exactly deterministic given the prompt set and immune to contention. Targets are aimed at the compute budget directly, which also puts all four methods on identical compute β€” the right basis for a later quality comparison. Correcting the target by a measured overhead ratio was tried and abandoned: that ratio is itself a timing measurement, one probe returned 1.110 (physically impossible), and it dragged a whole method's operating points below target.

  2. Paired A/B. Baseline and method run back to back on the same prompt, so drift between them is seconds rather than minutes.

  3. Minimum, not median, as the headline. Contention is one-sided β€” a neighbour can only make a run slower β€” so the fastest observed run is the best estimate of the uncontended time, and the headline speedup is min(baseline) / min(method) over the pairs. The paired median is reported alongside but is biased upward, because the baseline run is longer and so more exposed to being hit by a spike; that is why it sometimes exceeded a method's arithmetic ceiling.

  4. Compute fraction from the timed runs. A threshold tuned on one prompt subset and timed on another silently disagrees whenever it sits near a decision boundary. finalize.py reports the compute fraction of the very runs it timed, and re-tunes on those prompts if it misses.

finalize.py also re-checks every point on a disjoint prompt slice, so a threshold that only hit its target by sitting inside a narrow band of the indicator distribution shows up as held-out drift. verify.py does the same check standalone against a sweep file.

Reproducing

PY=/local/zoubin/cz/envs/self_forcing/bin/python

# 1. fit the rescale polynomial (once per base model)
CUDA_VISIBLE_DEVICES=0 $PY calibrate.py --base self_forcing \
    --num-prompts 12 --out results/coeff_self_forcing.json

# 2. sweep one method to the three target speedups
CUDA_VISIBLE_DEVICES=0 $PY sweep.py --base self_forcing --method teacache \
    --coefficients results/coeff_self_forcing.json \
    --targets 1.3,1.6,2.0 --save-video-dir videos \
    --out results/sweep_self_forcing_teacache.json

# 3. check the settings transfer to unseen prompts
CUDA_VISIBLE_DEVICES=0 $PY verify.py \
    --sweep results/sweep_self_forcing_teacache.json --prompt-offset 64 \
    --out results/verify_self_forcing_teacache.json

# 4. collect everything
$PY summarize.py

run_compare.py runs a single configuration if you just want one number:

CUDA_VISIBLE_DEVICES=0 $PY run_compare.py --base causal_forcing \
    --method taylorseer --interval 2.5 --num-prompts 8

Environment note

Sep 3: the machine was restored again -- /mnt/local_nvme became /local, both environments lost their executable bits a second time (1632 files restored), the vbench_eval venv's pyvenv.cfg/interpreter links pointed at the old path (now symlinks into /local/.../self_forcing), and the caches moved to /local/zoubin/cz/.cache/. Paths in eval/*.sh, eval/aggregate.py and the per-prompt records were updated accordingly. The node is shared: other tenants' launches have repeatedly killed GPU processes of ours, so long runs go through eval/retime_until_clean.sh-style loops that only use quiet GPUs and resume.

The conda environment at /mnt/local_nvme/zoubin/cz/envs/self_forcing had lost every executable bit (a bad restore left all files rw-rw-r--), so nothing in it could run β€” including Triton's bundled ptxas, which broke torch.compile. The exec bits were restored on files that are ELF binaries or start with #! (1772 files). No package contents were changed.

Results

Full tables: results/SUMMARY.md (regenerate with python summarize.py), machine readable in results/SUMMARY.csv and the per-stage JSONs.

Headline speedup is the denoise-DiT speedup implied by the measured compute fraction and the measured per-step cost model (see Measurement protocol). The stopwatch column is the plain paired whole-video ratio; it agrees within its (large) spread but is not the number to quote from this node.

Operating points found

base method knob 1.3x 1.6x 2.0x
Self-Forcing TeaCache thresh 0.6348 β†’ 1.30x 1.2507 β†’ 1.48x 1.2917 β†’ 1.96x
Self-Forcing FlowCache thresh 0.6360 β†’ 1.30x 1.2500 β†’ 1.51x 1.2917 β†’ 1.96x
Self-Forcing TaylorSeer interval 2.083 β†’ 1.31x 2.583 β†’ 1.59x 2.917 β†’ 1.89x
Self-Forcing MotionCache thresh 0.7560 β†’ 1.29x 1.2461 β†’ 1.60x 2.5000 β†’ 1.96x
Causal-Forcing TeaCache thresh 0.6491 β†’ 1.30x 1.2760 β†’ 1.51x 1.2917 β†’ 1.96x
Causal-Forcing FlowCache thresh 0.6488 β†’ 1.26x 1.2760 β†’ 1.61x 1.2917 β†’ 1.95x
Causal-Forcing TaylorSeer interval 2.083 β†’ 1.31x 2.583 β†’ 1.59x 2.917 β†’ 1.89x
Causal-Forcing MotionCache thresh 0.7695 β†’ 1.30x 1.2647 β†’ 1.57x 2.8750 β†’ 1.96x

What a cached step actually costs

Minima over ~1900 per-step samples per base, so these are solid:

method cached step, as a fraction of a full step ceiling at 14/28 forwards
TeaCache 1.9% 1.96x
MotionCache 2.0% 1.96x
FlowCache 2.3% 1.96x
TaylorSeer 5.6 - 5.8% 1.89x

A full denoising forward costs 76-77 ms uncontended (1 chunk = 3 latent frames, 4680 tokens, Wan2.1-1.3B, H100). 2.0x is the arithmetic ceiling of this schedule, and the achievable ceiling is ~1.96x for three of the methods. TaylorSeer pays more because its cached step is a Taylor forecast, not a copy: 30 layers x 3 module features have to be extrapolated and re-modulated. That is already with the torch.compile fusion upstream uses; without it the same step costs ~25% of a full one and the method tops out near 1.7x.

Three findings

  1. The TeaCache-family indicator carries almost no signal on a 4-step distilled model. Calibration over 252 (chunk, step) pairs: at a given denoising step the indicator's spread across chunks is a standard deviation of 0.0013 on a mean of 1.35 (0.1%), while the gaps between steps are 0.04-0.10. So the threshold is really choosing which step index to skip, not adapting per chunk. Worse, within a step the indicator is negatively correlated with the actual reuse error (r = -0.54 / -0.58 at steps 1 and 2); the +0.49 overall correlation is entirely a between-step effect, and a degree-4 fit only reaches R^2 = 0.26.

  2. That makes 1.6x hard to hit for the threshold methods and easy for the other two. TeaCache's reachable points on Self-Forcing are essentially quantised to {1.0, 1.33, ~1.48, 2.0} β€” the requested 1.6x falls in a gap, and the closest honest point is 1.48x. FlowCache, deciding per frame group, does slightly better (1.51x on Self-Forcing, 1.61x on Causal-Forcing). TaylorSeer (fractional interval) and MotionCache (per token) land on 1.59x / 1.60x directly, because both have a genuinely continuous knob.

  3. Generalisation splits the methods the same way. Re-running each setting on 64 prompts it was never tuned on, the compute fraction drifts by:

    method worst held-out drift
    TaylorSeer 0.000
    MotionCache 0.008
    TeaCache 0.079
    FlowCache 0.070

    TeaCache's and FlowCache's 1.3x and 1.6x settings sit inside a razor-thin band of the indicator distribution, so they do not transfer; their 2.0x settings (where everything skips regardless) transfer perfectly. TaylorSeer's schedule is content-independent by construction, and MotionCache's per-token thresholding averages over 4680 decisions per step instead of one.

Artifacts for a quality comparison

videos/<base>_<method>_x<target>/NNNN.mp4 holds 7-10 clips per operating point, and videos/<base>_baseline/NNNN.mp4 the matching all-full references (same prompts, same seed, indices align). The quality numbers are in Experiment.md (Extended-251 evaluation, next section).


Self-Forcing Extended-251 Full Evaluation

VBench-8 + FFFF-relative pixel metrics + matched-prompt policy latency, following self_forcing_extended_full_evaluation_protocol.md, with the protocol's paths remapped to this machine.

Path remapping

Protocol path Here
/data3/chenzhuo/anaconda3/envs/self_forcing /mnt/local_nvme/zoubin/cz/envs/self_forcing
$SELF_FORCING_REPO/.evaluation_env/vbench_site2 absent -> /mnt/local_nvme/zoubin/cz/projects/VBench on PYTHONPATH
/data3/chenzhuo/.cache/vbench /mnt/local_nvme/zoubin/cache/vbench
prompts/vbench/all_dimension{,_extended}.txt ../Self-Forcing/prompts/vbench/
assets/vbench8_extended_subset_mapping.json absent -> rebuilt by eval/build_mapping.py

Files

assets/vbench8_extended_subset_mapping.json   the 251-row mapping (built + validated here)
eval/build_mapping.py      946 -> 251 selection, with the protocol's validation gate
eval/strategies.py         FFFF + the 12 settled operating points, per base model
eval/generate_eval.py      video generation, latency split, pixel metrics
eval/pixel_metrics.py      PSNR / SSIM / LPIPS on decoded frames, pre-encode
eval/run_vbench.py         the eight dimensions, stock VBench metric code
eval/aggregate.py          normalize -> Quality/Semantic/Selected, latency, final tables
eval/test_protocol.py      the section 12 unit tests
eval/run_generation.sh / run_vbench_all.sh / run_all_phases.sh   8-GPU drivers
eval_out/                  generated_videos/ per_prompt/ vbench/ summaries/

Strategies

26 = 2 base models x (FFFF + 4 cache methods x 3 speed targets). FFFF is the all-full 4-step baseline and is each base model's own reference for both the pixel metrics and the latency speedup. Cache parameters are read back from results/final_<base>*.json, so the evaluation scores exactly the operating points the speed study settled on.

Two environments, on purpose

Generation runs in the base self_forcing env -- the same interpreter and package set the speed study used, so the evaluated videos come from exactly the benchmarked stack. The only addition to that env is lpips.

VBench runs in the vbench_eval venv. It has to: the base env carries torchao 0.17.0, which is unimportable under torch 2.5.1 (it references torch.int1), and transformers.modeling_utils pulls torchao in through its quantizer registry -- so BertModel, and therefore Tag2Text and the scene dimension, cannot load there. The venv pins torchao==0.7.0, transformers==4.46.3 and adds fairscale. This is a deviation from protocol section 2.1 (one environment for everything); it does not touch generation.

The venv itself needed repair first: a bad restore had flattened its bin/python symlinks into text files and stripped every executable bit.

Protocol conformance

Verified before the run:

  • mapping is 251 rows = 72 / 93 / 86, global_index unique, suite_index contiguous, all 86 scene rows keep auxiliary_info, and the 946 short prompts match VBench_full_info.json in order exactly (selection is by index -- two prompt texts are duplicated, so text matching is unsafe and is forbidden);
  • SHA256 recorded for the short prompts, extended prompts, VBench_full_info.json and the mapping;
  • videos are 81 frames, 480x832, 16 FPS, seed 0, one sample per prompt;
  • prompt_en handed to VBench is the extended prompt actually generated from;
  • each dimension scores only its own suite (72 / 93 / 86), driven by the mapping;
  • policy_latency_ms covers the denoise path only; the context/KV-cache DiT is timed separately into excluded_context_kv_latency_ms (~790 ms/video) and never added in;
  • pixel metrics run on the decoded RGB tensors before MP4 encoding, over all 81 frames, with FFFF-vs-itself pinned to 120 dB / 1.0 / 0.0;
  • eval/test_protocol.py covers the section 12 checklist and passes.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support