From d6b4a2fe607f99a02a604fa581e3c7b98ea04472 Mon Sep 17 00:00:00 2001 From: stevhliu Date: Wed, 30 Sep 2026 15:44:32 -0700 Subject: [PATCH 1/2] freeu --- docs/source/en/_toctree.yml | 12 +- docs/source/en/api/pipelines/kandinsky.md | 2 +- .../stable_diffusion/stable_diffusion_xl.md | 2 +- docs/source/en/optimization/deepcache.md | 62 ------ docs/source/en/optimization/tgate.md | 182 ------------------ docs/source/en/optimization/tome.md | 96 --------- docs/source/en/optimization/xformers.md | 29 --- docs/source/en/training/controlnet.md | 2 +- docs/source/en/training/dreambooth.md | 2 +- docs/source/en/training/kandinsky.md | 2 +- docs/source/en/training/lcm_distill.md | 2 +- docs/source/en/training/overview.md | 2 +- docs/source/en/training/sdxl.md | 2 +- docs/source/en/training/text2image.md | 2 +- docs/source/en/training/text_inversion.md | 2 +- .../en/using-diffusers/image_quality.md | 102 +--------- docs/source/en/using-diffusers/img2img.md | 2 +- docs/source/en/using-diffusers/inpaint.md | 2 +- 18 files changed, 24 insertions(+), 483 deletions(-) delete mode 100644 docs/source/en/optimization/deepcache.md delete mode 100644 docs/source/en/optimization/tgate.md delete mode 100644 docs/source/en/optimization/tome.md delete mode 100644 docs/source/en/optimization/xformers.md diff --git a/docs/source/en/_toctree.yml b/docs/source/en/_toctree.yml index 386dbbe302f4..344b5733b74e 100644 --- a/docs/source/en/_toctree.yml +++ b/docs/source/en/_toctree.yml @@ -24,6 +24,8 @@ title: Schedulers - local: using-diffusers/weighted_prompts title: Prompting + - local: using-diffusers/image_quality + title: FreeU - local: using-diffusers/reusing_seeds title: Reproducibility - local: using-diffusers/callback @@ -70,22 +72,12 @@ sections: - local: optimization/pruna title: Pruna - - local: optimization/xformers - title: xFormers - - local: optimization/tome - title: Token merging - - local: optimization/deepcache - title: DeepCache - local: optimization/cache_dit title: CacheDiT - - local: optimization/tgate - title: TGATE - local: optimization/xdit title: xDiT - local: optimization/para_attn title: ParaAttention - - local: using-diffusers/image_quality - title: FreeU title: Community methods - isExpanded: false sections: diff --git a/docs/source/en/api/pipelines/kandinsky.md b/docs/source/en/api/pipelines/kandinsky.md index a1965eb3cb72..5347c6f895e1 100644 --- a/docs/source/en/api/pipelines/kandinsky.md +++ b/docs/source/en/api/pipelines/kandinsky.md @@ -712,7 +712,7 @@ make_image_grid([img.resize((512, 512)), image.resize((512, 512))], rows=1, cols Kandinsky is unique because it requires a prior pipeline to generate the mappings, and a second pipeline to decode the latents into an image. Optimization efforts should be focused on the second pipeline because that is where the bulk of the computation is done. Here are some tips to improve Kandinsky during inference. -1. Enable [xFormers](../../optimization/xformers) if you're using PyTorch < 2.0: +1. Enable [xFormers](../../optimization/attention_backends) if you're using PyTorch < 2.0: ```diff from diffusers import DiffusionPipeline diff --git a/docs/source/en/api/pipelines/stable_diffusion/stable_diffusion_xl.md b/docs/source/en/api/pipelines/stable_diffusion/stable_diffusion_xl.md index 69f45576ca10..b87a442272aa 100644 --- a/docs/source/en/api/pipelines/stable_diffusion/stable_diffusion_xl.md +++ b/docs/source/en/api/pipelines/stable_diffusion/stable_diffusion_xl.md @@ -448,7 +448,7 @@ SDXL is a large model, and you may need to optimize memory to get it to run on y + refiner.unet = torch.compile(refiner.unet, mode="reduce-overhead", fullgraph=True) ``` -3. Enable [xFormers](../../../optimization/xformers) to run SDXL if `torch<2.0`: +3. Enable [xFormers](../../../optimization/attention_backends) to run SDXL if `torch<2.0`: ```diff + base.enable_xformers_memory_efficient_attention() diff --git a/docs/source/en/optimization/deepcache.md b/docs/source/en/optimization/deepcache.md deleted file mode 100644 index 2514867a504f..000000000000 --- a/docs/source/en/optimization/deepcache.md +++ /dev/null @@ -1,62 +0,0 @@ - - -# DeepCache -[DeepCache](https://huggingface.co/papers/2312.00858) accelerates [`StableDiffusionPipeline`] and [`StableDiffusionXLPipeline`] by strategically caching and reusing high-level features while efficiently updating low-level features by taking advantage of the U-Net architecture. - -Start by installing [DeepCache](https://github.com/horseee/DeepCache): -```bash -pip install DeepCache -``` - -Then load and enable the [`DeepCacheSDHelper`](https://github.com/horseee/DeepCache#usage): - -```diff - import torch - from diffusers import StableDiffusionPipeline - pipe = StableDiffusionPipeline.from_pretrained('stable-diffusion-v1-5/stable-diffusion-v1-5', dtype=torch.float16).to("cuda") # or "mps", "xpu", "cpu" - -+ from DeepCache import DeepCacheSDHelper -+ helper = DeepCacheSDHelper(pipe=pipe) -+ helper.set_params( -+ cache_interval=3, -+ cache_branch_id=0, -+ ) -+ helper.enable() - - image = pipe("a photo of an astronaut on a moon").images[0] -``` - -The `set_params` method accepts two arguments: `cache_interval` and `cache_branch_id`. `cache_interval` means the frequency of feature caching, specified as the number of steps between each cache operation. `cache_branch_id` identifies which branch of the network (ordered from the shallowest to the deepest layer) is responsible for executing the caching processes. -Opting for a lower `cache_branch_id` or a larger `cache_interval` can lead to faster inference speed at the expense of reduced image quality (ablation experiments of these two hyperparameters can be found in the [paper](https://huggingface.co/papers/2312.00858)). Once those arguments are set, use the `enable` or `disable` methods to activate or deactivate the `DeepCacheSDHelper`. - -
- -
- -You can find more generated samples (original pipeline vs DeepCache) and the corresponding inference latency in the [WandB report](https://wandb.ai/horseee/DeepCache/runs/jwlsqqgt?workspace=user-horseee). The prompts are randomly selected from the [MS-COCO 2017](https://cocodataset.org/#home) dataset. - -## Benchmark - -We tested how much faster DeepCache accelerates [Stable Diffusion v2.1](https://huggingface.co/stabilityai/stable-diffusion-2-1) with 50 inference steps on an NVIDIA RTX A5000, using different configurations for resolution, batch size, cache interval (I), and cache branch (B). - -| **Resolution** | **Batch size** | **Original** | **DeepCache(I=3, B=0)** | **DeepCache(I=5, B=0)** | **DeepCache(I=5, B=1)** | -|----------------|----------------|--------------|-------------------------|-------------------------|-------------------------| -| 512| 8| 15.96| 6.88(2.32x)| 5.03(3.18x)| 7.27(2.20x)| -| | 4| 8.39| 3.60(2.33x)| 2.62(3.21x)| 3.75(2.24x)| -| | 1| 2.61| 1.12(2.33x)| 0.81(3.24x)| 1.11(2.35x)| -| 768| 8| 43.58| 18.99(2.29x)| 13.96(3.12x)| 21.27(2.05x)| -| | 4| 22.24| 9.67(2.30x)| 7.10(3.13x)| 10.74(2.07x)| -| | 1| 6.33| 2.72(2.33x)| 1.97(3.21x)| 2.98(2.12x)| -| 1024| 8| 101.95| 45.57(2.24x)| 33.72(3.02x)| 53.00(1.92x)| -| | 4| 49.25| 21.86(2.25x)| 16.19(3.04x)| 25.78(1.91x)| -| | 1| 13.83| 6.07(2.28x)| 4.43(3.12x)| 7.15(1.93x)| diff --git a/docs/source/en/optimization/tgate.md b/docs/source/en/optimization/tgate.md deleted file mode 100644 index 57e18090c03d..000000000000 --- a/docs/source/en/optimization/tgate.md +++ /dev/null @@ -1,182 +0,0 @@ -# T-GATE - -[T-GATE](https://github.com/HaozheLiu-ST/T-GATE/tree/main) accelerates inference for [Stable Diffusion](../api/pipelines/stable_diffusion/overview), [PixArt](../api/pipelines/pixart), and [Latency Consistency Model](../api/pipelines/latent_consistency_models.md) pipelines by skipping the cross-attention calculation once it converges. This method doesn't require any additional training and it can speed up inference from 10-50%. T-GATE is also compatible with other optimization methods like [DeepCache](./deepcache). - -Before you begin, make sure you install T-GATE. - -```bash -pip install tgate -pip install -U torch diffusers transformers accelerate DeepCache -``` - - -To use T-GATE with a pipeline, you need to use its corresponding loader. - -| Pipeline | T-GATE Loader | -|---|---| -| PixArt | TgatePixArtLoader | -| Stable Diffusion XL | TgateSDXLLoader | -| Stable Diffusion XL + DeepCache | TgateSDXLDeepCacheLoader | -| Stable Diffusion | TgateSDLoader | -| Stable Diffusion + DeepCache | TgateSDDeepCacheLoader | - -Next, create a `TgateLoader` with a pipeline, the gate step (the time step to stop calculating the cross attention), and the number of inference steps. Then call the `tgate` method on the pipeline with a prompt, gate step, and the number of inference steps. - -Let's see how to enable this for several different pipelines. - - - - -Accelerate `PixArtAlphaPipeline` with T-GATE: - -```py -import torch -from diffusers import PixArtAlphaPipeline -from tgate import TgatePixArtLoader - -pipe = PixArtAlphaPipeline.from_pretrained("PixArt-alpha/PixArt-XL-2-1024-MS", dtype=torch.float16) - -gate_step = 8 -inference_step = 25 -pipe = TgatePixArtLoader( - pipe, - gate_step=gate_step, - num_inference_steps=inference_step, -).to("cuda") # or "mps", "xpu", "cpu" - -image = pipe.tgate( - "An alpaca made of colorful building blocks, cyberpunk.", - gate_step=gate_step, - num_inference_steps=inference_step, -).images[0] -``` - - - -Accelerate `StableDiffusionXLPipeline` with T-GATE: - -```py -import torch -from diffusers import StableDiffusionXLPipeline -from diffusers import DPMSolverMultistepScheduler -from tgate import TgateSDXLLoader - -pipe = StableDiffusionXLPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", - dtype=torch.float16, - variant="fp16", - use_safetensors=True, -) -pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config) - -gate_step = 10 -inference_step = 25 -pipe = TgateSDXLLoader( - pipe, - gate_step=gate_step, - num_inference_steps=inference_step, -).to("cuda") # or "mps", "xpu", "cpu" - -image = pipe.tgate( - "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k.", - gate_step=gate_step, - num_inference_steps=inference_step -).images[0] -``` - - - -Accelerate `StableDiffusionXLPipeline` with [DeepCache](https://github.com/horseee/DeepCache) and T-GATE: - -```py -import torch -from diffusers import StableDiffusionXLPipeline -from diffusers import DPMSolverMultistepScheduler -from tgate import TgateSDXLDeepCacheLoader - -pipe = StableDiffusionXLPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", - dtype=torch.float16, - variant="fp16", - use_safetensors=True, -) -pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config) - -gate_step = 10 -inference_step = 25 -pipe = TgateSDXLDeepCacheLoader( - pipe, - cache_interval=3, - cache_branch_id=0, -).to("cuda") # or "mps", "xpu", "cpu" - -image = pipe.tgate( - "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k.", - gate_step=gate_step, - num_inference_steps=inference_step -).images[0] -``` - - - -Accelerate `latent-consistency/lcm-sdxl` with T-GATE: - -```py -import torch -from diffusers import StableDiffusionXLPipeline -from diffusers import UNet2DConditionModel, LCMScheduler -from diffusers import DPMSolverMultistepScheduler -from tgate import TgateSDXLLoader - -unet = UNet2DConditionModel.from_pretrained( - "latent-consistency/lcm-sdxl", - dtype=torch.float16, - variant="fp16", -) -pipe = StableDiffusionXLPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", - unet=unet, - dtype=torch.float16, - variant="fp16", -) -pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config) - -gate_step = 1 -inference_step = 4 -pipe = TgateSDXLLoader( - pipe, - gate_step=gate_step, - num_inference_steps=inference_step, - lcm=True -).to("cuda") # or "mps", "xpu", "cpu" - -image = pipe.tgate( - "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k.", - gate_step=gate_step, - num_inference_steps=inference_step -).images[0] -``` - - - -T-GATE also supports [`StableDiffusionPipeline`] and [PixArt-alpha/PixArt-LCM-XL-2-1024-MS](https://hf.co/PixArt-alpha/PixArt-LCM-XL-2-1024-MS). - -## Benchmarks -| Model | MACs | Param | Latency | Zero-shot 10K-FID on MS-COCO | -|-----------------------|----------|-----------|---------|---------------------------| -| SD-1.5 | 16.938T | 859.520M | 7.032s | 23.927 | -| SD-1.5 w/ T-GATE | 9.875T | 815.557M | 4.313s | 20.789 | -| SD-2.1 | 38.041T | 865.785M | 16.121s | 22.609 | -| SD-2.1 w/ T-GATE | 22.208T | 815.433 M | 9.878s | 19.940 | -| SD-XL | 149.438T | 2.570B | 53.187s | 24.628 | -| SD-XL w/ T-GATE | 84.438T | 2.024B | 27.932s | 22.738 | -| Pixart-Alpha | 107.031T | 611.350M | 61.502s | 38.669 | -| Pixart-Alpha w/ T-GATE | 65.318T | 462.585M | 37.867s | 35.825 | -| DeepCache (SD-XL) | 57.888T | - | 19.931s | 23.755 | -| DeepCache w/ T-GATE | 43.868T | - | 14.666s | 23.999 | -| LCM (SD-XL) | 11.955T | 2.570B | 3.805s | 25.044 | -| LCM w/ T-GATE | 11.171T | 2.024B | 3.533s | 25.028 | -| LCM (Pixart-Alpha) | 8.563T | 611.350M | 4.733s | 36.086 | -| LCM w/ T-GATE | 7.623T | 462.585M | 4.543s | 37.048 | - -The latency is tested on an NVIDIA 1080TI, MACs and Params are calculated with [calflops](https://github.com/MrYxJ/calculate-flops.pytorch), and the FID is calculated with [PytorchFID](https://github.com/mseitzer/pytorch-fid). diff --git a/docs/source/en/optimization/tome.md b/docs/source/en/optimization/tome.md deleted file mode 100644 index 833bd69058a9..000000000000 --- a/docs/source/en/optimization/tome.md +++ /dev/null @@ -1,96 +0,0 @@ - - -# Token merging - -[Token merging](https://huggingface.co/papers/2303.17604) (ToMe) merges redundant tokens/patches progressively in the forward pass of a Transformer-based network which can speed-up the inference latency of [`StableDiffusionPipeline`]. - -Install ToMe from `pip`: - -```bash -pip install tomesd -``` - -You can use ToMe from the [`tomesd`](https://github.com/dbolya/tomesd) library with the [`apply_patch`](https://github.com/dbolya/tomesd?tab=readme-ov-file#usage) function: - -```diff - from diffusers import StableDiffusionPipeline - import torch - import tomesd - - pipeline = StableDiffusionPipeline.from_pretrained( - "stable-diffusion-v1-5/stable-diffusion-v1-5", dtype=torch.float16, use_safetensors=True, - ).to("cuda") # or "mps", "xpu", "cpu" -+ tomesd.apply_patch(pipeline, ratio=0.5) - - image = pipeline("a photo of an astronaut riding a horse on mars").images[0] -``` - -The `apply_patch` function exposes a number of [arguments](https://github.com/dbolya/tomesd#usage) to help strike a balance between pipeline inference speed and the quality of the generated tokens. The most important argument is `ratio` which controls the number of tokens that are merged during the forward pass. - -As reported in the [paper](https://huggingface.co/papers/2303.17604), ToMe can greatly preserve the quality of the generated images while boosting inference speed. By increasing the `ratio`, you can speed-up inference even further, but at the cost of some degraded image quality. - -To test the quality of the generated images, we sampled a few prompts from [Parti Prompts](https://parti.research.google/) and performed inference with the [`StableDiffusionPipeline`] with the following settings: - -
- -
- -We didn’t notice any significant decrease in the quality of the generated samples, and you can check out the generated samples in this [WandB report](https://wandb.ai/sayakpaul/tomesd-results/runs/23j4bj3i?workspace=). If you're interested in reproducing this experiment, use this [script](https://gist.github.com/sayakpaul/8cac98d7f22399085a060992f411ecbd). - -## Benchmarks - -We also benchmarked the impact of `tomesd` on the [`StableDiffusionPipeline`] with [xFormers](https://huggingface.co/docs/diffusers/optimization/xformers) enabled across several image resolutions. The results are obtained from A100 and V100 GPUs in the following development environment: - -```bash -- `diffusers` version: 0.15.1 -- Python version: 3.8.16 -- PyTorch version (GPU?): 1.13.1+cu116 (True) -- Huggingface_hub version: 0.13.2 -- Transformers version: 4.27.2 -- Accelerate version: 0.18.0 -- xFormers version: 0.0.16 -- tomesd version: 0.1.2 -``` - -To reproduce this benchmark, feel free to use this [script](https://gist.github.com/sayakpaul/27aec6bca7eb7b0e0aa4112205850335). The results are reported in seconds, and where applicable we report the speed-up percentage over the vanilla pipeline when using ToMe and ToMe + xFormers. - -| **GPU** | **Resolution** | **Batch size** | **Vanilla** | **ToMe** | **ToMe + xFormers** | -|----------|----------------|----------------|-------------|----------------|---------------------| -| **A100** | 512 | 10 | 6.88 | 5.26 (+23.55%) | 4.69 (+31.83%) | -| | 768 | 10 | OOM | 14.71 | 11 | -| | | 8 | OOM | 11.56 | 8.84 | -| | | 4 | OOM | 5.98 | 4.66 | -| | | 2 | 4.99 | 3.24 (+35.07%) | 2.1 (+37.88%) | -| | | 1 | 3.29 | 2.24 (+31.91%) | 2.03 (+38.3%) | -| | 1024 | 10 | OOM | OOM | OOM | -| | | 8 | OOM | OOM | OOM | -| | | 4 | OOM | 12.51 | 9.09 | -| | | 2 | OOM | 6.52 | 4.96 | -| | | 1 | 6.4 | 3.61 (+43.59%) | 2.81 (+56.09%) | -| **V100** | 512 | 10 | OOM | 10.03 | 9.29 | -| | | 8 | OOM | 8.05 | 7.47 | -| | | 4 | 5.7 | 4.3 (+24.56%) | 3.98 (+30.18%) | -| | | 2 | 3.14 | 2.43 (+22.61%) | 2.27 (+27.71%) | -| | | 1 | 1.88 | 1.57 (+16.49%) | 1.57 (+16.49%) | -| | 768 | 10 | OOM | OOM | 23.67 | -| | | 8 | OOM | OOM | 18.81 | -| | | 4 | OOM | 11.81 | 9.7 | -| | | 2 | OOM | 6.27 | 5.2 | -| | | 1 | 5.43 | 3.38 (+37.75%) | 2.82 (+48.07%) | -| | 1024 | 10 | OOM | OOM | OOM | -| | | 8 | OOM | OOM | OOM | -| | | 4 | OOM | OOM | 19.35 | -| | | 2 | OOM | 13 | 10.78 | -| | | 1 | OOM | 6.66 | 5.54 | - -As seen in the tables above, the speed-up from `tomesd` becomes more pronounced for larger image resolutions. It is also interesting to note that with `tomesd`, it is possible to run the pipeline on a higher resolution like 1024x1024. You may be able to speed-up inference even more with [`torch.compile`](fp16#torchcompile). diff --git a/docs/source/en/optimization/xformers.md b/docs/source/en/optimization/xformers.md deleted file mode 100644 index a5ef4c6fbdb9..000000000000 --- a/docs/source/en/optimization/xformers.md +++ /dev/null @@ -1,29 +0,0 @@ - - -# xFormers - -We recommend [xFormers](https://github.com/facebookresearch/xformers) for both inference and training. In our tests, the optimizations performed in the attention blocks allow for both faster speed and reduced memory consumption. - -Install xFormers from `pip`: - -```bash -pip install xformers -``` - -> [!TIP] -> The xFormers `pip` package requires the latest version of PyTorch. If you need to use a previous version of PyTorch, then we recommend [installing xFormers from the source](https://github.com/facebookresearch/xformers#installing-xformers). - -After xFormers is installed, you can use it with [`~ModelMixin.set_attention_backend`] as shown in the [Attention backends](./attention_backends) guide. - -> [!WARNING] -> According to this [issue](https://github.com/huggingface/diffusers/issues/2234#issuecomment-1416931212), xFormers `v0.0.16` cannot be used for training (fine-tune or DreamBooth) in some GPUs. If you observe this problem, please install a development version as indicated in the issue comments. diff --git a/docs/source/en/training/controlnet.md b/docs/source/en/training/controlnet.md index f52ba33f242f..9cb7526dd131 100644 --- a/docs/source/en/training/controlnet.md +++ b/docs/source/en/training/controlnet.md @@ -14,7 +14,7 @@ specific language governing permissions and limitations under the License. [ControlNet](https://hf.co/papers/2302.05543) models are adapters trained on top of another pretrained model. It allows for a greater degree of control over image generation by conditioning the model with an additional input image. The input image can be a canny edge, depth map, human pose, and many more. -If you're training on a GPU with limited vRAM, you should try enabling the `gradient_checkpointing`, `gradient_accumulation_steps`, and `mixed_precision` parameters in the training command. You can also reduce your memory footprint by using memory-efficient attention with [xFormers](../optimization/xformers). +If you're training on a GPU with limited vRAM, you should try enabling the `gradient_checkpointing`, `gradient_accumulation_steps`, and `mixed_precision` parameters in the training command. You can also reduce your memory footprint by using memory-efficient attention with [xFormers](../optimization/attention_backends). This guide will explore the [train_controlnet.py](https://github.com/huggingface/diffusers/blob/main/examples/controlnet/train_controlnet.py) training script to help you become familiar with it, and how you can adapt it for your own use-case. diff --git a/docs/source/en/training/dreambooth.md b/docs/source/en/training/dreambooth.md index ed2a79c8a889..87d8817055db 100644 --- a/docs/source/en/training/dreambooth.md +++ b/docs/source/en/training/dreambooth.md @@ -16,7 +16,7 @@ specific language governing permissions and limitations under the License. To load a trained checkpoint for inference, see [Load a DreamBooth model for inference](../using-diffusers/dreambooth). -If you're training on a GPU with limited vRAM, you should try enabling the `gradient_checkpointing` and `mixed_precision` parameters in the training command. You can also reduce your memory footprint by using memory-efficient attention with [xFormers](../optimization/xformers). +If you're training on a GPU with limited vRAM, you should try enabling the `gradient_checkpointing` and `mixed_precision` parameters in the training command. You can also reduce your memory footprint by using memory-efficient attention with [xFormers](../optimization/attention_backends). This guide will explore the [train_dreambooth.py](https://github.com/huggingface/diffusers/blob/main/examples/dreambooth/train_dreambooth.py) script to help you become more familiar with it, and how you can adapt it for your own use-case. diff --git a/docs/source/en/training/kandinsky.md b/docs/source/en/training/kandinsky.md index 89532c556031..3653b4c84b84 100644 --- a/docs/source/en/training/kandinsky.md +++ b/docs/source/en/training/kandinsky.md @@ -17,7 +17,7 @@ specific language governing permissions and limitations under the License. Kandinsky 2.2 is a multilingual text-to-image model capable of producing more photorealistic images. The model includes an image prior model for creating image embeddings from text prompts, and a decoder model that generates images based on the prior model's embeddings. That's why you'll find two separate scripts in Diffusers for Kandinsky 2.2, one for training the prior model and one for training the decoder model. You can train both models separately, but to get the best results, you should train both the prior and decoder models. -Depending on your GPU, you may need to enable `gradient_checkpointing` (⚠️ not supported for the prior model!), `mixed_precision`, and `gradient_accumulation_steps` to help fit the model into memory and to speedup training. You can reduce your memory-usage even more by enabling memory-efficient attention with [xFormers](../optimization/xformers) (version [v0.0.16](https://github.com/huggingface/diffusers/issues/2234#issuecomment-1416931212) fails for training on some GPUs so you may need to install a development version instead). +Depending on your GPU, you may need to enable `gradient_checkpointing` (⚠️ not supported for the prior model!), `mixed_precision`, and `gradient_accumulation_steps` to help fit the model into memory and to speedup training. You can reduce your memory-usage even more by enabling memory-efficient attention with [xFormers](../optimization/attention_backends) (version [v0.0.16](https://github.com/huggingface/diffusers/issues/2234#issuecomment-1416931212) fails for training on some GPUs so you may need to install a development version instead). This guide explores the [train_text_to_image_prior.py](https://github.com/huggingface/diffusers/blob/main/examples/kandinsky2_2/text_to_image/train_text_to_image_prior.py) and the [train_text_to_image_decoder.py](https://github.com/huggingface/diffusers/blob/main/examples/kandinsky2_2/text_to_image/train_text_to_image_decoder.py) scripts to help you become more familiar with it, and how you can adapt it for your own use-case. diff --git a/docs/source/en/training/lcm_distill.md b/docs/source/en/training/lcm_distill.md index cfe7d7e2fdce..9bdc3130ac59 100644 --- a/docs/source/en/training/lcm_distill.md +++ b/docs/source/en/training/lcm_distill.md @@ -14,7 +14,7 @@ specific language governing permissions and limitations under the License. [Latent Consistency Models (LCMs)](https://hf.co/papers/2310.04378) are able to generate high-quality images in just a few steps, representing a big leap forward because many pipelines require at least 25+ steps. LCMs are produced by applying the latent consistency distillation method to any Stable Diffusion model. This method works by applying *one-stage guided distillation* to the latent space, and incorporating a *skipping-step* method to consistently skip timesteps to accelerate the distillation process (refer to section 4.1, 4.2, and 4.3 of the paper for more details). -If you're training on a GPU with limited vRAM, try enabling `gradient_checkpointing`, `gradient_accumulation_steps`, and `mixed_precision` to reduce memory-usage and speedup training. You can reduce your memory-usage even more by enabling memory-efficient attention with [xFormers](../optimization/xformers) and [bitsandbytes'](https://github.com/TimDettmers/bitsandbytes) 8-bit optimizer. +If you're training on a GPU with limited vRAM, try enabling `gradient_checkpointing`, `gradient_accumulation_steps`, and `mixed_precision` to reduce memory-usage and speedup training. You can reduce your memory-usage even more by enabling memory-efficient attention with [xFormers](../optimization/attention_backends) and [bitsandbytes'](https://github.com/TimDettmers/bitsandbytes) 8-bit optimizer. This guide will explore the [train_lcm_distill_sd_wds.py](https://github.com/huggingface/diffusers/blob/main/examples/consistency_distillation/train_lcm_distill_sd_wds.py) script to help you become more familiar with it, and how you can adapt it for your own use-case. diff --git a/docs/source/en/training/overview.md b/docs/source/en/training/overview.md index ecd7e7780ccf..b33bee8730a5 100644 --- a/docs/source/en/training/overview.md +++ b/docs/source/en/training/overview.md @@ -59,4 +59,4 @@ pip install -r requirements_sdxl.txt To speedup training and reduce memory-usage, we recommend: - using PyTorch 2.0 or higher to automatically use [scaled dot product attention](../optimization/fp16#scaled-dot-product-attention) during training (you don't need to make any changes to the training code) -- installing [xFormers](../optimization/xformers) to enable memory-efficient attention +- installing [xFormers](../optimization/attention_backends) to enable memory-efficient attention diff --git a/docs/source/en/training/sdxl.md b/docs/source/en/training/sdxl.md index cdbea957e2b2..79ee1df9bdca 100644 --- a/docs/source/en/training/sdxl.md +++ b/docs/source/en/training/sdxl.md @@ -17,7 +17,7 @@ specific language governing permissions and limitations under the License. [Stable Diffusion XL (SDXL)](https://hf.co/papers/2307.01952) is a larger and more powerful iteration of the Stable Diffusion model, capable of producing higher resolution images. -SDXL's UNet is 3x larger and the model adds a second text encoder to the architecture. Depending on the hardware available to you, this can be very computationally intensive and it may not run on a consumer GPU like a Tesla T4. To help fit this larger model into memory and to speedup training, try enabling `gradient_checkpointing`, `mixed_precision`, and `gradient_accumulation_steps`. You can reduce your memory-usage even more by enabling memory-efficient attention with [xFormers](../optimization/xformers) and using [bitsandbytes'](https://github.com/TimDettmers/bitsandbytes) 8-bit optimizer. +SDXL's UNet is 3x larger and the model adds a second text encoder to the architecture. Depending on the hardware available to you, this can be very computationally intensive and it may not run on a consumer GPU like a Tesla T4. To help fit this larger model into memory and to speedup training, try enabling `gradient_checkpointing`, `mixed_precision`, and `gradient_accumulation_steps`. You can reduce your memory-usage even more by enabling memory-efficient attention with [xFormers](../optimization/attention_backends) and using [bitsandbytes'](https://github.com/TimDettmers/bitsandbytes) 8-bit optimizer. This guide will explore the [train_text_to_image_sdxl.py](https://github.com/huggingface/diffusers/blob/main/examples/text_to_image/train_text_to_image_sdxl.py) training script to help you become more familiar with it, and how you can adapt it for your own use-case. diff --git a/docs/source/en/training/text2image.md b/docs/source/en/training/text2image.md index 33df598ea16b..00ddd1214168 100644 --- a/docs/source/en/training/text2image.md +++ b/docs/source/en/training/text2image.md @@ -17,7 +17,7 @@ specific language governing permissions and limitations under the License. Text-to-image models like Stable Diffusion are conditioned to generate images given a text prompt. -Training a model can be taxing on your hardware, but if you enable `gradient_checkpointing` and `mixed_precision`, it is possible to train a model on a single 24GB GPU. If you're training with larger batch sizes or want to train faster, it's better to use GPUs with more than 30GB of memory. You can reduce your memory footprint by enabling memory-efficient attention with [xFormers](../optimization/xformers). +Training a model can be taxing on your hardware, but if you enable `gradient_checkpointing` and `mixed_precision`, it is possible to train a model on a single 24GB GPU. If you're training with larger batch sizes or want to train faster, it's better to use GPUs with more than 30GB of memory. You can reduce your memory footprint by enabling memory-efficient attention with [xFormers](../optimization/attention_backends). This guide will explore the [train_text_to_image.py](https://github.com/huggingface/diffusers/blob/main/examples/text_to_image/train_text_to_image.py) training script to help you become familiar with it, and how you can adapt it for your own use-case. diff --git a/docs/source/en/training/text_inversion.md b/docs/source/en/training/text_inversion.md index a9116b0a1b61..c1914de57785 100644 --- a/docs/source/en/training/text_inversion.md +++ b/docs/source/en/training/text_inversion.md @@ -16,7 +16,7 @@ specific language governing permissions and limitations under the License. For inference with trained embeddings, see [Textual inversion inference](../using-diffusers/legacy_adapters#textual-inversion). -If you're training on a GPU with limited vRAM, you should try enabling the `gradient_checkpointing` and `mixed_precision` parameters in the training command. You can also reduce your memory footprint by using memory-efficient attention with [xFormers](../optimization/xformers). +If you're training on a GPU with limited vRAM, you should try enabling the `gradient_checkpointing` and `mixed_precision` parameters in the training command. You can also reduce your memory footprint by using memory-efficient attention with [xFormers](../optimization/attention_backends). This guide will explore the [textual_inversion.py](https://github.com/huggingface/diffusers/blob/main/examples/textual_inversion/textual_inversion.py) script to help you become more familiar with it, and how you can adapt it for your own use-case. diff --git a/docs/source/en/using-diffusers/image_quality.md b/docs/source/en/using-diffusers/image_quality.md index 8cbf11ea5388..f7c53c6816bb 100644 --- a/docs/source/en/using-diffusers/image_quality.md +++ b/docs/source/en/using-diffusers/image_quality.md @@ -12,68 +12,14 @@ specific language governing permissions and limitations under the License. # FreeU -[FreeU](https://hf.co/papers/2309.11497) improves image details by rebalancing the UNet's backbone and skip connection weights. The skip connections can cause the model to overlook some of the backbone semantics which may lead to unnatural image details in the generated image. This technique does not require any additional training and can be applied on the fly during inference for tasks like image-to-image and text-to-video. +[FreeU](https://huggingface.co/papers/2309.11497) improves image detail by rebalancing how much the UNet decoder draws from backbone features versus skip-connection features. Skip connections can drown out the backbone's semantic features, which produces unnatural detail in the output. FreeU needs no training, and you can turn it on or off at inference time for text-to-image and text-to-video pipelines. -Use the [`~pipelines.StableDiffusionMixin.enable_freeu`] method on your pipeline and configure the scaling factors for the backbone (`b1` and `b2`) and skip connections (`s1` and `s2`). The number after each scaling factor corresponds to the stage in the UNet where the factor is applied. Take a look at the [FreeU](https://github.com/ChenyangSi/FreeU#parameters) repository for reference hyperparameters for different models. +> [!NOTE] +> FreeU only works with UNet-based pipelines like Stable Diffusion, SDXL, and AnimateDiff. It isn't supported by transformer-based pipelines like Flux or Qwen-Image. - - +Use the [`~pipelines.StableDiffusionMixin.enable_freeu`] method on your pipeline and configure the scaling factors. `b1` and `b2` amplify the backbone features, and `s1` and `s2` dampen the skip features. The `1` and `2` refer to the first two upsampling stages of the UNet decoder. See the [FreeU](https://github.com/ChenyangSi/FreeU#parameters) repository for reference hyperparameters for different models. -```py -import torch -from diffusers import DiffusionPipeline - -pipeline = DiffusionPipeline.from_pretrained( - "stable-diffusion-v1-5/stable-diffusion-v1-5", dtype=torch.float16, safety_checker=None -).to("cuda") # or "mps", "xpu", "cpu" -pipeline.enable_freeu(s1=0.9, s2=0.2, b1=1.5, b2=1.6) -generator = torch.Generator(device="cpu").manual_seed(33) -prompt = "" -image = pipeline(prompt, generator=generator).images[0] -image -``` - -
-
- -
FreeU disabled
-
-
- -
FreeU enabled
-
-
- -
- - -```py -import torch -from diffusers import DiffusionPipeline - -pipeline = DiffusionPipeline.from_pretrained( - "stabilityai/stable-diffusion-2-1", dtype=torch.float16, safety_checker=None -).to("cuda") # or "mps", "xpu", "cpu" -pipeline.enable_freeu(s1=0.9, s2=0.2, b1=1.4, b2=1.6) -generator = torch.Generator(device="cpu").manual_seed(80) -prompt = "A squirrel eating a burger" -image = pipeline(prompt, generator=generator).images[0] -image -``` - -
-
- -
FreeU disabled
-
-
- -
FreeU enabled
-
-
- -
- +Start with the repository values for a model. To tune for other models, keep `s1=0.9` and `s2=0.2` and adjust `b1` and `b2` first. Setting all four factors to `1.0` is the same as disabling FreeU. Larger `b` values strengthen the effect but can oversmooth fine texture, and lowering `s1` and `s2` counteracts that. ```py import torch @@ -100,41 +46,13 @@ image - - - -```py -import torch -from diffusers import DiffusionPipeline -from diffusers.utils import export_to_video - -pipeline = DiffusionPipeline.from_pretrained( - "damo-vilab/text-to-video-ms-1.7b", dtype=torch.float16 -).to("cuda") # or "mps", "xpu", "cpu" -# values come from https://github.com/lyn-rgb/FreeU_Diffusers#video-pipelines -pipeline.enable_freeu(b1=1.2, b2=1.4, s1=0.9, s2=0.2) -prompt = "Confident teddy bear surfer rides the wave in the tropics" -generator = torch.Generator(device="cpu").manual_seed(47) -video_frames = pipeline(prompt, generator=generator).frames[0] -export_to_video(video_frames, "teddy_bear.mp4", fps=10) -``` - -
-
- -
FreeU disabled
-
-
- -
FreeU enabled
-
-
- -
-
- Call the [`~pipelines.StableDiffusionMixin.disable_freeu`] method to disable FreeU. ```py pipeline.disable_freeu() ``` + +## Next steps + +- See the [`~pipelines.StableDiffusionMixin.enable_freeu`] API reference for the full parameter descriptions. +- Try FreeU on video with [AnimateDiff](../api/pipelines/animatediff). diff --git a/docs/source/en/using-diffusers/img2img.md b/docs/source/en/using-diffusers/img2img.md index 366be176bd86..e689f6d9fce6 100644 --- a/docs/source/en/using-diffusers/img2img.md +++ b/docs/source/en/using-diffusers/img2img.md @@ -580,7 +580,7 @@ make_image_grid([init_image, depth_image, image_control_net, image_elden_ring], ## Optimize -Running diffusion models is computationally expensive and intensive, but with a few optimization tricks, it is entirely possible to run them on consumer and free-tier GPUs. For example, you can use a more memory-efficient form of attention such as PyTorch 2.0's [scaled-dot product attention](../optimization/fp16#scaled-dot-product-attention) or [xFormers](../optimization/xformers) (you can use one or the other, but there's no need to use both). You can also offload the model to the GPU while the other pipeline components wait on the CPU. +Running diffusion models is computationally expensive and intensive, but with a few optimization tricks, it is entirely possible to run them on consumer and free-tier GPUs. For example, you can use a more memory-efficient form of attention such as PyTorch 2.0's [scaled-dot product attention](../optimization/fp16#scaled-dot-product-attention) or [xFormers](../optimization/attention_backends) (you can use one or the other, but there's no need to use both). You can also offload the model to the GPU while the other pipeline components wait on the CPU. ```diff + pipeline.enable_model_cpu_offload() diff --git a/docs/source/en/using-diffusers/inpaint.md b/docs/source/en/using-diffusers/inpaint.md index b0fb51bcdb89..6d7d174e2bd1 100644 --- a/docs/source/en/using-diffusers/inpaint.md +++ b/docs/source/en/using-diffusers/inpaint.md @@ -782,7 +782,7 @@ make_image_grid([init_image, mask_image, image, image_elden_ring], rows=2, cols= ## Optimize -It can be difficult and slow to run diffusion models if you're resource constrained, but it doesn't have to be with a few optimization tricks. One of the biggest (and easiest) optimizations you can enable is switching to memory-efficient attention. If you're using PyTorch 2.0, [scaled-dot product attention](../optimization/fp16#scaled-dot-product-attention) is automatically enabled and you don't need to do anything else. For non-PyTorch 2.0 users, you can install and use [xFormers](../optimization/xformers)'s implementation of memory-efficient attention. Both options reduce memory usage and accelerate inference. +It can be difficult and slow to run diffusion models if you're resource constrained, but it doesn't have to be with a few optimization tricks. One of the biggest (and easiest) optimizations you can enable is switching to memory-efficient attention. If you're using PyTorch 2.0, [scaled-dot product attention](../optimization/fp16#scaled-dot-product-attention) is automatically enabled and you don't need to do anything else. For non-PyTorch 2.0 users, you can install and use [xFormers](../optimization/attention_backends)'s implementation of memory-efficient attention. Both options reduce memory usage and accelerate inference. You can also offload the model to the CPU to save even more memory: From 7dcdcc9feb4dfaf53e136165ec6237d11b868392 Mon Sep 17 00:00:00 2001 From: stevhliu Date: Wed, 30 Sep 2026 22:00:22 -0700 Subject: [PATCH 2/2] cachedit --- docs/source/en/_toctree.yml | 8 +- docs/source/en/optimization/cache_dit.md | 276 ++++------------------- 2 files changed, 49 insertions(+), 235 deletions(-) diff --git a/docs/source/en/_toctree.yml b/docs/source/en/_toctree.yml index 344b5733b74e..4b4e4247195d 100644 --- a/docs/source/en/_toctree.yml +++ b/docs/source/en/_toctree.yml @@ -70,14 +70,14 @@ title: Inference - isExpanded: false sections: - - local: optimization/pruna - title: Pruna - local: optimization/cache_dit title: CacheDiT - - local: optimization/xdit - title: xDiT - local: optimization/para_attn title: ParaAttention + - local: optimization/pruna + title: Pruna + - local: optimization/xdit + title: xDiT title: Community methods - isExpanded: false sections: diff --git a/docs/source/en/optimization/cache_dit.md b/docs/source/en/optimization/cache_dit.md index 9c3edb23fdf7..16beb5ee5b3f 100644 --- a/docs/source/en/optimization/cache_dit.md +++ b/docs/source/en/optimization/cache_dit.md @@ -1,270 +1,84 @@ -## CacheDiT +# CacheDiT -CacheDiT is a unified, flexible, and training-free cache acceleration framework designed to support nearly all Diffusers' DiT-based pipelines. It provides a unified cache API that supports automatic block adapter, DBCache, and more. +[CacheDiT](https://github.com/vipshop/cache-dit) speeds up DiT pipelines by reusing transformer block outputs across denoising steps. It does not need training and supports most Diffusers DiT pipelines, including Flux, Qwen-Image, Wan, and HunyuanVideo. Diffusers also has [built-in caching](./cache) which doesn't require an extra dependency. -To learn more, refer to the [CacheDiT](https://github.com/vipshop/cache-dit) repository. - -Install a stable release of CacheDiT from PyPI or you can install the latest version from GitHub. - - - +Install CacheDiT from PyPI. ```bash -pip3 install -U cache-dit +pip install -U cache-dit ``` - - - -```bash -pip3 install git+https://github.com/vipshop/cache-dit.git -``` - - - - -Run the command below to view supported DiT pipelines. - -```python ->>> import cache_dit ->>> cache_dit.supported_pipelines() -(30, ['Flux*', 'Mochi*', 'CogVideoX*', 'Wan*', 'HunyuanVideo*', 'QwenImage*', 'LTX*', 'Allegro*', -'CogView3Plus*', 'CogView4*', 'Cosmos*', 'EasyAnimate*', 'SkyReelsV2*', 'StableDiffusion3*', -'ConsisID*', 'DiT*', 'Amused*', 'Bria*', 'Lumina*', 'OmniGen*', 'PixArt*', 'Sana*', 'StableAudio*', -'VisualCloze*', 'AuraFlow*', 'Chroma*', 'ShapE*', 'HiDream*', 'HunyuanDiT*', 'HunyuanDiTPAG*']) -``` +Call `cache_dit.supported_pipelines()` to list the pipeline families CacheDiT supports. -For a complete benchmark, please refer to [Benchmarks](https://github.com/vipshop/cache-dit/blob/main/bench/). - - -## Unified Cache API - -CacheDiT works by matching specific input/output patterns as shown below. - -![](https://github.com/vipshop/cache-dit/raw/main/assets/patterns-v1.png) - -Call the `enable_cache()` function on a pipeline to enable cache acceleration. This function is the entry point to many of CacheDiT's features. - -```python +```py import cache_dit -from diffusers import DiffusionPipeline - -# Can be any diffusion pipeline -pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image") - -# One-line code with default cache options. -cache_dit.enable_cache(pipe) -# Just call the pipe as normal. -output = pipe(...) - -# Disable cache and run original pipe. -cache_dit.disable_cache(pipe) +cache_dit.supported_pipelines() ``` -## Automatic Block Adapter - -For custom or modified pipelines or transformers not included in Diffusers, use the `BlockAdapter` in `auto` mode or via manual configuration. Please check the [BlockAdapter](https://github.com/vipshop/cache-dit/blob/main/docs/User_Guide.md#automatic-block-adapter) docs for more details. Refer to [Qwen-Image w/ BlockAdapter](https://github.com/vipshop/cache-dit/blob/main/examples/adapter/run_qwen_image_adapter.py) as an example. - - -```python -from cache_dit import ForwardPattern, BlockAdapter +## Enable caching -# Use 🔥BlockAdapter with `auto` mode. -cache_dit.enable_cache( - BlockAdapter( - # Any DiffusionPipeline, Qwen-Image, etc. - pipe=pipe, auto=True, - # Check `📚Forward Pattern Matching` documentation and hack the code of - # of Qwen-Image, you will find that it has satisfied `FORWARD_PATTERN_1`. - forward_pattern=ForwardPattern.Pattern_1, - ), -) +Call `cache_dit.enable_cache` on a pipeline to cache it with the default settings, then run the pipeline as usual. -# Or, manually setup transformer configurations. -cache_dit.enable_cache( - BlockAdapter( - pipe=pipe, # Qwen-Image, etc. - transformer=pipe.transformer, - blocks=pipe.transformer.transformer_blocks, - forward_pattern=ForwardPattern.Pattern_1, - ), -) -``` +```py +import torch +import cache_dit +from diffusers import FluxPipeline -Sometimes, a Transformer class will contain more than one transformer `blocks`. For example, FLUX.1 (HiDream, Chroma, etc) contains `transformer_blocks` and `single_transformer_blocks` (with different forward patterns). The BlockAdapter is able to detect this hybrid pattern type as well. -Refer to [FLUX.1](https://github.com/vipshop/cache-dit/blob/main/examples/adapter/run_flux_adapter.py) as an example. +pipeline = FluxPipeline.from_pretrained( + "black-forest-labs/FLUX.1-dev", dtype=torch.bfloat16 +).to("cuda") +cache_dit.enable_cache(pipeline) -```python -# For diffusers <= 0.34.0, FLUX.1 transformer_blocks and -# single_transformer_blocks have different forward patterns. -cache_dit.enable_cache( - BlockAdapter( - pipe=pipe, # FLUX.1, etc. - transformer=pipe.transformer, - blocks=[ - pipe.transformer.transformer_blocks, - pipe.transformer.single_transformer_blocks, - ], - forward_pattern=[ - ForwardPattern.Pattern_1, - ForwardPattern.Pattern_3, - ], - ), -) +image = pipeline( + "A cat holding a sign that says hello world", num_inference_steps=28 +).images[0] ``` -This also works if there is more than one transformer (namely `transformer` and `transformer_2`) in its structure. Refer to [Wan 2.2 MoE](https://github.com/vipshop/cache-dit/blob/main/examples/pipeline/run_wan_2.2.py) as an example. - -## Patch Functor - -For any pattern not included in CacheDiT, use the Patch Functor to convert the pattern into a known pattern. You need to subclass the Patch Functor and may also need to fuse the operations within the blocks for loop into block `forward`. After implementing a Patch Functor, set the `patch_functor` property in `BlockAdapter`. +CacheDiT also works with `torch.compile`. Compile the transformer after you call `enable_cache`. See the [compile](https://github.com/vipshop/cache-dit/blob/main/docs/user_guide/COMPILE.md) docs for settings that avoid recompilation with dynamic input shapes. -![](https://github.com/vipshop/cache-dit/raw/main/assets/patch-functor.png) - -Some Patch Functors are already provided in CacheDiT, [HiDreamPatchFunctor](https://github.com/vipshop/cache-dit/blob/main/src/cache_dit/cache_factory/patch_functors/functor_hidream.py), [ChromaPatchFunctor](https://github.com/vipshop/cache-dit/blob/main/src/cache_dit/cache_factory/patch_functors/functor_chroma.py), etc. - -```python -@BlockAdapterRegistry.register("HiDream") -def hidream_adapter(pipe, **kwargs) -> BlockAdapter: - from diffusers import HiDreamImageTransformer2DModel - from cache_dit.cache_factory.patch_functors import HiDreamPatchFunctor - - assert isinstance(pipe.transformer, HiDreamImageTransformer2DModel) - return BlockAdapter( - pipe=pipe, - transformer=pipe.transformer, - blocks=[ - pipe.transformer.double_stream_blocks, - pipe.transformer.single_stream_blocks, - ], - forward_pattern=[ - ForwardPattern.Pattern_0, - ForwardPattern.Pattern_3, - ], - # NOTE: Setup your custom patch functor here. - patch_functor=HiDreamPatchFunctor(), - **kwargs, - ) +```py +pipeline.transformer = torch.compile(pipeline.transformer) ``` -Finally, you can call the `cache_dit.summary()` function on a pipeline after its completed inference to get the cache acceleration details. +Call `cache_dit.summary` after inference to log how many steps were cached and the residual differences between steps. -```python -stats = cache_dit.summary(pipe) +```py +stats = cache_dit.summary(pipeline) ``` -```python -⚡️Cache Steps and Residual Diffs Statistics: QwenImagePipeline +Call `cache_dit.disable_cache` to restore the original pipeline. -| Cache Steps | Diffs Min | Diffs P25 | Diffs P50 | Diffs P75 | Diffs P95 | Diffs Max | -|-------------|-----------|-----------|-----------|-----------|-----------|-----------| -| 23 | 0.045 | 0.084 | 0.114 | 0.147 | 0.241 | 0.297 | +```py +cache_dit.disable_cache(pipeline) ``` -## DBCache: Dual Block Cache - -![](https://github.com/vipshop/cache-dit/raw/main/assets/dbcache-v1.png) - -DBCache (Dual Block Caching) supports different configurations of compute blocks (F8B12, etc.) to enable a balanced trade-off between performance and precision. -- Fn_compute_blocks: Specifies that DBCache uses the **first n** Transformer blocks to fit the information at time step t, enabling the calculation of a more stable L1 diff and delivering more accurate information to subsequent blocks. -- Bn_compute_blocks: Further fuses approximate information in the **last n** Transformer blocks to enhance prediction accuracy. These blocks act as an auto-scaler for approximate hidden states that use residual cache. - - -```python -import cache_dit -from diffusers import FluxPipeline +## Configure the cache -pipe_or_adapter = FluxPipeline.from_pretrained( - "black-forest-labs/FLUX.1-dev", - dtype=torch.bfloat16, -).to("cuda") # or "mps", "xpu", "cpu" +DBCache (Dual Block Cache) computes the first n blocks (Fn) at every step. When their output barely changes from the previous step, it reuses the cached output for the remaining blocks, and it can recompute the last n blocks (Bn) to correct it. The TaylorSeer calibrator predicts the cached output from earlier steps instead of reusing it as is. -# Default options, F8B0, 8 warmup steps, and unlimited cached -# steps for good balance between performance and precision -cache_dit.enable_cache(pipe_or_adapter) +`enable_cache` defaults to DBCache with the first 8 blocks always computed (F8B0) and 8 uncached warmup steps. To trade speed for quality, raise `Fn_compute_blocks` or lower `residual_diff_threshold` (default `0.08`). For the best quality at high cache rates, add the TaylorSeer calibrator. -# Custom options, F8B8, higher precision -from cache_dit import BasicCacheConfig +```py +from cache_dit import DBCacheConfig, TaylorSeerCalibratorConfig cache_dit.enable_cache( - pipe_or_adapter, - cache_config=BasicCacheConfig( - max_warmup_steps=8, # steps do not cache - max_cached_steps=-1, # -1 means no limit - Fn_compute_blocks=8, # Fn, F8, etc. - Bn_compute_blocks=8, # Bn, B8, etc. + pipeline, + cache_config=DBCacheConfig( + max_warmup_steps=8, + Fn_compute_blocks=8, + Bn_compute_blocks=0, residual_diff_threshold=0.12, ), -) -``` -Check the [DBCache](https://github.com/vipshop/cache-dit/blob/main/docs/DBCache.md) and [User Guide](https://github.com/vipshop/cache-dit/blob/main/docs/User_Guide.md#dbcache) docs for more design details. - -## TaylorSeer Calibrator - -The [TaylorSeers](https://huggingface.co/papers/2503.06923) algorithm further improves the precision of DBCache in cases where the cached steps are large (Hybrid TaylorSeer + DBCache). At timesteps with significant intervals, the feature similarity in diffusion models decreases substantially, significantly harming the generation quality. - -TaylorSeer employs a differential method to approximate the higher-order derivatives of features and predict features in future timesteps with Taylor series expansion. The TaylorSeer implemented in CacheDiT supports both hidden states and residual cache types. F_pred can be a residual cache or a hidden-state cache. - -```python -from cache_dit import BasicCacheConfig, TaylorSeerCalibratorConfig - -cache_dit.enable_cache( - pipe_or_adapter, - # Basic DBCache w/ FnBn configurations - cache_config=BasicCacheConfig( - max_warmup_steps=8, # steps do not cache - max_cached_steps=-1, # -1 means no limit - Fn_compute_blocks=8, # Fn, F8, etc. - Bn_compute_blocks=8, # Bn, B8, etc. - residual_diff_threshold=0.12, - ), - # Then, you can use the TaylorSeer Calibrator to approximate - # the values in cached steps, taylorseer_order default is 1. - calibrator_config=TaylorSeerCalibratorConfig( - taylorseer_order=1, - ), -) -``` - -> [!TIP] -> The `Bn_compute_blocks` parameter of DBCache can be set to `0` if you use TaylorSeer as the calibrator for approximate hidden states. DBCache's `Bn_compute_blocks` also acts as a calibrator, so you can choose either `Bn_compute_blocks` > 0 or TaylorSeer. We recommend using the configuration scheme of TaylorSeer + DBCache FnB0. - -## Hybrid Cache CFG - -CacheDiT supports caching for CFG (classifier-free guidance). For models that fuse CFG and non-CFG into a single forward step, or models that do not include CFG in the forward step, please set `enable_separate_cfg` parameter to `False (default, None)`. Otherwise, set it to `True`. - -```python -from cache_dit import BasicCacheConfig - -cache_dit.enable_cache( - pipe_or_adapter, - cache_config=BasicCacheConfig( - ..., - # For example, set it as True for Wan 2.1, Qwen-Image - # and set it as False for FLUX.1, HunyuanVideo, etc. - enable_separate_cfg=True, - ), + calibrator_config=TaylorSeerCalibratorConfig(taylorseer_order=1), ) ``` -## torch.compile - -CacheDiT is designed to work with torch.compile for even better performance. Call `torch.compile` after enabling the cache. - +For supported pipelines, CacheDiT already knows whether CFG runs as a separate forward pass. For other models, set `enable_separate_cfg=True` in `DBCacheConfig` if the model runs the conditional and unconditional passes separately, or `False` if it fuses them or doesn't use CFG. -```python -cache_dit.enable_cache(pipe) +See the [DBCache design](https://github.com/vipshop/cache-dit/blob/main/docs/user_guide/DBCACHE_DESIGN.md) docs for how the Fn and Bn blocks work, and the [cache benchmarks](https://github.com/vipshop/cache-dit/blob/main/bench/cache/README.md) for speed and quality numbers. -# Compile the Transformer module -pipe.transformer = torch.compile(pipe.transformer) -``` - -If you're using CacheDiT with dynamic input shapes, consider increasing the `recompile_limit` of `torch._dynamo`. Otherwise, the `recompile_limit` error may be triggered, causing the module to fall back to eager mode. - -```python -torch._dynamo.config.recompile_limit = 96 # default is 8 -torch._dynamo.config.accumulated_recompile_limit = 2048 # default is 256 -``` +## Next steps -Please check [perf.py](https://github.com/vipshop/cache-dit/blob/main/bench/perf.py) for more details. +- For pipelines CacheDiT doesn't support yet, see the [BlockAdapter](https://github.com/vipshop/cache-dit/blob/main/docs/user_guide/CACHE_API.md#automatic-block-adapter) docs. +- CacheDiT also supports [context parallelism](https://github.com/vipshop/cache-dit/blob/main/docs/user_guide/CONTEXT_PARALLEL.md) and [quantization](https://github.com/vipshop/cache-dit/blob/main/docs/user_guide/QUANTIZATION.md).