Consistency Models
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation. To overcome this limitation, we propose consistency models, a new family of models t...
Appears in
- Uploaded
- Uploaded Jul 9, 2026
- File type
- BLOG
- Queried
- 0
Full article
Showing the full article.
# Consistency Models Yang Song 1 Prafulla Dhariwal 1 Mark Chen 1 Ilya Sutskever 1 ## Abstract Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation. To overcome this limita-tion, we propose consistency models, a new fam-ily of models that generate high quality samples by directly mapping noise to data. They support fast one-step generation by design, while still al-lowing multistep sampling to trade compute for sample quality. They also support zero-shot data editing, such as image inpainting, colorization, and super-resolution, without requiring explicit training on these tasks.
Consistency models can be trained either by distilling pre-trained diffu-sion models, or as standalone generative models altogether. Through extensive experiments, we demonstrate that they outperform existing distilla-tion techniques for diffusion models in one- and few-step sampling, achieving the new state-of-the-art FID of 3.55 on CIFAR-10 and 6.20 on ImageNet 64 ˆ 64 for one-step generation. When trained in isolation, consistency models become a new family of generative models that can outper-form existing one-step, non-adversarial generative models on standard benchmarks such as CIFAR-10, ImageNet 64 ˆ 64 and LSUN 256 ˆ 256. ## 1. Introduction Diffusion models (Sohl-Dickstein et
, 2015; Song & Er-mon, 2019; 2020; Ho et, 2020; Song et, 2021), also known as score-based generative models, have achieved unprecedented success across multiple fields, including im-age generation (Dhariwal & Nichol, 2021; Nichol et, 2021; Ramesh et, 2022; Saharia et, 2022; Rombach et, 2022), audio synthesis (Kong et, 2020; Chen et, 2021; Popov et, 2021), and video generation (Ho et, > 1 OpenAI, San Francisco, CA 94110, USA. Correspondence to: Yang Song com >. Proceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA.
PMLR 202, 2023. Copyright 2023 by the author(s). Figure 1: Given a Probability Flow (PF) ODE that smoothly converts data to noise, we learn to map any point, xt, xt1, and xT) on the ODE trajectory to its origin, x0)for generative modeling. Models of these mappings are called consistency models, as their outputs are trained to be consistent for points on the same trajectory. 2022b;a). A key feature of diffusion models is the iterative sampling process which progressively removes noise from random initial vectors. This iterative process provides a flexible trade-off of compute and sample quality, as using extra compute for more iterations usually yields samples of better quality.
It is also the crux of many zero-shot data editing capabilities of diffusion models, enabling them to solve challenging inverse problems ranging from image inpainting, colorization, stroke-guided image editing, to Computed Tomography and Magnetic Resonance Imaging (Song & Ermon, 2019; Song et, 2021; 2022; 2023; Kawar et, 2021; 2022; Chung et, 2023; Meng et, 2021). However, compared to single-step generative models like GANs (Goodfellow et, 2014), VAEs (Kingma & Welling, 2014; Rezende et, 2014), or normalizing flows (Dinh et, 2015; 2017; Kingma & Dhariwal, 2018), the iterative generation procedure of diffusion models typically requires 10–2000 times more compute for sample generation (Song & Ermon, 2020; Ho et
, 2020; Song et, 2021; Zhang & Chen, 2022; Lu et, 2022), causing slow inference and limited real-time applications. Our objective is to create generative models that facilitate ef-ficient, single-step generation without sacrificing important advantages of iterative sampling, such as trading compute for sample quality when necessary, as well as performing zero-shot data editing tasks. As illustrated in Fig. 1, we build on top of the probability flow (PF) ordinary differen-tial equation (ODE) in continuous-time diffusion models (Song et, 2021), whose trajectories smoothly transition 1 > arXiv:2303.01469v2 LG] 31 May 2023 Consistency Models the data distribution into a tractable noise distribution.
We propose to learn a model that maps any point at any time step to the trajectory’s starting point. A notable property of our model is self-consistency: points on the same tra-jectory map to the same initial point. We therefore refer to such models as consistency models. Consistency models allow us to generate data samples (initial points of ODE trajectories,, x0 in Fig. 1) by converting random noise vectors (endpoints of ODE trajectories,, xT in Fig. 1) with only one network evaluation. Importantly, by chaining the outputs of consistency models at multiple time steps, we can
improve sample quality and perform zero-shot data editing at the cost of more compute, similar to what iterative sampling enables for diffusion models. To train a consistency model, we offer two methods based on enforcing the self-consistency property. The first method relies on using numerical ODE solvers and a pre-trained diffusion model to generate pairs of adjacent points on a PF ODE trajectory. By minimizing the difference between model outputs for these pairs, we can effectively distill a diffusion model into a consistency model, which allows gen-erating high-quality samples with one network evaluation. By contrast, our second method eliminates the need for a pre-trained diffusion model altogether, allowing us to train a consistency model in isolation.
This approach situates consistency models as an independent family of generative models. Importantly, neither approach necessitates adver-sarial training, and they both place minor constraints on the architecture, allowing the use of flexible neural networks for parameterizing consistency models. We demonstrate the efficacy of consistency models on sev-eral image datasets, including CIFAR-10 (Krizhevsky et, 2009), ImageNet 64 ˆ 64 (Deng et, 2009), and LSUN 256 ˆ 256 (Yu et, 2015). Empirically, we observe that as a distillation approach, consistency models outperform existing diffusion distillation methods like progressive dis-tillation (Salimans & Ho, 2022) across a variety of datasets in few-step generation: On CIFAR-10, consistency models reach new state-of-the-art FIDs of 3.55 and 2.93 for one-step and
two-step generation; on ImageNet 64 ˆ 64, it achieves record-breaking FIDs of 6.20 and 4.70 with one and two net-work evaluations respectively. When trained as standalone generative models, consistency models can match or surpass the quality of one-step samples from progressive distillation, despite having no access to pre-trained diffusion models. They are also able to outperform many GANs, and exist-ing non-adversarial, single-step generative models across multiple datasets. Furthermore, we show that consistency models can be used to perform a wide range of zero-shot data editing tasks, including image denoising, interpolation, inpainting, colorization, super-resolution, and stroke-guided image editing (SDEdit, Meng et al.
(2021)). ## 2. Diffusion Models Consistency models are heavily inspired by the theory of continuous-time diffusion models (Song et, 2021; Karras et, 2022). Diffusion models generate data by progres-sively perturbing data to noise via Gaussian perturbations, then creating samples from noise via sequential denoising steps. Let pdata pxq denote the data distribution. Diffusion models start by diffusing pdata pxq with a stochastic differen-tial equation (SDE) (Song et, 2021) dxt “ μpxt, t q dt ` σptq dwt, (1) where t P r 0, T s, T ą 0 is a fixed constant, μp¨, ¨q and σp¨q are the drift and diffusion coefficients respectively, and twtutPr 0,T s denotes the standard Brownian motion.
We denote the distribution of xt as ptpxq and as a result p0pxq ” pdata pxq. A remarkable property of this SDE is the existence of an ordinary differential equation (ODE), dubbed the Probability Flow (PF) ODE by Song et al. (2021), whose solution trajectories sampled at t are dis-tributed according to ptpxq: dxt “ „ μpxt, t q ´ 12 σptq2∇ log ptpxtq ȷ dt. (2) Here ∇ log ptpxq is the score function of ptpxq; hence dif-fusion models are also known as score-based generative models (Song & Ermon, 2019; 2020; Song et, 2021).
Typically, the SDE in Eq. (1) is designed such that pT pxq is close to a tractable Gaussian distribution πpxq. We hereafter adopt the settings in Karras et al. (2022), where μpx, t q “ 0 and σptq 2t. In this case, we have ptpxq “ pdata pxq b N p0, t 2Iq, where b denotes the convo-lution operation, and πpxq “ N p0, T 2Iq. For sampling, we first train a score model sϕpx, t q « ∇ log ptpxq via score matching (Hyv ¨arinen & Dayan, 2005; Vincent, 2011; Song et, 2019; Song & Ermon, 2019; Ho et
, 2020), then plug it into Eq. (2) to obtain an empirical estimate of the PF ODE, which takes the form of dxt dt “ ´ tsϕpxt, t q. (3) We call Eq. (3) the empirical PF ODE. Next, we sample ˆxT „ π “ N p0, T 2Iq to initialize the empirical PF ODE and solve it backwards in time with any numerical ODE solver, such as Euler (Song et, 2020; 2021) and Heun solvers (Karras et, 2022), to obtain the solution trajectory tˆxtutPr 0,T s. The resulting ˆx0 can then be viewed as an approximate sample from the data distribution pdata pxq.
To avoid numerical instability, one typically stops the solver at t “ ϵ, where ϵ is a fixed small positive number, and accepts ˆxϵ as the approximate sample. Following Karras et al. (2022), we rescale image pixel values to r´ 1, 1s, and set T “ 80, ϵ “ 0.002.2Consistency Models Figure 2: Consistency models are trained to map points on any trajectory of the PF ODE to the trajectory’s origin. Diffusion models are bottlenecked by their slow sampling speed. Clearly, using ODE solvers for sampling requires iterative evaluations of the score model sϕpx, t q, which is computationally costly.
Existing methods for fast sampling include faster numerical ODE solvers (Song et, 2020; Zhang & Chen, 2022; Lu et, 2022; Dockhorn et, 2022), and distillation techniques (Luhman & Luhman, 2021; Sali-mans & Ho, 2022; Meng et, 2022; Zheng et, 2022). However, ODE solvers still need more than 10 evaluation steps to generate competitive samples. Most distillation methods like Luhman & Luhman (2021) and Zheng et al. (2022) rely on collecting a large dataset of samples from the diffusion model prior to distillation, which itself is com-putationally expensive. To our best knowledge, the only distillation approach that does not suffer from this drawback is progressive distillation (PD, Salimans & Ho (2022)), with which we compare consistency models extensively in our experiments.
## 3. Consistency Models We propose consistency models, a new type of models that support single-step generation at the core of its design, while still allowing iterative generation for trade-offs between sam-ple quality and compute, and zero-shot data editing. Consis-tency models can be trained in either the distillation mode or the isolation mode. In the former case, consistency models distill the knowledge of pre-trained diffusion models into a single-step sampler, significantly improving other distilla-tion approaches in sample quality, while allowing zero-shot image editing applications. In the latter case, consistency models are trained in isolation, with no dependence on pre-trained diffusion models.
This makes them an independent new class of generative models. Below we introduce the definition, parameterization, and sampling of consistency models, plus a brief discussion on their applications to zero-shot data editing. Definition Given a solution trajectory txtutPr ϵ,T s of the PF ODE in Eq. (2), we define the consistency function as f: pxt, t q Þ Ñ xϵ. A consistency function has the property of self-consistency: its outputs are consistent for arbitrary pairs of pxt, t q that belong to the same PF ODE trajectory,, f pxt, t q “ f pxt1, t 1q for all t, t 1 P r ϵ, T s.
As illustrated in Fig. 2, the goal of a consistency model, symbolized as fθ, is to estimate this consistency function f from data by learning to enforce the self-consistency property (details in Sections 4 and 5). Note that a similar definition is used for neural flows (Bilo ˇs et, 2021) in the context of neural ODEs (Chen et, 2018). Compared to neural flows, how-ever, we do not enforce consistency models to be invertible. Parameterization For any consistency function f p¨, ¨q, we have f pxϵ, ϵ q “ xϵ,, f p¨, ϵ q is an identity function.
We call this constraint the boundary condition. All consistency models have to meet this boundary condition, as it plays a crucial role in the successful training of consistency models. This boundary condition is also the most confining archi-tectural constraint on consistency models. For consistency models based on deep neural networks, we discuss two ways to implement this boundary condition almost for Suppose we have a free-form deep neural network Fθ px, t q whose output has the same dimensionality as x. The first way is to simply parameterize the consistency model as fθ px, t q “ # x t “ ϵFθ px, t q t P p ϵ, T s.
(4) The second method is to parameterize the consistency model using skip connections, that is, fθ px, t q “ cskip ptqx ` cout ptqFθ px, t q, (5) where cskip ptq and cout ptq are differentiable functions such that cskip pϵq “ 1, and cout pϵq “ 0. This way, the consistency model is differentiable at t “ ϵ if Fθ px, t q, c skip ptq, c out ptq are all differentiable, which is criti-cal for training continuous-time consistency models (Appen-dices 1 and 2). The parameterization in Eq. (5) bears strong resemblance to many successful diffusion models (Karras et
, 2022; Balaji et, 2022), making it easier to borrow powerful diffusion model architectures for construct-ing consistency models. We therefore follow the second parameterization in all experiments. Sampling With a well-trained consistency model fθ p¨, ¨q,we can generate samples by sampling from the initial dis-tribution ˆxT „ N p0, T 2Iq and then evaluating the consis-tency model for ˆxϵ “ fθ pˆxT, T q. This involves only one forward pass through the consistency model and therefore generates samples in a single step. Importantly, one can also evaluate the consistency model multiple times by al-ternating denoising and noise injection steps for improved sample quality.
Summarized in Algorithm 1, this multistep sampling procedure provides the flexibility to trade com-pute for sample quality. It also has important applications in zero-shot data editing. In practice, we find time points 3Consistency Models Algorithm 1 Multistep Consistency Sampling Input: Consistency model fθ p¨, ¨q, sequence of time points τ1 ą τ2 ą ¨ ¨ ¨ ą τN ´1, initial noise ˆxT x Ð fθ pˆxT, T q for n “ 1 to N ´ 1 do Sample z „ N p0, Iq ˆxτn Ð x ` aτ 2 > n ´ ϵ2zx Ð fθ pˆxτn, τ nq end for Output: x tτ1, τ 2, ¨ ¨ ¨, τ N ´1u in Algorithm 1 with a greedy algorithm, where the time points are pinpointed one at a time using ternary search to optimize the FID of samples obtained from Algorithm 1.
This assumes that given prior time points, the FID is a unimodal function of the next time point. We find this assumption to hold empirically in our experiments, and leave the exploration of better strategies as future work. Zero-Shot Data Editing Similar to diffusion models, con-sistency models enable various data editing and manipu-lation applications in zero shot; they do not require ex-plicit training to perform these tasks. For example, consis-tency models define a one-to-one mapping from a Gaussian noise vector to a data sample. Similar to latent variable models like GANs, VAEs, and normalizing flows, consis-tency models can easily interpolate between samples by traversing the latent space (Fig.
11). As consistency models are trained to recover xϵ from any noisy input xt where t P r ϵ, T s, they can perform denoising for various noise levels (Fig. 12). Moreover, the multistep generation pro-cedure in Algorithm 1 is useful for solving certain inverse problems in zero shot by using an iterative replacement pro-cedure similar to that of diffusion models (Song & Ermon, 2019; Song et, 2021; Ho et, 2022b). This enables many applications in the context of image editing, including inpainting (Fig. 10), colorization (Fig. 8), super-resolution (Fig. 6b) and stroke-guided image editing (Fig.
13) as in SDEdit (Meng et, 2021). In Section 6.3, we empiri-cally demonstrate the power of consistency models on many zero-shot image editing tasks. ## 4. Training Consistency Models via Distillation We present our first method for training consistency mod-els based on distilling a pre-trained score model sϕpx, t Our discussion revolves around the empirical PF ODE in Eq. (3), obtained by plugging the score model sϕpx, t q into the PF ODE. Consider discretizing the time horizon rϵ, T s into N ´ 1 sub-intervals, with boundaries t1 “ ϵ ă t2 ă ¨ ¨ ¨ ă tN “ T.
In practice, we follow Karras et al. (2022) to determine the boundaries with the formula ti “ p ϵ1{ρ ` i´1{N ´1pT 1{ρ ´ ϵ1{ρqq ρ, where ρ “ 7. When N is sufficiently large, we can obtain an accurate estimate of xtn from xtn`1 by running one discretization step of a numerical ODE solver. This estimate, which we denote as ˆxϕ > tn, is defined by ˆxϕ > tn:“ xtn`1 ` p tn ´ tn`1qΦpxtn`1, t n`1; ϕq, (6) where Φp¨ ¨ ¨; ϕq represents the update function of a one-step ODE solver applied to the empirical PF ODE.
For example, when using the Euler solver, we have Φpx, t; ϕq “ ´tsϕpx, t q which corresponds to the following update rule ˆxϕ > tn “ xtn`1 ´ p tn ´ tn`1qtn`1sϕpxtn`1, t n`1q. For simplicity, we only consider one-step ODE solvers in this work. It is straightforward to generalize our framework to multistep ODE solvers and we leave it as future work. Due to the connection between the PF ODE in Eq. (2) and the SDE in Eq. (1) (see Section 2), one can sample along the distribution of ODE trajectories by first sampling x „ pdata,then adding Gaussian noise to x.
Specifically, given a data point x, we can generate a pair of adjacent data points pˆxϕ > tn, xtn`1 q on the PF ODE trajectory efficiently by sam-pling x from the dataset, followed by sampling xtn`1 from the transition density of the SDE N px, t 2 > n`1 Iq, and then computing ˆxϕ > tn using one discretization step of the numeri-cal ODE solver according to Eq. (6). Afterwards, we train the consistency model by minimizing its output differences on the pair pˆxϕ > tn, xtn`1 q. This motivates our following con-sistency distillation loss for training consistency models.
Definition 1. The consistency distillation loss is defined as LN > CD pθ, θ´; ϕq:“ Erλptnqdpfθ pxtn`1, t n`1q, fθ´ pˆxϕ > tn, t nqqs, (7) where the expectation is taken with respect to x „ pdata, n „ UJ1, N ´1K, and xtn`1 „ N px; t2 > n`1 Iq. Here UJ1, N ´1K denotes the uniform distribution over t1, 2, ¨ ¨ ¨, N ´ 1u, λp¨q P R` is a positive weighting function, ˆxϕ > tn is given by Eq. (6), θ´ denotes a running average of the past values of θ during the course of optimization, and dp¨, ¨q is a metric function that satisfies @x, y: dpx, yq ě 0 and dpx, yq “ 0 if and only if x “ y.
Unless otherwise stated, we adopt the notations in Defi-nition 1 throughout this paper, and use Er¨s to denote the expectation over all random variables. In our experiments, we consider the squared ℓ2 distance dpx, yq “} x ´ y}22, ℓ1 distance dpx, yq “} x ´ y}1, and the Learned Perceptual Image Patch Similarity (LPIPS, Zhang et al. (2018)). We find λptnq ” 1 performs well across all tasks and datasets. In practice, we minimize the objective by stochastic gradient descent on the model parameters θ, while updating θ´ with exponential moving average (EMA). That is, given a decay 4Consistency Models Algorithm 2 Consistency Distillation (CD) Input: dataset D, initial model parameter θ, learning rate η, ODE solver Φp¨, ¨; ϕq, dp¨, ¨q, λp¨q, and μ θ´ Ð θ repeat Sample x „ D and n „ UJ1, N ´ 1K Sample xtn`1 „ N px; t2 > n`1 Iq ˆxϕ > tn Ð xtn`1 ` p tn ´ tn`1qΦpxtn`1, t n`1; ϕq Lpθ, θ´; ϕq Ð λptnqdpfθ pxtn`1, t n`1q, fθ´ pˆxϕ > tn, t nqq θ Ð θ ´ η∇θ Lpθ, θ´; ϕq θ´ Ð stopgrad pμθ´ ` p 1 ´ μqθ) until convergence rate 0 ď μ ă 1, we perform the following update after each optimization step: θ´ Ð stopgrad pμθ´ ` p 1 ´ μqθq.
(8) The overall training procedure is summarized in Algo-rithm 2. In alignment with the convention in deep reinforce-ment learning (Mnih et, 2013; 2015; Lillicrap et, 2015) and momentum based contrastive learning (Grill et, 2020; He et, 2020), we refer to fθ´ as the “target network”, and fθ as the “online network”. We find that compared to simply setting θ´ “ θ, the EMA update and “stopgrad” operator in Eq. (8) can greatly stabilize the training process and improve the final performance of the consistency model. Below we provide a theoretical justification for consistency distillation based on asymptotic analysis.
Theorem 1. Let ∆t:“ max nPJ1,N ´1Kt| tn`1 ´ tn|u, and f p¨, ¨; ϕq be the consistency function of the empirical PF ODE in Eq. (3). Assume fθ satisfies the Lipschitz condition: there exists L ą 0 such that for all t P r ϵ, T s, x, and y,we have ∥fθ px, t q ´ fθ py, t q∥2 ď L ∥x ´ y∥2. Assume further that for all n P J1, N ´ 1K, the ODE solver called at tn`1 has local error uniformly bounded by Opp tn`1 ´ tnqp`1q with p ě 1.
Then, if LN > CD pθ, θ; ϕq “ 0, we have sup > n, x}fθ px, t nq ´ f px, t n; ϕq} 2 “ Opp ∆tqpq. Proof. The proof is based on induction and parallels the classic proof of global error bounds for numerical ODE solvers (S ¨uli & Mayers, 2003). We provide the full proof in Appendix 2. Since θ´ is a running average of the history of θ, we have θ´ “ θ when the optimization of Algorithm 2 converges. That is, the target and online consistency models will eventu-ally match each other.
If the consistency model additionally achieves zero consistency distillation loss, then Theorem 1 Algorithm 3 Consistency Training (CT) Input: dataset D, initial model parameter θ, learning rate η, step schedule N p¨q, EMA decay rate schedule μp¨q, dp¨, ¨q, and λp¨q θ´ Ð θ and k Ð 0 repeat Sample x „ D, and n „ UJ1, N pkq ´ 1K Sample z „ N p0, Iq Lpθ, θ´q Ð λptnqdpfθ px ` tn`1z, t n`1q, fθ´ px ` tnz, t nqq θ Ð θ ´ η∇θ Lpθ, θ´q θ´ Ð stopgrad pμpkqθ´ ` p 1 ´ μpkqq θq k Ð k ` 1 until convergence implies that, under some regularity conditions, the estimated consistency model can become arbitrarily accurate, as long as the step size of the ODE solver is sufficiently small.
Im-portantly, our boundary condition fθ px, ϵ q ” x precludes the trivial solution fθ px, t q ” 0 from arising in consistency model training. The consistency distillation loss LN > CD pθ, θ´; ϕq can be ex-tended to hold for infinitely many time steps (N Ñ 8) if θ´ “ θ or θ´ “ stopgrad pθq. The resulting continuous-time loss functions do not require specifying N nor the time steps tt1, t 2, ¨ ¨ ¨, t N u. Nonetheless, they involve Jacobian-vector products and require forward-mode automatic dif-ferentiation for efficient implementation, which may not be well-supported in some deep learning frameworks.
We provide these continuous-time distillation loss functions in Theorems 3 to 5, and relegate details to Appendix 1. ## 5. Training Consistency Models in Isolation Consistency models can be trained without relying on any pre-trained diffusion models. This differs from existing diffusion distillation techniques, making consistency models a new independent family of generative models. Recall that in consistency distillation, we rely on a pre-trained score model sϕpx, t q to approximate the ground truth score function ∇ log ptpxq. It turns out that we can avoid this pre-trained score model altogether by leveraging the following unbiased estimator (Lemma 1 in Appendix A): ∇ log ptpxtq “ ´ E „ xt ´ x t2 ˇˇˇˇ xt ȷ, where x „ pdata and xt „ N px; t2Iq.
That is, given x and xt, we can estimate ∇ log ptpxtq with ´p xt ´ This unbiased estimate suffices to replace the pre-trained diffusion model in consistency distillation when using the Euler method as the ODE solver in the limit of N Ñ 8, as 5Consistency Models justified by the following result. Theorem 2. Let ∆t:“ max nPJ1,N ´1Kt| tn`1 ´ tn|u. As-sume d and fθ´ are both twice continuously differentiable with bounded second derivatives, the weighting function λp¨q is bounded, and Er∥∇ log ptn pxtn q∥22s ă 8. As-sume further that we use the Euler ODE solver, and the pre-trained score model matches the ground truth,
, @t P r ϵ, T s: sϕpx, t q ” ∇ log ptpxq. Then, LN > CD pθ, θ´; ϕq “ LN > CT pθ, θ´q ` op∆tq, (9) where the expectation is taken with respect to x „ pdata, n „ UJ1, N ´ 1K, and xtn`1 „ N px; t2 > n`1 Iq. The consistency training objective, denoted by LN > CT pθ, θ´q, is defined as Erλptnqdpfθ px ` tn`1z, t n`1q, fθ´ px ` tnz, t nqqs, (10) where z „ N p0, Iq. Moreover, LN > CT pθ, θ´q ě Op∆tq if inf N LN > CD pθ, θ´; ϕq ą
Proof. The proof is based on Taylor series expansion and properties of score functions (Lemma 1). A complete proof is provided in Appendix 3. We refer to Eq. (10) as the consistency training (CT) loss. Crucially, Lpθ, θ´q only depends on the online network fθ, and the target network fθ´, while being completely agnostic to diffusion model parameters ϕ. The loss function Lpθ, θ´q ě Op∆tq decreases at a slower rate than the remainder op∆tq and thus will dominate the loss in Eq. (9) as N Ñ 8 and ∆t Ñ For improved practical performance, we propose to progres-sively increase N during training according to a schedule function N p¨q.
The intuition, Fig. 3d) is that the consis-tency training loss has less “variance” but more “bias” with respect to the underlying consistency distillation loss, the left-hand side of Eq. (9)) when N is small, ∆t is large), which facilitates faster convergence at the beginning of training. On the contrary, it has more “variance” but less “bias” when N is large, ∆t is small), which is desirable when closer to the end of training. For best performance, we also find that μ should change along with N, according to a schedule function μp¨q.
The full algorithm of consis-tency training is provided in Algorithm 3, and the schedule functions used in our experiments are given in Appendix C. Similar to consistency distillation, the consistency training loss LN > CT pθ, θ´q can be extended to hold in continuous time, N Ñ 8) if θ´ “ stopgrad pθq, as shown in Theo-rem 6. This continuous-time loss function does not require schedule functions for N or μ, but requires forward-mode automatic differentiation for efficient implementation. Un-like the discrete-time CT loss, there is no undesirable “bias” associated with the continuous-time objective, as we effec-tively take ∆t Ñ 0 in Theorem 2.
We relegate more details to Appendix 2. ## 6. Experiments We employ consistency distillation and consistency train-ing to learn consistency models on real image datasets, including CIFAR-10 (Krizhevsky et, 2009), ImageNet 64 ˆ 64 (Deng et, 2009), LSUN Bedroom 256 ˆ 256,and LSUN Cat 256 ˆ 256 (Yu et, 2015). Results are compared according to Fr ´echet Inception Distance (FID, Heusel et al. (2017), lower is better), Inception Score (IS, Salimans et al. (2016), higher is better), Precision, Kynk ¨a ¨anniemi et al. (2019), higher is better), and Recall, Kynk ¨a ¨anniemi et al.
(2019), higher is better). Addi-tional experimental details are provided in Appendix C. 6.1. Training Consistency Models We perform a series of experiments on CIFAR-10 to under-stand the effect of various hyperparameters on the perfor-mance of consistency models trained by consistency distil-lation (CD) and consistency training (CT). We first focus on the effect of the metric function dp¨, ¨q, the ODE solver, and the number of discretization steps N in CD, then investigate the effect of the schedule functions N p¨q and μp¨q in CT. To set up our experiments for CD, we consider the squared ℓ2 distance dpx, yq “} x ´ y}22, ℓ1 distance dpx, yq “}x ´ y}1, and the Learned Perceptual Image Patch Simi-larity (LPIPS, Zhang et al.
(2018)) as the metric function. For the ODE solver, we compare Euler’s forward method and Heun’s second order method as detailed in Karras et al. (2022). For the number of discretization steps N, we com-pare N P t 9, 12, 18, 36, 50, 60, 80, 120 u. All consistency models trained by CD in our experiments are initialized with the corresponding pre-trained diffusion models, whereas models trained by CT are randomly initialized. As visualized in Fig. 3a, the optimal metric for CD is LPIPS, which outperforms both ℓ1 and ℓ2 by a large margin over all training iterations.
This is expected as the outputs of consistency models are images on CIFAR-10, and LPIPS is specifically designed for measuring the similarity between natural images. Next, we investigate which ODE solver and which discretization step N work the best for CD. As shown in Figs. 3b and 3c, Heun ODE solver and N “ 18 are the best choices. Both are in line with the recommendation of Karras et al. (2022) despite the fact that we are train-ing consistency models, not diffusion models. Moreover, Fig. 3b shows that with the same N, Heun’s second order solver uniformly outperforms Euler’s first order solver.
This corroborates with Theorem 1, which states that the optimal consistency models trained by higher order ODE solvers have smaller estimation errors with the same N. The results of Fig. 3c also indicate that once N is sufficiently large, the performance of CD becomes insensitive to N. Given these insights, we hereafter use LPIPS and Heun ODE solver for CD unless otherwise stated. For N in CD, we follow the 6Consistency Models > (a) Metric functions in CD. (b) Solvers and Nin CD. (c) Nwith Heun solver in CD. (d) Adaptive Nand μin CT. Figure 3: Various factors that affect consistency distillation (CD) and consistency training (CT) on CIFAR-10.
The best configuration for CD is LPIPS, Heun ODE solver, and N “ 18. Our adaptive schedule functions for N and μ make CT converge significantly faster than fixing them to be constants during the course of optimization. > (a) CIFAR-10 (b) ImageNet 64 ˆ64 (c) Bedroom 256 ˆ256 (d) Cat 256 ˆ256 Figure 4: Multistep image generation with consistency distillation (CD). CD outperforms progressive distillation (PD) across all datasets and sampling steps. The only exception is single-step generation on Bedroom 256 ˆ suggestions in Karras et al. (2022) on CIFAR-10 and Im-ageNet 64 ˆ 64.
We tune N separately on other datasets (details in Appendix C). Due to the strong connection between CD and CT, we adopt LPIPS for our CT experiments throughout this paper. Unlike CD, there is no need for using Heun’s second order solver in CT as the loss function does not rely on any particular numerical ODE solver. As demonstrated in Fig. 3d, the con-vergence of CT is highly sensitive to N —smaller N leads to faster convergence but worse samples, whereas larger N leads to slower convergence but better samples upon convergence. This matches our analysis in Section 5, and motivates our practical choice of progressively growing N and μ for CT to balance the trade-off between convergence speed and sample quality.
As shown in Fig. 3d, adaptive schedules of N and μ significantly improve the convergence speed and sample quality of CT. In our experiments, we tune the schedules N p¨q and μp¨q separately for images of different resolutions, with more details in Appendix C. 6.2. Few-Step Image Generation Distillation In current literature, the most directly compara-ble approach to our consistency distillation (CD) is progres-sive distillation (PD, Salimans & Ho (2022)); both are thus far the only distillation approaches that do not construct synthetic data before distillation. In stark contrast, other dis-tillation techniques, such as knowledge distillation (Luhman & Luhman, 2021) and DFNO (Zheng et
, 2022), have to prepare a large synthetic dataset by generating numerous samples from the diffusion model with expensive numerical ODE/SDE solvers. We perform comprehensive comparison for PD and CD on CIFAR-10, ImageNet 64 ˆ64, and LSUN 256 ˆ 256, with all results reported in Fig. 4. All methods distill from an EDM (Karras et, 2022) model that we pre-trained in-house. We note that across all sampling iterations, using the LPIPS metric uniformly improves PD compared to the squared ℓ2 distance in the original paper of Salimans & Ho (2022). Both PD and CD improve as we take more sampling steps.
We find that CD uniformly outperforms PD across all datasets, sampling steps, and metric functions considered, except for single-step generation on Bedroom 256 ˆ 256, where CD with ℓ2 slightly underperforms PD with ℓ2. As shown in Table 1, CD even outperforms distilla-tion approaches that require synthetic dataset construction, such as Knowledge Distillation (Luhman & Luhman, 2021) and DFNO (Zheng et, 2022). Direct Generation In Tables 1 and 2, we compare the sample quality of consistency training (CT) with other gen-erative models using one-step and two-step generation. We also include PD and CD results for reference.
Both tables re-port PD results obtained from the ℓ2 metric function, as this is the default setting used in the original paper of Salimans 7Consistency Models Table 1: Sample quality on CIFAR-10. ˚Methods that require synthetic data construction for distillation. METHOD NFE (Ó) FID (Ó) IS (Ò) Diffusion + Samplers DDIM (Song et, 2020) 50 4.67 DDIM (Song et, 2020) 20 6.84 DDIM (Song et, 2020) 10 8.23 DPM-solver-2 (Lu et, 2022) 10 5.94 DPM-solver-fast (Lu et, 2022) 10 4.70 3-DEIS (Zhang & Chen, 2022) 10 4.17 Diffusion + Distillation Knowledge Distillation ˚ (Luhman & Luhman, 2021) 1 9.36 DFNO ˚ (Zheng et
, 2022) 1 4.12 1-Rectified Flow (+distill) ˚ (Liu et, 2022) 1 6.18 9.08 2-Rectified Flow (+distill) ˚ (Liu et, 2022) 1 4.85 9.01 3-Rectified Flow (+distill) ˚ (Liu et, 2022) 1 5.21 8.79 PD (Salimans & Ho, 2022) 1 8.34 8.69 CD 1 3.55 9.48 PD (Salimans & Ho, 2022) 2 5.58 9.05 CD 2 2.93 9.75 Direct Generation BigGAN (Brock et, 2019) 1 14.7 9.22 Diffusion GAN (Xiao et, 2022) 1 14.6 8.93 AutoGAN (Gong et, 2019) 1 12.4 8.55 E2GAN (Tian et, 2020) 1 11.3 8.51 ViTGAN (Lee et
, 2021) 1 6.66 9.30 TransGAN (Jiang et, 2021) 1 9.26 9.05 StyleGAN2-ADA (Karras et, 2020) 1 2.92 9.83 StyleGAN-XL (Sauer et, 2022) 1 1.85 Score SDE (Song et, 2021) 2000 2.20 9.89 DDPM (Ho et, 2020) 1000 3.17 9.46 LSGM (Vahdat et, 2021) 147 2.10 PFGM (Xu et, 2022) 110 2.35 9.68 EDM (Karras et, 2022) 35 2.04 9.84 1-Rectified Flow (Liu et, 2022) 1 378 1.13 Glow (Kingma & Dhariwal, 2018) 1 48.9 3.92 Residual Flow (Chen et, 2019) 1 46.4 GLFlow (Xiao et
, 2019) 1 44.6 DenseFlow (Grci´ c et, 2021) 1 34.9 DC-VAE (Parmar et, 2021) 1 17.9 8.20 CT 1 8.70 8.49 CT 2 5.83 8.85 Table 2: Sample quality on ImageNet 64 ˆ 64, and LSUN Bedroom & Cat 256 ˆ:Distillation techniques. METHOD NFE (Ó) FID (Ó) Prec. (Ò) Rec. (Ò) ImageNet 64 ˆ 64 PD: (Salimans & Ho, 2022) 1 15.39 0.59 0.62 DFNO: (Zheng et, 2022) 1 8.35 CD: 1 6.20 0.68 0.63 PD: (Salimans & Ho, 2022) 2 8.95 0.63 0.65 CD: 2 4.70 0.69 0.64 ADM (Dhariwal & Nichol, 2021) 250 2.07 0.74 0.63 EDM (Karras et
, 2022) 79 2.44 0.71 0.67 BigGAN-deep (Brock et, 2019) 1 4.06 0.79 0.48 CT 1 13.0 0.71 0.47 CT 2 11.1 0.69 0.56 LSUN Bedroom 256 ˆ 256 PD: (Salimans & Ho, 2022) 1 16.92 0.47 0.27 PD: (Salimans & Ho, 2022) 2 8.47 0.56 0.39 CD: 1 7.80 0.66 0.34 CD: 2 5.22 0.68 0.39 DDPM (Ho et, 2020) 1000 4.89 0.60 0.45 ADM (Dhariwal & Nichol, 2021) 1000 1.90 0.66 0.51 EDM (Karras et, 2022) 79 3.57 0.66 0.45 PGGAN (Karras et, 2018) 1 8.34 PG-SWGAN (Wu et
, 2019) 1 8.0 TDPM (GAN) (Zheng et, 2023) 1 5.24 StyleGAN2 (Karras et, 2020) 1 2.35 0.59 0.48 CT 1 16.0 0.60 0.17 CT 2 7.85 0.68 0.33 LSUN Cat 256 ˆ 256 PD: (Salimans & Ho, 2022) 1 29.6 0.51 0.25 PD: (Salimans & Ho, 2022) 2 15.5 0.59 0.36 CD: 1 11.0 0.65 0.36 CD: 2 8.84 0.66 0.40 DDPM (Ho et, 2020) 1000 17.1 0.53 0.48 ADM (Dhariwal & Nichol, 2021) 1000 5.57 0.63 0.52 EDM (Karras et, 2022) 79 6.69 0.70 0.43 PGGAN (Karras et, 2018) 1 37.5 StyleGAN2 (Karras et
, 2020) 1 7.25 0.58 0.43 CT 1 20.7 0.56 0.23 CT 2 11.7 0.63 0.36 Figure 5: Samples generated by EDM (top), CT + single-step generation (middle), and CT + 2-step generation (Bottom). All corresponding images are generated from the same initial noise. 8Consistency Models > (a) Left: The gray-scale image. Middle: Colorized images. Right: The ground-truth image. > (b) Left: The downsampled image (32 ˆ32). Middle: Full resolution images (256 ˆ256). Right: The ground-truth image (256 ˆ256). > (c) Left: A stroke input provided by users. Right: Stroke-guided image generation. Figure 6: Zero-shot image editing with a consistency model trained by consistency distillation on LSUN Bedroom 256
& Ho (2022). For fair comparison, we ensure PD and CD distill the same EDM models. In Tables 1 and 2, we observe that CT outperforms existing single-step, non-adversarial generative models,, VAEs and normalizing flows, by a significant margin on CIFAR-10. Moreover, CT achieves comparable quality to one-step samples from PD without relying on distillation. In Fig. 5, we provide EDM samples (top), single-step CT samples (middle), and two-step CT samples (bottom). In Appendix E, we show additional sam-ples for both CD and CT in Figs. 14 to 21. Importantly, all samples obtained from the same initial noise vector share significant structural similarity, even though CT and EDM models are trained independently from one another.
This indicates that CT is less likely to suffer from mode collapse, as EDMs do not. 6.3. Zero-Shot Image Editing Similar to diffusion models, consistency models allow zero-shot image editing by modifying the multistep sampling process in Algorithm 1. We demonstrate this capability with a consistency model trained on the LSUN bedroom dataset using consistency distillation. In Fig. 6a, we show such a consistency model can colorize gray-scale bedroom images at test time, even though it has never been trained on colorization tasks. In Fig. 6b, we show the same con-sistency model can generate high-resolution images from low-resolution inputs.
In Fig. 6c, we additionally demon-strate that it can generate images based on stroke inputs cre-ated by humans, as in SDEdit for diffusion models (Meng et, 2021). Again, this editing capability is zero-shot, as the model has not been trained on stroke inputs. In Appendix D, we additionally demonstrate the zero-shot capability of consistency models on inpainting (Fig. 10), interpolation (Fig. 11) and denoising (Fig. 12), with more examples on colorization (Fig. 8), super-resolution (Fig. 9) and stroke-guided image generation (Fig. 13). ## 7. Conclusion We have introduced consistency models, a type of generative models that are specifically designed to support one-step and few-step generation.
We have empirically demonstrated that our consistency distillation method outshines the exist-ing distillation techniques for diffusion models on multiple image benchmarks and small sampling iterations. Further-more, as a standalone generative model, consistency models generate better samples than existing single-step genera-tion models except for GANs. Similar to diffusion models, they also allow zero-shot image editing applications such as inpainting, colorization, super-resolution, denoising, inter-polation, and stroke-guided image generation. In addition, consistency models share striking similarities with techniques employed in other fields, including deep Q-learning (Mnih et, 2015) and momentum-based con-trastive learning (Grill et, 2020; He et
, 2020). This offers exciting prospects for cross-pollination of ideas and methods among these diverse fields. ## Acknowledgements We thank Alex Nichol for reviewing the manuscript and providing valuable feedback, Chenlin Meng for providing stroke inputs needed in our stroke-guided image generation experiments, and the OpenAI Algorithms team. 9Consistency Models ## References Balaji,, Nah,, Huang,, Vahdat,, Song,, Kreis,, Aittala,, Aila,, Laine,, Catanzaro,, Kar-ras,, and Liu, -Y. ediff-i: Text-to-image diffusion models with ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022. Bilo ˇs,, Sommer,, Rangapuram, S.
, Januschowski,, and G ¨unnemann, S. Neural flows: Efficient alternative to neural odes. Advances in Neural Information Processing Systems, 34:21325–21337, 2021. Brock,, Donahue,, and Simonyan, K. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations,2019. URL net/forum? Chen,, Zhang,, Zen,, Weiss, R., Norouzi,, and Chan, W. Wavegrad: Estimating gradients for waveform generation. In International Conference on Learning Representations (ICLR), 2021. Chen, R., Rubanova,, Bettencourt,, and Duvenaud, D. K. Neural Ordinary Differential Equations. In Ad-vances in neural information processing systems, pp.
6571–6583, 2018. Chen, R., Behrmann,, Duvenaud, D., and Jacobsen, -H. Residual flows for invertible generative modeling. In Advances in Neural Information Processing Systems,pp. 9916–9926, 2019. Chung,, Kim,, Mccann, M., Klasky, M., and Ye, J. C. Diffusion posterior sampling for general noisy in-verse problems. In International Conference on Learning Representations, 2023. URL https://openreview. Deng,, Dong,, Socher,, Li,, Li,, and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009. Dhariwal, P. and Nichol, A.
Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems (NeurIPS), 2021. Dinh,, Krueger,, and Bengio, Y. NICE: Non-linear independent components estimation. International Con-ference in Learning Representations Workshop Track,2015. Dinh,, Sohl-Dickstein,, and Bengio, S. Density es-timation using real NVP. In 5th International Confer-ence on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track net, 2017. URL https://openreview. Dockhorn,, Vahdat,, and Kreis, K. Genie: Higher-order denoising diffusion solvers. arXiv preprint arXiv:2210.05475, 2022. Gong,, Chang,, Jiang,, and Wang, Z. Autogan: Neural architecture search for generative adversarial net-works.
In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3224–3234, 2019. Goodfellow,, Pouget-Abadie,, Mirza,, Xu,, Warde-Farley,, Ozair,, Courville,, and Bengio, Y. Generative adversarial nets. In Advances in neural information processing systems, pp. 2672–2680, 2014. Grci ´c,, Grubi ˇsi ´c,, and ˇSegvi ´c, S. Densely connected normalizing flows. Advances in Neural Information Pro-cessing Systems, 34:23968–23982, 2021. Grill,, Strub,, Altch ´e,, Tallec,, Richemond,, Buchatskaya,, Doersch,, Avila Pires,, Guo,, Gheshlaghi Azar,, et al. Bootstrap your own latent-a new approach to self-supervised learning.
Advances in neural information processing systems, 33:21271–21284, 2020. He,, Fan,, Wu,, Xie,, and Girshick, R. Mo-mentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9729–9738, 2020. Heusel,, Ramsauer,, Unterthiner,, Nessler,, and Hochreiter, S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems, pp. 6626–6637, 2017. Ho,, Jain,, and Abbeel, P. Denoising Diffusion Proba-bilistic Models. Advances in Neural Information Process-ing Systems, 33, 2020.
Ho,, Chan,, Saharia,, Whang,, Gao,, Gritsenko,, Kingma, D., Poole,, Norouzi,, Fleet, D., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303,2022a. Ho,, Salimans,, Gritsenko, A., Chan,, Norouzi,, and Fleet, D. J. Video diffusion models. In ICLR Workshop on Deep Generative Models for Highly Struc-tured Data, 2022b. URL https://openreview. Hyv ¨arinen, A. and Dayan, P. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research (JMLR), 6(4), 2005. 10 Consistency Models Jiang,, Chang,
, and Wang, Z. Transgan: Two pure transformers can make one strong gan, and that can scale up. Advances in Neural Information Processing Systems,34:14745–14758, 2021. Karras,, Aila,, Laine,, and Lehtinen, J. Progres-sive growing of GANs for improved quality, stability, and variation. In International Conference on Learning Representations, 2018. URL https://openreview. Karras,, Laine,, Aittala,, Hellsten,, Lehtinen,, and Aila, T. Analyzing and improving the image quality of stylegan. 2020. Karras,, Aittala,, Aila,, and Laine, S. Elucidating the design space of diffusion-based generative models. In Proc. NeurIPS, 2022.
Kawar,, Vaksman,, and Elad, M. Snips: Solving noisy inverse problems stochastically. arXiv preprint arXiv:2105.14951, 2021. Kawar,, Elad,, Ermon,, and Song, J. Denoising diffusion restoration models. In Advances in Neural In-formation Processing Systems, 2022. Kingma, D. P. and Dhariwal, P. Glow: Generative flow with invertible 1x1 convolutions. In Bengio,, Wal-lach,, Larochelle,, Grauman,, Cesa-Bianchi,, and Garnett, R.), Advances in Neural Information Processing Systems 31, pp. 10215–10224. 2018. Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In International Conference on Learning Repre-sentations, 2014. Kong,, Ping,
, Huang,, Zhao,, and Catanzaro, B. DiffWave: A Versatile Diffusion Model for Audio Synthesis. arXiv preprint arXiv:2009.09761, 2020. Krizhevsky,, Hinton,, et al. Learning multiple layers of features from tiny images. 2009. Kynk ¨a ¨anniemi,, Karras,, Laine,, Lehtinen,, and Aila, T. Improved precision and recall metric for assess-ing generative models. Advances in Neural Information Processing Systems, 32, 2019. Lee,, Chang,, Jiang,, Zhang,, Tu,, and Liu, C. Vitgan: Training gans with vision transformers. arXiv preprint arXiv:2107.04589, 2021. Lillicrap, T., Hunt, J., Pritzel,
, Heess,, Erez,, Tassa,, Silver,, and Wierstra, D. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015. Liu,, Jiang,, He,, Chen,, Liu,, Gao,, and Han, J. On the variance of the adaptive learning rate and beyond. arXiv preprint arXiv:1908.03265, 2019. Liu,, Gong,, and Liu, Q. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022. Lu,, Zhou,, Bao,, Chen,, Li,, and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilis-tic model sampling in around 10 steps.
arXiv preprint arXiv:2206.00927, 2022. Luhman, E. and Luhman, T. Knowledge distillation in iterative generative models for improved sampling speed. arXiv preprint arXiv:2101.02388, 2021. Meng,, Song,, Song,, Wu,, Zhu,, and Ermon, S. Sdedit: Image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073,2021. Meng,, Gao,, Kingma, D., Ermon,, Ho,, and Salimans, T. On distillation of guided diffusion models. arXiv preprint arXiv:2210.03142, 2022. Mnih,, Kavukcuoglu,, Silver,, Graves,, Antonoglou,, Wierstra,, and Riedmiller, M. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
Mnih,, Kavukcuoglu,, Silver,, Rusu, A., Veness,, Bellemare, M., Graves,, Riedmiller,, Fidje-land, A., Ostrovski,, et al. Human-level control through deep reinforcement learning. nature, 518(7540): 529–533, 2015. Nichol,, Dhariwal,, Ramesh,, Shyam,, Mishkin,, McGrew,, Sutskever,, and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. Parmar,, Li,, Lee,, and Tu, Z. Dual contradistinctive generative autoencoder. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp. 823–832, 2021. Popov,
, Vovk,, Gogoryan,, Sadekova,, and Kudi-nov, M. Grad-TTS: A diffusion probabilistic model for text-to-speech. arXiv preprint arXiv:2105.06337, 2021. Ramesh,, Dhariwal,, Nichol,, Chu,, and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. Rezende, D., Mohamed,, and Wierstra, D. Stochastic backpropagation and approximate inference in deep gen-erative models. In Proceedings of the 31st International Conference on Machine Learning, pp. 1278–1286, 2014. 11 Consistency Models Rombach,, Blattmann,, Lorenz,, Esser,, and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Con-ference on Computer Vision and Pattern Recognition, pp.
10684–10695, 2022. Saharia,, Chan,, Saxena,, Li,, Whang,, Denton,, Ghasemipour, S. K., Ayan, B., Mahdavi, S., Lopes, R., et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022. Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. In International Confer-ence on Learning Representations, 2022. URL https: Salimans,, Goodfellow,, Zaremba,, Cheung,, Radford,, and Chen, X. Improved techniques for train-ing gans. In Advances in neural information processing systems, pp. 2234–2242, 2016. Sauer,, Schwarz,, and Geiger, A.
Stylegan-xl: Scaling stylegan to large diverse datasets. In ACM SIGGRAPH 2022 conference proceedings, pp. 1–10, 2022. Sohl-Dickstein,, Weiss,, Maheswaranathan,, and Ganguli, S. Deep Unsupervised Learning Using Nonequi-librium Thermodynamics. In International Conference on Machine Learning, pp. 2256–2265, 2015. Song,, Meng,, and Ermon, S. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
Want to learn more?
Ask about this article