What are the benefits of consistency loss in consistency model distillation?

What are the benefits of consistency loss in consistency model distillation?

Manage alerts

Loading saved threads...

Andrea Allais · External communityPost link
External question — Cross Validated Stack Exchange Author: Andrea Allais Original post: https://stats.stackexchange.com/questions/664916 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. When training consistency models with distillation, the loss is designed to drive the model to produce similar outputs on two consecutive points of the discretized probability flow ODE trajectory (eq. 7). Naively, it seems it would be easier to directly minimize the distance between the model output and the end point of the ODE trajectory, which is also available. After all, the defining property of the consistency function $f$ , as defined on page 3, is that it maps noisy data $x_t$ to clean data $x_\epsilon$ . Of course, there must be some reason why this naive approach does not work as well as the consistency loss, but I can't find any discussion of the trade-offs. Can someone help shed some light here?
Quote
Report
Andrea Allais · External communityPost link
External answer — Cross Validated Stack Exchange Author: Andrea Allais Original post: https://stats.stackexchange.com/a/667634 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. You can directly target the end point of the ODE trajectory, but it's very expensive to do so. For every step of the learning process, you need to integrate the ODE, which requires tens to hundreds of evaluations of the teacher model. In contrast, evaluating the consistency loss requires only one evaluation of the teacher model.
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: Andrea Allais Source score (net votes, not local likes): 1 Original post: https://stats.stackexchange.com/questions/664916 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. When training consistency models with distillation, the loss is designed to drive the model to produce similar outputs on two consecutive points of the discretized probability flow ODE trajectory (eq. 7). Naively, it seems it would be easier to directly minimize the distance between the model output and the end point of the ODE trajectory, which is also available. After all, the defining property of the consistency function $f$ , as defined on page 3, is that it maps noisy data $x_t$ to clean data $x_\epsilon$ . Of course, there must be some reason why this naive approach does not work as well as the consistency loss, but I can't find any discussion of the trade-offs. Can someone help shed some light here?

Cancel quote

Checking account access…