Discover how denoising irreversibility exposes vulnerabilities in diffusion language models and explore strategies to enhance AI safety and robustness.
Explore EMA's role and limits in sequence models, revealing the need for input-dependent selection to improve temporal structure encoding and reduce inform...