AIToday
Large Language ModelsImage GenerationApple Machine LearningPublished: Oct 9, 2026, 01:00 JST

Apple's NTM hits strong image quality in four sampling steps

Apple's NTM hits strong image quality in four sampling steps

3 Key Points

  1. What happened

    Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse step as a conditional normalizing flow with exact likelihood training. On text-to-image benchmarks, NTM matches or outperforms strong image generation baselines in just four sampling steps.

  2. Why it matters

    Existing few-step methods rely on distillation, consistency training, or adversarial objectives but sacrifice the likelihood framework, according to the authors. NTM keeps exact trajectory likelihood, which also enables self-distillation, a lightweight denoiser trained on the model's own score function.

WHO IT HITSAI researchers and teams working on generative image models who care about fast sampling with likelihood-based training could see a new option that avoids distillation trade-offs. Product teams evaluating text-to-image generation may find four-step sampling relevant for reducing generation cost, if the approach holds up in practice.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Diffusion-based models break generation into many small Gaussian denoising steps, an assumption that weakens when generation is compressed to a few coarse transitions. Existing few-step approaches, such as distillation, consistency training, and adversarial objectives, address this compression but give up the likelihood framework in the process.

NTM keeps that likelihood framework by modeling each reverse step as an expressive conditional normalizing flow with exact likelihood training. Architecturally, it pairs shallow invertible blocks within each step with a deep parallel predictor across the trajectory, so the whole network can be trained end-to-end from scratch or initialized from pretrained flow-matching models.

The exact trajectory likelihood also opens the door to self-distillation, where a lightweight denoiser is trained on the score function induced by the model itself. According to the authors, this produces high-quality samples in four steps, and on text-to-image benchmarks NTM matches or outperforms strong image generation baselines at that step count.

FAQ
What makes NTM different from existing few-step image generation methods?
Existing few-step methods use distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework. NTM instead models each reverse step as an expressive conditional normalizing flow with exact likelihood training.
How well does NTM perform on text-to-image benchmarks?
On text-to-image benchmarks, NTM matches or outperforms strong image generation baselines in just four sampling steps, while uniquely retaining exact likelihood over the generative trajectory.
Can NTM be trained from scratch or added to existing models?
NTM combines shallow invertible blocks within each step with a deep parallel predictor across the trajectory, forming an end-to-end network trainable from scratch or initializable from pretrained flow-matching models.
Apple Machine LearningRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI bans two influence ops, one hits Category 5