菜单

🤖 系统
📄 Abstract - NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation

Standard diffusion corrupts data using Gaussian noise whose Fourier coefficients have random magnitudes and random phases. While effective for unconditional or text-to-image generation, corrupting phase components destroys spatial structure, making it ill-suited for tasks requiring geometric consistency, such as re-rendering, simulation enhancement, and image-to-image translation. We introduce Phase-Preserving Diffusion {\phi}-PD, a model-agnostic reformulation of the diffusion process that preserves input phase while randomizing magnitude, enabling structure-aligned generation without architectural changes or additional parameters. We further propose Frequency-Selective Structured (FSS) noise, which provides continuous control over structural rigidity via a single frequency-cutoff parameter. {\phi}-PD adds no inference-time cost and is compatible with any diffusion model for images or videos. Across photorealistic and stylized re-rendering, as well as sim-to-real enhancement for driving planners, {\phi}-PD produces controllable, spatially aligned results. When applied to the CARLA simulator, {\phi}-PD improves CARLA-to-Waymo planner performance by 50\%. The method is complementary to existing conditioning approaches and broadly applicable to image-to-image and video-to-video generation. Videos, additional examples, and code are available on our \href{this https URL}{project page}.

顶级标签: computer vision model training multi-modal
详细标签: diffusion models image-to-image translation phase preservation structure alignment frequency-domain generation 或 搜索:

神经重制:用于结构对齐生成的相位保持扩散模型 / NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation


1️⃣ 一句话总结

这篇论文提出了一种新的扩散模型方法,它在生成新图像或视频时能保持原始输入的空间结构(如物体形状和位置),从而在图像重渲染、模拟器增强等需要几何一致性的任务上表现更优,且无需增加额外计算成本。


📄 打开原文 PDF