Qortora · Search · Indexed page

huggingface.coFetched 2026-08-13T19:37:05Z

Paper page - A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

Join the discussion on this paper page

Open original source · Full cached text

Paper page - A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples Code: <a href=\"https://github.com/zfu006/SSG\" rel=\"nofollow\">https://github.com/zfu006/SSG</a></p>\n","updatedAt":"2026-08-04T12:29:37.053Z","author":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","fullname":"Niels Rogge","name":"nielsr","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":1273,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6485579609870911},"editors":["nielsr"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg"],"reactions":[],"isReport":false}},{"id":"6a72b81c277c86a6e7b4f1dd","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-05T04:12:12.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion](https://huggingface.co/papers/2606.27760) (2026)\n* [PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation](https://huggingface.co/papers/2607.02515) (2026)\n* [Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation](https://huggingface.co/papers/2606.27978) (2026)\n* [Pixel-Space Diffusion Transformers](https://huggingface.co/papers/2607.17585) (2026)\n* [DiffusionBench: On Holistic Evaluation of Diffusion Transformers](https://huggingface.co/papers/2606.24888) (2026)\n* [DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer](https://huggingface.co/papers/2607.18510) (2026)\n* [Perceptual Flow Matching for Few-Step Generative Modeling](https://huggingface.co/papers/2607.03524) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.27760\">PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.02515\">PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.27978\">Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.17585\">Pixel-Space Diffusion Transformers</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.24888\">DiffusionBench: On Holistic Evaluation of Diffusion Transformers</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.18510\">DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.03524\">Perceptual Flow Matching for Few-Step Generative Modeling</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-05T04:12:12.397Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6520636081695557},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.29122","authors":[{"_id":"6a71d7f05067ac40957f7ce2","user":{"_id":"66309e099c604a44f60d46e8","avatarUrl":"/avatars/8a6df598de40e12cfff5c61aa2b3bf61.svg","isPro":false,"fullname":"Zixuan Fu","user":"ZXFu","type":"user","name":"ZXFu"},"name":"Zixuan Fu","status":"claimed_verified","statusLastChangedAt":"2026-08-05T08:45:04.512Z","hidden":false},{"_id":"6a71d7f05067ac40957f7ce3","name":"Chong Wang","hidden":false},{"_id":"6a71d7f05067ac40957f7ce4","name":"Lanqing Guo","hidden":false},{"_id":"6a71d7f05067ac40957f7ce5","name":"Kailai Zhou","hidden":false},{"_id":"6a71d7f05067ac40957f7ce6","name":"Jiahao Nie","hidden":false},{"_id":"6a71d7f05067ac40957f7ce7","name":"Bihan Wen","hidden":false}],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples","submittedOnDailyBy":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","isPro":false,"fullname":"Niels Rogge","user":"nielsr","type":"user","name":"nielsr"},"summary":"Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both global structure and local texture in the same high-dimensional space. While recent work improves pixel diffusion through alternative prediction targets, training objectives, and architectures, these advances typically require training a new model from scratch. We show there is a cheaper, complementary strategy: a frozen, pretrained pixel diffusion model can guide itself. Our key observation is that intermediate layers of a pretrained pixel diffusion transformer can be decoded into coarse predictions that capture the main low-frequency structure, while the final layers progressively refine local, high-frequency details. We therefore attach a lightweight prediction head to an intermediate layer, keep the backbone frozen, and use the discrepancy between the intermediate and final predictions as a self-guidance direction during sampling. To train this head, we further find that real images are not necessary. Instead, model-generated samples suffice and even outperform real images for training the head, especially in enhancing the high-frequency components that pixel diffusion tends to underfit. Across multiple pixel diffusion models on ImageNet, our Synthetic Self-Guidance (SSG) consistently improves generation while adapter training requires less than 1% of full-model training compute: it reduces FID by over 50% across the evaluated JiT variants without classifier-free guidance (CFG) and further improves strong baselines with CFG, e.g., JiT-H/16 from 1.86 to 1.67 and PixelREPA-H/16 from 1.81 to 1.59. Our code is available at https://github.com/zfu006/SSG.","upvotes":4,"discussionId":"6a71d7f15067ac40957f7ce8","githubRepo":"https://github.com/zfu006/SSG","githubRepoAddedBy":"user","ai_summary":"Freezing a pretrained pixel diffusion model and training a lightweight head on synthetic samples enables self-guidance that improves image generation with minimal compute.","ai_keywords":["pixel diffusion","self-guidance","synthetic self-guidance","intermediate layer decoding","prediction head","frozen backbone","classifier-free guidance","FID"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":10,"organization":{"_id":"6371470aafbe42caa5a76208","name":"nanyang-technological-university-singapore","fullname":"Nanyang Technological University Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637146c5afbe42caa5a75e1b/sZyHSA1AQaAS4nrGan682.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69bce12e696bd8657c81b55e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/heZLBpVqxp7IxvCLHmgLs.png","isPro":false,"fullname":"王浩然","user":"xuruilin36","type":"user"},{"_id":"66309e099c604a44f60d46e8","avatarUrl":"/avatars/8a6df598de40e12cfff5c61aa2b3bf61.svg","isPro":false,"fullname":"Zixuan Fu","user":"ZXFu","type":"user"},{"_id":"69783018f72a660840bf6cf4","avatarUrl":"/avatars/d24d25b55bb18a1d5b55e9a05f12019b.svg","isPro":false,"fullname":"Jiarui Zhang","user":"zhan0618","type":"user"},{"_id":"652b64f0680fa36b3b04c217","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/GCtF6U3pTGmUq4t98qo30.jpeg","isPro":false,"fullname":"mjorah7","user":"mjorah7","type":"user"}],"acceptLanguages":["en","es","pt","fr","de","ja","zh","ko","ar","hi","ru","*"],"dailyPaperRank":0,"organization":{"_id":"6371470aafbe42caa5a76208","name":"nanyang-technological-university-singapore","fullname":"Nanyang Technological University Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637146c5afbe42caa5a75e1b/sZyHSA1AQaAS4nrGan682.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.29122.md","query":{}}"> Papers arxiv:2607.29122 Copy markdown A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples Published on Jul 31 · Submitted by Niels Rogge on Aug 4 · Nanyang Technological University Singapore Upvote 4 Authors: Zixuan Fu , Chong Wang , Lanqing Guo , Kailai Zhou , Jiahao Nie , Bihan Wen Abstract Freezing a pretrained pixel diffusion model and training a lightweight head on synthetic samples enables self-guidance that improves image generation with minimal compute. Generated by thinkingmachines/Inkling-Small Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both global structure and local texture in the same high-dimensional space. While recent work improves pixel diffusion through alternative prediction targets, training objectives, and architectures, these advances typically require training a new model from scratch. We show there is a cheaper, complementary strategy: a frozen, pretrained pixel diffusion model can guide itself. Our key observation is that intermediate layers of a pretrained pixel diffusion transformer can be decoded into coarse predictions that capture the main low-frequency structure, while the final layers progressively refine local, high-frequency details. We therefore attach a lightweight prediction head to an intermediate layer, keep the backbone frozen, and use the discrepancy between the intermediate and final predictions as a self-guidance direction during sampling. To train this head, we further find that real images are not necessary. Instead, model-generated samples suffice and even outperform real images for training the head, especially in enhancing the high-frequency components that pixel diffusion tends to underfit. Across multiple pixel diffusion models on ImageNet, our Synthetic Self-Guidance (SSG) consistently improves generation while adapter training requires less than 1% of full-mo…