<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Nankai on SoloSoft</title><link>https://www.solosoft.dev/tags/nankai/</link><description>Recent content in Nankai on SoloSoft</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.solosoft.dev/tags/nankai/index.xml" rel="self" type="application/rss+xml"/><item><title>StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation</title><link>https://www.solosoft.dev/post/storydiffusion-image-video-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/storydiffusion-image-video-2026/</guid><description>&lt;p&gt;&lt;strong&gt;StoryDiffusion&lt;/strong&gt; is a research project from Nankai University and ByteDance that tackles one of the hardest problems in generative AI: maintaining visual consistency across long sequences of images and videos. Accepted as a major research contribution, it introduces a novel &lt;strong&gt;consistent self-attention (CSA)&lt;/strong&gt; mechanism that enables diffusion models to generate coherent comic strips, animations, and videos &amp;ndash; all without finetuning or per-sequence training.&lt;/p&gt;
&lt;p&gt;The core challenge StoryDiffusion addresses is simple to state but extremely difficult to solve: how do you generate a sequence of images where the same character looks consistently the same in every frame? Previous diffusion models could produce stunning single images, but when asked to generate a multi-panel comic or a video clip, characters would subtly change appearance between frames &amp;ndash; a different nose shape, a changed outfit, a shifted background style.&lt;/p&gt;</description></item></channel></rss>