A hierarchical architecture for text-to-video generation that separates coarse scene/motion planning from fine-grained per-frame rendering, aiming to spend heavy compute on visual detail only where needed while keeping full-clip temporal reasoning cheap.
Topics: text-to-video generation, efficient transformers, hierarchical reasoning, spatiotemporal modeling.
See the paper breakdown for a walkthrough of the coarse/fine split and its failure modes.