A hierarchical architecture for text-to-video generation that separates coarse scene/motion planning from fine-grained per-frame rendering, aiming to spend heavy compute on visual detail only where needed while keeping full-clip temporal reasoning cheap.

Topics: text-to-video generation, efficient transformers, hierarchical reasoning, spatiotemporal modeling.

See the paper breakdown for a walkthrough of the coarse/fine split and its failure modes.