a series of text-conditioned Diffusion Transformers (DiT) capable of transforming textual descriptions into vivid images, dynamic videos, detailed multi-view 3D images, and synthesized speech.
Code - https://github.com/Alpha-VLLM/Lumina-T2X
Reactions are currently unavailable
a series of text-conditioned Diffusion Transformers (DiT) capable of transforming textual descriptions into vivid images, dynamic videos, detailed multi-view 3D images, and synthesized speech.
Code - https://github.com/Alpha-VLLM/Lumina-T2X