| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Hello! This is the GitHub space for the FoundationVision @ ByteDance.
We are dedicated to exploring the frontiers of multimodal intelligence, with the ultimate goal of building Artificial General Intelligence systems (AGI).
Our research focuses on deep learning and multimodal intelligence. We are particularly interested in:
Our group strives to push the boundaries of multimodal intelligence and has produced highly influential works in the field, including:
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
[NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.
This organization has no public members. You must be a member to see who’s a part of this organization.
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |