| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this userโs behavior. Learn more about reporting abuse.
Report abuse๐ญ Iโm currently working on video understanding and large multimodal models
๐ซ How to reach me: shehan.munasinghe@mbzuai.ac.ae
[CVPR 2025 ๐ฅ]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
PG-Video-LLaVA: Pixel Grounding in Large Multimodal Video Models
๐ค Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
| Back | FazBrowse Home | New Git URL |