| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Official PyTorch implementation of the paper "VTimeLLM: Empower LLM to Grasp Video Moments".
VTimeLLM is a novel Video LLM designed for fine-grained video moment understanding and reasoning with respect to time boundary.
VTimeLLM adopts a boundary-aware three-stage training strategy, which respectively utilizes image-text pairs for feature alignment, multiple-event videos to increase temporal-boundary awareness, and high-quality video-instruction tuning to further improve temporal understanding ability as well as align with human intents.
We recommend setting up a conda environment for the project:
conda create --name=vtimellm python=3.10
conda activate vtimellm
git clone https://github.com/huangb23/VTimeLLM.git
cd VTimeLLM
pip install -r requirements.txtAdditionally, install additional packages for training cases.
pip install ninja
pip install flash-attn --no-build-isolationTo run the demo offline, please refer to the instructions in offline_demo.md.
For training instructions, check out train.md.
A Comprehensive Evaluation of VTimeLLM's Performance across Multiple Tasks.
We are grateful for the following awesome projects our VTimeLLM arising from:
If you're using VTimeLLM in your research or applications, please cite using this BibTeX:
@inproceedings{huang2024vtimellm,
title={Vtimellm: Empower llm to grasp video moments},
author={Huang, Bin and Wang, Xin and Chen, Hong and Song, Zihan and Zhu, Wenwu},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={14271--14280},
year={2024}
}This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License.
Looking forward to your feedback, contributions, and stars! 🌟
| Back | FazBrowse Home | New Git URL |