FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

vision-language · GitHub Topics · GitHub

#

vision-language

Here are 289 public repositories matching this topic...

[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"

  • Updated Aug 12, 2024
  • Python

Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.

  • Updated Mar 31, 2026
  • Jupyter Notebook

PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

  • Updated Mar 3, 2026
  • Jupyter Notebook

Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

  • Updated Apr 24, 2024
  • Python

An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

  • Updated Jul 25, 2026
  • Python

A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.

  • Updated Mar 17, 2026
  • C++

[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

  • Updated Aug 5, 2025
  • Python

[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering

  • Updated Jul 2, 2025
  • HTML

A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

  • Updated Oct 6, 2024
  • Python

A Framework of Small-scale Large Multimodal Models

  • Updated Jul 23, 2026
  • Python

CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

  • Updated Sep 8, 2025
  • Python

Pix2Seq codebase: multi-tasks with generative modeling (autoregressive and diffusion)

  • Updated Nov 7, 2023
  • Jupyter Notebook

[CVPR 2024] Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

  • Updated Jul 20, 2025
  • Jupyter Notebook

🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)

  • Updated Aug 5, 2025
  • Python

This is the third party implementation of the paper Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

  • Updated Jul 27, 2025
  • Python

[ICLR 2024] Controlling Vision-Language Models for Universal Image Restoration. 5th place in the NTIRE 2024 Restore Any Image Model in the Wild Challenge.

  • Updated Aug 7, 2024
  • Python

Official implementation of SEED-LLaMA (ICLR 2024).

  • Updated Sep 21, 2024
  • Python

🛰️ Official repository of paper "RemoteCLIP: A Vision Language Foundation Model for Remote Sensing" (IEEE TGRS)

  • Updated Jun 27, 2024
  • Jupyter Notebook

CLIPort: What and Where Pathways for Robotic Manipulation

  • Updated Nov 2, 2023
  • Jupyter Notebook

Improve this page

Add a description, image, and links to the vision-language topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the vision-language topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL