FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

multi-modal · GitHub Topics · GitHub

#

multi-modal

Here are 547 public repositories matching this topic...

Build and run agents you can see, understand and trust.

  • Updated Aug 19, 2026
  • Python

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

  • Updated Aug 12, 2026
  • Python

Open-source framework for conversational voice AI agents

  • Updated Aug 18, 2026
  • Python

[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型

  • Updated Sep 22, 2025
  • Python

ModelScope: bring the notion of Model-as-a-Service to life.

  • Updated Aug 19, 2026
  • Python

AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas, AGI functions, world-class Beam multi-model chats, text-to-image, voice, response streaming, code highlighting and execution, PDF import, presets for developers, much more. Deploy on-prem or in the cloud.

  • Updated Aug 18, 2026
  • TypeScript

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

  • Updated Aug 19, 2026
  • Python

a state-of-the-art-level open visual language model | 多模态预训练模型

  • Updated May 29, 2024
  • Python

Open Source Routing Engine for OpenStreetMap

  • Updated Aug 19, 2026
  • C++

Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.

  • Updated Mar 31, 2026
  • Jupyter Notebook

Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch

  • Updated Feb 17, 2024
  • Python

Ecommerce Search and Discovery - marqo.ai

  • Updated Aug 8, 2026
  • Python

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

  • Updated Aug 18, 2026
  • Python

OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340

  • Updated Dec 4, 2025
  • Jupyter Notebook

Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型

  • Updated Aug 23, 2024
  • Python

A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.

  • Updated Aug 16, 2026
  • C#

【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

  • Updated Dec 3, 2024
  • Python

Open-source AI assistant ecosystem with MCP integrations, multimodal workflows, IoT support, and cross-platform voice interaction.

  • Updated Aug 17, 2026
  • Python

Improve this page

Add a description, image, and links to the multi-modal topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the multi-modal topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL