FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Original HTTPS Page]

Om AI Lab ยท GitHub

Om AI Lab

Open Multimodal AGI Research

Building the foundational brains for the physical world.


๐ŸŒŒ About Us

At Om AI Lab, we believe the future of AI extends far beyond pure text. We are dedicated to building the "brains" for next-generation systems by focusing on the intersection of Spatial Intelligence, Visual Reasoning, and Embodied Agents.

Our research spans across open-vocabulary perception, reinforced vision-language models, and real-time inference. We aim to bridge the critical gap between high-level logical reasoning and fine-grained visual actionโ€”building models that don't just "see" the world, but intuitively understand and interact with it.


๐Ÿšข Flagship VLX Model Series

  • ๐Ÿ“น VLX-Flow: A real-time VLM for streaming video understanding .
  • ๐Ÿ” VLX-Seek: Fine-grained visual perception and grounding for physical AI.
  • ๐Ÿš— VLX-Go: Efficient general-purpose embodied navigation in the wild.

๐Ÿš€ Core Research Tracks

๐Ÿง  Reinforced & Advanced Visual Reasoning

Models that think, reason, and understand the visual world at a granular level.

  • ๐ŸŒŸ VLM-R1: Solving Visual Understanding with Reinforced VLMs. (Highly active)
  • ๐Ÿ”Ž ZoomEye: Enhancing Multimodal LLMs with human-like zooming capabilities through tree-based image exploration.
  • ๐ŸŒ ImageRAG: Enhancing ultrahigh-resolution remote sensing imagery analysis.

๐Ÿ‘๏ธ Real-Time Perception & Open-World Visual Detection

Foundational spatial understanding optimized for edge and on-premise speeds.

  • โšก OmDet: Real-time, highly accurate, open-vocabulary end-to-end object detection.
  • ๐Ÿ” VLM-FO1: Bridging the gap between high-level reasoning and fine-grained perception in Vision-Language Models.
  • ๐Ÿ“ GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training.

๐Ÿค– Multimodal Agents & Embodied AI

Action-oriented intelligence for physical and virtual environments.

  • ๐Ÿ› ๏ธ OmAgent: A comprehensive framework to build multimodal language agents for fast prototyping and production.
  • ๐ŸŽฏ OmTrackVLA: Open and reproducible research for tracking Vision-Language-Action (VLA) models.

๐Ÿ“Š Benchmarks & Evaluation

Rigorous standards for the open-source multimodal community.

  • ๐Ÿ“ OVDEval: A comprehensive evaluation benchmark for Open-Vocabulary Detection.
  • ๐Ÿ“ VL-CheckList: Evaluating Vision & Language Pretraining Models with Objects, Attributes, and Relations.

Pinned Loading

  1. VLM-R1 VLM-R1 Public

    Solve Visual Understanding with Reinforced VLMs

    Python 6k 384

  2. OmDet OmDet Public

    Real-time and accurate open-vocabulary end-to-end object detection

    Python 1.4k 118

  3. OmAgent OmAgent Public

    [EMNLP-2024] Build multimodal language agents for fast prototype and production

    Python 2.7k 292

  4. VLM-FO1 VLM-FO1 Public

    VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs

    Python 330 15

  5. OmTrackVLA OmTrackVLA Public

    Open & Reproducible Research for Tracking VLAs

    Python 273 15

  6. ZoomEye ZoomEye Public

    [EMNLP-2025 Oral] ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

    Python 91 10

Repositories

Loading
Type
Select type
All Public Sources Forks Archived Mirrors Templates
Language
Select language
All HTML Jupyter Notebook Python
Sort
Select order
Last updated Name Stars
Showing 10 of 26 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loadingโ€ฆ

Most used topics

Loadingโ€ฆ


Back | FazBrowse Home | New Git URL