OpenGVLab

community

https://github.com/opengvlab

opengvlab

OpenGVLab

Activity Feed Request to join this org

AI & ML interests

Computer Vision

Recent Activity

duanyuchen updated a model about 5 hours ago

OpenGVLab/Docopilot-8B

duanyuchen updated a model about 8 hours ago

OpenGVLab/Docopilot-2B

duanyuchen published a model about 8 hours ago

OpenGVLab/Docopilot-8B

View all activity

Organization Card

Community About org cards

OpenGVLab

Welcome to OpenGVLab! We are a research group from Shanghai AI Lab focused on Vision-Centric AI research. The GV in our name, OpenGVLab, means general vision, a general understanding of vision, so little effort is needed to adapt to new vision-based tasks.

Models

InternVL: a pioneering open-source alternative to GPT-4V.
InternImage: a large-scale vision foundation models with deformable convolutions.
InternVideo: large-scale video foundation models for multimodal understanding.
VideoChat: an end-to-end chat assistant for video comprehension.
All-Seeing-Project: towards panoptic visual recognition and understanding of the open world.

Datasets

ShareGPT4o: a groundbreaking large-scale resource that we plan to open-source with 200K meticulously annotated images, 10K videos with highly descriptive captions, and 10K audio files with detailed descriptions.
InternVid: a large-scale video-text dataset for multimodal understanding and generation.
MMPR: a high-quality, large-scale multimodal preference dataset.

Benchmarks

MVBench: a comprehensive benchmark for multimodal video understanding.
CRPE: a benchmark covering all elements of the relation triplets (subject, predicate, object), providing a systematic platform for the evaluation of relation comprehension ability.
MM-NIAH: a comprehensive benchmark for long multimodal documents comprehension.
GMAI-MMBench: a comprehensive multimodal evaluation benchmark towards general medical AI.

Collections 25

View 25 collections

spaces 11

InternVideo2.5

Hierarchical Compression for Long-Context Video Modeling

InternVL

Chat with an AI that understands text and images

MVBench Leaderboard

Submit model evaluation and view leaderboard

InternVideo2 Chat 8B HD

Upload a video to chat about its contents

ControlLLM

Display maintenance message for ControlLLM

VideoMamba

Identify actions and objects in videos and images

models 223

OpenGVLab/Docopilot-8B

8B • Updated about 5 hours ago • 1

OpenGVLab/Docopilot-2B

2B • Updated about 8 hours ago • 1

OpenGVLab/ZeroGUI-OSWorld-7B

Image-Text-to-Text • 8B • Updated 29 days ago • 32 • 4

OpenGVLab/InternVideo1.0

Video Classification • Updated Jun 10

OpenGVLab/ZeroGUI-AndroidLab-7B

Image-Text-to-Text • 8B • Updated May 30 • 14 • 4

OpenGVLab/InternVL3-78B-AWQ

Image-Text-to-Text • Updated May 29 • 1.2k • 8

OpenGVLab/InternVL3-38B-AWQ

Image-Text-to-Text • Updated May 29 • 20.4k • 2

OpenGVLab/InternVL3-14B-AWQ

Image-Text-to-Text • Updated May 29 • 4.24k • 4

OpenGVLab/InternVL3-8B-AWQ

Image-Text-to-Text • Updated May 29 • 11.9k • 7

OpenGVLab/InternVL3-78B-Instruct

Image-Text-to-Text • 78B • Updated May 29 • 11.9k • 6

View 223 models

datasets 43

OpenGVLab/MMBench-GUI

Preview • Updated 23 days ago • 235 • 32

OpenGVLab/VideoChat-Flash-Training-Data

Viewer • Updated 25 days ago • 87k • 127k • 10

OpenGVLab/VRBench

Updated Jun 12 • 74 • 2

OpenGVLab/VisualPRM400K-v1.1

Preview • Updated May 29 • 19.2k • 6

OpenGVLab/MMPR-v1.2-prompts

Updated May 29 • 9.26k • 1

OpenGVLab/MMPR-v1.2

Updated May 29 • 17.7k • 23

OpenGVLab/InternVL-Data

Updated May 8 • 117 • 8

OpenGVLab/VisualPRM400K-v1.1-Raw

Preview • Updated Apr 15 • 22 • 2

OpenGVLab/VisualPRM400K

Preview • Updated Apr 15 • 179 • 13

OpenGVLab/MMPR-v1.1

Preview • Updated Apr 13 • 83 • 44

View 43 datasets