ZeroGUI: Automating Online GUI Learning at Zero Human Cost
AI & ML interests
Computer Vision
Recent Activity
Organization Card
OpenGVLab
Welcome to OpenGVLab! We are a research group from Shanghai AI Lab focused on Vision-Centric AI research. The GV in our name, OpenGVLab, means general vision, a general understanding of vision, so little effort is needed to adapt to new vision-based tasks.
Models
- InternVL: a pioneering open-source alternative to GPT-4V.
- InternImage: a large-scale vision foundation models with deformable convolutions.
- InternVideo: large-scale video foundation models for multimodal understanding.
- VideoChat: an end-to-end chat assistant for video comprehension.
- All-Seeing-Project: towards panoptic visual recognition and understanding of the open world.
Datasets
- ShareGPT4o: a groundbreaking large-scale resource that we plan to open-source with 200K meticulously annotated images, 10K videos with highly descriptive captions, and 10K audio files with detailed descriptions.
- InternVid: a large-scale video-text dataset for multimodal understanding and generation.
- MMPR: a high-quality, large-scale multimodal preference dataset.
Benchmarks
- MVBench: a comprehensive benchmark for multimodal video understanding.
- CRPE: a benchmark covering all elements of the relation triplets (subject, predicate, object), providing a systematic platform for the evaluation of relation comprehension ability.
- MM-NIAH: a comprehensive benchmark for long multimodal documents comprehension.
- GMAI-MMBench: a comprehensive multimodal evaluation benchmark towards general medical AI.
-
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Paper • 2504.10479 • Published • 277 -
OpenGVLab/InternVL3-1B
Image-Text-to-Text • 0.9B • Updated • 74k • 67 -
OpenGVLab/InternVL3-2B
Image-Text-to-Text • 2B • Updated • 70k • 27 -
OpenGVLab/InternVL3-8B
Image-Text-to-Text • 8B • Updated • 346k • 85
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
-
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Paper • 2504.10479 • Published • 277 -
OpenGVLab/InternVL3-1B
Image-Text-to-Text • 0.9B • Updated • 74k • 67 -
OpenGVLab/InternVL3-2B
Image-Text-to-Text • 2B • Updated • 70k • 27 -
OpenGVLab/InternVL3-8B
Image-Text-to-Text • 8B • Updated • 346k • 85
spaces
11
Runtime error
InternVideo2.5
💬
Hierarchical Compression for Long-Context Video Modeling
Running
484
InternVL
⚡
Chat with an AI that understands text and images
Running
38
MVBench Leaderboard
🐨
Submit model evaluation and view leaderboard
Runtime error
18
InternVideo2 Chat 8B HD
👁
Upload a video to chat about its contents
Running
10
ControlLLM
🚀
Display maintenance message for ControlLLM
Runtime error
98
VideoMamba
🐍
Identify actions and objects in videos and images
models
223

OpenGVLab/Docopilot-8B
8B
•
Updated
•
1

OpenGVLab/Docopilot-2B
2B
•
Updated
•
1

OpenGVLab/ZeroGUI-OSWorld-7B
Image-Text-to-Text
•
8B
•
Updated
•
32
•
4

OpenGVLab/InternVideo1.0
Video Classification
•
Updated

OpenGVLab/ZeroGUI-AndroidLab-7B
Image-Text-to-Text
•
8B
•
Updated
•
14
•
4

OpenGVLab/InternVL3-78B-AWQ
Image-Text-to-Text
•
Updated
•
1.2k
•
8

OpenGVLab/InternVL3-38B-AWQ
Image-Text-to-Text
•
Updated
•
20.4k
•
2

OpenGVLab/InternVL3-14B-AWQ
Image-Text-to-Text
•
Updated
•
4.24k
•
4

OpenGVLab/InternVL3-8B-AWQ
Image-Text-to-Text
•
Updated
•
11.9k
•
7

OpenGVLab/InternVL3-78B-Instruct
Image-Text-to-Text
•
78B
•
Updated
•
11.9k
•
6
datasets
43
OpenGVLab/MMBench-GUI
Preview
•
Updated
•
235
•
32
OpenGVLab/VideoChat-Flash-Training-Data
Viewer
•
Updated
•
87k
•
127k
•
10
OpenGVLab/VRBench
Updated
•
74
•
2
OpenGVLab/VisualPRM400K-v1.1
Preview
•
Updated
•
19.2k
•
6
OpenGVLab/MMPR-v1.2-prompts
Updated
•
9.26k
•
1
OpenGVLab/MMPR-v1.2
Updated
•
17.7k
•
23
OpenGVLab/InternVL-Data
Updated
•
117
•
8
OpenGVLab/VisualPRM400K-v1.1-Raw
Preview
•
Updated
•
22
•
2
OpenGVLab/VisualPRM400K
Preview
•
Updated
•
179
•
13
OpenGVLab/MMPR-v1.1
Preview
•
Updated
•
83
•
44