AV LLMs
A collection of Audio, Video and Visual LLMs.
-
Text-to-Speech • Updated • 481 -
1.1k
OpenVoice
🤗Generate voice from text using a reference audio
-
dataautogpt3/ProteusV0.3
Text-to-Image • Updated • 307k • 94 -
ByteDance/SDXL-Lightning
Text-to-Image • Updated • 130k • • 2.1k -
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.15M • • 5.05k -
stabilityai/TripoSR
Image-to-3D • Updated • 15.2k • 565 -
Efficient-Large-Model/VILA-7b
Text Generation • 7B • Updated • 2.25k • 27 -
google/paligemma-3b-pt-896
Image-Text-to-Text • 3B • Updated • 79 • 121 -
microsoft/Phi-3-vision-128k-instruct
Text Generation • 4B • Updated • 34.1k • 968 -
stabilityai/stable-audio-open-1.0
Text-to-Audio • Updated • 33.9k • 1.33k -
OpenVLA: An Open-Source Vision-Language-Action Model
Paper • 2406.09246 • Published • 41 -
aiola/whisper-medusa-v1
2B • Updated • 28 • 178 -
merve/idefics3llama-vqav2
Updated • 8 -
black-forest-labs/FLUX.1-schnell
Text-to-Image • Updated • 961k • • 4.37k -
115
Llama3.1 S V0.2 Checkpoint 2024 08 20
😻Convert text to audio and vice versa
-
gpt-omni/mini-omni
Text-to-Speech • Updated • 429 -
fishaudio/fish-speech-1.4
Text-to-Speech • Updated • 182 • 451 -
179
Tonic's GOT OCR
📲GOT - OCR (from : UCAS, Beijing)
-
stepfun-ai/GOT-OCR2_0
Image-Text-to-Text • 0.7B • Updated • 57.1k • 1.52k -
apple/coreml-sam2-large
Mask Generation • Updated • 44 • 28 -
coreml-projects/sam-2-studio
Updated • 26 -
mistralai/Pixtral-12B-2409
Updated • 3.86k • 669 -
allenai/Molmo-72B-0924
Image-Text-to-Text • 73B • Updated • 2.54k • 294 -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 4.09M • • 2.67k -
Revai/reverb-asr
Automatic Speech Recognition • Updated • 13 • 90 -
359
GOT Online
💬Extract text from images using various OCR modes
-
facebook/vfusion3d
Image-to-3D • 0.5B • Updated • 17 • 65 -
facebook/cotracker
Updated • 647 • 36 -
rhymes-ai/Aria
Image-Text-to-Text • 25B • Updated • 39k • 636 -
SWivid/F5-TTS
Text-to-Speech • Updated • 619k • 1.12k -
64
Ichigo Llama3.1 S Instruct
🏢Generate text from audio recordings
-
kyutai/moshiko-mlx-q4
Updated • 136 • 28 -
kyutai/moshiko-mlx-q8
Updated • 746 • 5 -
128
Open VLM Video Leaderboard
🌎VLMEvalKit Eval Results in video understanding benchmark
-
jimmycarter/LibreFLUX
Text-to-Image • Updated • 102 • 171 -
microsoft/OmniParser
Image-Text-to-Text • Updated • 345 • 1.69k -
325
Aya Models
🌍Interact with the Aya family of models.
-
CohereLabs/aya-expanse-32b
Text Generation • 32B • Updated • 7.88k • • 278 -
stabilityai/stable-diffusion-3.5-medium
Text-to-Image • Updated • 202k • • 847 -
OuteAI/OuteTTS-0.1-350M
Text-to-Speech • 0.4B • Updated • 42 • 302 -
vidore/colpali
Visual Document Retrieval • Updated • 4.66k • 463 -
vidore/colpali-v1.2
Visual Document Retrieval • Updated • 34.1k • 112 -
si-pbc/hertz-dev
Audio-to-Audio • Updated • 214 -
38
Talk To Ultravox
⚡Talk to Fixie.ai's Ultravox with WebRTC ⚡️
-
LLaVA-o1: Let Vision Language Models Reason Step-by-Step
Paper • 2411.10440 • Published • 129 -
Xkev/Llama-3.2V-11B-cot
Image-Text-to-Text • 11B • Updated • 1.63k • 158 -
google/paligemma-3b-pt-224
Image-Text-to-Text • 3B • Updated • 41.2k • 366 -
apple/coreml-mobileclip
Updated • 760 • 47 -
InstantX/InstantIR
Image-to-Image • Updated • 2 • 180 -
86
InstantIR
🖼diffusion-based Image Restoration model
-
169
Flux IP Adapter
🖼Prompt with Images in flux[dev]
-
38
Image Preferences - Argilla annotation space
🖼A community project to create an image preferences dataset.
-
fishaudio/fish-speech-1.5
Text-to-Speech • Updated • 1.7k • 638 -
meta-llama/Llama-3.3-70B-Instruct
Text Generation • 71B • Updated • 755k • • 2.55k -
48
Paligemma2 Vqav2
🐨PaliGemma2 LoRA finetuned on VQAv2
-
VisionZip: Longer is Better but Not Necessary in Vision Language Models
Paper • 2412.04467 • Published • 118 -
fancyfeast/llama-joycaption-alpha-two-hf-llava
8B • Updated • 10.1k • 197 -
taohu/mask
Updated • 5 -
[MASK] is All You Need
Paper • 2412.06787 • Published • 2 -
924
Open VLM Leaderboard
🌎VLMEvalKit Evaluation Results Collection
-
microsoft/LLM2CLIP-Llama3.2-1B-EVA02-L-14-336
Zero-Shot Image Classification • Updated • 10 -
LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation
Paper • 2411.04997 • Published • 39 -
Generative Powers of Ten
Paper • 2312.02149 • Published • 8 -
24
StoryStar
💬Fantasy story generator
-
GoodiesHere/Apollo-LMMs-Apollo-7B-t32
Video-Text-to-Text • Updated • 2 • 57 -
Apollo: An Exploration of Video Understanding in Large Multimodal Models
Paper • 2412.10360 • Published • 147 -
Qwen/Qwen2-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 1.9M • • 1.24k -
XiaoduoAILab/Xmodel_VLM
Text Generation • 2B • Updated • 254 • 13 -
nvidia/Cosmos-1.0-Diffusion-14B-Text2World
Updated • 94 • 60 -
nvidia/Cosmos-1.0-Autoregressive-12B
Updated • 26 • 30 -
nvidia/Cosmos-1.0-Autoregressive-13B-Video2World
Updated • 27 • 32 -
nvidia/Cosmos-1.0-Diffusion-7B-Text2World
Text-to-Video • Updated • 945 • 228 -
nvidia/Cosmos-1.0-Diffusion-14B-Video2World
Updated • 49 • 56 -
457
Stable Point-Aware 3D
⚡Generate 3D models from images
-
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 4.33M • • 5.22k -
2.99k
Kokoro TTS
❤Upgraded to v1.0!
-
openbmb/MiniCPM-o-2_6
Any-to-Any • 9B • Updated • 103k • 1.26k -
445
TTS Spaces Arena
🤗Blind vote on HF TTS models!
-
google/paligemma2-10b-pt-896
Image-Text-to-Text • 10B • Updated • 98 • 32 -
NovaSky-AI/Sky-T1-32B-Preview
Text Generation • 33B • Updated • 72 • • 551 -
MiniMaxAI/MiniMax-VL-01
Image-Text-to-Text • 456B • Updated • 81.6k • 279 -
66
SmolVLM
📊Generate descriptions from images and text prompts
-
HKUSTAudio/Llasa-3B
Text-to-Speech • 4B • Updated • 769 • 523 -
HuggingFaceTB/SmolVLM-500M-Instruct
Image-Text-to-Text • 0.5B • Updated • 80.8k • 182 -
deepseek-ai/Janus-Pro-7B
Any-to-Any • Updated • 85.6k • 3.52k -
309
Kokoro TTS Zero
🎴✨[With v1.0.0] Accelerated TTS on Kokoro-82M
-
kyutai/hibiki-2b-mlx-bf16
Translation • Updated • 26 • 22 -
kyutai/hibiki-2b-pytorch-bf16
Translation • Updated • 78 • 55 -
ARTPARK-IISc/Vaani
Viewer • Updated • 20.5M • 9.23k • 76 -
Zyphra/Zonos-v0.1-hybrid
Text-to-Speech • Updated • 26.7k • 1.1k -
Zyphra/Zonos-v0.1-transformer
Text-to-Speech • Updated • 28.3k • 417 -
microsoft/OmniParser-v2.0
Updated • 35k • 1.3k -
95
Paligemma2 Mix
🌖Generate text and segment images using PaliGemma 2
-
google/paligemma2-3b-mix-448
Image-Text-to-Text • 3B • Updated • 4.98k • 50 -
google/paligemma2-3b-mix-224
Image-Text-to-Text • 3B • Updated • 10.3k • 38 -
google/paligemma2-28b-mix-224
Image-Text-to-Text • 28B • Updated • 360 • 4 -
google/paligemma2-28b-mix-448
Image-Text-to-Text • 28B • Updated • 167 • 27 -
google/paligemma2-10b-mix-224
Image-Text-to-Text • 10B • Updated • 1.04k • 9 -
google/paligemma2-10b-mix-448
Image-Text-to-Text • 10B • Updated • 1.75k • 34 -
stepfun-ai/stepvideo-t2v
Text-to-Video • Updated • 47 • 471 -
stepfun-ai/stepvideo-t2v-turbo
Updated • 97 -
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Paper • 2502.10248 • Published • 55 -
HuggingFaceTB/SmolVLM2-2.2B-Instruct
Image-Text-to-Text • 2B • Updated • 147k • 278 -
nvidia/canary-1b
Automatic Speech Recognition • Updated • 1.78k • 448 -
Wan-AI/Wan2.1-I2V-14B-720P
Image-to-Video • Updated • 16.4k • • 545 -
fastrtc/kokoro-onnx
Updated • 11 -
2
Fastphone
🐠Download and run a Hugging Face app
-
microsoft/Phi-4-multimodal-instruct
Automatic Speech Recognition • 6B • Updated • 402k • 1.52k -
microsoft/Magma-8B
Image-Text-to-Text • 9B • Updated • 3.35k • 411 -
46
Magma UI
📚Magma-8B model for UI Agents
-
645
Di♪♪Rhythm
🎶Blazingly Fast and Embarrassingly Simple Song Generation
-
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Paper • 2503.01183 • Published • 29 -
ASLP-lab/DiffRhythm-vae
Updated • 41 -
ASLP-lab/DiffRhythm-base
Updated • 43 • 168 -
Large Language Diffusion Models
Paper • 2502.09992 • Published • 122 -
GSAI-ML/LLaDA-8B-Instruct
Text Generation • 8B • Updated • 260k • 325 -
unsloth/gemma-3-12b-pt
Image-Text-to-Text • 12B • Updated • 839 • 5 -
google/gemma-3-27b-it
Image-Text-to-Text • 27B • Updated • 890k • • 1.67k -
sesame/csm-1b
Text-to-Speech • Updated • 20.5k • 2.25k -
unsloth/gemma-3-27b-it-GGUF
Image-Text-to-Text • 27B • Updated • 48.4k • 163 -
docling-project/SmolDocling-256M-preview
Image-Text-to-Text • 0.3B • Updated • 354k • 1.59k -
starvector/starvector-8b-im2svg
Text Generation • 8B • Updated • 1.98k • 512 -
starvector/starvector-1b-im2svg
Text Generation • 1B • Updated • 2.25k • 175 -
Tokenize Image as a Set
Paper • 2503.16425 • Published • 16 -
kyutai/moshika-vis-pytorch-bf16
Updated • 56 -
kyutai/Babillage
Viewer • Updated • 465k • 121 • 12 -
ByteDance/InfiniteYou
Text-to-Image • Updated • 5.64k • 635 -
InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
Paper • 2503.16418 • Published • 36 -
openfree/flux-chatgpt-ghibli-lora
Text-to-Image • Updated • 931 • • 320 -
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
Paper • 2504.00595 • Published • 36 -
weizhiwang/Open-Qwen2VL
Image-Text-to-Text • Updated • 14 • 20 -
ostris/Flex.1-alpha-Redux
Text-to-Image • Updated • 697 • 114 -
unsloth/Llama-4-Scout-17B-16E-Instruct-unsloth-bnb-4bit
Image-Text-to-Text • 57B • Updated • 2.78k • 80 -
unsloth/Llama-4-Scout-17B-16E-Instruct-unsloth-bnb-8bit
Image-Text-to-Text • 109B • Updated • 112 • 9 -
SmolVLM: Redefining small and efficient multimodal models
Paper • 2504.05299 • Published • 200 -
canopylabs/3b-hi-ft-research_release
Text-to-Speech • 3B • Updated • 2.2k • 26 -
canopylabs/3b-es_it-ft-research_release
Text-to-Speech • 3B • Updated • 930 • 15 -
nvidia/C-RADIOv2-g
Image Feature Extraction • 1B • Updated • 66 • 12 -
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Paper • 2504.10479 • Published • 300 -
OpenGVLab/InternVL3-1B
Image-Text-to-Text • 0.9B • Updated • 85.9k • 74 -
OpenGVLab/InternVL3-78B
Image-Text-to-Text • 78B • Updated • 4.45k • 222 -
InteractVLM: 3D Interaction Reasoning from 2D Foundational Models
Paper • 2504.05303 • Published • 5 -
1.7k
Dia 1.6B
👯Generate realistic dialogue from a script, using Dia!
-
nari-labs/Dia-1.6B
Text-to-Speech • Updated • 182k • • 2.8k -
Describe Anything: Detailed Localized Image and Video Captioning
Paper • 2504.16072 • Published • 63 -
nvidia/DAM-3B-Self-Contained
Image-Text-to-Text • Updated • 708 • 24 -
nvidia/DAM-3B-Video
Image-Text-to-Text • Updated • 375 • 56 -
nvidia/DAM-3B
Image-Text-to-Text • Updated • 4.75k • 127 -
Qwen/Qwen2.5-Omni-3B
Any-to-Any • 6B • Updated • 256k • 301 -
MMaDA: Multimodal Large Diffusion Language Models
Paper • 2505.15809 • Published • 96 -
One RL to See Them All: Visual Triple Unified Reinforcement Learning
Paper • 2505.18129 • Published • 59 -
117
PlayDiffusion
🎨Generate modified audio from text and voice
-
lerobot/smolvla_base
Robotics • Updated • 16.9k • 290 -
stockmark/Stockmark-2-VL-100B-beta
Image-Text-to-Text • 96B • Updated • 982 • 22 -
Qwen/Qwen2.5-Omni-7B
Any-to-Any • 11B • Updated • 221k • 1.81k -
Qwen2.5-Omni Technical Report
Paper • 2503.20215 • Published • 166 -
1.58k
Chatterbox TTS
🍿Expressive Zeroshot TTS
-
ResembleAI/chatterbox
Text-to-Speech • Updated • 857k • • 1.25k -
PrunaAI/FLUX.1-schnell-smashed
Text-to-Image • Updated • 75 • 5 -
ByteDance/Dolphin
Image-Text-to-Text • 0.4B • Updated • 6.46k • 504 -
nanonets/Nanonets-OCR-s
Image-Text-to-Text • 4B • Updated • 173k • 1.55k -
35
Nanonets Ocr S
👁https://nanonets.com/research/nanonets-ocr-s/
-
calcuis/cosmos-predict2-gguf
Text-to-Image • 14B • Updated • 10.5k • 31 -
Arrexel/pattern-diffusion
Text-to-Image • Updated • 203 • 105 -
numind/NuMarkdown-8B-Thinking
Image-to-Text • 8B • Updated • 2.08k • 215 -
Qwen/Qwen-Image
Text-to-Image • Updated • 196k • • 2.16k -
rednote-hilab/dots.ocr
Image-Text-to-Text • 3B • Updated • 1.13M • 1.11k -
Runware/Qwen-Image-Edit
Image-to-Image • Updated • 102 • 15 -
688
Qwen Image Edit
✒Edit images based on user instructions
-
Qwen/Qwen-Image-Edit
Image-to-Image • Updated • 199k • • 2.08k -
zju-community/matchanything_eloftr
16.1M • Updated • 2.75k • 75 -
234
MatchAnything
🏢Find matching images based on input criteria
-
microsoft/VibeVoice-1.5B
Text-to-Speech • 3B • Updated • 210k • 1.94k -
bytedance-research/USO
Text-to-Image • Updated • 228 • 175 -
398
FastVLM WebGPU
🍎Real-time video captioning powered by FastVLM
-
onnx-community/FastVLM-0.5B-ONNX
Image-Text-to-Text • Updated • 1.31k • 91 -
apple/FastVLM-0.5B
Text Generation • 0.8B • Updated • 12.6k • 342 -
Qwen/Qwen3-Omni-30B-A3B-Instruct
Any-to-Any • 35B • Updated • 307k • 695 -
smolagents/SmolVLM2-2.2B-Instruct-Agentic-GUI
Image-Text-to-Text • 2B • Updated • 1.38k • 52 -
facebook/sam2.1-hiera-large
Mask Generation • 0.2B • Updated • 54k • 108 -
PaddlePaddle/PaddleOCR-VL
Image-Text-to-Text • 1.0B • Updated • 28k • 1.2k -
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
Paper • 2510.14528 • Published • 80 -
164
PaddleOCR-VL Online Demo
📈Recognize text and elements in images
-
nanonets/Nanonets-OCR2-3B
Image-Text-to-Text • 4B • Updated • 67.5k • 431 -
nanonets/Nanonets-OCR2-1.5B-exp
Image-Text-to-Text • 2B • Updated • 21.4k • 44 -
deepseek-ai/DeepSeek-OCR
Image-Text-to-Text • 3B • Updated • 2.06M • 2.4k -
lightonai/LightOnOCR-1B-1025
Image-to-Text • Updated • 10.7k • 131