thomas-mayne 's Collections Vision Language Models
updated
Mini-Gemini: Mining the Potential of Multi-modality Vision Language
Models
Paper
• 2403.18814
• Published • 49
meta-llama/Llama-3.2-11B-Vision
Image-Text-to-Text
• 11B • Updated • 9.56k
• 593
google/paligemma-3b-pt-224
Image-Text-to-Text
• 3B • Updated • 168k
• 516
Qwen/Qwen2-VL-2B-Instruct
Image-Text-to-Text
• 2B • Updated • 3.31M
• 516
Qwen/Qwen2-VL-7B-Instruct
Image-Text-to-Text
• 8B • Updated • 1.44M
• 1.28k
Image-Text-to-Text
• 0.7B • Updated • 346k
• 1.55k
Image-Text-to-Text
• 25B • Updated • 132k
• 638
Salesforce/blip-image-captioning-large
Image-to-Text
• 0.5B • Updated • 739k
• 1.48k
black-forest-labs/FLUX.1-dev
Text-to-Image
• 12B • Updated • 595k
• • 13.7k
black-forest-labs/FLUX.1-schnell
Text-to-Image
• 12B • Updated • 245k
• • 5.4k
stabilityai/stable-diffusion-3.5-large
Text-to-Image
• 8B • Updated • 49.1k
• • 3.6k
Image-Text-to-Text
• 8B • Updated • 36.6k
• 1.06k
stabilityai/stable-diffusion-3.5-medium
Text-to-Image
• 2B • Updated • 223k
• • 1.02k
Image-Text-to-Text
• Updated • 393
• 1.71k