24 similar Multimodal you might want to consider.
Fuyu 2 is a Multimodal model developed by Adept AI. This page compares it with 24 other Multimodal models in the same category — 16 open source and 0 offering a free tier.
| Model | Developer | Release date | Context window | Max output tokens | Modalities (input → output) | Open source / licence | Free tier | API available |
|---|---|---|---|---|---|---|---|---|
| Fuyu 2 | Adept AI | Jul 1, 2025 | — | — | — | No | — | — |
| GPT-5 | OpenAI | Jun 1, 2025 | — | — | — | No | — | — |
| Gemini 2.5 Pro | Mar 1, 2025 | — | — | — | No | — | — | |
| Amazon Nova Premier | Amazon | Feb 1, 2025 | — | — | — | No | — | — |
| Cambrian-1 | NYU | Jun 1, 2024 | — | — | — | Yes | — | — |
| Claude 3.5 Sonnet Vision | Anthropic | Oct 1, 2024 | — | — | — | No | — | — |
| CogVLM 2.5 | Zhipu AI | Apr 1, 2025 | — | — | — | Yes | — | — |
| DeepSeek VL2 | DeepSeek | Jun 1, 2025 | — | — | — | Yes | — | — |
| Emu3 | Meta | Aug 1, 2025 | — | — | — | Yes | — | — |
| Florence 2.5 | Microsoft | Apr 1, 2025 | — | — | — | Yes | — | — |
| Gemini 2.5 Flash Vision | Mar 1, 2025 | — | — | — | No | — | — | |
| Grok 3 | xAI | Feb 1, 2025 | — | — | — | No | — | — |
| Idefics 3 | Hugging Face | May 1, 2025 | — | — | — | Yes | — | — |
| InternVL3-78B | Shanghai AI Lab | Apr 1, 2025 | — | — | — | Yes | — | — |
| Janus Pro | DeepSeek | Mar 1, 2025 | — | — | — | Yes | — | — |
| LLaMA 3.3 | Meta | Jan 1, 2025 | — | — | — | Yes | — | — |
| LLaVA-OneVision | UW/Microsoft | Aug 1, 2024 | — | — | — | Yes | — | — |
| Meta Chameleon 2 | Meta | — | — | — | — | Yes | — | — |
| Molmo | AI2 | Mar 1, 2025 | — | — | — | Yes | — | — |
| Moondream 2 | Vikhyat | Jan 1, 2025 | — | — | — | Yes | — | — |
| Pixtral Large | Mistral AI | Mar 1, 2025 | — | — | — | No | — | — |
| Qwen2.5-Omni | Alibaba | May 1, 2025 | — | — | — | Yes | — | — |
| Reka Flash | Reka AI | Apr 1, 2024 | — | — | — | No | — | — |
| Unified-IO 2 | Allen AI | Feb 1, 2024 | — | — | — | Yes | — | — |
| Unified-IO 3 | AI2 | Apr 1, 2025 | — | — | — | Yes | — | — |
OpenAI
OpenAI's next-generation model unifying text, image, and audio reasoning.
Google's thinking model with native multimodal reasoning across text, image, audio, video.
Amazon
Most capable Amazon Nova model for complex reasoning across text, image, and video modalities.
NYU
Open-source vision-centric multimodal model with strong visual grounding and understanding capabilities.
Anthropic
Vision-capable version of Claude 3.5 Sonnet with strong image understanding and analysis abilities.
Zhipu AI
Open-source visual language model with strong visual grounding.
DeepSeek
Cost-effective multimodal model with vision and language capabilities.
Meta
Meta's unified multimodal model generating text, images, and video.
Microsoft
Microsoft's vision foundation model for diverse visual tasks.
Multimodal version of Gemini 2.5 Flash with strong vision understanding at high speed.
xAI
xAI's multimodal model with real-time data and image generation.
Hugging Face
Community-built multimodal model accessible through Hugging Face.
Shanghai AI Lab
Strong open-source vision-language model excelling at visual tasks.
DeepSeek
Unified model for both multimodal understanding and generation.
Meta
Meta's efficient multimodal open model matching larger model performance.
UW/Microsoft
Open-source vision-language model excelling at single-image, multi-image, and video understanding tasks.
Meta
Second-generation multimodal foundation model from Meta that natively processes and generates text, images, and mixed-modal content in a unified architecture.
AI2
Fully open multimodal model with transparent training and pointing capability.
Vikhyat
Tiny but powerful open-source vision-language model that runs on edge devices.
Mistral AI
Mistral's vision-language model with strong visual reasoning.
Alibaba
Alibaba's omni-modal model processing text, image, audio, and video.
Reka AI
Multimodal model processing text, images, video and audio with competitive performance at efficient cost.
Allen AI
Unified autoregressive model handling text, images, audio, and actions in a single architecture.
AI2
Any-to-any multimodal model handling text, images, audio, and video in a unified architecture.