Multimodal
20 entities · 0 projects · 1 tools · 18 models · 1 news
DeepSeek V4.1 Flash
An MIT-licensed DeepSeek V4.1 checkpoint for fast multimodal text generation.
Voxtral Mini Realtime
Mistral AI's compact realtime speech model for low-latency transcription workflows.
Qwen Image 2.1
Qwen's Diffusers-compatible model for text-to-image generation and image editing.
Llama 4 Scout
Meta's open-weight Llama 4 Scout mixture-of-experts instruction model.
Llama 4 Maverick
Meta's larger Llama 4 Maverick mixture-of-experts instruction checkpoint.
Concat
Free and open-source CapCut replacement that supports MCP servers.
Holo4 27B
H Company's 27B agent model combines graphical interfaces, code, MCP and APIs; its downloadable checkpoint carries a noncommercial license.
ElevenLabs employee tender
ElevenLabs offered employees a chance to sell vested shares in a $300 million tender priced at a $22 billion valuation.
Claude Sonnet 5.5
Anthropic's September 28 Sonnet release targets faster coding and knowledge work at the same base token prices as Sonnet 5.
GPT-6.1 Sol
OpenAI's September 29 model for coding, computer use and professional work, with standard API pricing of $2 input and $10 output per million tokens.
Amazon Nova Pro
A higher-capability Amazon Nova model for complex multimodal enterprise workloads.
Amazon Nova 2 Sonic
Amazon's next-generation speech model for conversational and realtime voice experiences.
Command A Vision
Cohere's multimodal Command A model for enterprise image and document understanding.
Gemini 3.1 Pro
A Gemini Pro model for complex multimodal reasoning and production applications.
Gemini 3.5 Flash
A fast Gemini model for multimodal, high-throughput and interactive applications.
Gemini 3.5 Flash-Lite
A lightweight Gemini tier designed for cost-sensitive, high-volume inference.
GPT Realtime 2.1
OpenAI's realtime model family for low-latency voice and interactive multimodal sessions.
Gemini 3.8 Flash
Google DeepMind's Gemini 3.8 Flash and 3.8 Flash Cyber, announced September 2026.
Gemini 3.8 Live
Google DeepMind's Gemini 3.8 Live and 3.8 Live Extended Thinking.
Gemma 3
Google's open model family designed for capable, efficient multimodal applications.