Multimodal
Models and tools that work across text, images, audio, and video.
✦ Prompts
Full library →◈ News
Full timeline →Kuaishou's Kling AI Closes $2B Round at an $18B Valuation
Kling AI formally closed a $2 billion round led by General Atlantic on July 2, valuing the Kuaishou-backed video-generation company at $18 billion. The close confirms the process first disclosed in mid-June and ranks among the largest single rounds ever raised by a Chinese AI company ahead of an expected IPO.
Google Ships Gemini 3.5 Live Translate for Real-Time Speech-to-Speech Translation
Google's June AI roundup introduced Gemini 3.5 Live Translate, an audio model for live speech-to-speech translation that auto-detects more than 70 languages while preserving the speaker's natural intonation, targeting multilingual calls, meetings, and travel. The month also brought Nano Banana 2 Lite as Google's fastest image model and Gemini Omni Flash reaching APIs in public preview.
Gemini 3.5 Pro Enters Limited Vertex Preview with 2M-Token Context and Deep Think Reasoning
Google opened Gemini 3.5 Pro to select Vertex AI enterprise customers in limited preview in early June 2026, ahead of a broader GA Sundar Pichai described as 'next month' at Google I/O. The model targets a 2-million-token context window, Deep Think multi-step reasoning, and frontier multimodal understanding — absorbing the use cases previously routed to Gemini Ultra.
Ideogram Releases 4.0: First Open-Weight 9.3B Text-to-Image Model with Native 2K Output and JSON Layout Control
Ideogram released its first open-weight model — a 9.3-billion-parameter text-to-image diffusion transformer using a frozen Qwen3-VL-8B-Instruct text encoder. Unlike prompt-based models, Ideogram 4.0 uses structured JSON bounding-box placement for pixel-precise layout control and generates natively at 2K resolution. Weights and Apache 2.0 inference code are on Hugging Face and GitHub; commercial deployment requires a separate license. It claimed #1 among all open-weight image models on DesignArena on launch day.
Cognition Launches Devin Desktop — Agent-Neutral IDE Built on Windsurf Rebrand with ACP Support
Cognition rebranded Windsurf as Devin Desktop, a full code editor and agent orchestration layer for engineering teams. The update introduces Devin Local (rewritten in Rust with 30% greater token efficiency), native support for the open Agent Client Protocol (ACP), and out-of-the-box integrations with OpenAI Codex, Claude Agent, and OpenCode. The move positions Cognition — which raised $1B at a $25B valuation in May 2026 — to compete directly with GitHub Copilot and Cursor at the IDE layer.
Alibaba Launches Qwen3.7-Plus — Multimodal GUI Agent Unifying Vision and Language
Alibaba released Qwen3.7-Plus, a multimodal agent model that unifies GUI and CLI operation across visual and text tasks in a single set of weights. On ScreenSpot Pro it scores 79.0% and Terminal-Bench 70.3%, leading the open-API GUI agent field. Available via Alibaba Cloud at $0.40/M input tokens with native vision for images and video up to 5 minutes.
⬡ Tools
All tools →OpenAI's AI assistant used by over 200 million people worldwide
Google's AI assistant with real-time Search and up to 2M-token context
xAI's assistant with real-time X/Twitter data and Aurora image generation
Meta's free AI assistant built into WhatsApp, Instagram, and Facebook
AI-first code editor built on VS Code with deep codebase context
Agentic AI coding IDE with Cascade for autonomous task completion
OpenAI's image generator with best-in-class prompt adherence
Leading AI video generation platform trusted by professional filmmakers
OpenAI's text-to-video model generating cinematic clips up to 60 seconds
Kuaishou's high-quality video generation with up to 2-minute clips
AI video creation and editing with scene modification features
AI avatar video creation for marketing, training, and localization
Luma AI's fast, high-quality text-to-video and image-to-video generator
AI voiceover studio with 200+ voices for creators and businesses
AI podcast and video editor that lets you edit media by editing text
AI audio enhancement that makes any microphone sound studio-quality
Google's AI research assistant grounded exclusively in your own documents
Apple's single-photo → 3D model with view-dependent lighting and reflections
2B video model — dense captions and plain-English timestamp search
Foundation ASR built for real-world audio where Whisper breaks down
Cognition's autonomous AI software engineer — plans, codes, debugs end-to-end
AWS's AI coding assistant with deep AWS service knowledge
Real-time AI image generation and enhancement with live canvas editing
Hollywood-grade AI video generation and editing
AI video generation and editing with scene-modify and Pikaffects
Enterprise AI video platform — professional training and comms videos at scale
△ Concepts
All concepts →AI that understands and generates more than one type of media — text, images, audio, or video.
The AI behind image generation: learning to reverse the process of adding noise.