Release notes
Google's official release notes. Antigravity: Google's AI coding app. Gemini CLI: Gemini in a terminal window. Gemini app: the website and phone app. Gemini API: what developers build with.
New model·Sep 22, 2026
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS generally available (GA): Released our next-generation text-to-speech (TTS) audio models and the Gemini API Voices endpoint (/v1beta/voices): Gemini 3.8 Flash TTS (gemini-3.8-flash-tts): Flagship creative TTS model engineered for studio-grade voice fidelity, nuanced acting, regional dialects, and long-form multi-turn stability. Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts): Fast, cost-efficient TTS model built to replace gemini-3.1-flash-tts-preview for high-throughput production and real-time voice agent cascades. Voice design, Voice replication, and the Extended Voice Library: Create persistent custom vocal personas from text prompts, replicate voices with consent verification, and query 150+ prebuilt and custom voices.
New model·Sep 15, 2026
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking generally available (GA): Released two new audio-to-audio models for real-time voice applications using the Live API: Gemini 3.8 Live (gemini-3.8-live): The default option for most low-latency voice agent experiences and real-time dialogue without reasoning delays. Features interleaved reasoning, default asynchronous function calling, and full session client content updates. Gemini 3.8 Live Extended Thinking (gemini-3.8-live-extended-thinking): High-reasoning audio-to-audio model supporting background reasoning during live audio interactions, recommended when higher background reasoning is required. To get started, see the Live API guide, the Capabilities guide, and the Thinking guide.
New model·Sep 2, 2026
Gemini 3.8 Flash generally available (GA): Released gemini-3.8-flash, our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. To get started, see the Gemini 3.8 Flash model page and the Latest model guide.
New model·Aug 27, 2026
Gemini Omni Flash generally available (GA): Released gemini-omni-1.1-flash, the GA version of our fast, conversational video generation and editing model. This release includes significant new capabilities: Video extension: Seamlessly extend existing videos by generating continuations at the end of a clip using the extend task or directly with a prompt. Interpolation (first + last frame): Generate a video transitioning between two images using the image_to_video task with up to 2 images. Resolution control: New resolution parameter in video_config supports 360p, 720p (default), 1080p, and 4k outputs. 1080p and 4K outputs are generated using upscaling. The existing gemini-omni-flash-preview endpoint will be deprecated on September 30, 2026. To get started, see the Gemini Omni Flash model page and the omni guide.
New model·Aug 26, 2026
Gemini 3.5 Transcribe generally available (GA): Released two dedicated speech-to-text models based on Gemini's audio understanding: Gemini 3.5 Transcribe (gemini-3.5-transcribe): High-accuracy, low-latency non-streaming speech-to-text with utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps, and custom vocabulary biasing (up to 1,000 terms). Gemini 3.5 Transcribe Live (gemini-3.5-transcribe-live): Low-latency, bidirectional streaming speech-to-text over WebSockets using the Live API, supporting interim and finalized transcription events, Smart transcription mode, and multiple Voice Activity Detection (VAD) strategies. To get started, see the Audio transcription guide, the Live transcription guide, and the Gemini 3.5 Transcribe model page.
New model·Aug 13, 2026
Gemini 3.7 Flash generally available (GA): Released our most intelligent workhorse model yet for coding and agents: Gemini 3.7 Flash (gemini-3.7-flash): Substantial improvements across software engineering, web development, and agentic workflows, available at an introductory price through December 31, 2026. To get started, see the Gemini 3.7 Flash model page and the Latest model guide.
New model·Jul 30, 2026
Gemini Robotics ER 2 in public preview: Released two new embodied reasoning model endpoints for robotics: gemini-robotics-er-2-preview: Advanced spatial reasoning, agentic code execution, multi-step tool orchestration, video moment finding, progress classification, and multi-robot coordination. gemini-robotics-er-2-streaming-preview: Optimized for real-time text streaming using the Live API, enabling low-latency robot agents with bidirectional audio and video input. Both model endpoints accept text, image, video, and audio inputs and support function calling with blocking behavior for physical robot actions. To get started, see the Gemini Robotics ER overview. For real-time streaming use cases, see Robotics with streaming.
New model·Jul 21, 2026
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available (GA): Released stable, production-ready versions of our latest 3.x Flash models: Gemini 3.6 Flash (gemini-3.6-flash): Features improved token efficiency and code/agentic planning capabilities at a lower price point than 3.5 Flash, resolving developer feedback around output verbosity. Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite): Offers a low-latency, highly cost-effective subagent option designed for high-volume automation. To learn more, see the Latest Gemini model guide.
New model·Jun 30, 2026
Gemini Omni Flash in public preview: Released gemini-omni-flash-preview, a high-performance multimodal model designed for high-speed video generation and conversational video editing. Using the Interactions API, you can generate 3–10 second videos at 720p from text descriptions or animate still images, and then conversationally edit and refine the outputs. To get started, see the Gemini Omni Flash guide and the Gemini Omni Flash model card.