Google held its annual developer conference today and shipped two things worth understanding: a smarter, cheaper AI model for running complex tasks, and a video creation tool that takes whatever you have and turns it into something new. Start with Gemini 3.5 Flash. The short version is that it performs at the level of the best AI models currently available, but runs four times faster and costs roughly half to a third as much to use. Google's CEO Sundar Pichai acknowledged the cost problem directly: companies are already burning through their annual AI budgets by May, and that is a real operational issue. The model is designed specifically for tasks where AI needs to work through a long sequence of steps on its own, things like drafting and refining documents, running research across multiple sources, or building and iterating on software, without needing a human to approve every single step. That kind of autonomous, multi-step work is where AI is headed, and the cost per step matters enormously at any real scale. Gemini 3.5 Flash is priced at $1.50 per million input tokens and $9.00 per million output tokens. For comparison, the previous flagship model, 3.1 Pro, costs $2 to $18 per million tokens depending on context length. The speed difference is also significant: at close to 280 output tokens per second, versus around 60 to 70 for comparable models from OpenAI and Anthropic, tasks that used to feel sluggish now happen fast enough to feel real. Now for Omni. Most current video tools do one thing: you give them a text prompt, they give you a video. Omni works differently. You can give it a photo of a product, a voice recording, a clip you shot on your phone, and a written description, and it reasons across all of them to produce a single, coherent output. Characters stay consistent from scene to scene. Text rendered in the video, like a slogan or a label, is accurate. Physics behaves correctly. The practical uses are wide. A marketing team can produce variant ads without an agency cycle. A training team can build explainer videos from a document and a product photo. A retail group can put a product into a lifestyle scene it never filmed. These are not hypothetical futures; they are the specific use cases Google named at the launch. Timing matters here. OpenAI shut down its consumer-facing Sora app in April because the costs were unsustainable: running costs reportedly ran between $8 to $12 million a month while subscription revenue never covered it. Google, which has its own hardware infrastructure and an existing base of paying Gemini subscribers, is absorbing those economics differently. The gap left by Sora's closure is real, and Omni Flash is squarely positioned to fill it. Omni Flash currently produces clips up to 10 seconds long, which Google says is a deliberate choice based on what most users actually need, not a technical ceiling. Longer durations are coming. A more capable Pro version is also in development, aimed at professional and enterprise users who need higher quality output. API access for businesses arrives in the coming weeks. Two caveats are worth noting. First, early reports from developers who received early access say the content restrictions in Omni are fairly strict, which could limit some business use cases. Second, Omni Flash is only available to paid Gemini subscribers for now, not on the free tier. For anyone managing teams that produce content, train staff, or run marketing operations: the production chain for short video is compressing fast. What required a brief, a shoot, an edit, a voiceover session, and a round of revisions is becoming a single prompt-driven workflow. That does not eliminate judgment about what to make or why, but it does change how much you need to spend to make it.
Product Launch3 min read
Google Launches Gemini 3.5 Flash and Omni Video Model
June 2, 2026Synthesized from 3 sources: TechCrunch, Ars Technica, Google DeepMind
At Google I/O 2026, Google launched Gemini 3.5 Flash, a fast and cheap AI model built to run long, autonomous tasks, and Gemini Omni Flash, a tool that turns any mix of text, images, audio, and video into a new video, with implications for anyone who produces content, trains staff, or manages marketing.
Related Coverage
Stay informed
Get AI intelligence like this delivered to your inbox.