Models

Google's Omni 1.1 Flash and Gemini 3.5 Transcribe: Two Models, One Strategy

CRAZE CRAZE Summary 3 things to know
  • The 360p draft mode is the real news: generate video previews 60% faster at one-third cost, then upscale only final takes.
  • Gemini 3.5 Transcribe treats speech as meaning, not sound—removing filler words and resolving corrections, with a 2.6% non-streaming Word Error Rate.
  • Same-day releases show a multimodal stack: Omni for video, Transcribe for voice, distributed through Adobe, Figma, Runway, Gboard, and Chrome.
Jeff Editorial | · 5 min read
Google's Omni 1.1 Flash and Gemini 3.5 Transcribe: Two Models, One Strategy

On August 27, Google released two Gemini models on the same day. Gemini Omni 1.1 Flash is a generative video model that went from preview to general availability. It creates and edits video from text, images, and other video clips. Gemini 3.5 Transcribe is a speech-to-text model described as Google's "most precise" yet, converting speech into formatted, cleaned-up text.

Two models are not a coincidence. They are a suite.

Draft at 360p. Export at 4K. The Cost Problem Is Solved.

Video generation has been expensive and unpredictable. Omni 1.1 Flash's new 360p draft mode is the quietest but most important improvement. Previews can be generated up to 60% faster than 720p at one-third of the cost. For developers iterating on prompts, this is the difference between a production budget and a demo budget.

The pricing structure: 360p video costs $0.03 per second, 720p is $0.10 per second, 1080p is $0.15 per second, and 4K is $0.30 per second. Developers can draft at 360p until the prompt is right, then upscale the final version to 1080p or 4K.

The model also analyzes up to 10 seconds of prior context when extending videos, compared to just one second in the preview version. Extensions happen in 10-second increments, up to 40 seconds total. The result is improved visual consistency.

Other additions: first-and-last-frame control (specify the start and end frames, and the model generates the footage between them) and video references (up to three seconds of reference clips to maintain character consistency).

2.6% WER, 70% Faster, and It Removes Your "Ums"

The traditional speech-to-text model is a tape recorder. It writes down whatever it hears. Gemini 3.5 Transcribe is being positioned as something different: a model that understands what you meant to say and cleans it up for you.

If you say "Let's meet Tuesday—no, Wednesday," the model recognizes the correction and outputs "Wednesday." It removes filler words like "um" and "uh." It auto-formats text with punctuation and capitalization. It can also delegate complex tasks like image generation or file analysis to other Gemini models through function calling.

The performance numbers: 4.0% Word Error Rate for streaming and 2.6% for non-streaming use cases. Time to final transcription is 70% faster than Google's previous Chirp 3 model. It supports over 85 languages and can handle regional accents.

Google's Omni 1.1 Flash and Gemini 3.5 Transcribe: Two Models, One Strategy
Google's Gemini Omni 1.1 Flash and Gemini 3.5 Transcribe—video generation and speech-to-text, released on the same day.

Video + Voice + Text = Google's Real Strategy

Omni 1.1 Flash and Transcribe are not separate releases. They are two parts of a larger strategy: Gemini Audio (including Gemini 3.5 Live and Gemini 3.5 Transcribe) is the voice layer; Omni is the visual generation layer. Together, they form a multimodal stack.

The distribution channels are the real story. Omni 1.1 Flash is available through Google AI Studio, the Gemini API, and Google Flow. Adobe has already integrated Omni Flash into Firefly, with Figma Weave, GMI Cloud, and Runway following. Gemini 3.5 Transcribe is being embedded into Gboard's Rambler feature on Android, the Gemini app on macOS, and is coming to Chrome.

Google Is Building a Multimodal Stack—and Selling It Piece by Piece

Google is not trying to win with a single model. It is trying to make AI frictionless across every input and output modality—text, voice, image, video—at prices that developers can afford. Two releases on the same day are not a coincidence. They are a statement.


P.S. The quietest signal in the Omni 1.1 Flash release is the 360p draft mode. Video generation is still expensive and unpredictable. Google's bet is that if you can iterate cheaply and upscale only the keepers, developers will build on its platform. That is not a technical insight. It is a business insight.


Frequently Asked Questions

Q: What is Gemini Omni 1.1 Flash?

A: Gemini Omni 1.1 Flash is Google's generative video model, released on August 27, 2026. It creates and edits video from text, images, and other video clips. Features include 360p draft mode (60% faster, 1/3 the cost), 4K output, first-and-last-frame control, and video references for character consistency.

Q: What is Gemini 3.5 Transcribe?

A: Gemini 3.5 Transcribe is Google's speech-to-text model, also released on August 27. It converts speech into formatted, cleaned-up text—removing filler words, auto-correcting errors, and adding punctuation. It achieves 2.6% Word Error Rate (non-streaming) and is 70% faster than Google's previous Chirp 3 model.

Q: How does the 360p draft mode work?

A: Developers can generate low-resolution previews at 360p (up to 60% faster, one-third the cost of 720p) to iterate on prompts. Once the prompt is right, they can upscale the final version to 1080p or 4K. This solves the cost problem in video generation.

Q: What are the pricing tiers for Omni 1.1 Flash?

A: 360p costs $0.03 per second, 720p is $0.10 per second, 1080p is $0.15 per second, and 4K is $0.30 per second. Previews are generated at 360p or 720p to keep iteration costs low.

Q: What are the new video features in Omni 1.1 Flash?

A: Extended context (analyzes up to 10 seconds of prior video), first-and-last-frame control (specify start and end frames), video references (up to 3 seconds to maintain character consistency), and draft-then-upscale workflow.

Q: How does Gemini 3.5 Transcribe handle corrections?

A: If you say "Let's meet Tuesday—no, Wednesday," the model recognizes the correction and outputs "Wednesday." It also removes filler words like "um" and "uh," auto-formats with punctuation and capitalization, and can delegate tasks to other Gemini models.

Q: What are the key performance metrics for Gemini 3.5 Transcribe?

A: 4.0% Word Error Rate for streaming and 2.6% for non-streaming use cases. Time to final transcription is 70% faster than Google's previous Chirp 3 model.

Q: What languages does Gemini 3.5 Transcribe support?

A: It supports over 85 languages and can handle regional accents. It is now the default voice model for Gboard's Rambler feature on Android and the Gemini app on macOS.

Q: Where is Omni 1.1 Flash available?

A: Through Google AI Studio, the Gemini API, and Google Flow. Adobe has integrated it into Firefly, with Figma Weave, GMI Cloud, and Runway following.

Q: What is the strategic significance of releasing both models on the same day?

A: The two models are parts of a larger multimodal strategy—Gemini Audio (voice) and Gemini Omni (video). Together with text and image capabilities, they form a complete multimodal stack. Google is betting on making AI frictionless across every input and output modality at prices developers can afford.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article