
google/gemini-omni-flash/text-to-videoA natively multimodal Google DeepMind model that generates cinematic videos with synchronized native audio from a text prompt alone, grounded in real-world physics for controllable, high-speed video generation.
Set up the model and run your first generation.
Ctrl ⏎Runs published with this model. Load one into the playground and change whatever you like.
A solitary man leans against a vintage red coupe on an empty seaside road at dusk. The wind gently moves the grass and his coat while a lonely cloud hangs motionless above him. The camera slowly tracks sideways, capturing reflections on the car body and the endless blue ocean beyond.
Where this model can run, priced for your account. Auto routing picks the top row and moves down it if a provider fails.
| Provider | Price | Speed | Completed |
|---|---|---|---|
| Atlas CloudLowest price | $0.125 /sec | — | — |
| fal.ai | 1 unit | — | — |
| OpenRouter | — | — | — |
| Replicate | — | — | — |
Speed and reliability appear once this platform has enough recent runs on a provider.