
nvidia/cosmos-3-super/text-to-imageCosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.
Set up the model and run your first generation.
Ctrl ⏎Runs published with this model. Load one into the playground and change whatever you like.

A symphony orchestra performing inside a grand European train station after the last train has departed. Empty platforms stretch into darkness while warm lantern light illuminates the musicians. Steam drifts through the air, creating a dreamlike cinematic atmosphere. Emotional storytelling, movie still, ultra realistic.
Where this model can run, priced for your account. Auto routing picks the top row and moves down it if a provider fails.
| Provider | Price | Speed | Completed |
|---|---|---|---|
| fal.aiLowest price | $0.040 /gen | — | — |
| Atlas Cloud | $0.044 /gen | — | — |
| OpenRouter | — | — | — |
| Replicate | — | — | — |
Speed and reliability appear once this platform has enough recent runs on a provider.