| Native outputResolution and frame rate | 720p demonstratedBFL used 720p clips in its preliminary evaluation. Maximum native resolution and FPS are not published yet. | 720p / 1080p · 24fpsPublished native output options on Vertex AI; some reference and extension modes have exceptions. | 720p landscape or portrait1280×720 and 720×1280 are listed for Sora 2. A native FPS is not stated on the current model card. | 720p · 24/25fpsPublished Gen-4.5 output specification for text-to-video and image-to-video workflows. |
|---|
| Single-clip durationHow long one generation can hold | Up to 20 secondsVideo and audio are generated together in one clip—the longest published single generation in this comparison. | 4, 6, or 8 secondsVertex AI supports fixed durations; reference-image workflows are limited to 8 seconds where available. | Not stated on model cardThe current Sora 2 model page publishes output size and per-second pricing, but not a duration limit. | 2–10 secondsCreators select a duration between two and ten seconds in the Gen-4.5 workflow. |
|---|
| Input & controlT2V, I2V, V2V, frames, references | Broad multimodal controlText-to-video, image animation and references, video-to-video, keyframes, plus video-and-audio continuation. | Text, image, first + last frameStrong shot framing controls; exact reference-image and extension support depends on the Veo variant. | Text + image inputThe current API model card lists natural-language and image inputs for video generation. | Text-to-video + image-to-videoCamera motion and scene behavior are directed through prompts; an input image anchors appearance. |
|---|
| Native audioDialogue, ambience, SFX, lip sync | Joint video + audio generationAnnounced multilingual dialogue, ambience, physical sound effects, and speech synchronized with lip movement. | Native audio generationVeo 3.1 generates video with audio, including dialogue, sound effects, and environmental sound. | Synchronized audio outputSora 2 produces video with synchronized audio from text or image inputs. | Not native to Gen-4.5The published Gen-4.5 specification covers video output; audio uses separate Runway tools or workflows. |
|---|
| Continuity workflowCharacter, style, and multi-shot cohesion | References + agentic chainingVideo-to-video can carry central elements into a new scene; chained clips target longer multi-shot sequences. | First / last frame controlDefined boundary frames help control transitions and shot endpoints within supported variants. | No comparable API specThe model card does not publish a standardized character-reference or multi-shot continuity control. | Image-anchored clipsInput images help establish a shot, while longer scene continuity requires an assembled workflow. |
|---|
| Aspect ratiosLandscape, vertical, square, cinematic | Broad range announcedBFL states support extends beyond conventional cinematic output; the exact production ratio list is pending. | 16:9 · 9:16Landscape and vertical output are documented, with mode-specific exceptions. | 16:9 · 9:16The model card lists 1280×720 landscape and 720×1280 portrait output. | 16:9 · 9:16 · 1:1 · 4:3 · 3:4 · 21:9Gen-4.5 image-to-video publishes the widest exact ratio list in this comparison. |
|---|