OpenAI Image to Video Generator

Sora 2 Image to Video: From Still to Story

Give Sora 2 a single frame — a product shot, a generated keyframe, a campaign visual — and it builds the scene around it: motion, camera, and sound for up to 12 seconds.

10 credits per second of video.By OpenAI

Video Generation Input

0/2000 characters
OpenAI10 credits per second of video.

A high-intent Sora variant for turning a still image into motion.

Your generated video will appear here.

Enter a prompt and click 'Generate Video' — or start from an example:

Examples

Sora 2 Image to Video video examples

Hover a card to preview, then open the player to watch with sound and review the full prompt.

Skincare serum ad in motion

Static serum ad animated with sliding droplets and a slow push-in, image to video by Sora 2 Image to Video

View prompt

Animate this skincare serum advertisement into a premium product video, keeping the bottle, pedestal, eucalyptus leaves, and all text exactly as they are. From 0s to 3s: the water droplets around the base shimmer and one droplet slides slowly down the green glass bottle, the headline "DEWY, NOT GREASY" staying perfectly sharp and static. From 3s to 6.5s: a soft breeze moves the eucalyptus leaves gently, light flares subtly across the glass shoulder of the bottle, faint morning ambience with distant birds. From 6.5s to 10s: the camera pushes in very slowly toward the bottle while the light warms slightly, ending on a gentle settle, a single soft piano note fading in underneath. No warping of the typography, no changes to the layout, photorealistic glass and liquid motion only.

LinkedIn headshot becomes a profile clip

Corporate headshot animated into a subtle professional profile clip, image to video by Sora 2 Image to Video

View prompt

Animate this professional corporate headshot into a natural, restrained portrait video while keeping the woman's face, hair, blazer, and office background completely consistent. From 0s to 2.5s: she holds her gentle confident smile, perfectly still except for one natural blink, soft office air moving a few strands of hair. From 2.5s to 5s: her smile widens slightly and warms as if greeting a colleague, her head tilting a few degrees, the blurred glass-wall background staying fixed. From 5s to 7s: she settles back to the original composed expression with a small nod. Quiet office ambience, a faint HVAC hum and distant keyboard taps. Locked camera, photorealistic skin texture, no change to her identity, clothing, or the lighting.

Sci-fi one-sheet becomes a motion poster

Sci-fi movie poster animated into a living one-sheet with locked typography, image to video by Sora 2 Image to Video

View prompt

Animate this science fiction movie poster into a theatrical motion poster, keeping the astronaut, the ring-shaped structure, both moons, and the title "THE QUIET ORBIT" and its billing block exactly in place. From 0s to 3s: atmospheric dust drifts slowly across the rust-red desert, the lavender sky darkening almost imperceptibly, a low ambient synth drone fading in with deep rumbling. From 3s to 6.5s: the ring structure emits a faint slow pulse of light traveling along its curve, the astronaut's suit fabric stirring in the wind, sand trickling off a nearby dune. From 6.5s to 10s: the two moons brighten slightly as the pulse completes one full circuit of the ring, a deep cinematic bass swell rising and cutting to silence on the final frame. The title serif lettering stays razor sharp and static the entire time, no layout drift, epic and still.

Fashion editorial coat in the wind

Fashion editorial still animated with billowing coat fabric, image to video by Sora 2 Image to Video

View prompt

Animate this fashion editorial photograph into a high-fashion film moment, keeping the model's face, pose, and the concrete parking structure exactly consistent. From 0s to 3s: the hem of the oversized crimson trench coat begins to lift and ripple in a slow gust, fabric waves rolling naturally through the heavy wool, distant city rumble and wind. From 3s to 6.5s: the gust peaks — the coat billows sculpturally around her while she holds her pose completely still, one lock of hair crossing her face, dust motes swirling through the shaft of side light. From 6.5s to 10s: the wind dies down and the fabric settles back with realistic weight, the long shadows holding steady, the wind fading to a low airy hum. Locked camera, medium-format film look, muted gray environment with the red coat dominant, photorealistic fabric motion.

What is Sora 2 Image to Video?

Sora 2 Image to Video uses your image as the opening frame and generates the scene that follows: 4, 8, or 12 seconds of 720p motion at 10 credits per second on RenderFlow AI.

The still fixes the look — subject, styling, lighting — which frees the model to spend its effort on what Sora 2 does best: coherent action, camera language, and synchronized audio. It is the difference between animating a picture and directing a scene that happens to start from one.

Why start Sora 2 from a still instead of text

A scene, not just movement

A scene, not just movement for stills with a story in them. Where simpler image-to-video models add drift and zoom, Sora 2 builds events: a barista finishes the pour, steam rises, the customer reaches in. Your frame becomes the first second of something.

Sound generated from one image

Sound generated from one image when the clip needs to feel alive. The model infers the scene's ambience from the still — a rainy street gets rain — and your prompt can quote dialogue for the people in frame. Few image-to-video workflows even attempt this.

Twelve seconds from a single frame

Twelve seconds from a single frame for beats that need room. Product reveals with a slow orbit, portraits that settle into a smile, establishing frames that come alive — the 4/8/12 second options let the clip match the idea instead of the other way round.

Stills that turn into stories with Sora 2

  • Animating campaign key visuals into scenes
  • Product reveals from packshots
  • AI keyframes extended into full clips
  • Portrait and character moments with sound

Building a Sora 2 clip around one frame

Choose a frame with a scene inside it

The best inputs imply what happens next: a hand reaching, a glass half-filled, a door ajar. Flat, sealed compositions give the model less story to work with.

Script the following seconds

Your prompt describes what happens after the first frame, in order: the action, the camera move, and how the clip ends. At 12 seconds, beats need timing — name what happens early and late.

Set duration to the arc

Pick 4 seconds for a gesture, 8 for a reveal, 12 for a full two-beat scene. Each second costs 10 credits, so match runtime to the idea rather than maxing out by default.

Grade the audio on review

Listen as much as you watch: check that inferred ambience fits the scene and quoted dialogue lands on the right character. If sound misses, adjust the audio direction before touching the visual notes.

Frequently Asked Questions

What does the input image control?

The still anchors subject identity, styling, lighting, and the opening composition. Everything after frame one — action, camera movement, sound — comes from your prompt and the model's reading of the scene.

What durations and output sizes are available?

Clips run 4, 8, or 12 seconds at 720p, billed at 10 credits per second. Framing follows your source image, so upload the crop you want on screen.

Does Sora 2 image-to-video include audio?

Yes — it generates sound in the same render, inferring ambience from the image and following quoted dialogue in your prompt. Undirected sound is the more common surprise, so describe the audio you want.

Standard or Pro image-to-video — how do I choose?

This standard tier is 720p at 10 credits per second, right for drafting and most social placements. Sora 2 Image to Video Pro adds a 1080p option at 30 to 50 credits per second for clips headed to premium placements.

What inputs work poorly?

Dense collages, heavy text overlays, and very low-resolution images give the model an unstable foundation — expect drift or warped details. Clean, single-scene stills with clear subjects animate most predictably.