Text or image to video

AI Video Generator

Create clips from a prompt or starting frame with three current model families and five purpose-built generation modes.

  • 35 free monthly credits
  • No provider API keys
  • Credit price shown by model

Video Generation Input

0/2000 characters
ByteDanceFrom 40 credits for 4 seconds, based on duration and resolution.

Fast text-to-video generation with native audio, reference images, web search, and output up to 4K.

Enable
Enable
Enable
Enable
Enable
Enable
5

Your generated video will appear here.

Enter a prompt and click 'Generate Video' — or start from an example:

Model guide

Choose the right AI video model

Select by input type first, then compare duration, resolution, native audio, motion quality, and credit cost.

Model Input Best for Credit cost
Seedance 2.0 ByteDance text to video native audio, reference-guided video, high-resolution video From 40 credits for 4 seconds, based on duration and resolution.
Wan 2.7 Text to Video Alibaba text to video longer clips, 1080p video, prompt expansion From 20 credits for 2 seconds, based on duration and resolution.
Wan 2.7 Image to Video Alibaba image to video image animation, start and end frames, 1080p video From 20 credits for 2 seconds, based on duration and resolution.
Kling 3.0 Text to Video Kuaishou text to video cinematic motion, social video, fast iteration From 34 credits for 3 seconds.
Kling 3.0 Image to Video Kuaishou image to video premium image animation, character motion, cinematic shots From 42 credits for 3 seconds.

ByteDance · text to video

Seedance 2.0

Fast text-to-video generation with native audio, reference images, web search, and output up to 4K.

  • native audio
  • reference-guided video
  • high-resolution video

From 40 credits for 4 seconds, based on duration and resolution.

Alibaba · text to video

Wan 2.7 Text to Video

Text-to-video generation with 720p and 1080p output, optional audio guidance, and clips up to 15 seconds.

  • longer clips
  • 1080p video
  • prompt expansion

From 20 credits for 2 seconds, based on duration and resolution.

Alibaba · image to video

Wan 2.7 Image to Video

Animate a first frame with an optional ending frame, audio guidance, 1080p output, and clips up to 15 seconds.

  • image animation
  • start and end frames
  • 1080p video

From 20 credits for 2 seconds, based on duration and resolution.

Kuaishou · text to video

Kling 3.0 Text to Video

Fast Kling 3.0 text-to-video generation with flexible aspect ratios and clips from 3 to 15 seconds.

  • cinematic motion
  • social video
  • fast iteration

From 34 credits for 3 seconds.

Kuaishou · image to video

Kling 3.0 Image to Video

Higher-quality Kling 3.0 image animation with controlled motion and clips from 3 to 15 seconds.

  • premium image animation
  • character motion
  • cinematic shots

From 42 credits for 3 seconds.

Use cases

Create motion for the work you already do

Start from words when the scene is open-ended. Start from an image when visual consistency matters more.

Cinematic scenes and storyboards

Turn a shot description into moving concept footage with deliberate camera motion, lighting, pacing, and atmosphere.

Product and campaign videos

Create short product reveals, social ads, launch clips, and visual variations sized for the channel where they will run.

Animate a still image

Use a product photo, illustration, character, or generated image as the first frame and describe the motion you want.

Social video and B-roll

Generate vertical clips, establishing shots, transitions, and supporting footage without a traditional production setup.

Examples

AI video examples from real prompts

Compare text-to-video, image-to-video, native audio, camera movement, and start-frame consistency across the three featured model families. Hover a card to preview, click to play with sound in the viewer.

Seedance 2.0

Text-to-video with native audio, physical motion, and connected shots

Espresso craft commercial

Espresso commercial with four-shot structure and synced foley, text to video by Seedance 2.0

View prompt

A cinematic 10-second espresso commercial in a warm artisan cafe, four connected shots with realistic physics and synchronized sound. From 0s to 2s: extreme macro of roasted coffee beans tumbling in slow motion onto a wooden counter, each bean bouncing naturally, soft morning light, gentle rattling sound. From 2s to 4s: hands tamp ground coffee into a steel portafilter, fine dust drifting in the light, a firm thud as the tamper sets down. From 4s to 7s: rich espresso streams spiral out of the portafilter spouts into a white cup, crema swirling, thin steam rising, the liquid sound shifting from trickle to flow. From 7s to 10s: the finished cup slides slowly toward the camera across the counter, cafe ambience and soft milk-steamer hiss in the background, empty space in the upper third for a logo. Warm color grade, shallow depth of field, photorealistic.

Domino-and-marble physics run

Chain-reaction physics run with collisions, rolling, and a water splash, text to video by Seedance 2.0

View prompt

A 10-second continuous shot of a physics chain reaction on a sunlit workshop bench, realistic motion throughout. From 0s to 2s: a fingertip tips the first wooden domino and a line of dominoes falls left to right with crisp clacking sounds, each impact transferring momentum believably. From 2s to 4s: the last domino nudges a glass marble that rolls down a sloped brass track, accelerating naturally with a rolling hum. From 4s to 6s: the marble drops off the track and lands on a small drum skin, bouncing once with a deep boom that makes a ring of paper scraps jump. From 6s to 8s: the marble rolls off the drum into a tilted funnel, spirals around it with increasing speed, and drops out the bottom. From 8s to 10s: the marble splashes into a glass of water, realistic splash crown and ripples, water droplets catching the sunlight, gentle splash and settling sounds. Macro photography style, warm side light, dust motes in the air.

Street food vendor dialogue

Night-market vendor with English lip-synced dialogue and layered ambience, text to video by Seedance 2.0

View prompt

A 9-second street food scene at a night market, one continuous shot. From 0s to 3s: a middle-aged vendor in a blue apron tosses noodles in a blazing wok, flames leaping up, sparks and steam rising, loud sizzling and the roar of the wok burner. From 3s to 6s: he turns to the camera with a big smile and says clearly "Best noodles in the whole night market!", his lip movements matching the English words precisely, the sizzle continuing underneath. From 6s to 9s: he slides the noodles onto a plate and pushes it toward the camera, steam curling into the cool night air, neon stall signs glowing in the background, crowd murmur and a distant scooter passing. Handheld documentary camera feel, warm stall lighting against blue night, photorealistic.

Wan 2.7 Image to Video

Start-frame fidelity, style consistency, and restrained character motion

Golden-hour skyline breathing

Golden-hour skyline animated with drifting clouds and shifting light from a single still, image to video by Wan 2.7

View prompt

Animate this golden-hour city skyline into a calm late-afternoon moment, keeping every building, window, and rooftop exactly as it is. From 0s to 3s: thin clouds drift slowly across the sky and the sunlight on the glass towers warms slightly, the water in the foreground shimmering with small moving ripples. From 3s to 6s: a few tower windows catch brief sun glints as the light angle shifts, a tiny boat wake trails slowly across the lower-left water, and haze drifts gently between the buildings. From 6s to 9s: the light begins to cool toward dusk blue, the tallest tower's glass catching one last warm reflection, clouds continuing their slow drift. The camera holds a locked position the entire time, no cuts, no new objects appearing, photorealistic natural motion only.

The Great Wave as a woodblock animation

Hokusai woodblock print animated in strict ukiyo-e style, image to video by Wan 2.7

View prompt

Animate this Japanese woodblock print while strictly preserving the ukiyo-e style: flat color areas, bold dark outlines, and visible print texture. From 0s to 3s: the foam claws at the crest of the great wave curl and drip slowly, small spray dots drifting off the wave tips, the flat blue water swelling gently. From 3s to 6s: the three boats rock and pitch over the swells, the rowers bracing in unison, the wave rising higher with stylized layered motion. From 6s to 9s: the crest of the great wave crashes forward in flat-color, contour-lined animation, foam patterns swirling, while Mount Fuji stays perfectly still in the background. No camera movement, loop-friendly rhythm, hand-drawn frame-by-frame feel, no photorealistic blending.

Studio portrait micro-expressions

Studio portrait animated with natural micro-expressions, image to video by Wan 2.7

View prompt

Animate this studio portrait into a natural, restrained portrait video while keeping the man's face, hair, sweater, and background completely consistent. From 0s to 2s: he holds his relaxed expression, perfectly still except for one natural blink. From 2s to 5s: a warm, subtle smile forms slowly, his eyebrows lifting slightly, a few hairs moving in soft studio air, and he blinks again naturally. From 5s to 7s: he gives a small, friendly nod as if greeting someone, then settles back to the original calm expression. The gray background and lighting stay completely unchanged. Locked camera, photorealistic skin texture, no change to his identity or clothing.

Kling 3.0 Image to Video

Cinematic camera movement, atmosphere, and realistic subject motion

Coastal cliffs aerial push

Coastal cliffs animated into an aerial push with crashing surf, image to video by Kling 3.0

View prompt

Animate this coastal cliff photograph into a cinematic aerial shot with realistic ocean motion. From 0s to 4s: the camera pushes slowly forward toward the sea stack, waves rolling in from the horizon and breaking white against the rocks, the sound of surf building gradually. From 4s to 7s: a larger swell crashes over the foreground rock ledge, spray bursting upward and blowing sideways in the wind, a few gulls crossing the frame and crying faintly. From 7s to 10s: the camera continues gliding past the sea stack, revealing more of the lighthouse headland, the wave patterns shimmering under the pale light, wind and distant surf underneath. A sparse cinematic score of slow strings and soft piano swells gently as the big wave crashes, then settles back with the camera. One continuous shot, no cuts, photorealistic water motion, natural coastal ambience.

Racehorse over the fence

Racehorse still animated into a slow-motion fence jump, image to video by Kling 3.0

View prompt

Animate this racehorse photograph with completely realistic weight and motion, keeping the horse, fence, and track consistent. From 0s to 3s: the gray horse stretches over the brush fence in slow motion, mane and tail flowing, muscles working under the skin, hooves tucked tight, the sound of hoofbeats and rushing air. From 3s to 6s: it lands hard on the far side, turf clods kicking up from its hooves, ears pinned, and drives forward into a gallop. From 6s to 8s: the camera pans smoothly to follow as the horse accelerates down the track, mane whipping, the white rails and hedges passing steadily behind. A driving rhythmic score with low strings and percussion builds tension through the jump and releases as the horse gallops away. Realistic anatomy and momentum throughout, no floating or sliding, natural trackside sound with a faint distant crowd.

Dancer in the spotlight

Spotlit dancer animated with slow cinematic motion, image to video by Kling 3.0

View prompt

Animate this stage photograph into a cinematic dance moment, keeping the dancer, spotlight, and dark stage consistent. From 0s to 3s: she holds her pose as fine dust drifts slowly through the spotlight beam, her breathing subtle, the stage floor creaking softly once. From 3s to 6s: she sweeps her raised arm down and turns, her hair swinging with the motion, fabric catching the light as her shadow sweeps across the worn stage floor. From 6s to 9s: she rises into a slow, controlled leg extension, perfectly balanced, the spotlight holding on her while the theater stays dark and silent around her. A lone Spanish guitar plays a slow flamenco phrase, joined by soft palmas hand claps that build gently as she moves. The camera pushes in very slowly the entire time. Graceful, weighted, realistic human motion, quiet theater ambience with faint echoing footsteps.

Source images by dolbinator1000, mypubliclands, Paolo Camera, and OmarOBorO under CC BY 2.0. The Great Wave off Kanagawa is public domain.

Workflow

From shot idea to generated video

A controlled video prompt separates what moves, how the camera moves, and what must stay visually consistent.

  1. 01

    Choose text or image input

    Start from a prompt, or select an image-to-video model and upload the frame that should anchor the shot.

  2. 02

    Describe action and camera

    Name the subject movement, environmental motion, shot size, camera direction, and pace in concrete terms.

  3. 03

    Set format and duration

    Choose the aspect ratio, resolution, and clip length supported by the selected model.

  4. 04

    Generate and refine

    Review the motion and continuity, then adjust one variable at a time before generating the next version.

Video prompt formula

[Subject] performs [specific action] in [environment]. [Camera movement and shot size]. [Environmental motion]. [Lighting, pace, and mood]. Keep [identity, product, or composition] consistent.

Practical details

Frequently asked questions

Which AI video model should I use?

Seedance 2.0 is a strong fit for native audio and reference-guided video. Wan 2.7 Text to Video is a strong fit for longer clips and 1080p video. Wan 2.7 Image to Video is a strong fit for image animation and start and end frames. Kling 3.0 Text to Video is a strong fit for cinematic motion and social video. Kling 3.0 Image to Video is a strong fit for premium image animation and character motion.

What is the difference between text-to-video and image-to-video?

Text-to-video creates a shot from a written description. Image-to-video uses an uploaded image as the visual starting point, which is better when the subject, character, product, or composition must remain recognizable.

How should I write an AI video prompt?

Describe the subject and action first, then add camera movement, environment motion, lighting, pacing, and mood. For image-to-video, explicitly state which parts of the starting frame should remain consistent.

How long can generated videos be?

The featured models support different ranges. Current duration choices are shown after you select a model, with the listed endpoints supporting clips up to 15 seconds.

Can I generate video with sound?

Seedance 2.0 can generate synchronized native audio. Other models may accept audio guidance or generate silent video, depending on the selected endpoint.

Can I try the AI video generator for free?

Yes. New RenderFlow AI accounts receive 35 monthly credits on the Free plan, with no credit card required.

How many credits does an AI video cost?

Seedance 2.0: From 40 credits for 4 seconds, based on duration and resolution. Wan 2.7 Text to Video: From 20 credits for 2 seconds, based on duration and resolution. Wan 2.7 Image to Video: From 20 credits for 2 seconds, based on duration and resolution. Kling 3.0 Text to Video: From 34 credits for 3 seconds. Kling 3.0 Image to Video: From 42 credits for 3 seconds.

Do I need provider API keys?

No. RenderFlow AI handles provider access, so you can use the available video models through one account and credit balance.