Veo 3.1
AI Video Generator

Turn a prompt or a photo into an 8-second clip with its own dialogue and sound in 720p, 1080p or 4K, using the Veo 3.1 AI video generator.

The Veo 3.1 AI video generator renders picture and sound in the same pass, so voices, footsteps and room tone land on the right frame, in landscape or vertical.

A woman in a black turtleneck with sunglasses on her head stands in a gilded hotel hallway with a checkered marble floor.

Characters from Reference Images Give Veo 3.1 a face, an outfit and a location, and it places the same person in that setting with the styling intact.

A chimpanzee in a flat cap and denim overalls sits in a flower meadow in front of a circus tent and carousel.

Native Vertical Video Vertical clips are composed for a 9:16 screen from the first frame, ready for Reels, TikTok and Shorts.

A motocross rider in red climbs a hill of glittering purple crystals under floating golden lights.

Physical Motion in Imagined Worlds Spray, dust and suspension react to speed and weight, even when the hill is made of crystals.

A smiling woman in a grey sweater lifts a brown felt hat onto her head against a plain wall.

First and Last Frame in Portrait Upload a start frame and an end frame, and Veo 3.1 fills the moment between them, here a hat going on in one natural gesture.

Key Features of Veo 3.1

Direct speech, sound, framing and transitions with the Veo 3.1 AI video generator. Start from text, a photo, a pair of frames or up to three reference images.

Native Audio and Dialogue

Dialogue, sound effects and ambience are generated with the picture, and lips move with the words. Put each line in quotes and name the speaker, and the right character delivers it.
Create a video with dialogue
A detective in a trench coat and fedora looks up from his desk at a woman standing in his black-and-white office.

Up to Three Reference Images

Add up to three images of a person, a product or a place, and Veo 3.1 keeps their look as the camera moves. A portrait, a white tiger and a skyline become one scene of the emperor walking his tiger.
Create from references
An emperor in teal silk robes and a jeweled crown walks beside a white tiger on a terrace above a futuristic mountain city.

First and Last Frame Control

Set the opening and closing image, and Veo 3.1 builds the camera move between them with matching sound. Start inside a barn, end on a rider in the field, and the shot travels through the doors for you.
Create a frame-to-frame video
A cowboy rides a horse across a golden prairie at sunset after the camera leaves a dark wooden barn.

Camera and Lens Direction

Veo 3.1 reads film language such as dolly, crane, close-up and shallow depth of field. Name the shot, the lens and the light, like a shallow-focus close-up through a rainy bus window, and the frame follows your direction.
Create a cinematic shot
A young woman looks out a rain-streaked bus window at night, city lights blurred behind her reflection.

Veo 3.1 Use Cases

Use the Veo 3.1 AI video generator for ads, social clips, short scenes and music content. Every clip arrives with its own sound, so it drops straight into an edit.

Product Ads and Commercials

Product Ads and Commercials

Add your packshot as a reference and describe the camera path, such as a fly-through from a window to the cans on a kitchen counter. The label stays readable, ready for product pages and paid social.

UGC and Social Clips

UGC and Social Clips

Describe the speakers, the street and what each one says, and you get a vox-pop clip with natural speech and matching lip movement. Switch to 9:16 for Reels, TikTok and Shorts, without booking talent.

Short Films and Drama Scenes

Short Films and Drama Scenes

Write the setting, the mood and a line of dialogue, and Veo 3.1 stages the scene with performance, music and city ambience. The result holds up in pitch reels, storyboards and short-form series.

Music Videos and Live Performance

Music Videos and Live Performance

Use a close-up of the singer and a crowd shot as first and last frames, then write the lyric in quotes. The camera arcs around the stage while she sings the line.

Fashion and Lookbook Films

Fashion and Lookbook Films

Place a model and an outfit on a black-sand beach or any location you describe. Wind, fabric and overcast light behave like an editorial shoot, ready for campaign teasers and lookbooks.

Food and Restaurant Content

Food and Restaurant Content

Ask for a handheld shot of a wok toss with a metallic clank and a sharp whoosh. Flame, steam and sizzle follow the motion, ready for menus, delivery listings and restaurant socials.

Veo 3.1 Prompting Techniques

The Veo 3.1 AI video generator responds to film vocabulary, quoted speech and timed instructions.

Quoted Dialogue

Put spoken lines in quotation marks and tie each one to a character, for example: A woman says, "We have to leave now." Keep lines short enough to fit the clip. English is the language Google has fully tested.

Sound Effects and Ambience

Write each sound as its own cue, such as SFX: thunder cracks in the distance, or Ambient: the quiet hum of a starship bridge. Veo 3.1 places each one against the action.

Five-Part Prompt Formula

Build prompts from five parts: cinematography, subject, action, context, and style and ambiance. Medium shot, a tired office worker, rubbing his temples, a cluttered 1980s office at night, grainy color film.

Timestamped Shots

Split an 8-second prompt into segments such as [00:00-00:02] and [00:02-00:04], each with its own shot and sound. Veo 3.1 moves through them in order within one clip.

Roles for Reference Images

Say what each reference image is for, such as the detective, the woman and the office. Generate the next shot from the same set and faces, clothing and rooms stay consistent.

Length, Resolution and Framing

Choose 4, 6 or 8 seconds, 720p, 1080p or 4K, and 16:9, 9:16 or Auto. Turn sound off when you plan to score the clip yourself.

Made with Veo 3.1

The Veo 3.1 AI video generator follows long, specific prompts, from shot type and lighting to timed actions and spoken lines.

A tired office worker in a suit sits beside a bulky beige computer with a green monochrome screen in a dim 1980s office.

Late night in a 1980s office

Medium shot of an exhausted office worker rubbing his temples beside a bulky 1980s computer in a cluttered office late at night. Harsh fluorescent overhead light and the green glow of a monochrome monitor. Retro look, shot on grainy 1980s color film.

A woman in a veiled hat and dark suit smiles slightly at a man in a fedora in a black-and-white office.

Film noir reply

Using the reference images of the detective, the woman and the office, a shot focused on the woman. A faint, mysterious smile crosses her lips as she replies, "You were highly recommended." Black-and-white film noir.

A small wax figure holds a flame aloft as it walks across a landscape of molten wax beside a dripping candle.

Wax figure in timed beats

A small pale-yellow wax figure holds a flame in a landscape of molten wax, a dripping candle to its left. 0-1s: eye-level tracking shot as it starts to walk, feet rippling the wax. 1-7s: slow, deliberate steps, arm raised to guard the flame. 7-8s: the camera eases back to reveal the wider wax world.

A woman in a dark red trench coat leaps through the air down a green-tiled corridor.

Mid-air leap with a dolly move

The camera dollies around a woman suspended mid-leap in a long, sterile, monochrome green corridor. Her dark trench coat and trousers billow, arms outstretched, her profile set with focus. High-tension cinematic scene.

A chimpanzee in a banana hat and a polar bear in headphones host a podcast in front of jungle and glacier backdrops.

Two-host animal podcast

A monkey and a polar bear host a casual podcast, a tropical jungle behind one and an arctic glacier behind the other. Monkey: "Welcome back to Bananas and Ice! I am Banana." Polar bear: "And I'm Ice!"

A smiling bartender stands behind an iridescent tiled bar as a black cat steps past a cocktail.

Tracking shot into a bar

Slow track along iridescent tiled walls above still water, then a push-in and a turn that reveals a curved bar and a finished cocktail. Soft jazz, clinking ice. "Welcome," a smooth voice says. "Care for a taste?" A black cat jumps onto the bar.

How to generate

Generate a Veo 3.1 clip in three steps

  1. Add an image

    Step 01

    Add an image

    Upload a photo to animate, two frames to connect or up to three reference images.

  2. Describe the scene

    Step 02

    Describe the scene

    Write the shot, the action and the sound. Put dialogue in quotes.

  3. Generate and download

    Step 03

    Generate and download

    Pick 4, 6 or 8 seconds, a resolution and a ratio, then generate and download.

Candles burn across a sprawling miniature city carved from wax at night.

Try the Veo 3.1 AI Video Generator

Start with a sentence, a photo or two frames. Generate a clip with its own dialogue and sound, run as many takes as you need and download the one you want in 720p, 1080p or 4K.

Pricing

Upgrade or cancel anytime. Choose a plan for clean output, fast generation, full model access and priority support.

Monthly

Annual

36% OFF

Veo 3.1 AI Video Generator FAQs

Quick answers about Veo 3.1, plus the plans, credits and billing behind it on ImgVid.

Veo 3.1 is Google DeepMind's video model, released on October 15, 2025. It generates picture and sound together in one pass and builds on Veo 3 with closer prompt adherence and better sound and image quality when you start from a photo. On ImgVid we've set it up for text to video, image to video, first and last frame to video and reference to video, all in the same workspace.

Three things stand out. Sound now works in every mode, including reference images and first and last frames, with richer dialogue and effects. Prompts are followed more closely, and image to video looks and sounds better. A January 2026 update then added native vertical output for reference-based videos, sharper 1080p and a new 4K option.

Each clip runs 4, 6 or 8 seconds, in 720p, 1080p or 4K. Aspect ratios cover 16:9 and 9:16, plus Auto, which lets the model pick the fit. For a longer scene, generate it shot by shot and use the last frame of one clip as the first frame of the next. Our tip: draft at 720p and save 4K for the take you plan to publish.

Yes. Dialogue, sound effects, ambience and music are generated with the picture and timed to what happens on screen. Put spoken lines in quotes and name the speaker, and describe effects and background sound as separate cues. English is the language Google has fully tested, so results in other languages can vary. Prefer a silent clip? Just switch the sound off.

Add up to three images of a person, a character, a product or a location, and Veo 3.1 keeps their appearance in the video. Say in your prompt what each image is, for example which one is the main character and which one is the set. Reuse the same images across shots to keep a character consistent through a sequence. Only upload material you have the rights to use.

Upload a starting image and an ending image, and Veo 3.1 generates the motion between them with matching audio. It suits transitions, reveals and camera moves you want to land exactly, such as a 180-degree arc from a singer's face to the crowd behind her. Describe the move and the sound in your prompt so the model knows how to get from one frame to the other.

Veo 3.1 is included in the Plus and Premium plans, which unlock every model on ImgVid. Free and Starter include Seedance 2.0 Fast and Seedance 2.0 Mini, which are a good way to test ideas before you upgrade.

It depends on length, resolution and sound. With sound on, 720p and 1080p use 16 credits per second, so an 8-second clip costs 128 credits and a 4-second draft costs 64. The same 8 seconds without sound costs 86. 4K with sound uses 37.5 credits per second, 300 for an 8-second clip. The exact number appears in the generator before you confirm, and if a generation fails the credits go straight back to your balance.

Not on the Free plan itself, since Veo 3.1 comes with Plus and Premium. What you can do for free is sign up with no credit card, get a monthly batch of credits and try Seedance 2.0 Mini or image models such as GPT Image 2 and Seedream 5 Lite to see how the workspace feels. Free exports carry a watermark. Any paid plan removes it, and everything you generate on a paid plan is yours to use commercially.

Plus comes with 2,000 credits a month, enough for 15 eight-second Veo 3.1 clips with sound at 720p or 1080p, or 6 in 4K. Premium comes with 6,000, three times as much. Subscription credits stay valid for two months, so an unused balance carries into the next cycle. If you run short mid-month, a credit pack tops you up instantly and never expires.

Yes to both. Cancel from the Membership tab in your account settings and you keep every benefit until the end of the period you've paid for, after which the account returns to the Free plan. Refunds are available on monthly and annual orders, based on how much of the plan you've used, and our team handles requests within 24 hours.