MakeAIVideo

How to Use Veo 3 in 2026 (Beginner's Guide)

How to use Google Veo 3 in 2026: how to access it, how to write your first prompt with audio cues, and where beginners should actually start.

Jamie Partridge
Jamie Partridge
Founder··20 min read
Veo 3 2026 beginner-guide hero, dark indigo gradient with a stylised prompt-plus-sound-wave glyph

How to use Veo 3 in 2026 starts with the one detail that separates it from every other AI video model on the market: the audio comes out of the same generation as the picture. Prompt Veo 3 for "a barista tamping espresso, morning light, close-up," and it returns eight seconds of footage with the tamp click, the machine hum, and the room tone rendered in-model. Rival foundation models still ask you to add sound in a second tool. This beginner guide walks through what Veo 3 is, which surface to sign in through, how to write your first prompt with audio cues in it, what you get back, and where a pipeline like the MakeAIVideo route fits when your deliverable is a finished narrated video rather than eight raw seconds. If you already have a script and want the whole piece assembled around a Veo shot, skip to our narrated-video route at the bottom.

The 30-second call. Veo 3 is the pick when you need one photoreal eight-second clip with the sound baked in. It is not a finished-video tool. Beginners typically get the most value by using Veo 3 for a hero beat and a pipeline like MakeAIVideo for the surrounding script, voice, captions, and export.

What Veo 3 actually is

Veo 3 is Google DeepMind's third-generation text-to-video model, announced at Google I/O on May 20, 2025 and rolled out through the Gemini app, Google Flow (the filmmaking tool that replaced VideoFX), Google Vids inside Workspace, and Vertex AI for developers. Give it a written prompt (or a prompt plus a reference image), and it returns up to eight seconds of video per generation at 1080p or 4K, with synchronised audio produced in the same pass.

The version history matters because "Veo 3" is not the newest label you will see in a Google product menu. Veo 1 landed in May 2024 as a research preview and rendered up to 1080p. Veo 2 shipped in December 2024, added 4K, and pushed physics realism forward. Veo 3 arrived in May 2025 with the native audio pass that became the model's signature feature. Veo 3.1, released on October 15, 2025, is the current stable variant and the version you interact with when you generate today. Everything below applies to Veo 3.1 unless noted, and beginners can safely use the phrase "Veo 3" for the whole family.

Veo is not a full video product; it is a foundation model. There is no timeline, no scene manager, no caption pass, and no cross-clip character memory. It renders one clip, with one baked-in audio track, per generation. For anything longer than eight seconds, you stitch multiple clips or wrap Veo in an editor or pipeline. Our full Veo 3 review covers the strengths and limits in more depth than this walkthrough needs.

If your first instinct is "give it a still image and animate it," Veo does support that path. So does a pipeline layer, and the animate a still route inside MakeAIVideo covers the finished-video shape of that same job.

Who Veo 3 is for, and who it isn't

Setting expectations before you invest a subscription cost or hours of iteration:

Veo 3 suits beginners who want the highest-fidelity short cinematic clip available today. A product shot with realistic ambient sound. A single line of dialogue delivered in-frame by a character. A moody atmospheric beat where room tone carries the vibe. On any of those, Veo 3 currently outputs the most photoreal short-clip result of any mainstream model, and the audio pass saves an entire sound-design step.

Veo 3 does not suit beginners shipping finished long-form videos. Every generation caps at eight seconds. A 60-second YouTube Short, a two-minute explainer, or a scripted testimonial needs multiple Veo takes plus voiceover, captions, music, and assembly. Character continuity between separate Veo generations is imperfect, which is why longer projects usually reach for a pipeline layer instead. For that shape of job, the finished-video comparison covers the workflow tools that sit on top.

Veo 3 suits creators who want to iterate quickly on the same beat. Google's own filmmaking tool (Flow) is built around Veo and supports rapid iteration with visual prompt refinements, image-anchored takes, and side-by-side generation review. Beginners who like a photoshoot-style workflow (shoot, review, adjust, shoot again) find Flow more comfortable than the raw API.

Veo 3 does not suit creators optimising strictly for cost per clip. Vertex AI billing runs around $0.75 per second of generated video when audio is included. An eight-second clip with sound costs roughly $6 in compute alone, before any wrap-up work. If cost per finished clip is the binding constraint, our Sora alternatives roundup covers the value picks (Kling, Pika) that trade some quality for a fraction of the per-second rate.

If you sit inside the "one striking clip per beat" profile, work through the rest of this guide. If you sit inside the "cadence of finished videos" profile, jump straight to the pipeline section further down.

How to access Veo 3 in 2026

Four public surfaces expose Veo 3 today, each aimed at a different type of user.

Gemini app (consumer path). The Gemini web app and mobile apps include Veo 3 for subscribers on the Google AI Pro or Google AI Ultra plans. Type a prompt into the Gemini chat, hit the video button, and the same subscription that powers your Gemini Advanced text sessions covers the generation. This is the cheapest way for a beginner to try Veo 3 seriously; the Ultra tier is where most repeat users end up because the Pro tier caps monthly video quotas quickly.

Google Flow (creator path). Flow is the DeepMind filmmaking tool that replaced VideoFX, purpose-built around Veo. It supports image-anchored takes, prompt refinement panels, and side-by-side clip review. Access is included in the same Google AI Pro and Ultra plans; the Ultra tier unlocks higher per-project generation counts. If you plan to iterate on a single shot more than three or four times, Flow is a noticeably better working environment than the Gemini chat interface.

Google Vids (productivity path). Google Vids is the Workspace video-editor product with Veo baked in as a menu-level scene generator. If your organisation already pays for Workspace and you want to add a Veo shot to a slide-deck-style narrated video, Vids is the shortest path and no separate signup is needed.

Vertex AI (developer path). Vertex AI charges per-second of generated video, currently around $0.50 per second on silent Veo 3 output and $0.75 per second when audio is included. A typical eight-second clip with sound costs roughly $6 in raw compute. There is a cheaper Veo 3.1 Lite variant for high-volume silent workflows. This is the surface to use if you are wiring Veo into an application or automating renders at scale, and it is the only path that removes the monthly quota cap entirely.

Some YouTube Shorts creators also see Veo inside the Shorts creator flow as a "Dream Screen" background generator. If you already publish Shorts, check that first before signing up for anything else.

For the shape of a workflow subscription rather than a raw-render subscription, you can compare MakeAIVideo plans against a Google AI Ultra plan on the same-dollars basis. Different products, different problems.

Signing up and your first login walkthrough

The cleanest beginner path is Gemini plus a Google AI Pro subscription. Three steps:

  1. Sign into your Google account at gemini.google.com. If you already use Gmail, Drive, or YouTube, the same account works.
  2. Subscribe to Google AI Pro (or AI Ultra). From the Gemini sidebar, choose Upgrade and pick a plan. Pro is the entry tier and includes Veo 3 with a monthly video quota. Ultra sits above it with higher quotas and priority access to newer Veo variants; that is where serious users end up.
  3. Open the video button in a Gemini chat. From any Gemini chat, type a prompt and select the video icon in the prompt bar. Your first generation kicks off; the returned clip appears in the chat with a Download button.

That is the entire signup surface for a beginner. No API key, no billing dashboard, no rate-limit tuning. If you then want to iterate more comfortably on the same shot, open Google Flow from the same Google AI account; your subscription carries across.

For the developer path, the flow is longer: enable the Vertex AI API in a Google Cloud project, add a billing account, generate an API key or set up application-default credentials, then call the Veo generation endpoint from a script. If you have never touched Google Cloud before, budget an afternoon for the account plumbing before your first render. For the beginner-friendly shape of the same pipeline, the shorts route inside MakeAIVideo produces the finished vertical video without any of the API scaffolding.

Your first Veo 3 prompt, step by step

Walk through one concrete example: a product hero shot of a coffee cup, this time with audio cues baked into the prompt. This is where Veo differs sharply from Sora or Kling as a first-time experience.

Weak prompt (what beginners write):

A cup of coffee on a table.

Veo will return an eight-second silent-feeling clip with generic ambient noise. Nothing about the shot reads as intentional.

Strong prompt with audio cues (the same generation with structure):

A macro dolly-in shot of a ceramic latte cup on a dark walnut table, steam rising from the surface, warm golden-hour light from a window on the left, shallow depth of field, 35mm lens. Ambient sound: distant café chatter, the soft hiss of the espresso machine, a spoon clinking on porcelain. Cinematic colour grade, 8 seconds.

The strong prompt names the subject, the camera movement, the specific material, the lighting direction and quality, the lens focal length, and (critically) the sound design that should render alongside the visuals. Every experienced Veo user writes prompts in two blocks: the visual block and the audio block. Skipping the audio block leaves the model to guess, and Veo's guesses on ambient sound are competent but generic.

Six-part visual frame most experienced users adopt: subject, action, camera, lighting, style, duration. Then a separate audio frame naming ambient sound, sound effects, and dialogue if any. Thirty templates already written in that two-block shape will save you the first dozen wasted generations. Missing one or two parts is fine; missing four means the model has to guess. Our free script writer tool covers the pre-generation writing step for anyone drafting the narration before hitting a video model.

Two practical points on your first Veo prompt. First, dialogue in-frame is Veo's most novel capability. Ask for "a barista saying: 'we open at seven,' in a warm Melbourne accent," and Veo will render lip-synced speech from the character inside the clip. Rival models need a Wav2Lip pass or an ElevenLabs voiceover layered afterwards; Veo does it inside the model. Second, camera direction language matters more than colour language. "35mm anamorphic dolly-in" reliably changed the output; "warm colour" reliably did nothing on its own.

Do not skip the audio block on your first prompt. The single fastest way to waste a Veo 3 generation is to prompt only the visuals and let the model guess the sound. Name the ambient sound, the sound effects, and any dialogue explicitly. The audio pass is Veo's whole point; using it as a silent video generator wastes what you are paying for.

Understanding the output you get back

Every Veo 3 generation returns an MP4 file with these properties:

  • Duration: up to 8 seconds per generation. No longer clips available in a single call.
  • Resolution: 1080p or 4K, selected at generation time.
  • Frame rate: 24 fps by default.
  • Aspect ratio: 16:9 landscape, 9:16 vertical, or 1:1 square, set at generation time (Veo does not reframe after the fact).
  • Audio: baked into the MP4 track. Cannot be requested separately from the video pass on Veo 3.
  • Watermark: an invisible SynthID watermark, detectable by Google's classifier but not visible in the frame.

The SynthID watermark is the detail that most surprises first-time users who compared Veo to Sora during the app's active window. Sora burned a visible moving watermark into every generation; Veo does not. That makes Veo output usable for client and commercial work without cropping or licensing tricks.

Vertex AI users should still confirm the specific commercial licence attached to their contract, and note that a model licence is not an advertising clearance: if the finished piece carries a testimonial or endorsement, the FTC's Consumer Reviews and Testimonials Rule applies on top, and its guidance speaks directly to AI-generated avatars used in marketing. Google DeepMind's SynthID watermarking documentation describes the mark as embedded "directly into the pixels of every video frame, making it imperceptible to the human eye, but detectable for identification."

Two more things worth flagging on output review. Physics realism is best on plausible scenes and weakest on deliberately surreal ones. And multi-shot continuity between separate Veo generations is imperfect; a character rendered in one eight-second clip will not look identical to the same character in the next generation without careful reference-image anchoring. For vertical short-form output specifically, the vertical shorts pipeline handles the assembly and caption pass that Veo itself does not.

Veo 3-specific features to actually use

Three capabilities separate Veo 3 from every other mainstream video model. Beginners who ignore them are paying a premium for a silent-clip generator that does nothing special.

Native audio generation. Dialogue, sound effects, and ambient sound in the same generation. This is the marquee feature. Use it on any prompt where sound would sell the shot: product shots (ambient room tone), dialogue clips (in-frame character lines), and mood-driven pieces (rain on a window, footsteps on gravel). TechCrunch's coverage of the Veo 3 launch framed the audio pass as the most consequential capability in the model.

Image-to-video from a reference frame. Upload a still and ask Veo to animate forward from that frame. This is the closest Veo path to character consistency across multiple takes: use the same reference image for each generation, and the resulting clips share a consistent starting character. Not identical (the model still varies), but close enough to cut together in a two-or-three-clip sequence. For the finished-video equivalent of that same job, the prompt-only route inside a pipeline handles the surrounding assembly.

Prompt-driven camera control. Veo interprets camera direction language literally: "35mm anamorphic dolly-in," "orbit shot around the subject," "whip pan to the left," "low-angle wide shot." Rival models need more prompt iteration to hit the same camera intent; Veo often gets it first try. That predictability is worth as much to a beginner as raw quality is, because iteration budget on the Ultra tier (or the Vertex bill) burns down faster than beginners expect.

Google Flow adds a fourth capability worth naming even though it sits outside the model itself: rapid iteration UI. Side-by-side clip review, image-anchored takes, and prompt refinement panels are all built into the Flow surface. If you plan to iterate more than three times on a single beat, Flow is where to do it.

Beyond the raw clip: what to do with a Veo 3 output

A Veo clip is a raw asset, not a finished video. Everything that turns it into something you can publish sits downstream:

Longer runtime. Eight seconds is a hook, a beat, or a single shot. It is not a full YouTube Short, Reel, or explainer. Stitching multiple Veo takes into a longer sequence loses character continuity between generations.

Captions. Short-form platforms weight caption presence heavily. Veo does not caption the dialogue it renders; the frame stays clean, and burned-in captions are a separate downstream step. Reels, TikTok, and Shorts audiences skim on mute during silent autoplay, and non-caption video reads as effort-light to the platform algorithms.

Music beds. Veo renders ambient sound and dialogue well; scored music is possible but less predictable than the sound-design side. A dedicated music bed under the whole video is usually a separate track.

Scene assembly. A finished 60-second video is typically 6 to 12 short clips cut together on a timeline. Veo returns single clips. You need an editor or a pipeline to stitch them.

Multi-ratio export. Veo generates in the aspect ratio you pick; it does not reframe. If you need 9:16 for Reels and 16:9 for YouTube long-form from the same source, you generate twice or crop downstream.

The two shapes of downstream tool: a general video editor (DaVinci Resolve, Premiere, Google Vids) if you want to do this by hand, or a pipeline (like MakeAIVideo) if you want the assembly automated with script, voice, captions, and music built in. For the specific longer-explainer format that dominates B2B landing pages, the explainer pipeline covers the 60-second to 3-minute narrated shape.

Where MakeAIVideo fits after the clip

MakeAIVideo is not a Veo 3 competitor and never was. Veo 3 is a foundation model that returns raw clips (with baked audio). MakeAIVideo is the pipeline that turns clips, prompts, or scripts into a finished narrated MP4 with voiceover, captions, music, and platform-native export in multiple aspect ratios.

The split by production surface:

  • The prompt-first surface covers typing a prompt and receiving a finished narrated video, with the underlying model choice handled behind the scenes.
  • The script-first flow covers pasting an existing script and getting a matched video with scenes, voice, and captions rendered per beat, ideal for wrapping a single Veo hero shot inside a longer piece.
  • The anonymous channel workflow covers the faceless YouTube format for creators who never wanted to be on camera in the first place, using a pipeline for the whole show.
  • The short-form workflow for TikTok outputs vertical TikTok-native videos with auto-burned captions, voiceover, and platform-native ratio in one render.

Every surface uses the same $9 per month Lite starter tier, with a 7-day free trial, $0 today, cancel anytime. That collapses the per-second Vertex bill into a flat monthly cost for the wrap-around workflow, and it wraps Veo (or a successor model) inside a finished-video pipeline rather than leaving you with eight-second raw clips.

The specific handoff pattern that keeps working in 2026: use Veo 3 to render the hero shots that need photoreal quality and in-frame dialogue. Use a pipeline for the surrounding script, voiceover on any narration Veo did not handle, captioning the whole thing for silent autoplay, and exporting the aspect ratios each platform actually wants. Veo gives you five to ten seconds of "this is the moment." The pipeline gives you the surrounding two minutes that turn a moment into a story.

Pipeline plus Veo, not either or. Beginners who treat Veo 3 as their whole video toolkit end up shipping eight-second clips that need work everywhere else. Beginners who treat a pipeline as their whole toolkit miss out on Veo's audio-plus-video advantage on the hero beat. The creators shipping the most in 2026 run both, and the pipeline trial is 7 days at $0 today, cancel anytime.

Five beginner mistakes that waste Veo 3 credits

1. Skipping the audio block in the prompt. Veo's whole differentiator is native audio. A prompt that names only visuals leaves the model guessing at the sound, and its guesses are competent but generic. Always name ambient sound, sound effects, and dialogue explicitly.

2. Iterating past four tries on the same prompt. Diminishing returns kick in fast. If a prompt is not working by the fourth generation, the prompt structure itself is broken; a fifth pass will not fix it. Rewrite the whole thing, do not tune a single word.

3. Prompting for clips longer than eight seconds. The model refuses or truncates. Veo's cap is a hard cap, and longer output means multiple generations stitched together. Beginners waste an early week trying to force one long generation instead of accepting the eight-second reality.

4. Generating in the wrong aspect ratio and cropping downstream. Aspect ratio is set at generation time and Veo does not reframe after the fact. Cropping a 16:9 clip down to 9:16 costs the sides of the frame. If your destination is TikTok, Reels, or Shorts, generate in 9:16 from the first take.

5. Treating the eight-second Veo clip as a finished video. A raw Veo output is one input to a longer pipeline. Voiceover on any surrounding narration, captions, music bed, brand outro, and multi-ratio export are all separate jobs. Beginners who publish raw Veo clips to short-form platforms get low retention because the clip lacks the typographic and pacing cues short-form audiences expect. For the full retrospective on which tools handle each stage, the finished-video comparison covers the pipeline options in more depth.

Read the pattern above, ship the first attempt, and revise once. Iterating on individual generations is a distraction; iterating on the workflow shape is where the real improvement lives.

How to use Veo 3 FAQs

Is Veo 3 free to use?

Not directly. Consumer access sits inside Google AI Pro and Google AI Ultra subscriptions; both are paid plans that include Gemini Advanced alongside a Veo monthly quota. Vertex AI charges per-second for developers. Some Google Workspace and YouTube Shorts users see limited Veo access inside existing subscriptions without a dedicated video plan.

How long can Veo 3 videos be?

Each Veo 3 generation caps at eight seconds. Longer sequences require multiple generations stitched together in Google Flow, Google Vids, or a downstream pipeline. Character continuity between separate Veo generations is imperfect, which is why longer projects usually benefit from an assembly layer on top.

Can Veo 3 make audio?

Yes, that is Veo 3's marquee capability. Every generation renders synchronised audio (dialogue, ambient sound, sound effects) in the same pass as the video. It is the only mainstream foundation model that does so as of 2026; Sora, Runway, Kling, Luma, and Pika all render silent clips that need a separate audio pipeline.

How do I access Veo 3 in 2026?

Four main surfaces. The Gemini app for consumers (paid AI plan required). Google Flow for creators (same subscription). Google Vids for Workspace users (bundled inside the office suite). Vertex AI for developers (billed per-second). Some YouTube creators also see Veo inside the Shorts creator flow as a background-generation feature.

What is the current version of Veo?

Veo 3.1, released on October 15, 2025, is the current stable release. It is the version you interact with today when you generate through Gemini, Flow, Vids, or Vertex AI. Beginners can safely refer to the whole family as "Veo 3" without loss of accuracy.

What resolutions does Veo 3 support?

Veo 3 outputs 1080p and 4K at the point of generation. Aspect ratios are set at generation time rather than in post, so if you need 9:16 vertical output for Reels, Shorts, or TikTok, generate in that ratio from the first take rather than cropping downstream.

Is Veo 3 better than Sora?

On raw physics realism and on the audio pass, Veo 3 is currently ahead of any surviving Sora tier. Sora 2 pushes further on clip length (up to 25 seconds on the API) and on stylised looks. Sora is also being wound down as a product, with the API sunsetting on September 24, 2026; our full Sora review covers the shutdown timeline.

How does Veo 3 compare to Runway?

Different products, different jobs. Our Runway review covers the model in depth. Runway wins on editor UX, multi-model access in one seat (Gen-4.5, Aleph, Veo 3.1, Nano Banana Pro), and timeline editing. Veo 3 wins on raw clip quality and native audio. Working editorial teams often prefer Runway; solo creators shipping single-clip hero shots often prefer Veo.

Can I use Veo 3 for commercial work?

Yes on the paid tiers, with the standard usage terms Google publishes for each channel. Enterprise buyers should confirm the specific licence attached to their Vertex AI or Workspace agreement. Every Veo output carries an invisible SynthID watermark, which does not affect commercial use.

Does MakeAIVideo replace Veo 3?

No. Veo 3 is a foundation model that generates raw eight-second clips (with audio). MakeAIVideo is the pipeline that turns any model's clips, prompts, or scripts into a finished narrated MP4 with voice, captions, music, and platform-native export. Many creators use both. For value-tier successor picks (Kling, Pika) at a fraction of Veo's per-second Vertex cost, the Kling roundup is the sibling piece.

Where to start now

If you have never touched Veo 3, sign up for Google AI Pro through the Gemini app and run three test prompts with explicit audio cues in each one. That gets you a real feel for the model within an hour and under the cost of a single Vertex API render. If you are not sure Veo is the right engine at all, the twelve-model field comparison is the cheaper way to rule it in or out. If you already pay for a Google AI plan, open Google Flow and iterate on a single beat until you have a keeper worth wrapping into a longer piece.

For the surrounding piece itself, the shortest honest path is a workflow pipeline that handles the script, voice, captions, music, and multi-ratio export around a Veo hero shot. MakeAIVideo starts at $9 per month with a 7-day free trial, $0 today, cancel anytime in the trial window for no charge. Pick the model for the hero beat, pick the pipeline for the surrounding piece, and ship your first finished video with a Veo clip inside it this week. If your project is closer to a single prompt than to a full script, the prompt-first surface is the fastest way in.

Jamie Partridge

Written by Jamie Partridge

Founder at MakeAIVideo. Writing about AI video generation, scripting, scenes, and shipping content faster.

More about the author

Free AI video tools

More AI video guides

Comparisons

Best AI Video Generators in 2026 (12 Tested)

12 AI video generators tested in 2026: best for multi-scene, raw clips, avatars, repurposing, suite integration. Real pricing. Honest comparison.

21 min read
Comparisons

Best AI Video Maker in 2026 (10 Tested)

10 AI video makers tested on the same brief: best for prompt-to-finished-video, script-to-MP4, talking avatars, and short-form social. Real sourced pricing.

20 min read
Comparisons

Best AI Marketing Tools in 2026 (10 Tested)

10 AI marketing tools tested across video, copy, SEO, email, and ad creative. Real sourced pricing, honest verdicts, MakeAIVideo at #1 for video.

20 min read
Guides

How to Use Runway AI: A Beginner's Guide (2026)

How to use Runway AI from first signup to first Gen-4.5 clip. Pricing, credits, prompts, motion brush, Act-One and where Runway fits in a video workflow.

17 min read

Ship AI video in under 90 seconds

Prompt or script in, finished narrated MP4 out. Plans from $9/month. Plans start at $9/month with a 7-day free trial. $0 today, cancel anytime.