Google launches Veo 3.1 to compete with Sora 2 with AI videos with sound
Summarize this article with:

The artificial intelligence war is reaching a peak. With each announcement, a new model emerges, bolder, more immersive, more… expensive. In this battle of innovations, Google did not want to remain a spectator. By releasing Veo 3.1, it unveils a video AI armed with sounds, dialogues, and new editing capabilities. Faced with the viral popularity of Sora 2, the Mountain View firm is playing another card: that of narrative precision and creative control.

Two androids symbolizing Google and OpenAI face off in a futuristic city, separated by an orange energy sphere.

In brief

  • Veo 3.1 integrates audio, dialogue and sound effects to enrich AI-generated scenes.
  • The tool targets serious creators, with professional editing options and formats.
  • Three key modules: image composition, creative transitions and smooth clip expansion.
  • Google's AI values ​​visual consistency, sometimes to the detriment of speed of action.

Technological duel: Google takes on the queens of AI video

When OpenAI, valued at $500 billion without an IPO, launched Sora 2 on September 30, it was an immediate success. The app was downloaded over a million times in just five days, climbing to the top of the App Store. His approach? A “TikTokized” interface, designed for sharing and remixing.

Google did not choose this path. With Veo 3.1the objective is clear: to address creators, not influencers. The model allows you to generate videos with 1080p resolution, in horizontal or vertical format, integrating sound ambiance, synchronized voices and realistic effects. Accessible via Flow, Vertex AI and Gemini API, it offers two plans: a fast version at $0.15/second, and a standard version at $0.40/second.

The firm emphasizes audio capabilities, now present in all modules. It promises a unique result: the lip synchronization of Veo 3.1 exceeds that of all other models.

Where Sora favors visual dynamism, Veo chooses coherence. Movements are slower, but the elements remain stable. This is the price of precision. A positioning that contrasts with the ambitions of Meta or Luma Labs, which focus more on speed and the wow effect.

Stories that speak: Google's AI wants to tell

One of the major bets of Veo 3.1 is narrative immersion. The addition of sound allows Google to take a step forward: no longer just illustrate, but tell stories with images and voices. Three features stand out:

  • Ingredients to Video: you combine several reference images, and the AI ​​generates a scene with objects and characters;
  • Frames to Video: you give a starting and ending frame, and the AI ​​produces a consistent transition;
  • Extend: the AI ​​extends a clip by generating the continuation from the last second.

The tool also allows you to add or remove elements, taking into account light and shadow. This level of detail is the strength of the approach: a film studio in an artificial intelligence interface.

Start your crypto adventure safely with Coinhouse
This link uses an affiliate program

But not everything is perfect. When instructions stray too far from visual logic, AI goes off the rails. Some scenes jump from one shot to another, lose characters or completely change the mood. This remains a technology under construction.

As Google explained in its official blog:

We're also introducing Veo 3.1, which brings richer sound, better narrative control and increased realism capturing textures close to reality.

Veo 3.1 does not want to entertain: it wants to move. And this is undoubtedly where it differs radically from its competitors.

Demanding UX, stunning result: when artificial intelligence becomes a creative tool

The user experience offered by Veo 3.1 is not that of a social network. It is not a product to consume, but a tool to master. Creators must learn to speak the language of AI. A poorly written prompt or one that is too far from the reference images can produce an inconsistent result.

Some tips are already circulating among users. For example, using Seedream to generate a faithful initial image, before importing it into Veo. Or use an audio-aware construction, explicitly mentioning the desired sounds in the prompts.

In this regard, here are some concrete facts:

  • Veo has generated over 275 million videos since the launch of Flow;
  • Three creative modules are available: Ingredients, Frames, Extend;
  • The usage cost is up to 2 times lower than that of Sora 2 Pro;
  • Videos can be up to a minute long, with built-in sound;
  • Only three models support spoken voices: Sora, Grok, and now Veo.

The tool is not easy to tame. But once understood, it delivers videos of rare realism, with accurate intonations and credible characters. You just need patience, skill… and a few credits.

Google no longer hides its ambition to dominate generative AI. Veo 3.1 shows that the firm does not just want to follow. She wants to impose her tempo. And as if to confirm this thirst for prowess, one of its robots has just solved a mathematical problem deemed impossible. The message is clear: the AI ​​giant is only beginning to speak.

Maximize your Tremplin.io experience with our 'Read to Earn' program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.

Similar Posts