Vivold Consulting
Product & Tech Updates

Introducing Gemini Omni

Gemini Omni is Google's any-modality-in, any-modality-out model, starting with video

Key Insights

Google announced Gemini Omni, a new model family capable of generating output in any modality from any input - combining Gemini's intelligence with Google's generative media models in what it frames as a leap in world understanding. The first release, Gemini Omni Flash, starts with video outputs (with image and text to follow) and is available today in the Gemini app, Google Flow, and YouTube Shorts. API access for developers and enterprises follows in the coming weeks.

Stay Updated

Get the latest insights delivered to your inbox

From predicting text to simulating reality

Google introduced Gemini Omni, a model designed to generate samples in any output modality from any input - part of a broader shift the company describes as AI moving from predicting text to simulating reality through world models.

What it does

  • Omni combines Gemini's reasoning with Google's generative media models, which Google frames as a significant step forward in world understanding.
  • The first model in the family, Gemini Omni Flash, starts with video outputs, with image and text generation to be enabled over time.
  • It's available starting now in the Gemini app, Google Flow, and YouTube Shorts, with rollout to developers and enterprise customers via APIs in the coming weeks.

The bigger picture

Omni sits alongside Google's other world-simulation work shown at I/O - including Project Genie, which generates explorable real-world places - and reflects a strategic bet that unifying intelligence with generative media is the next frontier. Combined with the breakout success of Google's Nano Banana image models (more than 50 billion images generated to date), it underscores how central generative media has become to Google's roadmap. The natural question Omni raises, as with any high-quality generative video, is provenance - which is why Google paired its I/O media news with an expansion of SynthID watermarking and Content Credentials, now joined by partners including OpenAI, Kakao, and ElevenLabs.

More in Product & Tech Updates

All Product & Tech Updates stories

Claude gets a 'Reflect' dashboard: Spotify Wrapped for your AI habit - and a masterclass in retention design

Anthropic launched Reflect, a beta dashboard (Free, Pro, and Max users with Memory on) that visualises Claude usage over 1-12 months - top topics, task types, peak hours - and coaches you via its 4D AI Fluency Framework (delegation, description, discernment, diligence), suggesting features like Projects or custom skills based on your patterns. It ships wellbeing controls (quiet hours, break nudges, reflection prompts) built with MIT Media Lab and Boston Children's Hospital experts, and excludes health-linked conversations entirely. TechCrunch's sharp read: beneath the mindfulness framing, Reflect showcases how much of your work runs through Claude - a retention play as much as a wellness one.

Gemini Spark lands on the Mac: Google's 24/7 agent starts working your local files

Gemini Spark, Google's 24/7 agentic assistant, is now available on Mac (beta, US-only, Google AI Ultra subscribers), where it can work directly with files on the computer - sorting and organising them, or turning a folder of invoices into a budgeting worksheet in Google Workspace. The update adds long-requested Google Tasks and Keep integrations plus third-party hooks into Canva, Dropbox, Instacart, OpenTable, and Zillow Rentals, real-time tracking of topics like stocks and breaking news, and - notably for builders - custom MCP support for wiring in your own apps. It puts Spark in direct competition with Claude Desktop, Microsoft Copilot, and OpenClaw for the desktop, where the real productivity (and governance) questions live.

L'Oreal brings Maybelline virtual try-on to ChatGPT

L'Oreal has announced a wide-ranging collaboration with OpenAI, unveiled at VivaTech 2026, that brings Maybelline's virtual makeup try-on directly into ChatGPT via L'Oreal's ModiFace AR technology. The deal spans consumer shopping tools, product discovery for brands like Lancome and Kerastase, advertising pilots (SkinCeuticals, CeraVe, Garnier), and R&D - including using OpenAI's GPT-Rosalind life-sciences model for skin-microbiome research. It lands as OpenAI reports ChatGPT at more than 900 million weekly users.