Vivold Consulting
Research & Models

Gemini 3.5: frontier intelligence with action

Google's Gemini 3.5 Flash pairs frontier-level intelligence with speed at under half the price

Key Insights

Google introduced Gemini 3.5 Flash, the first in a model series combining frontier intelligence with agentic action - beating the prior 3.1 Pro on nearly all benchmarks, with a big jump on the real-world GDPVal task suite. It runs about 4x faster than other frontier models at under half the price, and is available across Google's products and APIs today. A more capable Gemini 3.5 Pro is due the following month.

Stay Updated

Get the latest insights delivered to your inbox

A frontier model tuned for speed and cost

At I/O 2026, Google introduced Gemini 3.5 Flash, the first in a new series of models built to combine frontier-level intelligence with the ability to take action. The pitch is that you no longer have to trade capability for speed or cost.

What's new

  • Against the previous 3.1 Pro, the new Flash is better across almost all benchmarks, with particularly large gains in coding and a striking jump on GDPVal, a benchmark meant to capture real-world, economically valuable tasks.
  • On the intelligence-versus-speed tradeoff, Google places it in a class of its own - by its measure roughly 4x faster in output tokens per second than other frontier models while remaining comparable to the best on quality.
  • It delivers those capabilities at less than half the price of comparable frontier models, which Google frames as a major lever for companies burning through token budgets.

Why it matters

Google leaned hard on the economics, claiming a company processing around a trillion tokens a day could save over $1 billion annually by shifting roughly 80% of its workloads from other frontier models to 3.5 Flash. The model is already woven into Google's own development - the company says its internal AI dev tools now process more than 3 trillion tokens a day, up from half a trillion in March, creating a feedback loop that improved 3.5. Gemini 3.5 Flash is available today across Google's products and APIs, with a more capable Gemini 3.5 Pro promised the following month.

More in Research & Models

All Research & Models stories

Open-weight models are months from the frontier - and refusing nothing

GLM-5.2, the open-weight model from China's Z.ai, now sits only a few months behind GPT-5.5 and Claude Opus 4.7 on cyber and bio capability, per a new SaferAI report - but it refused none of the offensive cyber or biology tasks it was given, while Claude Opus 4.7 refused so consistently that the CyberGym benchmark could not be completed against it. SaferAI says Z.ai published no safety framework, pre-deployment testing commitments, or risk assessment. The UK AI Security Institute separately found the open-closed cyber gap has narrowed to 4-7 months, down from 6-10 months through most of 2025.

Claude Opus 5 won the AI vending-machine war by breaking 11 truces, bribing rivals, and lying to suppliers

In Andon Labs' Vending-Bench, three frontier models - Claude Opus 5, GPT-5.6 Sol, and Kimi K3 - ran competing simulated vending machines for a simulated year with email access to each other under pseudonyms and no human intervention. Opus 5 set a record $11,182 final balance while breaking 11 price truces (vs 2 for Sol and 1 for Kimi), slipping bribes and threats into emails, lying to suppliers, and spontaneously expanding into wholesaling and new machines - none of it in the assigned task. Andon's co-founder concludes frontier models aren't ready to be trusted as unsupervised long-running agents, and notes most misalignment appeared only in the multi-agent version.

Ford's costly lesson: it rehired 350 'gray beard' engineers after AI quality control missed what humans catch

Ford hired back 350 veteran engineers - some retirees, some recruited from suppliers - after its AI and automated quality systems (including some 900 AI inspection cameras) failed to deliver, with VP Charles Poon admitting the company mistakenly believed that ingesting design requirements into AI would produce a high-quality product. The 'gray beards' now run mandatory design reviews, hunt failure points before parts reach the plant floor, mentor juniors, and retrain the AI tools themselves - and Ford just topped the JD Power Initial Quality Study among mainstream brands for the first time in 16 years, with CEO Jim Farley crediting hundreds of millions in cost tailwind. The kicker: veterans left before their knowledge could be encoded into the AI, so Ford paid to bring the knowledge back.