Vivold Consulting
Product & Tech Updates

Introducing Claude Opus 4.8

Claude Opus 4.8 lands with sharper judgment, effort controls, and cheaper fast mode

Key Insights

Anthropic released Claude Opus 4.8, an upgrade to its Opus class with stronger coding, agentic, and knowledge-work performance at the same price ($5/$25 per million input/output tokens). New features include user-controllable effort levels, a Claude Code "dynamic workflows" mode that runs hundreds of parallel subagents, and fast mode now 3x cheaper. Anthropic highlights improved honesty - the model is around 4x less likely than its predecessor to let flaws in its own code pass unremarked.

Stay Updated

Get the latest insights delivered to your inbox

A steady, useful upgrade - plus new dials for developers

Opus 4.8 isn't a reinvention; Anthropic itself calls it a modest-but-tangible step up from Opus 4.7. The gains show up across coding, agentic tasks, and reasoning, and - notably - it arrives at the same price as its predecessor, with a batch of features that matter more in day-to-day use than any single benchmark.

Better judgment, and a real push on honesty

The theme early testers kept hitting was judgment: Opus 4.8 asks better questions, catches its own mistakes, and pushes back when a plan is shaky before charging ahead. Anthropic leaned hard into honesty - a model that flags uncertainty instead of confidently claiming progress it hasn't made. Its evaluations show Opus 4.8 is roughly four times less likely than Opus 4.7 to let flaws in code it wrote slip by unremarked. The alignment team also reported lower rates of misaligned behavior, similar to its best-aligned model.

The features that change how you work

Three launches landed alongside the model:

  • A new effort control in claude.ai and Cowork lets you choose how hard Claude works on a response - think more deeply for quality, or answer faster and burn through rate limits more slowly. It's available on all plans.

  • In Claude Code, dynamic workflows (research preview) lets Claude plan a big job, spin up hundreds of parallel subagents, and verify its own outputs before reporting back - enough to run codebase-scale migrations across hundreds of thousands of lines from kickoff to merge.

  • And fast mode, which runs at 2.5x speed, is now three times cheaper than on previous models.
There's also a quietly useful developer change: the Messages API now accepts system entries inside the messages array, so you can update Claude's instructions mid-task - permissions, token budgets, environment context - without breaking the prompt cache.

What the early adopters are seeing

The testimonials skew technical, but the pattern is consistent: more reliable agentic runs, cleaner tool calls using fewer steps, and stronger performance on specialized benchmarks spanning coding, legal, finance, and computer use. Several testers flagged better citation precision and more token-efficient retrieval on dense documents - the unglamorous stuff that quietly makes production workloads cheaper to run.

Reading the tea leaves

Two forward hints stand out. First, Anthropic says it's working on models that deliver Opus-level capability at lower cost. Second, it teased a new class of model above Opus - Mythos-class - noting a small group was already using Claude Mythos Preview for cybersecurity, with broader release pending stronger safeguards. That tease became real days later with Fable 5 and Mythos 5. For most users, though, the headline is simpler: a better Opus, the same price, with new controls worth turning on.

More in Product & Tech Updates

All Product & Tech Updates stories

Claude gets a 'Reflect' dashboard: Spotify Wrapped for your AI habit - and a masterclass in retention design

Anthropic launched Reflect, a beta dashboard (Free, Pro, and Max users with Memory on) that visualises Claude usage over 1-12 months - top topics, task types, peak hours - and coaches you via its 4D AI Fluency Framework (delegation, description, discernment, diligence), suggesting features like Projects or custom skills based on your patterns. It ships wellbeing controls (quiet hours, break nudges, reflection prompts) built with MIT Media Lab and Boston Children's Hospital experts, and excludes health-linked conversations entirely. TechCrunch's sharp read: beneath the mindfulness framing, Reflect showcases how much of your work runs through Claude - a retention play as much as a wellness one.

Gemini Spark lands on the Mac: Google's 24/7 agent starts working your local files

Gemini Spark, Google's 24/7 agentic assistant, is now available on Mac (beta, US-only, Google AI Ultra subscribers), where it can work directly with files on the computer - sorting and organising them, or turning a folder of invoices into a budgeting worksheet in Google Workspace. The update adds long-requested Google Tasks and Keep integrations plus third-party hooks into Canva, Dropbox, Instacart, OpenTable, and Zillow Rentals, real-time tracking of topics like stocks and breaking news, and - notably for builders - custom MCP support for wiring in your own apps. It puts Spark in direct competition with Claude Desktop, Microsoft Copilot, and OpenClaw for the desktop, where the real productivity (and governance) questions live.

L'Oreal brings Maybelline virtual try-on to ChatGPT

L'Oreal has announced a wide-ranging collaboration with OpenAI, unveiled at VivaTech 2026, that brings Maybelline's virtual makeup try-on directly into ChatGPT via L'Oreal's ModiFace AR technology. The deal spans consumer shopping tools, product discovery for brands like Lancome and Kerastase, advertising pilots (SkinCeuticals, CeraVe, Garnier), and R&D - including using OpenAI's GPT-Rosalind life-sciences model for skin-microbiome research. It lands as OpenAI reports ChatGPT at more than 900 million weekly users.