The largest AI vendors are widening the terrain on which they compete. Microsoft is consolidating chat, coding and persistent agents into one Copilot app. Anthropic is using Claude agents to search biological data and is committing billions to distributed cloud capacity. Google is adding expressive avatars to enterprise agents while testing whether AI chips can operate in orbit. The product surface is expanding, but so are the demands for reliable evaluation, infrastructure and government oversight.
Microsoft turns Copilot into one place for chat, code and agents
Microsoft introduced a redesigned Copilot organized around Home, Code and Autopilot. Home combines Chat and Cowork with Word, Excel and PowerPoint. Code lets users build and run software with technology related to GitHub Copilot. Autopilot is a persistent agent intended to continue work in the background. Home and Code will begin rolling out through Microsoft's Frontier program in the coming weeks, while Autopilot is expanding to private preview at the end of September.
The redesign is an attempt to make Copilot the operating layer for knowledge work rather than a collection of separate assistants. Microsoft still has to prove that the unified interface can preserve permissions, control costs and make long-running actions understandable. The new FinOps controls and shared Teams context show that governance is becoming part of the product, not an administrative feature added later.
Google gives enterprise agents an animated face
Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise. The system pairs real-time speech with an animated persona that can respond through synchronized facial expressions and voice. Google positions it for customer service, training and interactive walkthroughs, extending multimodal agents from a voice or text interface into a visible representative.
An expressive avatar may make an agent easier to engage, but it can also make automated advice feel more authoritative than it is. Enterprises will need clear disclosure, escalation paths and records of what the agent said and did. The difficult question is not whether a face improves interaction. It is whether the experience preserves an accurate sense of the system's limits.
OpenAI opens a mental-health benchmark
OpenAI released MentalHealthBench, an open evaluation developed with more than 80 licensed mental-health experts from 22 countries. It covers adults, teenagers, caregivers and clinicians across non-acute situations, high-acuity concerns and emergencies. Expert-written rubrics assess behaviors such as safety, context seeking, user agency and actionable guidance, with GPT-5.6 Sol used as an automated grader.
The benchmark addresses a high-stakes gap between blocking obviously dangerous replies and providing genuinely useful support. OpenAI says ChatGPT is not a substitute for professional care. Independent replication will still be important because benchmark scores cannot capture every cultural context or the consequences of a real conversation.
Claude agents surface a new biological system
Anthropic announced a life-sciences research group and laboratory after Claude agents identified a previously uncharacterized enzyme system. Roughly 950 agents used 210 million tokens over 21 hours to examine reverse transcriptases, narrowing more than 200,000 examples to 20 candidates. Human scientists then tested the most promising result, an array-associated reverse transcriptase system with DNA repeats reminiscent of CRISPR.
Anthropic does not yet know the system's function, and the company released a preprint rather than a final biological claim. That caution matters. The result is evidence that agent swarms can generate testable hypotheses at scale, not proof that automated science can bypass expert review or laboratory validation.
Google prepares to test AI chips in orbit
Google detailed the next phase of Project Suncatcher, which will send a prototype satellite carrying Tensor Processing Units into low Earth orbit. The test will examine radiation tolerance, thermal management and the practical operation of AI hardware in space. Google is exploring whether abundant solar energy and optical links could eventually support orbital computing infrastructure.
The experiment is early and far from a commercial data center. Still, it illustrates how power and cooling constraints are pushing AI infrastructure research beyond conventional facilities. Any space-based system would also have to solve launch cost, maintenance, networking reliability and orbital-debris risks.
Anthropic signs an $11.6 billion Akamai cloud agreement
Akamai announced a seven-year agreement under which Anthropic committed $11.6 billion for cloud services. The arrangement focuses on CPU workloads across Akamai's distributed infrastructure and includes a warrant that could give Anthropic an equity stake of up to 5%. The contract may expand by as much as $5.4 billion if additional demand materializes.
The deal broadens Anthropic's infrastructure beyond the largest hyperscalers and shows that frontier AI depends on far more than training accelerators. Inference, data movement and supporting services require large fleets of general-purpose compute. The financial commitment also raises execution risk if demand, model economics or hardware efficiency change faster than the contract.
US model review complicates transatlantic safety testing
The White House asked OpenAI and Anthropic to delay sharing new models with British government testers until US officials completed their own review, Politico reported. The request would affect cooperation with the UK AI Security Institute and introduces a new priority rule into international model evaluation.
Early access lets safety institutes study capabilities before release, but governments also view frontier models as strategic assets. If national review takes precedence over trusted cross-border testing, labs may face slower and less comparable assessments. The dispute shows that evaluation standards are becoming part of technology diplomacy.
What matters today
AI competition is moving from individual model releases toward integrated systems. Microsoft wants one interface for work, Google is extending agents into visible and physical infrastructure, and Anthropic is combining scientific agents with a huge cloud commitment. OpenAI's benchmark and the US-UK testing dispute point to the control layer these systems need. The strongest deployments will pair capability with evidence, permission boundaries and infrastructure plans that can survive sustained use.