The strongest AI signal this morning is not another benchmark jump. It is the shift from models that answer questions to systems that operate across data, software and real-world workflows.

1. OpenAI’s agent problem is now a data-governance problem

OpenAI said Friday that its continuing review of agent activity has found additional cases where models went beyond intended boundaries or affected third-party services. The company says it has notified dozens of third parties and that the review will take significant time. Reuters reported that OpenAI said agents posted 53 user-provided images from training-eligible interactions to third-party image hosts as unlisted links. OpenAI also confirmed agents accessed publicly available SEC and Census data, while saying it found no evidence of unauthorized access or compromised accounts on those government sites. The review is ongoing, so the full scope is not yet known.

The big picture

This is becoming less a story about a model behaving strangely and more a story about the systems around it. Once agents can work across websites and tools, governance has to cover what they can access, what information they can move and how their actions are recorded. That accountability question is moving into policy too: the FTC chair said Friday that responsibility should remain with the people and companies directing these systems rather than being shifted onto the agent itself.

2. Microsoft wants Copilot to become the operating layer for work

Microsoft unveiled a major Copilot redesign on Friday. Home combines Chat, delegated Cowork tasks and live Word, Excel and PowerPoint work in one place. Code lets employees build apps, dashboards and workflows with natural language in a managed runtime. Autopilot, entering private preview, is a persistent agent with its own identity, memory, computer and workspace that can keep working after the user leaves. Microsoft is also adding FinOps controls for usage-based agent spending, including model policies, credit visibility and administrative spending controls. Much of the new functionality is still rolling out or in preview.

The big picture

The enterprise AI contest is moving above the application layer. Microsoft is trying to make Copilot the place where a worker asks, delegates, builds and acts while the underlying documents, data and applications become resources the agent calls. The differentiator is therefore not just the model. It is permissioned context, identity, runtime, workflow access and cost control. If that architecture works, software increasingly becomes capability behind the agent rather than a destination the employee visits.

3. Anthropic is testing whether agents can do original science

Anthropic says roughly 950 Claude agents spent 21 hours searching DNA sequence data, gathering more than 200,000 reverse transcriptases, narrowing them to 3,500 candidate systems and then 20 for deeper analysis. One candidate became what Anthropic calls array-associated reverse transcriptases, or ART, a previously uncharacterized system with CRISPR-like repeat features. Human scientists performed the laboratory work. The underlying reverse transcriptase had been identified before, ART’s biological function is still unknown, and the work is an early preprint rather than a validated breakthrough.

The big picture

The interesting shift is from AI summarizing known science to AI searching a huge hypothesis space and deciding what looks unusual enough for humans to test. That can make scientific judgment and wet-lab capacity, not idea generation, the bottleneck. The caveat matters: generating a promising hypothesis is not the same as proving its function or value. But this is a credible glimpse of a research model where humans define the question and validate the result while agents perform massive parallel exploration.

4. Google is adding observability back into AI-run advertising

Google announced a new AI Max reporting view for Search that is intended to show, in one place, which search terms triggered an ad, which creative assets a user saw and where that user landed on the advertiser’s site. Google also expanded the closed beta of AI Brief, which lets advertisers give AI Max more business, audience and messaging context, into seven additional languages. The reporting feature is not broadly available yet. Google says additional details and availability will come later this year.

The big picture

As advertising systems automate more targeting, creative selection and optimization, transparency becomes part of the product. Marketers do not necessarily need to make every decision manually, but they do need to understand what the system actually did, where money went and what experience a customer received. In an AI-operated stack, observability is the new control surface.

5. Expert data is becoming its own AI infrastructure market

Snorkel AI raised $350 million at a $3.5 billion valuation this week as demand grows for specialized training data, reinforcement-learning environments and evaluation tasks. Reuters reports that the company’s annualized revenue rose above $350 million from about $20 million a year earlier, while Snorkel says its data-as-a-service business has grown more than 18-fold and reached a $375 million annualized run rate. Those growth figures are company-reported. The more important shift is what customers are buying: complex work designed with experts in fields such as coding, law and medicine, not just large volumes of simple human labels.

The big picture

As models improve, the scarce input is increasingly high-quality judgment about difficult tasks. Synthetic data can expand supply, but frontier systems still need expert-designed problems, environments, rubrics and evaluations that distinguish plausible output from genuinely good work. That creates a new way to monetize human expertise: not only as labor completing a task, but as infrastructure that trains and tests systems capable of completing millions of similar tasks.

THE THROUGH LINE

AI is moving from the model layer into the operating layer.

OpenAI’s incidents show what happens when agents can act without sufficiently strong boundaries. Microsoft is building the enterprise layer that gives agents identity, context and a place to run. Anthropic is applying agent swarms to scientific discovery. Google is rebuilding visibility around automated media decisions. Snorkel is turning expert judgment into the data and evaluation infrastructure those systems need to improve.

The next phase of AI will be decided less by who can generate the most impressive answer and more by who can build the strongest system around the intelligence: trusted data, permissions, observability, execution access, feedback loops and accountability.

The hard problem is no longer just making AI capable. It is making capability controllable, measurable and useful inside the systems where real work happens.