OpenAI is packaging a managed agents product for DevDay, and Meta launched Muse, a consumer agent that runs on its own isolated machine and spends real money on your behalf. Then the counterweight: an Anthropic researcher quit, warning that the labs are racing toward self-improving systems with no plan to control them. Around that, OpenAI shipped a sharper image model, Google put a satellite-trained weather model into Search, Arm went after robotics standards, and six months of ChatGPT Ads data still gave advertisers nothing to measure against.
Listen to the AI-Powered Audio Recap
This AI-generated podcast is based on our editor team’s AI This Week posts. We use advanced tools like Google NotebookLM, Descript, and ElevenLabs to turn written insights into an engaging audio experience. While the process is AI-assisted, our team ensures each episode meets our quality standards. We’d love your feedback—let us know how we can make it even better.
TL;DR
- OpenAI is expected to unveil a managed agents product at DevDay on September 29, closely tracking Anthropic’s Claude Managed Agents, with a possible path from ChatGPT ads straight into an agent.
- Meta launched Muse, a consumer AI agent that runs on a dedicated secure VM, pays through Stripe’s Link with one-time card numbers, and acts across your apps under approval.
- Anthropic researcher Jacob Coxon resigned, warning that the labs are gambling with humanity’s future; alignment lead Evan Hubinger agreed and put the odds of catastrophe above 10% within the decade.
- OpenAI released ChatGPT Images 2.5 with Sketch input, better multi-turn edit consistency, and two API tiers split by speed versus control.
- Google’s WeatherNext 3 trains on live satellite data, refreshes hourly at 5-kilometre resolution, and is already live across Search, Gemini, and Maps.
- Microsoft’s MAI-Transcribe-2 claims the top spot on speed, accuracy, and price at $0.10 per hour, a rate that expires at year-end.
- Arm launched a Robotics Capability Framework and physical-AI design program with 80-plus partners to standardize how the industry describes robotic systems.
- Six months into ChatGPT Ads, published CPCs run from under $3 to nearly $23 with no benchmarks to explain the spread, and the ad-eligible audience excludes every paid tier above Go.
🤖 Agentic AI
OpenAI Is Lining Up Managed Agents for DevDay
OpenAI’s DevDay lands September 29 at Fort Mason in San Francisco, with the keynote livestreamed and in-person registration already closed. Ahead of it, a report from Alexey Shabanov, based on code found in OpenAI’s platform, points to the headline announcement: a managed agents product that looks a lot like what Anthropic already ships with Claude Managed Agents.
The unreleased UI reportedly lets users create separate agents and environments, manage them, and switch on specific skills and plugins. Self-hosted environments appear to be supported, and an OpenAI Developers plugin would let people spin up managed agents from the start. The expected audience is businesses as much as developers, which would mark a shift from the developer-first framing of past DevDays.
This would be OpenAI’s third pass at an agent ecosystem. Custom GPTs came in November 2023, Agent Builder followed in October 2025, and in June 2026 OpenAI said Agent Builder and Evals would wind down after November 30, with the Agents SDK and Workspace Agents named as successors. The visual pipeline-builder approach lasted barely a year before model capabilities made manual wiring unnecessary.
Shabanov also flags something further out: functionality that would let ChatGPT ads resolve directly to an agent, so a user clicking an ad gets handed to an agent that walks them through a purchase rather than a static link. OpenAI already runs ChatGPT Ads with campaign goals, bidding, and conversion measurement, and said in August it would explore more native ways for businesses to reach consumers. Whether the agent-ad piece surfaces at DevDay or later is unclear, and none of this is confirmed by OpenAI.
Why it matters: OpenAI and Anthropic have converged on the same product shape, which suggests agent orchestration is becoming a built-in layer of the model platform. Everything that sat between the model and the workflow, including OpenAI’s own Agent Builder, is getting absorbed upward.
Meta Launches Muse, a Personal Agent That Runs on Its Own Secure Machine
Meta introduced Muse, a personal AI agent aimed at consumers rather than developers. The pitch is that Muse does the work instead of answering questions: set it a goal and it builds a plan, opens a browser, fills forms, and advances the task on its own, checking back before anything sensitive like sending an email or making a purchase. It runs on Muse Spark, described as Meta’s most capable model, and you talk to it through a dedicated app or inside WhatsApp.

Each Muse runs on its own isolated cloud VM, the Muse Secure VM, where the agent and the user’s connected credentials live, walled off from anyone else’s agent. A separate Sentinel agent sits on the same machine at the system level, and nothing Muse does reaches the internet unless Sentinel approves it. Muse never sees passwords or card numbers directly; credentials go into secure storage it can use but not read. Meta says conversations and VM data stay out of its ad systems, users control which apps connect and how much access each gets, and a Confidential VM with user-held encryption keys is promised later this year.
Payments run through Stripe’s Link, which generates a one-time card number at checkout and, Meta says, makes Muse the first AI agent covered by Link’s purchase protections. Shop Pay and 1Password support are coming. Muse is live in the US on iOS, Android, and the web, free for most uses with paid tiers above that, and headed to Meta’s AI glasses.
Why it matters: Everyone shipping a consumer agent has hit the same wall, which is that letting software act on your email, calendar, and credit card is only acceptable if it demonstrably cannot be hijacked or quietly siphon your data. Meta’s answer, an isolated VM with a separate gatekeeper agent and credentials the agent cannot see, is a bet that trust infrastructure is what unlocks mass adoption. Whether people believe Meta, of all companies, when it promises the ad systems stay out is the open question, and the user-held encryption keys arriving later read as an attempt to answer it.
🧠 New Models
OpenAI Ships Images 2.5 With Sketch Input and Two API Tiers
OpenAI released ChatGPT Images 2.5 on the back of what it says is more than 3 billion images generated weekly across ChatGPT and the API. The model produces sharper detail and more natural lighting, holds reference-photo subjects more faithfully, and follows editing instructions more reliably across multiple turns. OpenAI puts generation latency at up to 50 percent lower than Images 2.0. It is available today to all ChatGPT, ChatGPT Work, and Codex tiers across desktop, mobile, and web.

The product additions are aimed at directing the model rather than prompting it blind. Sketch lets you draw a rough guide inside ChatGPT and have the model render from it, triggered by typing “@Sketch.” Templates give starting points for common formats like posters and merch, comments allow focused edits pinned to a spot on the image, and shared images can now carry the prompt that made them so others can reuse the idea with their own inputs.
For developers, there are two API models. GPT-Image-2.5 Flare is the default, matching the quality and editing gains at 50 percent lower latency than GPT-Image-2, pitched at high-volume and social workloads. GPT-Image-2.5 Sunburst trades speed for tighter control across edits, aimed at campaign and product imagery. OpenAI says it continues to tag outputs with C2PA metadata and invisible watermarking.
Why it matters: The upgrade that counts is edit consistency. Every image model can make a nice picture now, so the constraint on actual production work has been that each new edit degrades the last one and the asset drifts away from the brief. Holding composition and identity stable across many turns is what makes the model usable in a real creative pipeline instead of a one-shot toy.
Google’s WeatherNext 3 Forecasts From Live Satellite Data Instead of Simulations
Google DeepMind and Google Research released WeatherNext 3 on September 3, calling it the most accurate global weather model available based on independent live evaluations by Brightband. It is already running weather in Search, Gemini, Google Maps, the Maps Platform Weather API, and Earth Engine, with forecast data available through BigQuery and Cloud Storage.
The main change is what the model learns from. Most AI forecasters, including last year’s WeatherNext 2, train on the output of physics-based simulation systems that carry a six-hour lag. WeatherNext 3 ingests live geostationary satellite imagery and direct weather station readings, which lets it issue a new forecast every hour at 5-kilometre resolution for surface temperature and moisture, roughly five times sharper than its predecessor’s 25-kilometre, six-hourly grid.

Precipitation is where Google claims the largest gains, citing CRPS improvements of up to 60 percent against NASA’s IMERG dataset and consumer forecasts a day or more out that are up to 50 percent more accurate. It also adds turbine-height wind and solar radiation outputs for renewable operators, and pitches the model at Latin America, Africa, and Asia-Pacific, where high-resolution regional forecasting has been too expensive to run.
Why it matters: Training on observations rather than on another model’s output removes a ceiling every AI forecaster has shared, and hourly refreshes shift the use case from next week’s outlook to storms developing right now. But distribution is the real edge. Google is putting this in front of billions of people and into cloud data products on day one, a reach no national weather service can match. The disclaimer pointing users to their local meteorological agency for official warnings reads like a preview of who is about to be sidelined.
Microsoft Ships a Transcription Model Built to Undercut Everyone
Microsoft AI released MAI-Transcribe-2, claiming it is the fastest, most accurate, and cheapest speech recognition model available. The company says it tops the FLEURS multilingual benchmark across 60 languages with an average word error rate of 5.2 percent, sits second on the Artificial Analysis word error rate leaderboard, and defines that site’s accuracy-latency frontier. Per Artificial Analysis evals, it processes audio 10 times faster than OpenAI’s GPT-Transcribe, 7 times faster than ElevenLabs’ Scribe v2, and 5 times faster than Gemini 3.5 Transcribe.

The feature list targets production workloads rather than demos: speaker diarization, word-level timestamps, keyword biasing for domain terms, automatic language detection, and code-switching for mixed-language speech like Hinglish and Spanglish. Developers can also choose between a verbatim mode that keeps filler words and false starts for compliance use, and a clean mode for readable captions and notes.
Pricing is $0.10 per hour of audio, which Microsoft calls the lowest on the market. That rate is a limited-time offer running until the end of the year.
Why it matters: Transcription is close to a solved problem, so the competition has moved to speed, cost, and coverage, and Microsoft is pressing all three at once. A single model that handles 60 languages, noisy audio, and mixed-language speech removes the reason to stitch together specialist vendors, and a price at or below the cheapest rivals makes switching easy. The year-end expiry on that price is the tell: this is a land grab for the API integrations that get built now, with the assumption that few developers will rip them out when the rate goes up.
Ready to explore what AI can do for your organization?
🧯 Risk and Control
Anthropic Researcher Resigns Over the Race to Superintelligence
Jacob Coxon, a researcher at Anthropic, announced his resignation with a public warning that the AI labs are racing toward self-improving superintelligence without acting responsibly, naming both Anthropic and OpenAI. He wrote that people building these systems earnestly believe the technology could kill everyone by the end of the decade, and that advanced models would soon be capable of hacking, overturning entire fields overnight, and acquiring real-world resources and power. He floated a temporary ban on improving model capabilities as one possible step, while citing the July OpenAI breach of Hugging Face as a “warning shot” that has made coordination between US labs more plausible.
What makes this more than a departure post is who agreed with him. Evan Hubinger, an alignment science lead still at Anthropic, endorsed the concern and went further, putting the odds of a catastrophic outcome within the next decade above 10%. He said he believes Anthropic is trying its best but that the company does not yet have a plan to solve alignment for superintelligence and is not clearly on track to.
Why it matters: Departure warnings from AI safety people are common enough to be background noise now. What lifts this above the genre is Hubinger, who did not resign, agreeing from inside the safety team and attaching a number to it. Once a company’s own alignment lead says on the record that there is no plan yet and no clear path to one, “we take safety seriously” stops being a sufficient answer, because the objection is coming from the people the phrase is supposed to reassure you about.
🔧 Infrastructure and Hardware
Arm Wants to Set the Standards for Physical AI
Arm launched Arm Total Design for Physical AI on September 8, along with a Robotics Capability Framework meant to give the robotics industry a shared way to describe what a system can do. The framework, modelled on the SAE levels used for driving automation, sorts machines into tiers running from reactive systems up to context-aware, cognitive, and self-improving ones. Each tier ties use cases to expected behaviours and hardware constraints, including latency, compute placement, memory, power, determinism, and safety requirements.

The program brings together more than 80 partners across software, hardware, and AI, among them AWS, Hugging Face, NXP, Qwen, Siemens, QNX, and Unitree Robotics, while ANYbotics, Lenovo, McKinsey, and others contributed to the framework itself. Arm has applied the same collaborative model before in cloud infrastructure and, more recently, in automotive, where partners built a digital cockpit reference on its Zena platform and tested code before the silicon existed. The company puts the compute opportunity across mining, agriculture, manufacturing, and transport at roughly $200 billion a year by the 2030s.
Why it matters: Robotics is stuck at the stage where every integrator rebuilds the same perception, control, and safety stack by hand, which is a big part of why so much physical AI never gets past the demo. A shared capability framework and pre-silicon tooling attack that friction directly, and Arm benefits from anything that moves the field toward volume deployment, because deployed robots run on chips and its architecture already dominates low-power edge hardware. The open question is adoption. Standards like this only matter if the 80 partners actually build to them, and plenty of well-backed industry frameworks have gone nowhere.
📊 The Business of AI
Six Months In, Nobody Knows What a Good ChatGPT Ads Result Looks Like
Search Engine Journal’s Brooke Osmundson pulled together the advertiser data that has surfaced since OpenAI started testing ads in February, and the range is wide. Published CPCs run from under $3 to nearly $23, with geography explaining some of it: one agency saw $5.10 in the UK, $10.62 in the US, and $22.89 in New Zealand. Hostinger spent close to $70,000 and found CPMs above $65, with specific use cases converting and broad messaging falling flat. An ecommerce advertiser tracked by Common Thread Collective reported ROAS between 3.3x and 6.8x depending on attribution model. A B2B campaign that spent about $7,000 CAD used visitor deanonymization to identify 146 companies behind its clicks, and only five fit its customer profile.
The reason nobody can interpret those numbers is that OpenAI has not published benchmarks by industry, advertiser, or campaign type, and Ads Manager offers nothing like auction insights or impression share. Ads are placed through a relevance-weighted second-price auction driven by conversation intent rather than keywords, so advertisers cannot tell whether a rising CPC reflects competition, relevance scoring, or the conversations they happened to match. OpenAI’s widely repeated $3 to $5 figure is a recommended starting bid, not an average, and since August the default strategy sets bids automatically anyway.
The audience is also narrower than ChatGPT’s headline user count. Ads reach only Free and Go accounts, with Plus, Pro, Business, Enterprise, and Edu users excluded, and there is little public data on who the ad-eligible group actually is. Self-service access expanded to 52 countries on August 31.
Why it matters: OpenAI is not selling an ad platform yet so much as an experiment advertisers are funding. The auction is opaque, the audience is unprofiled, and there is no benchmark for what a click should cost, which means every reported CPC is a data point without a scale. The advertisers getting real signal are the ones measuring outside Ads Manager, so OpenAI’s own reporting is not yet the thing people trust.
Keep ahead of the curve – join our community today!
Follow us for the latest discoveries, innovations, and discussions that shape the world of artificial intelligence.
