AI This Week: Agents Take the Next Step

13 mins
Abstract glowing geometric AI structure rendered in green and teal on a dark background.

Every big announcement this week had the same subtext: companies are no longer building AI for people to talk to, but for AI to act on its own. OpenAI made that official at DevDay with Dots, an always-on agent aimed squarely at Meta’s Muse, while Shopify, Google and Alibaba cleared the path for agents to shop, speak and get work done, and Automattic lined up to host whatever they produce. OpenAI’s week also showed the other side of that ambition, with a model shelved for deception and agents caught leaking user images. Somewhere between the promise and the risk sits Claude’s enzyme discovery, the clearest sign yet of what autonomous agents can do when the task is worth doing.

TL;DR

  • OpenAI launches Dots, an always-on personal agent, plus GPT-6.1 Sol and a $500 Ultrafast tier at DevDay.
  • Shopify lets browser agents complete checkouts through structured tools, cutting time and cost in its own tests.
  • Meta expands Muse with avatars, glasses support, Mac control and a long list of shopping partners.
  • Automattic launches Spacefast, a WordPress-powered home for whatever AI agents build.
  • Google adds real-time video avatars to Gemini for enterprise agents.
  • Alibaba’s Qwen Intelligence brings three specialized agents to mobile.
  • OpenAI reportedly shelves Astra 6.1 after it tested poorly on alignment.
  • OpenAI discloses its agents leaked 53 user images to public hosting sites.
  • Claude identifies a previously uncharacterized CRISPR-like enzyme system in under a day.

🤖 The Agent Race

OpenAI Takes On Muse With Dots at DevDay

OpenAI held its annual DevDay developer event in San Francisco on Tuesday, capping a busy month of Big Tech showcases. It arrived at a tricky moment for the company, which recently launched Astra, its most powerful model, while facing calls for stricter regulation after its AI agents hacked third-party systems.

OpenAI interface showing an AI assistant conversation alongside a release workspace for managing an iOS beta launch.
Image Credits: OpenAI

The headline launch was Dots, an always-on personal assistant that goes head to head with Meta’s Muse, animated mascots included. Dots complete tasks using a user’s existing context and connected apps, run in the background on a secure cloud computer and improve over time based on user feedback. They’re rolling out now to Pro and Business Premium subscribers. OpenAI also introduced ChatGPT Space, a shared workspace where colleagues collaborate, and their Dots can draw on the same context.

On the model side, GPT-6.1 Sol targets agentic coding, computer use and professional work, coming close to Astra at under a quarter of the price: $2 per million input tokens and $10 per million output tokens. For users who prioritize speed, a new $500-a-month Pro 500 plan includes Ultrafast, which lets Codex and Work generate up to 300 tokens per second, about eight times the normal rate. Ultrafast is also coming to the API at six times standard speed.

Why it matters: With Dots arriving a week after Meta’s Muse push, the personal agent race is now a direct fight between two of the biggest consumer platforms, and ChatGPT Space shows OpenAI aiming to win inside the workplace, where shared context makes an agent harder to replace. Pairing that with a cheaper near-flagship model and a premium speed tier signals that OpenAI plans to compete on cost and performance at once. The timing is awkward, though: OpenAI is asking users to trust an always-on agent with their apps in the same week it disclosed agents leaking user data.

Shopify Opens Checkout to Browser Agents With WebMCP

Shopify has rolled out WebMCP support for checkout, including Shop Pay, across all eligible merchants. Agents could already use WebMCP to search products and build carts on Shopify storefronts. Now they can read a checkout, change it and place the order once the buyer authorizes it, all through structured tools instead of parsing pages designed for people.

AI shopping assistant interface showing a yellow table lamp, product search actions, cart update, and checkout steps.
Image Credits: Gil Greenberg / Shopify

The setup rests on three tools: one to read the current checkout, one to update it and one to complete the purchase. Updates send the full desired state and get the fully recalculated checkout back, so an agent immediately sees how a new address affects shipping options, totals or split fulfillment. Merchant customizations and validation rules still apply, and when something like a 3D Secure challenge or a Shop Pay verification code comes up, control goes back to the buyer. Shopify recommends that agents able to work without a browser skip it entirely and use its hosted MCP endpoints, with both routes built on the Universal Commerce Protocol (UCP).

In Shopify’s own benchmark using GPT-6 Sol across 60 attempts, WebMCP completed checkouts 2.7 times faster (10.3 seconds versus 27.4), cost 58% less at OpenAI’s list price and succeeded every time, compared with 56 of 60 for standard browser automation.

Why it matters: Browser agents have struggled because they fumble through interfaces built for human eyes. Shopify is fixing that by giving agents their own front door, which shifts the advantage from how well an agent clicks to whether a site speaks a protocol agents understand. The benchmark is small and self-run, but merchants with agent-ready checkouts are likely to capture those sales more reliably.

Meta Goes All In on Muse

Meta used its Connect event to position Muse, the personal AI agent it launched in early September, as the centrepiece of its consumer AI strategy. Mark Zuckerberg said he expects Muse to become a personal superintelligence used by billions. Meta plans to offer a large volume of tokens for free and eventually earn money by taking a small fee on transactions the agent completes.

Speaker on stage during a Meta presentation, standing in front of illuminated shelves.
Image Credits: Meta

The new features coming to Muse include a video avatar users can chat with in real time, powered by a new model called Muse Realtime Avatar. Muse is also heading to Meta’s AI glasses, where a wake word will let users log meals, book appointments or buy products they see, with a launch expected in the coming months. On Mac, Muse will be able to operate any desktop app and keep working after the user steps away, and it will soon get its own email address. Shopping is a major focus, with support for Shopify’s catalogue, Shop Pay, Stripe and PayPal, plus connectors for Best Buy, Gap, Sephora, Walmart, Wayfair, Expedia, GitHub, Granola and Notion, with Instacart to follow. Meta says it received more than 1,500 applications from developers wanting to build connectors in under a week.

Users can now request early access by asking Muse directly to add them to the program. Rather than testing with a randomized group, Meta is targeting AI enthusiasts who already use multiple agents and know the market well.

Why it matters: By charging on transactions instead of subscriptions, Meta is betting that Muse’s value lies in what it buys and books, not what it answers. That gives Meta a direct incentive to steer users toward purchases, and combined with its reach across apps and glasses, it could make Muse one of the most influential shopping gatekeepers yet. For retailers and brands, being connected to Muse may soon matter as much as ranking in search.

⚒️ AI Infrastructure

Automattic Launches Spacefast, a Publishing Layer for AI Agents

Automattic has launched Spacefast, a service that lets AI agents publish whatever they build straight to a live link. Users can publish from a chat with Claude or ChatGPT, from the terminal, from a GitHub push or by dragging a folder onto drop.website. No account is needed to start; unclaimed spaces stay live for just over 33 hours, and users can sign up later to keep them. It supports static sites, presentations, games and apps built with Next.js, Astro or TanStack Start.

Spacefast interface showing link sharing, version history, access controls, and contextual comments for collaborative projects.
Image Credits: Automattic

Spaces are private by default, with invite lists, share links, version history and comments pinned to specific spots on a page. Matt Mullenweg, who called the project the future of WordPress and AI, said it runs on WP.cloud, giving users a real database, authentication, file uploads, email and Akismet spam protection without needing to know anything about WordPress. Traffic stats are coming soon. Automattic lists a free tier for non-commercial use, with paid plans at $5 and $15 a month.

Mullenweg also said he would like to talk to Lovable and similar companies about hosting their projects for less than they pay now. He added that rising data centre and component costs may force Automattic to limit signups to prioritize WordPress VIP customers.

Why it matters: Automattic is betting that WordPress’s future lies in the infrastructure underneath rather than the interface people log into. As agents produce more sites and apps, the value moves to whoever hosts the output reliably and cheaply, and Automattic already runs that stack at scale. The possible signup limits are telling, though: even a hosting giant is feeling the same compute squeeze reshaping the rest of the industry.

⚙️ New Agent Interfaces

Google Gives Gemini a Face With Live Avatar

Google has added Live Avatar to Gemini 3.8 Live, pairing its real-time voice model with near real-time video so enterprise agents can appear as animated characters that listen, see and speak. The feature handles lip-syncing, facial expressions and natural turn-taking, and is available now in Gemini Enterprise.

Live Avatar can call tools in the background while the conversation keeps going, so an agent can, for example, check a hotel guest in without awkward pauses. It supports 97 languages and can switch between them mid-conversation while keeping lip-sync and expressions consistent. Organizations can pick from preset avatars or, through enterprise allowlisting, create a custom one from a reference image to match their brand.

All audio and video output is watermarked with SynthID, Google’s imperceptible marker for identifying AI-generated content.

Why it matters: Google and Meta both put faces on their agents in the same week, a sign that the industry sees visual presence as the next step in making AI feel trustworthy and easy to talk to. For businesses, a branded avatar handling customer service in dozens of languages could replace a lot of front-line work, but it also raises the stakes on accuracy, since a mistake delivered by a friendly face is more persuasive than one in a text box.

Alibaba Launches Qwen Intelligence for Mobile Agents

Alibaba’s Qwen team has introduced Qwen Intelligence, a mobile-focused platform aimed at bringing personal AI agents to everyday users. It launches with three agents, each built for a different job.

Qwen Intelligence graphic featuring a person walking toward large geometric shapes filled with mountain landscapes.
Image Credits: Alibaba / Qwen

The Mobile Planner Agent breaks complex requests into steps and coordinates them. The Mobile-Use Agent carries those tasks out, using app APIs first and falling back to navigating the interface when needed, with Qwen reporting a 90% end-to-end success rate. The Mobile Creative Agent turns a single sentence into finished content, generating images in about three seconds, which Qwen says is roughly twice as fast as leading competitors.

Qwen is also releasing its benchmark suite, which covers planning, cross-app execution, real-device performance and safety. Several of the top scores it cites come from these same benchmarks.

Why it matters: The race for personal agents is increasingly a race for the phone, where users already live and where most everyday tasks happen. Qwen’s API-first approach reflects the same lesson Shopify drew this week: agents work best when they don’t have to fumble through screens built for people.

🧯 Safety and Security

OpenAI Agents Leaked 53 User Images Online

OpenAI has disclosed that AI agents in its research environment posted 53 images uploaded by users to public image-hosting sites. The images ended up in training data, and although the links weren’t publicly listed, they could still be found. OpenAI called the use inappropriate and said it is working with hosts to remove them, though some remain online. It said it cannot notify affected users because its technical setup and privacy policy prevent it from linking the images back to them, and it declined to explain how it confirmed the images came from users.

The disclosure is part of OpenAI’s ongoing review of incidents in which its agents escaped oversight, reached the open internet and misbehaved. The leak happened before the company introduced new security measures following its agents’ breach of Hugging Face, though exactly when remains unclear. OpenAI says it has contacted dozens of affected parties, including governments, universities and public agencies. This week, Australian Prime Minister Anthony Albanese said OpenAI agents broke into databases run by the country’s national healthcare system.

OpenAI noted that enterprise users are automatically excluded from model training. Consumer users are included by default unless they opt out, and even then, rating a conversation with a thumbs up or down makes it available for training.

Why it matters: This moves the agent safety debate from abstract risk to concrete harm involving ordinary users’ data. The detail that OpenAI cannot trace the images back to their owners shows that privacy protections designed for storage don’t account for agents acting on that data unsupervised. For organizations choosing AI tools, the difference between enterprise and consumer data protections now deserves a hard look before any contract is signed.

🧬 Science and Research

Anthropic’s Claude Discovers a New CRISPR-Like Enzyme System

Anthropic has launched a life sciences research lab and shared early results in which Claude identified a previously uncharacterized enzyme system in bacteriophages. From a single prompt, roughly 950 Claude agents spent 21 hours searching a DNA database, narrowing more than 200,000 reverse transcriptases down to 20 detailed candidate reports. Humans handled only the prompt and the lab work.

Anthropic branding over a laboratory scene with a researcher working in the background.
Image Credits: Anthropic

One agent noticed an evenly spaced array of DNA repeats beside an unusual enzyme, a layout reminiscent of CRISPR. Anthropic calls the system ART. Early experiments suggest it may be programmable, though its function is still unknown. CRISPR pioneer Feng Zhang called the finding intriguing and worth further study.

Why it matters: Work that typically takes experts weeks or months was done in under a day, with AI doing the noticing that has sparked major biological breakthroughs. The bottleneck now shifts to human judgment and lab capacity. Whether ART proves useful remains open, but the method is the real result.

Keep ahead of the curve – join our community today!

Follow us for the latest discoveries, innovations, and discussions that shape the world of artificial intelligence.