AI This Week: Agents Get a Browser, Commerce, and Creative Tools

15 mins
Abstract black-and-white visualization of a connected geometric sphere surrounded by glowing data points and flowing network lines.

Two browsers, seventy tools, and a thousand bots for every person. Cloudflare built a browser that no human will ever use, then forecast a web where humans barely register. Meta gave away an agent model that runs on a laptop while keeping the frontier model it was distilled from behind closed doors. OpenAI flagged that its next model may cross its own top safety threshold, then widened access to offensive security capability for vetted defenders. Google Maps started taking food orders, and Adobe moved seventy of its tools into a chat window it does not own. The common thread is agents arriving in places built for people, and the infrastructure quietly being rebuilt to suit them instead.

TL;DR

  • Meta released Muse Glimmer, a 30B agent model under Apache 2.0 that runs on a single consumer GPU, while keeping the larger Muse Spark closed.
  • Cloudflare built Kitesurf, a browser engine for AI agents that uses three to seven times fewer resources than Chromium and runs entirely on Workers.
  • Google Maps added agentic food ordering with Square and Toast, live hotel pricing, and a Gmail connection that reads a user’s existing plans.
  • Adobe consolidated its ChatGPT connectors into one plugin exposing more than 70 tools across Photoshop, Premiere, Acrobat and Firefly.
  • OpenAI said it cannot rule out Critical cyber capability in its upcoming Astra model, then expanded Daybreak access and shipped GPT-5.6-Cyber, which completes 95 percent of advanced security requests versus 1.5 percent for the standard model.
  • Cloudflare’s CFO forecast non-human traffic reaching 1,000 times human traffic within five years, driven by agents rather than crawlers.

⚙️ Model Releases

Meta Open Sources a 30B Agent Model That Runs on a Laptop

Meta Superintelligence Labs has released Muse Glimmer, a 30-billion-parameter model built for always-on local agent workflows, with open weights on Hugging Face under a permissive Apache 2.0 licence. The target is a Mac or PC with a single consumer GPU, covering local agents, function calling, coding and LLM-as-a-judge evaluation. Glimmer is effectively an open counterpart to Muse Spark, the closed model Meta debuted in April, and was distilled from it.

Getting a 30B model onto consumer hardware took two main optimizations. At full precision, the weights would need over 55 GB, so quantization to roughly 4-bit brings the language model under 20 GB, leaving room for the KV cache, the perception encoder and the speculative decoding drafter inside a 24 or 32 GB envelope. Meta says the compression causes minimal to no degradation on agentic tasks. A lightweight drafter based on DFlash proposes blocks of tokens that the main model verifies in parallel, reported as a 3.1 times decode speedup on an RTX 5090, 1.8 times on an M5 Max and 1.5 times on an M4 Max. The model handles interleaved text and images, is trained on more than 100 languages, supports adjustable reasoning effort, and is designed to diagnose and retry failed tool calls rather than stop. It runs through Ollama, LM Studio, llama.cpp, ExecuTorch, MLX, vLLM and SGLang.

Benchmark table comparing Muse Glimmer-30B, Gemma4-31B, and Qwen3.6-27B across several agentic AI evaluations.
Featured Image: Meta

The release arrived alongside a letter from Mark Zuckerberg restating his case for distributing superintelligence widely, which he frames as the start of an era of personal empowerment, with an agent working continuously on a person’s relationships, health, career, finances and home management, and free or affordable access for everyone. Glimmer is the first concrete expression of that argument. It also marks where the line falls. Spark stays closed. The smaller model is the one people can download, fine-tune and own, which makes this release as much a statement about what Meta will not open as about what it will.

Why it matters: The pitch depends on access to personal context, and that is exactly the access most people and most organizations will not route through a cloud API. Local capability, not benchmark position, is what makes the personal agent argument plausible at all, and a competent model under a permissive licence running on hardware people already own removes both the per-token cost and the data residency objection at once. But the split between Spark and Glimmer is the part that defines the strategy. Meta is distributing the tier that runs on a laptop while retaining the tier that runs the frontier, so users own the agent and Meta owns the intelligence the agent was distilled from. The security ledger is separate and less comfortable. Weights that run offline cannot be rate limited, monitored or revoked, which is the same property that makes them worth having.

🌐 Agent Infrastructure

Cloudflare Builds a Browser for Agents, Not People

Cloudflare has released Kitesurf, a browser engine built specifically for AI agents and running entirely on Workers. It is available free during beta through Browser Run, the company’s existing headless automation product, which until now ran on Chromium, the open-source engine behind Chrome and Edge and the default for nearly all automated browsing. The reasoning behind Kitesurf is an argument about cost: Chromium was designed for humans, and it carries overhead that models do not need. Memory and compute consumption are high enough that giving every agent its own instance becomes prohibitively expensive, which in practice restricts large parts of the web to only the most capable and costly models.

Stylized illustration of a kitesurfer riding across an orange and red wave within a framed interface.
Image Credits: Claudflare

So Kitesurf drops what models do not use. No tabs, no extensions, no pixel-perfect rendering. Most of the engine is Rust compiled to WebAssembly, using Blitz for parsing and painting, Firefox’s Stylo for CSS and Parley for text shaping. A single component handles all network access, page sessions run in isolated Dynamic Workers, and the renderer holds no state, so a stalled request can be killed and retried. The project is twelve weeks old and passes more than 215,000 Web Platform Tests.

The benchmarks tell the story. Across a fourteen-URL corpus, Kitesurf used three to four times less CPU than Chromium and four to seven times less memory, while running roughly 1.7 times slower on wall time. Cloudflare optimized for the metrics that drive the bill, not the stopwatch. It cannot yet handle video, WebGL, bot challenges or long authenticated sessions, and the company plans to open source it.

Why it matters: Cloudflare already sits in front of a large share of web traffic and decides which automated requests get through. Building the browser those requests run in extends that position from gatekeeper to supplier, and a company that meters bot access selling the cheapest way to be a bot is a strange and powerful place to stand. The economics point the same direction. If agent browsing costs a fraction of what it did, the volume of it rises accordingly, and publishers already struggling with crawler load face a new wave from applications that previously could not justify the compute. Cheaper agents are not a neutral efficiency gain for the sites on the receiving end.

🛒 Agentic Commerce

Google Maps Starts Placing Orders

Google has expanded Ask Maps, the Gemini-powered layer it describes as the biggest change to Maps in over a decade, with agentic capabilities that complete tasks rather than just answer questions. The headline addition is food ordering: a request for a specific dish to pick up on the way home returns open restaurants along the route, accounting for saved places and dietary needs, and adds the item to a cart for review. It is launching with Square and Toast, with Uber Eats to follow, and Google says it is co-developing something called the Universal Commerce Protocol for Food with its partners. Hotel search now compares live prices and availability, and event listings arrive with ticket links.

Illustration of a person using a smartphone alongside Google Maps and search interfaces, featuring an “Ask Maps” route preview.
Image Credits: Google

The other additions are about context. Personal Intelligence lets users connect Ask Maps to Gmail, with Calendar coming, so the assistant can factor in flight times and dinner reservations without being told about them. The connection is off by default. Conversation history persists, so earlier trip planning can be recalled. A transit widget adds live delay data for buses, trains, subways and ferries on a virtual departure board. Contributions have gone conversational as well, including uploading a photo of a storefront sign so Maps can read the hours off the image and submit an edit for confirmation.

Ask Maps is now rolling out in Australia, Brazil, Canada, Indonesia, Japan and Mexico, plus more than 150 countries in English. Ordering, hotels, events and conversational contributions are US only for now.

Why it matters: Maps has spent twenty years as the layer that sends people to businesses. Taking the order collapses that referral into a transaction, and the local restaurant that once competed on its Maps listing now competes for placement inside an assistant’s shortlist, with no page to optimize and no obvious way to see why it was passed over. The protocol work is the part worth watching. Google writing a commerce standard for food ordering, with point-of-sale vendors as launch partners, is an attempt to make Maps the interface layer above the systems restaurants already run on. Meanwhile, the Gmail connection quietly moves Maps from a place people go to a service that reads their plans, which is a different privacy proposition than looking up directions, whatever the default setting says.

🎨 Creative AI

Adobe Puts 70 Tools Inside ChatGPT

Adobe has replaced its separate ChatGPT connectors with a single unified plugin, pulling more than 70 tools from Photoshop, Illustrator, InDesign, Express, Premiere, Lightroom, Acrobat, Firefly and Adobe Stock into one conversational surface across ChatGPT Work and Codex. Rather than switching between app-specific connectors, the plugin routes to the relevant tool based on the prompt, so image editing, video work, layout and document export can happen in one continuous exchange.

Adobe creative tools shown inside a ChatGPT interface, with a prompt asking to change a pickleball graphic’s background colour to pink.
Image Credits: Adobe

The functionality spans batch style adjustments across campaign photo sets, turning images into product mockups, cutting long footage into vertical highlight reels for Reels and Shorts, template search with copy and colour edits, and print-ready PDF export. Uploaded spreadsheets can be turned into formatted output like event badges, catalogues and infographics. Assets save back to Creative Cloud. The plugin is available globally on web and mobile, with guest mode for basic use and an Adobe sign-in required for generative features and file access.

Forest Key, who leads Firefly and agentic AI at Adobe, framed the move as meeting more than a billion ChatGPT users where they already work, and was explicit that the plugin is not a replacement for the flagship desktop apps. The positioning is a conversational entry point for ideation and batch work, with precision tasks moving into the desktop software. Adobe has already shipped comparable integrations with Claude and Copilot, and Slack and Gemini are reportedly next.

Why it matters: Adobe is doing the arithmetic that every software company with a mature install base will eventually have to do. Its tools have always been where the work happened, but the ideation now starts somewhere else, and defending the application boundary would mean ceding the moment a project actually begins. The insistence that the plugin complements rather than replaces the desktop apps is the interesting tension, because the tasks it covers, resizing for social, batch edits, template tweaks, PDF export, are precisely the routine production work that fills a lot of junior creative time. Distributing across ChatGPT, Claude, Copilot, Slack and Gemini simultaneously also tells you Adobe has no intention of picking a winner among the assistants. It would rather be the capability layer that every one of them calls, which keeps the subscription intact even if the interface people open each morning is no longer Adobe’s own.

🔓 Safety and Security

OpenAI Flags Its Own Ceiling, Then Arms Defenders

OpenAI disclosed that internal evaluations of Astra, an upcoming model, show enough advancement in agentic coding and cybersecurity that the company cannot rule out the Critical threshold under its Preparedness Framework. Critical has a specific definition: a model that can find and develop working zero-day exploits across many hardened real-world systems without human intervention, or devise and execute novel end-to-end attacks against hardened targets given only a high-level goal. Previous models, including GPT-5.6 Sol, were assessed at High. OpenAI has paused internal Astra work that does not meet strengthened security requirements, and imposed isolated testing environments, restricted network and tool access, additional weight encryption and sandboxed execution. Universal monitoring now runs across all agentic use of the model, including training and evaluation, with monitors reading chain of thought to interrupt high-risk activity.

Against that backdrop, OpenAI expanded Daybreak, its access programme for approved security professionals, into two tiers. Daybreak Blue removes the production guardrails that screen cybersecurity requests, giving vetted defenders fuller use of GPT-5.6 Sol for vulnerability management, secure code review, malware analysis and incident response. Daybreak Red provides purpose-trained models for authorized vulnerability research, exploit validation and security testing, and introduces GPT-5.6-Cyber, built on Sol and trained both to perform better on zero-day discovery and exploit chain development, and to refuse less often on dual-use requests.

The refusal numbers are stark. On OpenAI’s internal Advanced Cybersecurity Completion Rate evaluation, covering exploit chain development, authentication bypass and privilege escalation, GPT-5.6-Cyber completes 95 percent of requests. Sol completes 1.5 percent with safeguards on and 2 percent through Daybreak Blue. The previous Cyber model managed 57.3 percent, and OpenAI cites researcher frustration with persistent refusals as the reason for the change. Capability results are more mixed: the Cyber model leads on ExploitGym and zero-day discovery quality but scores lower than Sol on vulnerability report writing, and Sol remains more token-efficient on ExploitBench at the standard 300-turn limit. GPT-5.6-Cyber was assessed at High, not Critical.

Field results back the capability claim. OpenAI used the model to find two previously unknown V8 vulnerabilities that could be chained to corrupt memory and escape the heap sandbox, disclosed to Google and fixed as CVE-2026-15903. It also reports at least five vulnerabilities in a widely used mobile operating system, three critical issues in a popular database including a remote path to code execution, and more than 400 privilege escalation vulnerabilities in a widely used OS kernel. Access runs through identity verification, monitoring, approved-use restrictions and legal attestations, with hardware security keys mandatory for individual accounts from September 1, 2026.

Why it matters: The sequence is the argument. OpenAI names a capability it may not be able to contain, then moves within days to put a weaker version of that capability into defenders’ hands under contract and monitoring. The reasoning is that the window closes either way, so the only variable is who is prepared when it does. What the Daybreak numbers make explicit is that the safety layer on a production model was always a policy choice rather than a capability limit: the same weights go from 1.5 percent completion to 95 percent depending on which agreement the customer signed. That turns trust and safety into an access control problem, and access control has a long record of being defeated by insiders, stolen credentials and scope creep.

📉 The Machine Web

Cloudflare Forecasts Humans as a Rounding Error

On its Q2 earnings call, Cloudflare CFO Thomas Seifert projected that if current trends hold, non-human traffic will be as much as 1,000 times human traffic within five years. His framing was that humans become a rounding error on the internet, not because human traffic falls but because of how quickly the other kind is growing. He caveated that his predictions have been wrong before, which is a fair caution given that CEO Matthew Prince had forecast bots overtaking humans at the end of 2027 and the crossover arrived earlier this year.

The driver is agentic AI rather than conventional crawlers or fraud bots. Agent traffic resembles ordinary browsing but runs at machine speed and repeats at scale, so a single prompt can generate thousands of requests. Prince’s illustration was shopping: a person comparing cameras might visit five retailers, while an agent acting on their behalf could hit 5,000 sites. Multiply that across millions of people asking assistants to compare prices, check airfares, research topics or summarize articles and the arithmetic gets away quickly. Seifert’s conclusion was that the industry has to get considerably more efficient, and he flagged the security implications of a web composed almost entirely of automated clients.

Why it matters: A thousand to one is a ratio at which the economics of the open web stop working. Advertising, analytics and capacity planning all assume a person on the other end of a request, and none of that survives an audience that is overwhelmingly machine. The cost side inverts first: publishers pay to serve traffic that will never see an ad or return, while the value accrues to whoever operates the assistant. Independent sites feel it soonest, lacking both the leverage to set terms and the budget to absorb the volume.

Keep ahead of the curve – join our community today!

Follow us for the latest discoveries, innovations, and discussions that shape the world of artificial intelligence.