This week, the frontier split in two directions: some labs are opening their most capable models to everyone, while others are choosing carefully who gets them first. Mistral and Reflection AI are promising open weights for large models that organizations can run on their own infrastructure, while Google handed Gemini 4 Argon to cyber defenders before anyone else. Anthropic built Claude for Government around the budget and oversight rules agencies already follow, and Google Docs made room for the Markdown files AI tools produce. On the consumer side, OpenAI started putting ads beside generated images, and Pinterest turned saved inspiration into salon plans, a reminder that the AI people use most is often the kind they barely notice.
Listen to the AI-Powered Audio Recap
This AI-generated podcast is based on our editor team’s AI This Week posts. We use advanced tools like Google NotebookLM, Descript, and ElevenLabs to turn written insights into an engaging audio experience. While the process is AI-assisted, our team ensures each episode meets our quality standards. We’d love your feedback—let us know how we can make it even better.
TL;DR
- Mistral previews Large 4, a 1 trillion parameter open model trained and hosted in Europe, with weights due by month’s end.
- Reflection AI debuts Beam, a 501B open-weight model pitched on efficiency, with Apache 2.0 weights coming this month.
- Google unveils Gemini 4 Argon, releasing it first to cyber defenders without cyber guardrails before a wider rollout.
- Google Docs now opens, edits and shares Markdown files natively, no conversion required.
- Claude for Government is generally available to U.S. federal and state agencies, with hard spending caps and built-in oversight.
- OpenAI tests visual ads during ChatGPT image generation and expands its measurement and brand safety tools.
- Pinterest’s Beauty Guides turn hair and nail Pins into salon-ready plans with terminology, pricing and upkeep.
🧠 The Model Race
Mistral Previews Large 4, a 1 Trillion Parameter Open Model
Mistral has opened a public preview of Mistral Large 4, its largest and most capable model to date. Nicknamed “Le Chonk,” ML4 is a natively multimodal mixture-of-experts model with 1 trillion total parameters and 49 billion active parameters, combining instruction following, reasoning and agentic capabilities. The API is available now through Mistral Studio at $1.36 per million input tokens and $4.18 per million output tokens, with model weights expected by the end of the month.

Mistral is targeting software engineering, cybersecurity, finance, law, science and manufacturing. ML4 scored 61.7% on DeepSWE v1.1 and 59.9% on AutomationBench, and Mistral says third-party evaluations place it above GPT-6 Astra on legal and financial tasks. It can work across spreadsheets, PDFs, charts, technical drawings and large geospatial images. Before releasing the weights, Mistral is red-teaming the model with cybersecurity leaders, vetted partners and state authorities.
ML4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral’s own European data centres, which also serve the preview. It will be offered in multiple regions, including a European deployment run entirely by Mistral under European law, with private cloud and on-premises options planned. Training data covered more than 160 languages, including every official EU language. The release is the first milestone funded by Mistral’s €3 billion Series D and will serve as the base for a family of specialized models.
Why it matters: Mistral isn’t just shipping a big open model; it’s selling independence. Training in its own European facilities, offering a deployment governed by European law and planning on-premises options gives governments and regulated industries a frontier-class model they can run without depending on U.S. or Chinese infrastructure. Arriving the same week as Reflection’s Beam, it also shows the open-weight race is no longer a two-player contest between Chinese labs and everyone else, with Western developers now competing on cost, control and jurisdiction as much as benchmarks. And like Google with Argon, Mistral is gating its cyber capabilities behind a red-teaming phase before the weights go public, a sign that staged release is becoming standard practice for powerful models, open or closed.
Reflection AI Debuts Beam, Its First Open-Weight Model
Brooklyn-based Reflection AI unveiled Beam, its first open-weight model and a direct bid to compete with leading Chinese open models like those from DeepSeek, Qwen and Z.ai. Beam is a text-only mixture-of-experts model with 501 billion total parameters, 23 billion of them active at a time, built for coding, reasoning and agentic work. It was pretrained on 23.8 trillion tokens, supports a 1M token context window and went through a large reinforcement learning run that generated more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks.

Reflection isn’t claiming Beam beats the top open models outright. It says Beam is competitive with Z.ai’s GLM 5.2 on advanced reasoning benchmarks while using three to four times less inference compute, and acknowledges that models like Kimi K3 remain ahead on raw capability. The pitch is efficiency: strong performance at lower cost, positioned as a “workhorse” for enterprise coding and agentic workloads. As TechCrunch notes, these benchmark claims haven’t been independently verified.
Beam is currently in final red-teaming with early access by waitlist. Reflection plans to release the weights under an Apache 2.0 licence later this month, along with a technical report, model card and tooling for running and fine-tuning the model. The two-year-old company, founded by former Google DeepMind researchers, has raised roughly $4.7 billion and is targeting enterprises and sovereign nations with “AI factories” that let institutions train customized local models on their own data.
Why it matters: The best open-weight models have largely come from Chinese labs, which leaves Western enterprises and governments that want to run AI on their own infrastructure with an uncomfortable choice. Beam is an attempt to fill that gap, and Reflection is smart not to compete on raw capability alone. By leading with inference efficiency and a permissive licence, it’s selling something organizations actually budget for: the cost of running a model at scale, on their own terms. The open weights are the entry point; the real business is the AI factory model, where Reflection helps customers build sovereign systems on top of it. Whether that works depends on independent benchmarks confirming the efficiency claims once the weights ship.
Google Unveils Gemini 4 Argon, Starting With Cyber Defenders
Google DeepMind announced Gemini 4 Argon on September 30, a new frontier model built for long, complex workflows in software engineering, enterprise knowledge work like legal and finance, and cybersecurity defence. Rather than a broad launch, Argon is rolling out first to a group of trusted cyber defenders through Google’s Fairwind Program, and Google says it’s taking part in the U.S. government’s voluntary pre-release model access process as it expands availability. Paid API customers and Google AI Ultra subscribers are next in line, with introductory pricing set at $2 per million input tokens and $10 per million output tokens.
The headline technical change is output capacity: Argon can generate up to 1M tokens in a single response, up from 64K, giving it room to work through hard problems in one pass. Google reports state-of-the-art results on DeepSWE v1.1 (77.9%) for long-horizon coding and LVBench (91.7%) for long video understanding, along with the top spot on the Vals Index and Zapier’s AutomationBench. Internally, Argon agents have freed more than 300 TiB of memory across Google’s data centres and are migrating C/C++ codebases to Rust, including the 800K+ line Fuchsia Zircon kernel.
On security, trusted defenders and Google’s own teams will get Argon without cyber guardrails. Wiz has already used it through its Scan for Good initiative to uncover a critical vulnerability in healthcare software used by hospitals worldwide, one that earlier frontier models had missed. For the wider release, Google is strengthening misuse safeguards, prompt injection defences and monitoring of the model’s chain-of-thought to catch it acting beyond what users intended.
Why it matters: The rollout order is the real story. Google is giving its most capable cyber tooling to defenders first, unrestricted, before attackers can get comparable capability elsewhere, and treating general availability as something earned after safeguards are tested in the field. That staged, access-controlled approach is starting to look like the template for frontier releases rather than an exception. Meanwhile, the jump to 1M output tokens and the internal examples Google chose to highlight, like fleet-wide memory savings and kernel-scale rewrites, point to a model designed less for conversation and more for running long, autonomous jobs that used to require entire engineering teams.
Get AI This Week, along with industry news, delivered straight to your inbox
💼 AI at Work
Google Docs Now Edits Markdown Files Natively
Google is adding native Markdown support to Docs and Drive. Users can now open, edit and collaborate on .md and .markdown files directly in Google Docs without converting them, while Drive now displays rendered previews with formatted tables and clickable links.

Until now, working with Markdown in Docs meant importing and converting the file into Google’s own format. That process could change formatting, remove comments and leave teams juggling multiple versions of the same file. Google says the new approach keeps the file type intact while preserving Docs’ real-time editing and commenting tools, so people and AI agents can work on the same file without anyone touching raw syntax or converting manually.
The feature is rolling out gradually to Rapid Release and Scheduled Release domains, with full visibility expected within 15 days. There’s no admin control or user setting to manage, and it’s available to all Google Workspace customers and personal Google account holders.
Why it matters: Google’s own announcement frames Markdown through the lens of LLMs, which says a lot about where document workflows are heading. As more drafts start life as AI output, the format those tools produce is becoming the format teams actually work in, and Google is choosing to meet that content where it is rather than pull it into a proprietary format. Keeping the file as Markdown means the same document can move between a person editing in Docs, an agent generating or revising content, and other tools downstream without losing structure along the way. For content teams, that removes a small but constant point of friction in human-plus-AI collaboration.
Claude for Government Is Now Generally Available
Anthropic made Claude for Government generally available to U.S. federal and state agencies, following a public beta that began in July. The platform delivers Claude’s coding and agentic capabilities through a FedRAMP High authorized environment, with features comparable to what commercial customers get and new capabilities generally arriving on the same release schedule. Agency staff can work with files on their desktops for tasks like memos, RFP reviews and casework, while Claude Code helps teams build and modernize the software behind public services. The Claude Code CLI and Claude for Microsoft 365 are also entering early access in the same environment.
Much of the launch focuses on how agencies buy and manage the tool. There are no seat fees: agencies pay for usage in fixed increments with a hard not-to-exceed cap, so spending can’t go past what’s been obligated. Department-level administrators can allocate prepaid usage to sub-agencies, set spend and model limits by user tier, and receive alerts before balances run low. Procurement officers can also contract with Anthropic directly, without a separate cloud-provider relationship.
On oversight, administrative actions are captured in audit logs, sensitive operations on Anthropic’s side require two-person approval, and usage exports contain only metering data. Conversation history stays on the agency-managed device, and documentation is provided to support each agency’s Authority to Operate (ATO) process.
Why it matters: For government, the hurdle to AI adoption has rarely been the technology itself. It’s been procurement rules, budget controls and the paperwork needed to prove a system is safe to run. This launch is built around those constraints: hard spending caps that align with how public money is obligated, direct contracting that removes an intermediary, and oversight features designed to answer auditors and inspectors general. Just as significant is the promise of near-commercial release timing, since public sector teams have often worked with tools a generation behind the private sector. If agencies get current models inside a compliant wrapper, the gap between government and industry AI use could narrow quickly.
🛍️ Consumer AI
OpenAI Brings Visual Ads to ChatGPT Image Generation
OpenAI is adding a visual ad format to ChatGPT, starting with ads that appear during image generation. The image-based ads are meant to show products in use or the experiences they enable, and OpenAI says they’ll be clearly labelled, kept separate from the image being created, and won’t influence ChatGPT’s answers. Testing begins later this month in the U.S. with an initial group of advertisers. OpenAI introduced ads earlier this year to support its free and low-cost tiers and expanded them to India in August.

The bigger part of the announcement is measurement. OpenAI added integrations with Hightouch, Tealium and LiveRamp so advertisers can send conversion data from their existing systems, and now supports a long list of attribution partners including AppsFlyer, Triple Whale, Adjust and Northbeam. It’s also working with Haus, Measured and WorkMagic on geo-based experiments to measure incrementality. OpenAI shared early partner results, including DV Rockerbox data showing WeightWatchers’ attributed cost per acquisition on ChatGPT Ads was 15.3% lower than its blended paid search benchmark.
To address brand safety, OpenAI is piloting suitability evaluations with DoubleVerify and Integral Ad Science, letting them test how its standards are applied without accessing real user conversations. Qualifying advertisers can also use Negative Phrases to keep ads away from topics that conflict with their own policies.
Why it matters: The format is new, but the measurement push is the real signal. Advertisers won’t shift meaningful budget to a platform they can’t track with their existing tools, so OpenAI is plugging ChatGPT into the attribution stack marketers already use and pointedly benchmarking results against paid search. That’s the budget it’s after. Placing the first visual ads inside image generation also makes sense commercially, since people creating images are often already imagining a product, a space or an outcome. The open question is brand safety in a channel where context is a private conversation rather than a page, which is why independent verification without exposing user data could matter as much as the ads themselves.
Pinterest Turns Beauty Pins Into Salon-Ready Guides
Pinterest has introduced Beauty Guides, a feature that uses AI to turn saved hair and nail Pins into practical plans users can bring to an appointment. Powered by Pinterest Intelligence, the company’s visual and generative AI technology, the guides translate a look into the terms stylists and nail artists actually use, like “balayage,” “root melt” or “almond nails.”

Tapping “Get the Guide” on a Pin shows what to ask for to recreate the look, along with how long it takes, a price range and maintenance needs, such as how often it will require trims, toning or touch-ups. At launch, the feature covers Pins related to hairstyle, hair colour, nail shape and finish.
Why it matters: This is AI applied to something people already do rather than a new behaviour they have to learn. Bringing a photo to the salon is common, but the gap between “I want this” and knowing what to ask for, what it costs and what upkeep it needs is where expectations break down. By closing that gap, Pinterest moves from being where inspiration is saved to where decisions get made, a more valuable position for a platform built on discovery. It’s a useful reminder that some of the most practical AI features aren’t assistants or agents, but quiet translations between how people describe what they want and how professionals deliver it.
Keep ahead of the curve – join our community today!
Follow us for the latest discoveries, innovations, and discussions that shape the world of artificial intelligence.
