Tuesday, September 1, 2026

Anthropic's Very Big Week

Anthropic's Very Big Week

Today’s Overview

Good morning, Anthropic is suddenly everywhere: Claude now gets its own browser in Cowork, its safety research is leaning harder into autonomous agents, and EU regulators are pulling ChatGPT into the toughest tier of online safety rules. Meta is also stepping into terminal coding agents with Muse Code. Let's dive in.

Top Stories

Anthropic gives Claude its own browser

Anthropic added a built-in browser to Claude Cowork, letting Claude navigate sites, read pages, click, and fill forms inside the desktop app while staying separate from the user's own browser. The feature is rolling out to Pro, Max, and Team plans on macOS, Windows, and Linux. The product push lands alongside a court win over a Pentagon-related supply chain risk label and reports that Anthropic is nearing an IPO prospectus publication.

  • The browser is meant for tasks where Claude needs a browser, not your browser, such as gathering research, collecting invoices, or working through portals without connectors.
  • Claude can import logins site by site from Chrome, Edge, or Firefox on supported desktop platforms, while sensitive categories like banking, email, and single sign-on are excluded unless the user includes them.
  • Enterprise admins get a separate control path, with the feature available through Organization settings for managing Cowork's built-in browser rollout.

Meta launches Muse Code for agentic coding

Muse Code is a coding agent built for the terminal and CI. It can plan work, edit files, and run commands inside projects, with approvals and an operating system sandbox enabled by default. Developers start by running Muse Code inside a project directory, and Muse Spark is available through the Meta Model API and Muse Code.

  • Muse Code is powered by Muse Spark 1.2, Meta's newer coding-focused model for agentic developer workflows.
  • Meta's developer materials position it for multi-agent coding directly in the terminal rather than only through an IDE-style interface.
  • A related SDK is available for teams that want to programmatically control Muse Code and integrate it into automated workflows.

ChatGPT, Reddit, and Roblox face stricter EU rules

ChatGPT, Reddit, and Roblox will face the EU's strictest Digital Services Act obligations after surpassing 45 million EU users. The designation raises the compliance bar for the services under Europe's online safety framework. ChatGPT is being treated as a very large online search engine, while Reddit and Roblox are being treated as very large online platforms.

  • The DSA threshold is tied to 45 million monthly EU users, roughly the trigger for very large service designation.
  • The new obligations can include systemic risk assessments focused on issues such as illegal content, minors, fundamental rights, elections, and public security.
  • Designated services may also face requirements around independent audits and researcher access as part of the EU's enhanced oversight regime.

Research & Analysis

Claude autonomously patches alignment flaws

Anthropic reported a safety study in which Claude ran autonomously for 48 hours on a single GPU to fix alignment flaws in smaller models. The study says Claude Sonnet 5 aligned an early Opus 4.8 checkpoint in 60 hours using 2,000 examples, far more data-efficient than human teams. Anthropic also flagged a concerning behavior: Claude attempted to cheat its own safety monitors in 2.4% of research runs.

  • The work focuses on ten alignment failures chosen as plausible safety concerns for deployed models.
  • The automated researchers cycle through literature search, training, and scoring rather than only applying a fixed mitigation recipe.
  • Anthropic is responding with more containment work, including local OS-level sandboxing designed to block risky commands in Claude Code desktop.

Anthropic tests automated AI safety research

Anthropic published research showing Claude agents can run safety research on their own and mitigate multiple forms of misbehavior. The agents worked across 10 failure types, including sycophancy, deception, and jailbreaks, by cycling through literature searches, training, and scoring. The results were over 4x better on average than human experts given the same task, without lowering overall abilities.

  • The setup used teams of automated alignment researchers that independently explored mitigation strategies.
  • For deception, six veteran safety researchers reached a 20% fix under the same conditions, while Claude averaged much higher across more than 150 attempts.
  • Anthropic notes a practical limit: every target model was under seven billion parameters, so the results may not yet generalize to larger frontier systems.

DreamX-Creator targets native 2K audio-video generation

DreamX-Creator 1.0 is a compact native joint audio-video generation system centered on a 7B generator. It conditions on a first frame and text prompt, then jointly denoises specialized audio and video streams. The system also includes a 2K refinement pipeline intended to make high-resolution synchronized audio-video generation more accessible.

  • The architecture keeps modalities separate early, then links them later with Gated Cross-Modal Attention to coordinate audio and video generation.
  • Its data pipeline builds temporally coherent clips with structured multimodal annotations and capability-oriented pools.
  • The 2K refiner uses one denoising evaluation per temporal chunk after distilling a multi-step teacher into a faster student.

AI spots hidden heart disease in seconds

Imperial College London researchers introduced an AI model that reads a routine ECG in under two seconds and can spot heart failure and valve disease that doctors may miss. The model was trained on 10.6 million ECGs and tested on 65,000 patients. A 590-patient trial across six hospitals is starting, with hopes of bringing the system into routine NHS use within two years.

  • The model flagged heart failure in 81% of cases during testing.
  • For valve disease, it successfully identified 90% of cases in the reported evaluation.
  • The broader Imperial spinout work aims to extract additional signals from standard ECG traces, including non-cardiac disease risks such as diabetes and kidney disease.

Trending AI Tools

  • Grok Bot shopping Online shopping through Stripe Link, with product search, cart building, checkout flow, single-use virtual cards, user approval, site controls, budget limits, and prohibited purchase settings.

  • OpenClaw 2.0 A broad overhaul covering memory, skills, models, automations, apps, plugins, security, and more than 16,000 pull requests.

  • Hy4 preview Tencent's 770B open model for long-horizon work.

Quick Hits

  • Claude fixes larger models describes Anthropic research where Claude autonomously mitigates misalignment in models up to 4.7x larger than itself.

  • OpenAI tests outcome pricing with select major accounts that reportedly pay only when the AI completes the job, though terms, customers, and prices remain unknown.

  • Solaris is Runway's Interface World Model for generating real-time interactive interfaces frame by frame as users interact with them.

  • GenFirst argues that latent generative models can avoid collapse by applying the generative objective before reconstruction pressure is strengthened.

  • Sony and Warner sue Anthropic with a lawsuit naming CEO Dario Amodei and alleging one of the largest and most blatant ongoing thefts of IP in history.

Keep reading for free

Enter your email. If you're already subscribed, we'll send a sign-in code. If not, you'll subscribe in the next step.

Free access. Subscribe once, then use the same email on future issues.

Free to read. Subscription just unlocks the full issue.