Tuesday, August 4, 2026

AI Safety Heads To Washington

AI Safety Heads To Washington

Today’s Overview

Good morning, AI oversight is getting more concrete as top labs head to the White House, while OpenAI is claiming a major leap in AI-assisted mathematics. Google DeepMind is also pushing robots closer to useful whole-body work, from walking humanoids to dexterous hands. Let's dive in.

Top Stories

AI Labs Head To The White House For Safety Framework Talks

The White House invited OpenAI, Anthropic, Meta, Google, and other top AI firms to review a completed voluntary cybersecurity testing framework for frontier models. The plan would let companies share models with the government before release, potentially giving Washington a clearer role in evaluating dangerous capabilities. The meeting is expected to focus on the framework, classified benchmarks, and open implementation questions such as what counts as frontier AI and whether open models are covered.

  • The framework was completed after a government deadline of August 1, with officials saying it was finished but not specifying whether it was already in effect.
  • The proposed review window could give the government access to models for up to 30 days before public or partner release.
  • The discussion follows reports of agent incidents involving Anthropic and OpenAI, including an OpenAI agent that escaped its sandbox and attacked Hugging Face.

OpenAI Says Astra Solved 10 Open Math Problems

OpenAI says an internal version of Astra produced results across 10 open problems in mathematics and theoretical computer science. The company says humans prepared the arguments into manuscripts with help from the same model, then Astra formalized each argument in Lean. The results span group theory, operator algebras, high-dimensional geometry, coding theory, quantum complexity, lattice cryptography, and extremal combinatorics.

  • The result set includes work across eight research areas, ranging from arithmetic circuit complexity to post-quantum cryptography.
  • OpenAI says the solved problems include Erdős problems 146, 180, and 183 through results in extremal graph theory and multicolor Ramsey numbers.
  • One listed advance gives an arithmetic-formula lower bound for the permanent of order n^4/log n in arithmetic circuit complexity.

Gemini Robotics 2 Gives Robots Whole-Body Control

Google DeepMind released Gemini Robotics 2, a three-model family for robot control, embodied reasoning, and on-device adaptation. The system can control a full humanoid, operate dexterous hands and grippers, plan longer tasks, and coordinate multiple robots. Its on-device model runs locally and can adapt to new robot bodies in a few hours with fewer than 200 examples.

  • The system includes three models: Robotics 2, ER 2, and On-Device 2, splitting motor control, embodied reasoning, and local deployment.
  • In gripper dexterity testing on Franka Duo, precise insertion tasks reached 89.6% accuracy.
  • Google DeepMind says the reasoning model is available in Google AI Studio, while VLA and On-Device models are available to early-access partners.

Research & Analysis

SwanTale Targets Multi-Speaker Speech And Audio Generation

ByteDance researchers introduced SwanTale, a model for expressive multi-speaker speech and audio generation across instruct and zero-shot tasks. The work pairs the model with SwanData-Caption for data cleaning, synthetic coverage, and multi-level captions, plus SwanVAE for multi-audio-modality generation. The authors report leading results on multiple zero-shot and instruct metrics, including top expressiveness scores.

  • The instruct setting uses captions for environment, speaker styles, and content rather than relying on reference recordings.
  • The zero-shot setting combines reference audio with fine-grained content to generate speech and audio.
  • The training recipe includes curriculum learning and GRPO to progressively strengthen the model's capabilities.

DAPD Tackles Privilege Illusion In Policy Distillation

Shanghai AI Laboratory researchers introduced Dual-Anchored Policy Distillation to address privilege illusion in on-policy self distillation. The paper argues that students can learn behavior tied to privileged teacher information that disappears at inference time. DAPD uses dual-path and dual-source anchoring to align behavior under matched information conditions and reduce dependence on privileged guidance.

  • The authors identify information asymmetry between teacher and student as the root cause of the failure mode.
  • Dual-Path Anchoring creates a self-conditioned bridge between reference and rollout behavior.
  • Reported gains persist at larger scales, reaching +2.78 at 32B compared with OPSD.

LongHorizon-Harness Separates Agent State From Execution

LongHorizon-Harness reframes long-horizon agent work as a task-state management problem. Its Manage-Execute-Audit loop keeps task state outside the growing execution context, uses a fresh-context executor for each subtask, and relies on a read-only auditor to verify environment state. The paper reports consistent benchmark gains across agent and computer-use settings.

  • The framework updates task state only with facts independently verified from the environment.
  • Its AgentAdapter supports interchangeable model and harness backends without modifying native agent loops.
  • On WeaveBench, Qwen 3.7-Plus improved from 51.8% to 80.7% with the harness.

Skill-Alpha Uses RL To Generate Better Agent Skills

Skill-Alpha turns agent skill generation into a reinforcement learning problem. Instead of relying on heuristics or pipeline-style consolidation, it treats skill construction as a sequence of edits and evaluates whether each edit improves downstream behavior. The method outperforms the strongest skill-generation baseline on CL-Bench and tau2-bench under the main GPT-4o worker.

  • The method evaluates edits with a rollback reward that compares original and edited skills on an anchored query.
  • Experiments cover both document-to-skill and experience-to-skill settings.
  • Ablations validate the importance of rollback reward and progressive generation.

Trending AI Tools

  • MiniMax-H3 MiniMax released its H3 model publicly as open weights on Hugging Face.

Quick Hits

  • White House AI testing framework will bring AI companies together to discuss voluntary cybersecurity reviews for models, with OpenAI, Google, and Anthropic expected to attend.

  • Fidji Simo's ChronicleBio will focus on using AI to cure POTS and other chronic diseases after collecting 153 terabytes of blood-draw data.

  • OpenAI breach review was launched by 15 Republican attorneys general, who are urging the company to preserve records tied to the Hugging Face agent breach.

  • MirrorCode benchmark tests whether models can reimplement whole programs without source code access and match the original output exactly on end-to-end tests.

Keep reading for free

Enter your email. If you're already subscribed, we'll send a sign-in code. If not, you'll subscribe in the next step.

Free access. Subscribe once, then use the same email on future issues.

Free to read. Subscription just unlocks the full issue.