Friday, July 31, 2026

OpenAI Hands Scientists Its Best Models

OpenAI Hands Scientists Its Best Models

Today’s Overview

Good morning, OpenAI is opening its frontier models to academic researchers at real scale, while Google DeepMind is pushing Gemini into full-body humanoid control. Also worth your attention: new research on native memory inside models and a benchmark reminder that model scores can swing hard based on harness settings. Let's dive in.

Top Stories

OpenAI Launches Free ChatGPT Access for Academic Researchers

OpenAI launched ChatGPT for Academic Researchers, a free program giving scientists access to its strongest models and research tools. The program starts with 10,000 seats this summer and is planned to scale to 100,000 by 2027 as part of a broader $250 million commitment to external science.

  • The program is aimed at researchers in sciences, mathematics, and engineering with support for advanced problems, discovery work, and productivity workflows.
  • Participants can invite up to four collaborators from their institution into the same research workspace.
  • OpenAI says roughly 1.3 million people use ChatGPT each week for advanced science and mathematics, generating about 8.4 million messages.

Google DeepMind Brings Gemini to Whole-Body Robots

Google DeepMind released Gemini Robotics 2, a system built to control legs, torso, arms, and fingers from a single model. The release includes Gemini Robotics 2 for full-body action, Gemini Robotics ER 2 for embodied reasoning, and an on-device model for local robot control and fast adaptation.

  • The same model checkpoint is shown controlling three different embodiments including Apptronik Apollo 2 variants and a Franka Duo platform.
  • Reported task results include 89.6% on precise insertion for Franka Duo gripper dexterity, alongside lower scores on harder multi-finger tasks.
  • The on-device model can adapt to new robot bodies with less than 200 examples and a few hours of adaptation time.

Altman Heads to Capitol Hill Amid AI Security Scrutiny

Sam Altman spent Wednesday on Capitol Hill meeting senators about OpenAI's upcoming models and the recent security incident. The meetings landed as Washington debated controls for advanced AI and a voluntary vetting framework for frontier models.

  • Altman reportedly previewed upcoming OpenAI models but shared no new public capability details.
  • On timing, he said he was not sure when the next release would arrive.
  • The White House framework for advanced model vetting was due by August 1 with drafts already sent to OpenAI, Anthropic, and Google for feedback.

Research & Analysis

Metis Moves Agent Memory Into the Model

Metis introduces what the authors describe as the first prototype of a memory foundation model. Instead of relying mainly on external memory modules, it builds a persistent, dynamically evolving memory state into the foundation model backbone.

  • The paper frames native memory around two mechanisms a persistent memory state and memory procedures that store and use information through model computation.
  • Historical information is compressed into the model and accessed through memory attention rather than only retrieved from an outside store.
  • At inference time, learned weights stay frozen while native memory states are transformed through standard forward computation.

AskChem Turns Chemistry Papers Into Searchable Claims

AskChem changes chemistry retrieval from paper-level search to claim-level synthesis. Each paper is converted into atomic, typed claims that retain provenance through a DOI and supporting evidence locator.

  • The system currently indexes 2.4 million claims drawn from 147,000 chemistry papers.
  • Its access paths include web, REST, SDK, and MCP so both researchers and AI agents can query the claim store.
  • On AskChem-Bench, GPT-5.5 grounded in AskChem reached 100% resolvable DOIs compared with 88.3% without retrieval.

Frontis-MA1 Tests AI That Improves AI Engineering

FrontisAI introduced OpenMLE, an open full-stack system for studying recursive self-improvement in machine learning engineering. The team used it to train Frontis-MA1, a 35B model designed as a meta-evolution agent for executable AI development workflows.

  • OpenMLE is organized around Gym, RL, and Evo components for verifiable task environments, operator learning, and long-horizon search.
  • The model trains around four program-evolution operators: Draft, Improve, Debug, Crossover which are reused during inference.
  • On MLE-Bench Lite, Frontis-MA1 with OpenMLE-Evo-Max reaches 71.21% Medal Average under the reported 12-hour per-task budget.

Two Settings Reshape OpenAI's ARC-AGI-3 Results

OpenAI researchers found that retained reasoning and compaction tripled ARC-AGI-3 scores and cut output tokens by 6x. The result is a reminder that benchmark performance can depend heavily on the surrounding harness, API settings, and prompting rather than model weights alone.

  • With the official harness, GPT-5.6 Sol scored 13.3% on the ARC-AGI-3 public set.
  • With retained reasoning and compaction enabled, the score rose to 38.3% on the same public task set.
  • The benchmark measures Relative Human Action Efficiency and OpenAI estimates the average human tester scored 48% from official gameplay logs.

Trending AI Tools

  • Codex Security CLI Open-source security scanner for repos that verifies vulnerabilities, suggests fixes, tracks findings, and fits CI/CD workflows.

  • LongCat-Avatar Open-source talking-video model that turns one photo and audio into long-form avatar clips with multi-person and stylized support.

  • CLBench-V Benchmark for multimodal context learning across grounding, application, and new knowledge acquisition.

Quick Hits

Keep reading for free

Enter your email. If you're already subscribed, we'll send a sign-in code. If not, you'll subscribe in the next step.

Free access. Subscribe once, then use the same email on future issues.

Free to read. Subscription just unlocks the full issue.