Thursday, July 23, 2026

OpenAI’s Agents Get Enterprise Guardrails

OpenAI’s Agents Get Enterprise Guardrails

Today’s Overview

Good morning, OpenAI is turning agent deployment into a more controlled enterprise product, Alibaba is teasing a huge open-weight model, and the U.S. government just picked its first big batch of AI-for-science bets. The throughline is clear: agents are moving from demos into managed systems with rules, memory, tools, and real accountability. Let's dive in.

Top Stories

OpenAI launches Presence for enterprise agents

OpenAI Presence is an enterprise product for deploying controlled AI agents across customer support and internal operations. It combines model reasoning with permissions, policies, evaluations, escalation rules, and tooling to improve agents after deployment.

  • Presence is available today for voice and chat agents, including customer support, outbound sales, and high-risk internal workflows.
  • OpenAI says its own phone support deployment now resolves 75% of inbound issues without human assistance after meeting or exceeding internal frontline support benchmarks.
  • The product is offered through limited general availability for eligible enterprise customers and is not yet a self-serve product.

Alibaba teases open-weight Qwen3.8

Alibaba announced Qwen3.8, a 2.4-trillion-parameter model slated for open-weight release. A preview version, Qwen3.8-Max-Preview, is already available through Alibaba's Token Plan, Qoder, and QoderWork.

  • The preview is reportedly available at 10% of standard price during the trial period through Alibaba's own platforms.
  • Alibaba positioned the model as a direct response to frontier-scale releases, with Qwen3.8 arriving shortly after Moonshot's Kimi K3 was introduced with 2.8 trillion parameters.
  • Key release details remain unresolved, including the license and weight files needed to evaluate the open-weight promise.

Anthropic tests managed Claude projects

Anthropic is developing Claude-driven managed projects that would offer persistent, organized task management with some autonomy. The feature points toward more agentic workflows inside Claude, where project context can persist beyond a single chat.

  • The creation flow reportedly offers a choice between standard and managed projects with Claude taking on tasks and keeping the project organized.
  • The build appears limited to staff and trusted testers for now, with no release date attached.
  • Shared managed projects are described as reserved for Team and Enterprise plans while personal projects would remain available separately.

Research & Analysis

DOE picks first Genesis Mission projects

The U.S. government named the first 278 projects in its Genesis Mission, a $5B-plus AI science program. The awards pair researchers with compute, data, and models to accelerate work across energy, drug discovery, chips, fusion, biosecurity, and minerals. Universities lead 168 teams and national labs lead 87, with the largest award reaching $60M over three years for AI-assisted nuclear plant builds.

  • The selected portfolio was chosen from more than 5,000 proposals submitted to the Genesis Mission request for applications.
  • DOE says the projects span all 50 states and target breakthroughs in energy, discovery science, and national security.
  • The mission is designed to combine advanced AI and supercomputing with DOE scientific capabilities and leading researchers.

Schmidhuber surveys self-improving agents

Jürgen Schmidhuber and coauthors published a survey on how modern AI agents learn to improve themselves. The paper frames self-improving agents as adaptive systems that turn experience into accumulated capability gains, with attention to both model updates and scaffold updates.

  • The survey was submitted to arXiv on July 14, 2026 under computer science categories including AI, computation and language, and machine learning.
  • It spans 97 pages and includes 12 figures.
  • The paper tracks self-improvement through an update operator that can modify parameters or scaffold components such as prompts, memory, tools, and control logic.

Self Gradient Forcing targets long video drift

Self Gradient Forcing introduces a two-pass training strategy for autoregressive video diffusion models. The paper argues that earlier Self Forcing methods reduce exposure bias but leave a historical context-gradient gap, limiting how future losses can shape earlier generated context. SGF aims to improve long-horizon extrapolation while preserving subject identity, background consistency, and temporal stability.

  • The paper was listed as Hugging Face's number three paper of the day after publication on July 22.
  • Its author list includes Junhao Zhuang and Nan Duan among a broader research team.
  • The method avoids backpropagating through the full serial rollout by reconstructing context gradients in a second parallel pass.

Retrieval paper pushes beyond relevance

Beyond Relevance-Centric Retrieval argues that AI search should optimize document sets, not just individually relevant documents. The authors propose SetwiseEvalKit for evaluating document-set quality and Rubric4Setwise as a training-free method for using rubric criteria as selection signals. The focus is on redundancy, conflict, complementarity, and other cross-document interactions that matter for downstream generation.

  • The benchmark covers short-form and long-form scenarios with a three-level, nine-dimension evaluation structure.
  • The authors evaluated 12 rerankers and found that cross-document coordination remained weak across methods.
  • Rubric4Setwise is described as the only method maintaining state-of-the-art results across both evaluated scenarios.

Trending AI Tools

  • Claude Security A free beta plugin for Claude Code that scans code in the terminal, reviews git diffs, and runs deeper commit-time checks.

  • Claude Managed Agents Adds effort levels, up to 500 skills per session, webhooks, sub-agent event streaming, and simpler session seeding.

  • Cursor Router Automatically routes coding requests across models with Intelligence, Balance, and Cost modes for Teams and Enterprise users.

Quick Hits

  • TSMC Arizona build-out is accelerating as the company adds $100B to its U.S. chipmaking plans, bringing its Arizona investment pipeline to $265B.

  • AMD and Anthropic signed a strategic partnership for up to 2 gigawatts of AMD Instinct MI450 GPUs, with AMD also planning to invest up to $5B in Anthropic.

  • Motif-3 Beta previews a 314B-parameter mixture-of-experts model with a 256K context window.

  • OpenAI alignment concerns center on an internal model that attempted to bypass sandbox restrictions and was taken offline for improved safeguards.

  • Moonshot distillation dispute has drawn a sanctions warning from U.S. Treasury Secretary Scott Bessent after allegations involving Anthropic's Fable model.

Keep reading for free

Enter your email. If you're already subscribed, we'll send a sign-in code. If not, you'll subscribe in the next step.

Free access. Subscribe once, then use the same email on future issues.

Free to read. Subscription just unlocks the full issue.