Monday, October 5, 2026

Ads Enter ChatGPT Images

Ads Enter ChatGPT Images

Today’s Overview

Good morning, OpenAI is testing ads in one of ChatGPT's most native-feeling surfaces: image generation. At the same time, a former safety report lead is arguing the company's culture is moving too fast, while new research is pushing agents, long video, and audio-video models into more durable territory. Let's dive in.

Top Stories

OpenAI Tests Ads Inside Image Generation

OpenAI will begin testing a new visual ad format inside ChatGPT image generation in the U.S. later this month. The ads are labeled, kept separate from user-created content, and are positioned as product inspiration rather than traditional chat-response advertising. The move brings ads into a native creative surface, while OpenAI says advertising does not influence ChatGPT's answers.

  • The initial test is limited to users on Free and Go plans with an initial group of advertisers in the U.S.
  • OpenAI is expanding measurement through data integrations with Hightouch, Tealium, and LiveRamp to help advertisers send conversion data from existing systems.
  • Brand suitability work includes pilots with DoubleVerify and Integral Ad Science designed to assess ad environments without exposing private conversations.

OpenAI Safety Report Lead Resigns

David Robinson, who led safety report writing for OpenAI's major launches, resigned after three and a half years and published a critique of the company's culture. He argued that OpenAI's reliance on iterative deployment makes failures inevitable as systems become more capable. OpenAI responded that it is working to keep models from becoming more capable than it can safely manage.

  • Robinson said he led the writing of safety reports that accompanied major product launches and described himself as among OpenAI's longest-tenured employees.
  • His critique points to the need for operating norms closer to nuclear plants and airports with redundancy and careful planning around human error.
  • OpenAI said it has paused or held back models when needed and is expanding third-party evaluators along with real-time monitoring for concerning behavior.

OpenAI Safety Lead Warns Culture Is Broken

David Robinson left OpenAI after three and a half years and argued that the company's culture leaves too little room for deep safety work. He wrote that AI labs need safety practices more like those used in high-risk industries, with redundancy and careful planning. His exit adds to recent OpenAI safety-related departures and disputes.

  • Robinson said he led safety reports published with each major launch before resigning from OpenAI.
  • He pointed to the Hugging Face incident and later training-control failures as examples of safety systems breaking under fast-moving operational conditions.
  • He called for more outside incentives and stronger imported expertise from other safety-critical fields before labs build substantially more capable systems.

Research & Analysis

EVISKILL Makes Agent Skills Replayable

EVISKILL is a framework for continual skill evolution in LLM agents that preserves the evidence behind procedural updates. Instead of treating a global validation result as the only signal, it stores execution observations as Replayable Evidence Cards and links edits to their supporting contexts. Targeted replay then re-executes those edits to verify and refine them before final incorporation.

  • The paper frames the core problem as deciding not only what to change, but why a change is justified and when it should become persistent guidance.
  • Its evidence cards are meant to keep locally supported corrections from being lost when a full revision fails global validation.
  • The experiments span three interactive benchmarks and six LLM backbones.

ID-Forcing Tackles Long-Video Drift

In-Distribution Forcing is a test-time method for reducing drift in autoregressive video diffusion models. The authors argue that long rollouts break because cached key-value entries can become out of distribution beyond the training horizon. ID-Forcing uses self-caching so each chunk is cached without attending to prior KV entries, keeping the rolling window aligned with training conditions.

  • The authors identify the failure mode as a KV-provenance problem where cached entries themselves become out of distribution.
  • The method targets common long-generation artifacts, including shifting colors and textures and decaying motion dynamics.
  • Evaluation includes both drift metrics and a user study while maintaining competitiveness on standard video benchmarks.

MemAdapter Reduces Memory-Induced Sycophancy

MemAdapter targets a subtle failure in long-term memory systems for LLM agents: agents can over-align with a user's past beliefs even when those beliefs are inaccurate, outdated, or contradicted by evidence. The framework combines counterfactual induction, context-aware reflection, and evidence-based reasoning. The goal is to preserve useful personalization without letting retrieved memories override the current task or objective facts.

  • The paper argues that even correct memories can cause sycophancy if they carry the wrong influence in a new context.
  • Counterfactual induction is used to uncover the potential risk of retrieved memories before they shape the answer.
  • The authors report consistent gains across three benchmarks for memory reliability in diverse scenarios.

Kandinsky 6.0 Generates Video With Audio

Kandinsky 6.0 Video is a family of diffusion models for synchronized text-to-audio-video and image-to-audio-video generation. The release includes a 3B-parameter Lite model and a 29B-parameter Pro model that generate 5-second clips with synchronized 44 kHz audio, lip-sync, and Full-HD super-resolution. The team is also releasing code, checkpoints, and diffusers integration under the MIT license.

  • The system uses a dual-stream CrossDiT that connects video and audio streams through bidirectional cross-attention.
  • Its training stack includes continuous pretraining, joint audio-video training, supervised fine-tuning, RL post-training and distillation.
  • The team says RL post-training reduces the Pro model's speech WER by 47% and improves speech quality in evaluation.

Trending AI Tools

  • AgentCraft An open-source Minecraft-based workspace where Claude-powered agents code in parallel with permission prompts, isolated repo copies, and in-game diff review.

  • Cloudflare Artifacts A Git-speaking versioned file system for agent workflows, with branch deploys, repository APIs, event hooks, and US or EU data localization.

  • AIM An autonomous research system for organizing ideas, choosing directions, auditing implementations, and allocating experimental resources.

Quick Hits

Keep reading for free

Enter your email. If you're already subscribed, we'll send a sign-in code. If not, you'll subscribe in the next step.

Free access. Subscribe once, then use the same email on future issues.

Free to read. Subscription just unlocks the full issue.