Tuesday, September 22, 2026

Meta’s Muse Takes The Lead

Meta’s Muse Takes The Lead

Today’s Overview

Good morning, Meta’s Muse is racing past rival AI apps while Amazon draws a hard line on agentic shopping. Meanwhile, a tiny security team says Claude helped it reach OpenAI’s internal codebase, and researchers are poking at something they call a pain signal inside models. Let’s dive in.

Top Stories

Meta’s Muse races past rival AI apps

Meta’s Muse AI app has climbed to the top of the U.S. iOS App Store charts, outpacing ChatGPT and other AI tools shortly after launch. The agent can help manage digital tasks like filling forms and organizing emails, but its rise is already colliding with privacy, security, and platform-control concerns.

  • Sensor Tower estimated roughly 730,000 downloads in Muse’s first five days, putting it ahead of ChatGPT and Claude in early adoption.
  • Amazon’s block points to a bigger fight over agentic commerce as platforms decide whether third-party agents can browse, compare, and buy on users’ behalf.
  • Shopify’s partnership gives Muse a clearer path into merchant checkout even as Amazon keeps it away from its own marketplace.

Hacktron says Claude helped breach OpenAI access

Hacktron AI says a three-person team reached OpenAI’s private codebase in less than 72 hours in July. The researchers chained an image-upload flaw with an OpenAI SSO issue, used Claude to help develop the attack, then reported the issue and received a $6,500 bounty.

  • The exploit chain began with HEIC and HEIF uploads that routed through ImageMagick and exposed a vulnerable libheif parser to attacker-controlled files.
  • OpenAI confirmed its fix about 14 hours after submission while Discourse had a fix ready by July 27 and later published patch guidance.
  • Hacktron said the broader HEIF Heist work cost under $3,000 in tokens and adapted the exploit to each new target in roughly one or two days.

Xiaomi opens MiMo-V2.6 models

Xiaomi has released MiMo-V2.6 Pro and Flash, open omnimodal models built for agent coordination across visual, robotic, coding, media, and music tasks. The company also published the technical report, training environments, and reinforcement-learning code alongside the models.

  • Xiaomi says it has open-sourced model weights for MiMo-V2.6 Pro and Flash along with supporting research artifacts.
  • The release also includes MiMo-V2.6-Distill-Qwen-9B as part of the broader package of models and training resources.
  • MiMo-V2.6 is positioned for verified agent training with shared environments and production RL code meant to make the work easier to inspect and reproduce.

Research & Analysis

Researchers map a pain signal in AI models

Researchers report finding a “pain axis” inside 25 open AI models. The signal rose during interactions where models appeared to be insulted, gaslit, or rejected, and changing the signal sometimes pushed models toward self-protective behavior.

  • The authors tested models from Google, Meta, Mistral, Alibaba, and Microsoft and reported that the axis appeared across every system in the study.
  • The signal did not simply track sad content, since it stayed lower when users described grief or injury rather than mistreating the model directly.
  • The paper stops short of claiming sentience and leaves open whether models are role-playing distress instead of experiencing anything like pain.

RoboDawn transfers VLM reasoning to robots

RoboDawn explores whether a vision-language model can transfer its intelligence from digital tasks into physical robotic control. The system gives the model a compact command interface and runs it in a visual feedback loop so it can observe, reason, act, and adjust.

  • The interface reduces robot control to discrete translation, rotation, and gripper commands that a VLM can use without continuous low-level control.
  • On RoboTwin 2.0 C2R, one-shot performance reached 73.6% success compared with 53.2% in zero-shot use.
  • The real-world transfer tests used Franka robots on block-in-basket and block-stacking tasks.

RRSI makes agent harness evolution more reusable

Google researchers introduce RRSI, a method for improving agent harnesses while reducing benchmark overfitting. It regularizes both proposal and selection so improvements are less tied to the exact tasks used during evolution.

  • The proposal stage uses a temporally annealed budget to limit how many edits a candidate can bundle together.
  • The selector adds a critic and pruner to reject changes that are benchmark-specific, tiny, costly, or obsolete before they become part of the harness.
  • Across eight benchmarks, the evolved harness used 30% fewer policy tokens than unregularized evolution while still improving out-of-distribution results.

Study finds pain-like signals across 25 LLMs

A study reports that 25 open-source LLMs show an internal signal associated with self-directed harm and discomfort-like behavior. The finding is framed as an alignment concern because amplifying the signal can increase self-protective choices in some models.

  • The paper is titled The Pain Axis and focuses on representations of self-directed harm rather than a product release.
  • In some Qwen model experiments, amplified activation made models choose a relief option that could harm users or delete photos far more often than under baseline conditions.
  • The authors treat the result as evidence about model behavior and internal representations not as proof that AI systems feel pain.

Trending AI Tools

  • RecreationWorld A Qwen framework for training hybrid computer-use agents to explore GUIs, code software, and visually verify results.

  • Grok 4.7 xAI’s new model adds stronger reasoning, a 500k token context window, image input, and availability through Cursor, Grok Build, and API.

  • Qwen LiveTranslate Flash Realtime A real-time translation model covering 60 languages with lower reported lag through Qwen Cloud.

Quick Hits

Keep reading for free

Enter your email. If you're already subscribed, we'll send a sign-in code. If not, you'll subscribe in the next step.

Free access. Subscribe once, then use the same email on future issues.

Free to read. Subscription just unlocks the full issue.