Wednesday, August 26, 2026

Perplexity’s Local Agent Bet

Perplexity’s Local Agent Bet

Today’s Overview

Good morning, AI agents are getting a serious local-first push, with Perplexity and NVIDIA moving agent work onto machines users control. OpenAI is also showing off its own inference silicon, while a new physics startup is betting that models can replace parts of the lab loop. Let’s dive in.

Top Stories

Perplexity and NVIDIA launch a local AI agent

Perplexity’s Portable Computer brings its agentic Computer platform onto hardware users already own. The model, files, and work can stay on local machines, with no billing credits for local tasks and step-by-step permission before any cloud escalation. It is available for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support planned for September.

  • The product packages the local models, agent harness, inference engine, tools, app connectors, and security sandbox into one bundled system rather than asking users to assemble a local AI stack themselves.
  • Perplexity’s internal Local Knowledge Work Bench includes 53 tasks spanning deep research, financial analysis, and document creation.
  • The launch requires an RTX GPU with at least 24GB of VRAM which VentureBeat describes as roughly a GeForce RTX 3090 or newer.

OpenAI shows off its Jalapeño inference chip

OpenAI revealed benchmark results for Jalapeño, its first custom chip built specifically for inference rather than training. The company says the chip was co-built with Broadcom, designed partly with help from OpenAI models, and targets lower cost per response than leading NVIDIA chips. OpenAI says the chip is for its own infrastructure and should make ChatGPT, Codex, and other products faster and more resilient as demand grows.

  • OpenAI evaluated Jalapeño on InferenceX, a public SemiAnalysis benchmark that measures full request serving across throughput, power efficiency, and latency.
  • Across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5, Jalapeño delivered 1.5x to 1.9x more AI work per watt at peak throughput.
  • The chip is rated at 700 watts, but OpenAI says measured sustained power stayed at or below 550 watts on the tested workloads.

Accelerated Understanding targets physics world models

Caltech professor Anima Anandkumar and engineer Benedikt Jenik launched Accelerated Understanding to train AI that forecasts how physical systems evolve. Instead of predicting words, the company is building physics-grounded models that follow events through 3D space and time. Its early enterprise targets include chip materials, extreme-weather prediction, and robotics.

  • The company argues that real-world experiments reveal what happened, but often fail to provide directional feedback for how to improve the next design.
  • Its approach favors broad physical models over narrow surrogates because industrial datasets are often too sparse to support a standalone model for each use case.
  • Accelerated Understanding says its architecture has reached more than 5 trillion context at inference without patching or sub-sampling.

Research & Analysis

WeChat’s multimodal embedding model hits a new high

Tencent’s WeMM-Embedding report introduces a family of universal multimodal embedding models for text, images, videos, visual documents, and interleaved inputs. The models come in 2B, 4B, and 9B variants and use a two-stage training process. The 9B model reaches a reported overall score of 80.6, and the system is already deployed across WeChat recommendation and search products.

  • The report frames universal multimodal embeddings as a foundation for retrieval and recommendation as well as classification and agentic systems.
  • Training combines large-scale multimodal alignment with refinement based on curated data, fine-grained relevance supervision, and cross-scale knowledge transfer across model sizes.
  • The released Hugging Face page lists model weights for 2B, 4B, and 9B variants, plus code intended to support future research.

Outer Bio launches living-skin AI discovery loop

Outer Bio emerged with Yuna, a platform that keeps full-thickness ex vivo human skin alive for four weeks on a 3D-printed scaffold. The company feeds experimental skin data into an AI system that proposes compounds for specific skin processes, then uses each experiment to refine the next prediction. CEO Michael Polansky says the loop has cut lead discovery from two leads in roughly 18 months to a new candidate about every six weeks.

  • Outer Bio says topical skincare has produced exactly one FDA-approved mechanism for visible skin aging in thirty years: retinoids.
  • Yuna keeps epidermis, dermis, multiple resident cell types, and resident immune signaling intact across four weeks of measurement.
  • In its preprint, the company tested psoriasis-like inflammation, aged donor tissue, and UVB injury using interventions with known clinical answers.

MobilePA-Bench tests agents on real mobile planning

MobilePA-Bench is an interactive benchmark for mobile planning agents that goes beyond GUI manipulation and static API matching. It runs in an executable sandbox with live application databases and structured feedback across realistic mobile tools. The paper says frontier LLMs remain unreliable when strict tool ordering, permission limits, and runtime errors enter the loop.

  • The benchmark spans 13 functional domains to better reflect the variety of tasks a mobile copilot may face.
  • Its tool environment includes 212 realistic mobile tools rather than relying on offline function-call matching alone.
  • The benchmark evaluates sub-agent collaboration, memory usage, and composite skill invocation as advanced dimensions beyond basic tool use.

Ox Alpha sends model watchers hunting for clues

Ox Alpha appeared on OpenRouter as a free anonymous model with a 1M-token context window and multimodal input. Early testing made it interesting for coding, sustained agent work, and production workloads, but the creator has not been confirmed. Speculation has centered on Chinese labs such as Zhipu AI, while Microsoft’s MAI family has also been floated as a possible source.

  • OpenRouter describes Ox Alpha as a stealth model operated by a third-party provider that chose to stay anonymous during preview.
  • The model was presented as designed for coding, sustained agentic work, and production workload rather than only chat-style prompting.
  • Public speculation remains unsettled, with discussion pointing in different directions, including GLM models, Microsoft MAI, and China-linked forensics without confirmed attribution.

Trending AI Tools

  • Portable Computer A fully local Perplexity agent for Linux subscribers that runs the orchestrator, subagent, and tool harness on the user’s own hardware.

  • Portable Computer on DGX Spark A local-first Perplexity agent built with NVIDIA for DGX Spark, with optional cloud model calls that require user approval.

  • Thomson Thomson Reuters’ first homegrown legal AI model, adapted from Qwen and trained on decades of company content.

Quick Hits

  • Perplexity funding talks could bring in NVIDIA at a valuation as high as $30B, with reported annual revenue now above $750M.

  • OpenAI data center exit adds Chris Malone, its head of data centers, to a recent run of executive departures.

  • Hugging Face sale report says the company is exploring a deal at $13B or more, up from its $4.5B valuation in 2023.

  • Anthropic expands Mythos 5 to more cyber defenders beyond Project Glasswing, while bankers reportedly floated a future $2T IPO valuation.

  • Keenable’s agent web index emerged from stealth with a claimed 100B-plus document index, an API for AI labs, and a planned evidence-combining query language.

Keep reading for free

Enter your email. If you're already subscribed, we'll send a sign-in code. If not, you'll subscribe in the next step.

Free access. Subscribe once, then use the same email on future issues.

Free to read. Subscription just unlocks the full issue.