OpenAI published benchmark details for its custom inference chip, Jalapeño, claiming superior efficiency and latency compared to NVIDIA GB200/GB300 systems. Simultaneously, Perplexity launched the Portable Computer on NVIDIA DGX Spark hardware, enabling fully local agent orchestration without cloud dependency.

  • OpenAI's Jalapeño delivered 1.5–1.9× more work per watt and 1.7–3.6× lower end-to-end latency than NVIDIA systems in tests.
  • The chip is rated at 700W but stayed below 550W during testing, with deployment planned by year-end.
  • Perplexity's Portable Computer runs the orchestrator LLM, subagent LLM, and agent harness locally using PPLX 27B or Qwen 3.8 27B.
  • Microsoft-led AutoSaddler reported gains of +9.0 on GAIA2 and +10.0 on Terminal-Bench 2.0 by treating the agent harness as code.
  • SWE Refactor Bench found only a 5.4% survival rate for whole-repository migration tasks across 520 runs.

These developments highlight a shift toward specialized inference hardware and local-first agent architectures that reduce reliance on cloud infrastructure.