Perplexity has released Photon, an in-house Rust-based retrieval and ranking engine that now handles all production traffic, replacing a previously forked open-source component. The new system enables a "Fast Search" mode in the Perplexity Search API, offering significantly lower latency for agentic workflows.

  • Production p99 latency dropped from approximately 800 ms to 65 ms, with Fast Search reporting 160 ms at p50 and 230 ms at p95.
  • Photon operates on about 20% fewer serving machines while storing 2.5x more data per document compared to the previous setup.
  • Fast Search reduces estimated agent task costs by roughly 68% (from $187.60 to $59.73 across 3,554 tasks) but lowers relevance scores on internal benchmarks.
  • The engine uses adaptive posting lists, budgeted traversal, and batched async reads via io_uring to minimize disk I/O and cache misses.
  • Fast Search is priced at $1 per 1,000 requests and is recommended for day-to-day agent loops, while the default preset remains better for hard or ambiguous queries.

Photon allows Perplexity to scale its search infrastructure more efficiently and provides developers with a cost-effective, low-latency option for building AI agents.