tech

Perplexity splits AI inference between PCs and cloud to cut costs

Perplexity AI announced a platform at Computex that dynamically routes AI inference between PCs and cloud servers in real time, acting as an “air-traffic controller” for AI tasks. The chip-agnostic system targets the cost crisis of centralised inference as Perplexity’s revenue hits $500 million.

Perplexity splits AI inference between PCs and cloud to cut costs

TL;DR

  • Perplexity AI developed a platform that dynamically splits AI workloads between PCs and cloud servers.
  • The system acts as an "air-traffic controller for AI tasks" to reduce inference costs.
  • Simple tasks run locally on PCs, while complex tasks are routed to cloud servers.
  • This hybrid approach offloads inference work to existing PCs, reducing strain on data centers.
  • The platform is "chip agnostic" and works with processors from Intel and Nvidia.
  • Perplexity's revenue grew fivefold to $500 million, with significant growth per employee added.
  • By using user hardware, Perplexity can reduce marginal cost per query and improve response latency.