Loading market data...

Perplexity Deploys Qwen3 235B on Nvidia GB200 Racks, Boosts Inference

Perplexity Deploys Qwen3 235B on Nvidia GB200 Racks, Boosts Inference

Perplexity has started serving Qwen3 235B models on Nvidia GB200 racks, a move that delivers major inference performance gains. The deployment, first reported by Crypto Briefing, underscores Nvidia's continued lead in AI hardware and could shift the competitive dynamics of large-model serving.

Inference gains on GB200

The switch to GB200 racks gives Perplexity a noticeable lift in how fast it can run the 235-billion-parameter Qwen3 model. Inference throughput and latency both improved, though the company hasn't released specific benchmarks. The gains come from the tight integration of Nvidia's Grace CPU and Blackwell GPU, which cuts data-transfer bottlenecks.

Nvidia's hardware edge

This deployment is another example of Nvidia pulling ahead in the AI chip race. GB200 racks are designed for exactly these kinds of high-parameter workloads, and Perplexity's choice suggests the hardware delivers where it counts. Competitors like AMD and Intel face an uphill climb to match that performance at scale.

Accelerating model deployment

With better inference, Perplexity can roll out updates and new models faster. The Qwen3 235B is a dense, powerful model, and running it efficiently means less time between training and production. That speed matters as the race to deploy ever-larger language models heats up.

The move could pressure other inference providers to upgrade their hardware or risk falling behind. If Perplexity maintains this edge, it may attract more AI developers who need high-throughput, low-latency serving. The next few months will show whether rivals can close the gap or if Nvidia's GB200 becomes the de facto standard for heavy models.