Cloud inference provider DeepInfra claims NVIDIA’s upcoming Vera CPU significantly outperformed competing processors from AMD and Intel in production-scale AI performance testing.

Rather than using generic benchmarks, which don’t often reflect real-world usage, DeepInfra used its own production AI infrastructure for testing. It showed the Arm-based Vera CPU delivered up to 2.2 times faster agent orchestration than x86-based competitors while supporting up to 1.6 times more concurrent AI agents under the same quality-of-service requirements.

The results are notable because Vera is NVIDIA’s first custom server CPU designed specifically for AI infrastructure rather than traditional enterprise computing. The processor is specifically designed to handle the CPU demands of AI agents, which require constant orchestration, scheduling, networking, memory management and communication with GPUs.

DeepInfra, which operates a cloud platform optimized for AI inference, said it received early access to Vera hardware through NVIDIA’s ecosystem program and evaluated the processor using real production traffic instead of synthetic benchmarks. The company compared Vera against three leading CPUs from AMD and Intel under identical conditions.

DeepInfra also published a technical analysis with the full benchmark methodology and complete results here.

According to the benchmark, Vera was the fastest processor across every workload tested. On identical CPU partitions, the processor sustained 256 concurrent AI agents while maintaining response-time and error-rate targets, compared with between 160 and 192 agents on competing processors.

DeepInfra also reported that unused Vera CPU cores were able to simultaneously serve a 20-billion-parameter open-source large language model faster than an entire previous-generation CPU socket dedicated exclusively to that task, without degrading agent performance.

“We built DeepInfra’s infrastructure from the ground up for inference,” Nikola Borisov, co-founder and CEO of DeepInfra, said in a statement. “We have tuned every layer of it for cost, latency, and throughput at production scale. NVIDIA is building for where AI workloads are headed, and we believe Vera is exactly the kind of hardware solution the next iteration demands.”

GPUs get all the attention when it comes to AI processing, but the CPU plays a necessary role as well. Whereas the GPU does the heavy lifting of processing the models, CPUs coordinate requests, manage memory, schedule workloads, handle networking and storage operations, and orchestrate fleets of AI agents.

“Bringing Vera to market is an incredible opportunity, but performance is ultimately proven in production,” Ian Finder, director of Data Center CPU Products at NVIDIA, said in a statement. “DeepInfra is pushing Vera with demanding, real-world agentic workloads that reflect the needs of production environments. These results demonstrate exactly what Vera was designed to deliver.”