Stop Measuring GPU Speed! Why Agentic Workloads Require New Performance Metrics
Enterprise data centers are being redesigned to support the growing demands of agentic artificial intelligence (AI). However, as organizations deploy AI agents capable of reasoning, retrieving information and interacting with external tools, traditional data center performance metrics are becoming less useful.
Conventional benchmarks often emphasize processor speed, core density and accelerator throughput. While these figures indicate theoretical hardware capability, they do not always reflect how efficiently a complete system performs under sustained, concurrent workloads.

For agentic AI, the more important question is not simply how much computing power a data center has, but how much useful work it can complete.
Agentic AI Creates New Data Center Bottlenecks
Unlike conventional workloads, agentic AI systems can operate through continuous loops involving reasoning, database queries, information retrieval and tool execution. When hundreds of agents run simultaneously, they compete for CPU resources, memory, network capacity and other infrastructure components.
As concurrency increases, these shared resources can become bottlenecks even when processors and GPUs have substantial unused theoretical capacity.
Engineering tests indicate that at 32 concurrent agent sessions, CPU-side processing delays can rise from less than 1% of total latency to more than 15%. These queueing delays can prevent tasks from progressing and increase overall workflow latency.
Idle GPUs Can Reduce AI Infrastructure Efficiency
The problem becomes more significant when expensive AI accelerators are involved. An agent may pause while waiting for a database query, network request or sandboxed tool to complete. During that time, the GPU may have little useful work to perform.
Some enterprise testing suggests GPUs can perform useful work for only around 50% of active runtime in certain agentic workloads. This demonstrates why hardware utilisation alone is not an adequate measure of AI infrastructure efficiency.
A data center can contain large amounts of compute capacity while still delivering relatively little completed work if other components cannot keep pace.
Measuring Performance Through Delivered Agents
These challenges point to a need for new AI data center benchmarks. One approach is to measure delivered agents per rack rather than focusing exclusively on processor or accelerator throughput.
The metric would measure how many complete agent workflows a rack or cluster can successfully execute within defined latency, power and cost limits.
Such a benchmark could account for:
- CPU scheduling and queueing delays
- Memory capacity and bandwidth
- Network latency
- Storage and database performance
- Tool execution time
- Accelerator utilisation
- Power consumption
- Cost per completed workflow
- Performance at different concurrency levels
This provides a more practical measure of infrastructure capacity because it focuses on completed work rather than theoretical compute performance.
Memory and Hardware Balance Matter
Memory architecture will also play an important role as AI agents maintain larger amounts of context and state. Combining CXL-attached memory with DDR5 can provide additional capacity, although the cost and performance benefits depend heavily on workload and system design.
Hardware selection also needs to reflect the different stages of an agentic workflow. Orchestration, reasoning, data movement and tool execution can have very different processing requirements. Some tasks benefit from strong per-thread performance, while others benefit from higher core density and parallel throughput.
Infrastructure tasks such as compression and certain storage operations can also be offloaded to dedicated processing engines, freeing general-purpose CPU resources for latency-sensitive workloads.
Why Traditional AI Benchmarks Need to Evolve
Processor specifications and accelerator throughput will remain important, but they cannot fully describe the performance of agentic AI infrastructure.
Future benchmarks should report factors such as target concurrency, end-to-end latency, queueing overhead, accelerator utilisation, memory performance, network delays, power consumption and cost per completed workflow.
As AI agents become more autonomous and concurrent, the primary limitation may no longer be the speed of an individual processor or GPU. Instead, overall performance will increasingly depend on how efficiently CPUs, GPUs, memory, storage and networks operate together.
For enterprise AI infrastructure, the most meaningful measure of capacity may ultimately be how much useful work a system can reliably complete under real-world conditions.


Click it and Unblock the Notifications