• Dell Technologies believes that enterprise AI costs continue to rise despite a sharp decline in token prices because Agentic AI systems consume increasingly large amounts of tokens and context.
  • According to Goldman Sachs, the cost of generating tokens has decreased by about 60–70% annually due to more efficient models, more powerful hardware, and market competition.
  • However, FinOps Foundation stated that nearly three-quarters of enterprises have exceeded their AI cost budgets over the past year.
  • Goldman Sachs forecasts that global AI token volume will increase 24-fold by 2030.
  • Dell notes that the primary cause is no longer computing power, but storage infrastructure that is not designed for AI Agents and long context windows.
  • In large language models, GPU memory must simultaneously perform two tasks: computation and KV cache storage, which degrades performance as the workload increases.
  • Dell proposes shifting the KV cache to a dedicated storage system to free up GPU memory for computation.
  • According to Dell’s testing results on 4 NVIDIA H100 GPUs, the Time to First Token was reduced from 17 seconds to approximately 1 second when using the KV cache transfer mechanism.
  • The throughput for handling multi-turn conversations also nearly tripled after storage optimization.
  • Dell stated that its storage system achieves latency that is nearly half that of its closest competitor and up to 14 times faster than a configuration without KV cache transfer.
  • The article emphasizes that enterprises should deploy servers, storage, and networking synchronously from a single vendor to reduce integration time and accelerate AI operationalization.
  • Dell concludes that in the era of Agentic AI, competitive advantage will depend on the ability to optimize the entire infrastructure, particularly storage, rather than just investing in more GPUs.
  • 📌 Conclusion: The bottleneck for enterprise AI is shifting from GPU power to context management and data storage capabilities. As Agentic AI handles long sessions and multiple concurrent agents, the KV cache becomes the deciding factor for performance and cost. According to Dell, optimizing the storage architecture can significantly reduce latency, increase throughput, and improve AI investment efficiency without merely continuing to add more GPUs.
Share.
VIET NAM CONSULTING AND MEASUREMENT JOINT STOCK COMPANY
Contact

Email: info@vietmetric.vn
Address: No. 34, Alley 91, Tran Duy Hung Street, Yen Hoa Ward, Hanoi City

© 2026 Vietmetric
Exit mobile version