Test report DSG-9770 · Rev A · tested October 10, 2026

AI Datacenter InfrastructureDevice under test

HPE Cuts AI Support Token Spend 30x, Expands On-Premises Play

HPE's Fidelma Russo said running AI on-premises cut token costs more than 30x, saving nearly $100,000 a month, as the vendor expands GreenLake Intelligence, Morpheus 9, Vera-equipped ProLiant servers and Alletra X10000 storage for agentic workloads.

Read
4 min
Words
821
Node
14nm
Operator
Priya Raman

Spec summary

  1. HPE reduced AI support token costs by more than 30x, saving nearly $100,000 a month by moving workloads to its Private Cloud AI stack
  2. OpenClaw processed more than 600 billion tokens in a single month for ~100 coding agents, equating to about $13,000 per agent per month
  3. Dell Technologies study cited at Dell Technologies World 2026 found 67% of AI workloads run outside the cloud and 88% of respondents run at least one AI workload on-prem
  4. HPE Alletra Storage X10000 delivers 20x faster time-to-first-token and 17x higher throughput in HPE tests when paired with KV cache
  5. HPE's Juniper Networks acquisition closed in 2025 at $14 billion, supplying networking technology cross-pollinated with the Aruba lineup

Hewlett Packard Enterprise has reduced the cost of running its internal AI support infrastructure by more than 30x, saving nearly $100,000 a month, after moving workloads from public clouds onto an in-house stack built on GreenLake Intelligence and Private Cloud AI, an on-premises system co-engineered with Nvidia.

The figures came from Fidelma Russo, executive vice president, CTO and general manager of HPE's hybrid cloud business unit, speaking at HPE Discovery 2026 in Las Vegas. The savings show how sharply "tokenomics" — the per-token economics of generative AI — can move once an enterprise controls its own inference layer.

"It allowed us to govern that really important customer data and it gave us better performance," Russo said. "It also helped us significantly minimize the token spend associated with operating AI at scale. We stopped being consumers of AI and we became producers of intelligence."

What does tokenomics look like at scale?

The numbers behind AI agent consumption explain why HPE built its own stack. HPE support systems process billions of operational signals every day, and as those systems grew more autonomous, token use scaled with them.

OpenClaw, a widely deployed virtual personal AI agent, processed more than 600 billion tokens in a single month to support roughly 100 continuously operating coding agents, Russo said. That works out to about $13,000 per agent per month. HPE positions the implication bluntly: a single agent prompt can fan out into thousands or millions of model interactions, and inference becomes a continuous operational workload rather than a one-time request.

Validation firm Signal65 puts the inflation factor in range: agents consume 4x to 15x as many tokens as standard chat interactions, and as workloads mature, autonomous agents could push 1,000x the inference demand of reasoning-only AI.

Why are workloads moving back on-prem?

The shift cuts against the "cloud for everything" posture of 2023–2024. Steve McDowell, founder and chief analyst at NAND Research, framed the reversal directly:

"Training might happen in the cloud, and that part of the story remains largely true. But something unexpected is happening with inference. It's moving back on-prem and out to the edge, a quiet reversal that's forcing a fundamental rethink of enterprise AI architecture."

Dell Technologies founder and CEO Michael Dell cited a company study at Dell Technologies World 2026 in May, where 67 percent of AI workloads ran outside the cloud and 88 percent of respondents ran at least one AI workload in their own datacenter. Cisco Systems made a similar case at Cisco Live 2026, leaning on its networking portfolio to package AI infrastructure for enterprises.

Cheri Williams, SVP and GM of HPE's private cloud and flex solutions, drew the line between training and production:

"There's still a place for experimentation and model training in the public cloud, and you still see customers doing that," she said during a panel. "But when it comes to production, on-prem is where most of the enterprise customers are going. The economics don't make sense to be in production in the public cloud."

What is HPE shipping for the on-prem AI datacenter?

The product slate HPE outlined at Discovery 2026 spans software, storage and compute:

  • GreenLake Intelligence: A framework of AI agents with a central agent registry that catalogs what agents an organization runs, where they reside and what permissions they hold. OpsRamp Copilot monitors token-based consumption and cost across agents, AI factories and workloads.

  • Morpheus 9: The latest release in HPE's infrastructure automation line. New components include Morpheus Central for federated management across distributed instances, a Morpheus Orchestration Copilot accepting natural-language provisioning, and integrated software-defined networking.

  • Private Cloud AI + Alletra MPX 10000: HPE is integrating the MPX 10000 storage array with its two-year-old Private Cloud AI stack to apply governance and metadata policies automatically.

  • Alletra Storage X10000: HPE positions the array as "active memory" for AI via KV cache, which stores key-value context so systems do not rebuild it on every prompt. Test data from HPE: 20x faster time-to-first-token and 17x higher throughput.

  • Nvidia Agent Toolkit: Ships with Private Cloud AI servers and bundles OpenShell secure runtime, NemoClaw blueprints and Nemotron models for multi-agent design, build and orchestration.

  • ProLiant DL394 Gen12: A server due in early 2027 that pairs Nvidia's new Vera CPUs with GPU accelerators. HPE positions Vera for rapid tool calls, orchestration and real-time data processing in agent workloads.

  • Networking ties the strategy together. HPE continues to merge its Aruba campus and branch lineup with technology from Juniper Networks, an acquisition that closed last year at $14 billion.

    The cross-stack convergence matters because storage and memory, not raw GPU count, now decide the cost of an AI interaction. "In AI, memory is no longer a technical detail or a supply chain challenge," Russo said. "It is a strategic resource."

    via finops.org (Original)

    Filed under

    • hpe
    • nvidia
    • on-premises-ai
    • ai-inference
    • greenlake
    Share this article:

    More from Priya Raman

    Priya Raman

    Show full bio

    Correspondent covering business strategy at Die Signal.

    243 articles

    Same lot · LOT-C1C6

    « Previous article