Weka cuts AI inference load by caching 100% of pre-calculated tokens
NeuralMesh 6 turns NAND flash into GPU memory so teams stop redoing attention work every chat turn.

Weka co-founder and CEO Liran Zvibel says NeuralMesh 6 caches 100% of an AI model's pre-calculated tokens using its Augmented Memory Grid approach. For decision-makers, it aims to reduce GPU memory and compute pressure, lower inference costs, and speed up new AI deployments without waiting for more GPUs.
Weka’s pitch is simple, and the consequence is not: NeuralMesh 6 is designed to cache 100% of an AI model's pre-calculated tokens, so inference does not repeatedly redo the expensive work hidden inside every prompt. Weka co-founder and CEO Liran Zvibel told VentureBeat that the payoff shows up hardest in multi-turn sessions, where each new turn can force models to re-prefill everything that came before it. In Weka’s framing, that means you can “overcalculate” again and again unless that prefill work is cached.
And Zvibel is very explicit about the math. In multi-turn chat, he said, “If you have 10 turns, you may overcalculate 100 times because you're redoing all of them. If you have 20, you'll overcalculate 400 times.” His point ties directly to what NeuralMesh 6 targets: the attention prefill stage, which is computationally expensive, while decode is comparatively lightweight. Weka’s claim is that by using much cheaper NAND flash as an extension of GPU memory, it can cache that pre-calculated token work so you “never need to redo it.”
This is the “stop adding more GPUs” argument in practical form. GPU memory is the most expensive resource in production AI, and it is also the one running out fastest. Long context windows and multi-turn conversations amplify the problem because models repeatedly recompute information they already processed. The question for operators is always the same: do you treat GPU memory as the limiting resource, or do you extend it with lower-cost storage? Weka is betting on the second option, using flash storage technologies to close the gap. Its NeuralMesh 6 software platform launches alongside its first self-designed hardware line, Wekapod 3, and it positions both as part of an Augmented Memory Grid that aggregates NAND flash to behave like GPU memory at a fraction of the cost.
This matters because the market is already noisy. VentureBeat notes that this is an active and increasingly crowded category, and big storage players have repositioned toward AI infrastructure over the past two years. Dell, NetApp, Pure Storage, and VAST have all moved in this direction. The reason the competition is intense is also why Weka is leaning hard on “AI-native” credibility rather than general-purpose storage adaptation. In the short term, buyers are evaluating whether the vendor is solving real inference bottlenecks or merely updating their messaging to fit the trend. Steve McDowell, chief analyst at NAND Research, told VentureBeat that the storage world is shifting from serving bits to managing data at the speed of AI, and he singled out Weka as a true AI-native data company in that context.
NeuralMesh 6 is built to close multiple gaps at once, not just the caching angle. Weka said it adds four capabilities aimed at functionality gaps Zvibel claims have cost the company deals in competitive evaluations. First is composable and virtual multi-tenancy. Composable clusters give an anchor tenant full hardware-level isolation, including dedicated CPU, memory, and storage. Virtual multi-tenancy runs through Weka’s RDMA fabric, and Weka claims network-level isolation that scales past 1,000 tenants per cluster. Provisioning is listed as under 30 minutes. Weka also states that a single cluster running 50 composable clusters can support up to 50,000 tenants. Second is unified file and object storage. In many systems, file and object paths are separate, meaning a gateway translates and data effectively exists twice. Weka’s claim is that the same physical data on disk is directly readable through either path at once, without a translation layer and without a second copy.
Third is metadata-first replication, which targets another operational pain point: waiting for full data movement before compute can start working. Zvibel said destination environments become browsable before a full data copy arrives, with data hydrating only when accessed. He contrasted this with waiting days or weeks, and in extreme cases “a month,” for data to fully arrive. Weka says customers can grab some allocation of new GPUs and get up and running within an hour.
Fourth is AlloyFlash and Always-On data reduction. AlloyFlash mixes TLC and QLC NAND within a single cluster, routing latency-sensitive work to TLC and bulk-capacity workloads to QLC. Weka’s stated goal is cost reduction per terabyte without a performance penalty for the workloads that need speed. In addition, data reduction runs by default rather than as an optional feature.
So where does the caching claim actually land in real deployments? Weka frames its technology as most relevant for organizations already operating AI at scale or expecting rapid growth in usage, especially enterprises building internal copilots, customer service agents, software engineering assistants, or retrieval systems with long context windows. Smaller deployments may see less immediate benefit, particularly if GPU utilization is not yet the bottleneck. Weka also names specific non-AWS GPU clouds as targets for its object storage approach, including Lambda, Nebius, G42, and CoreWeave. Weka describes roughly two orders of magnitude higher performance than conventional S3, plus capacity-based pricing instead of per-API charges.
The strategic stakes are why this story belongs on an executive briefing deck, not just in a storage blog. If Weka’s approach can truly cache 100% of pre-calculated tokens and translate that into “level of GPU efficiency” that saves money on GPUs and memory, it changes the economics of capacity planning. It can also change how quickly customers can turn budget into live inference. McDowell flagged Weka’s contractual guarantee on its data reduction claims as underappreciated, saying the company is “putting its money where its mouth is” with contractual guarantees. For boards and leadership teams, that is a subtle but meaningful signal: it moves from marketing claims to enforceable commitments, which becomes important when buyers are under pressure to control inference costs in memory and GPU constrained markets.
Meanwhile, the broader second-order implication for competitors is blunt. If storage vendors cannot demonstrate real-world improvements in GPU efficiency, latency, and tenant provisioning speed, buyers may view the category as mostly about repositioning. McDowell advised enterprise buyers to look at what vendors are promising versus what they are actually delivering by talking to organizations running similar workloads at similar scale. In a world where GPU availability and inference costs can decide whether product teams scale or stall, the difference between “bits plus slides” and “attention-aware caching plus operational speed” is not theoretical. It is measurable, and it is starting to show up in the way deals are evaluated.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Nvidia and Wistron will build Blackwell AI servers in Texas, Nikkei Asia reports
A Texas manufacturing plan for Blackwell AI servers ties Nvidia's next platform rollout to Wistron's local capacity and supply chain risk.

Meta tests StoryKit bedtime stories in select regions to measure parent response
The experiment is regional, and the real question is how quickly parents adopt AI storytelling for kids.

Range Rover GT is not a Velar EV replacement, spy tests at Arctic Circle confirm
The EV plan is real, but the direction was misread for months. Here is the actual story.

