Subquadratic claims it eliminated a decade-old LLM bottleneck with far less compute
The AI hardware bottleneck fight just got a new entrant, plus BCI trials accelerating toward the market.

Subquadratic, an AI startup that came out of stealth last month, claims it solved a mathematical bottleneck that has held back large language models for almost a decade by slashing transformer computations. For decision-makers, the question is whether this changes model cost and energy assumptions enough to shift buy-versus-build and evaluation priorities.
Subquadratic says it cracked a mathematical bottleneck that has held back large language models for almost a decade. The company came out of stealth last month with a big claim: its approach slashes the number of computations transformers need to generate answers, making the resulting LLM faster, cheaper, and far less energy-hungry than other models on the market.
That headline matters because compute and energy are not just “ops” details anymore. They are the limiting factors that shape pricing, deployment timelines, and which models companies can afford to run at scale. Subquadratic’s founders have started to share “receipts,” and many experts who were skeptical initially have not been fully won over, but the conversation has moved from pure skepticism to a more serious debate about whether there is a real efficiency step-change here.
So what exactly is the “bottleneck” fight about? In plain English, transformers do a lot of math to produce each answer. Even if a model is smart, the cost to run it can dominate everything from inference latency to unit economics to whether a product can stay on when you turn on the lights. Subquadratic’s stated path is to reduce the computations required for generation rather than just scaling up brute-force. The purported result is an LLM that uses far less energy than any other model on the market, and that is the kind of claim that instantly forces buyers and builders to ask: does this show up in measurable benchmarks that hold up outside a lab demo?
MIT Technology Review notes that “many experts remained skeptical,” but that Subquadratic has started to share the receipts. That detail is important, because in this industry, efficiency claims can sound like marketing until they are supported by evidence. When receipts appear, the evaluation shifts from “can you imagine a better method?” to “can you reproduce the savings, the performance, and the energy reductions under realistic workloads?” For investors and operators, this is where diligence gets sharper. You stop asking only whether a method exists and start asking how it behaves across the full stack: model behavior, runtime, infrastructure fit, and whether the energy claim is actually comparable to what “other models on the market” do.
At the same time, there is a second storyline in today’s tech world that is less about model compute and more about biological compute. Brain-computer interface (BCI) trials are taking off, and MIT Technology Review frames the acceleration with a very specific human reference: Casey Harrell, a man with ALS, described as “the first power user” of a brain implant. Harrell told the reporter that the device enabled him to maintain an income, reconnect with friends and family, and read to his daughter, and he called it “nothing short of revolutionary.” Whether or not you track BCI weekly, that’s the real-world scoreboard: if people can do more because of these systems, the market pull and regulatory attention follow.
The regulatory angle here is not abstract either. MIT Technology Review says that over the past couple of years the number of BCI trial volunteers has soared, and that this year China became the first country to approve a BCI for medical use. That approval signals a broader shift from experimental trials to at least some regulated medical deployment. It also raises the stakes for BCI developers and partners because approval changes the funding and commercial math. Once a device is approved for medical use, the market stops being “maybe someday” and starts being “what exactly is reimbursable, supportable, and scalable?”
In both stories, you can see a shared pattern: technical bottlenecks are being challenged, but proof and regulatory framing determine who gets to cash in. In AI, the bottleneck is computations transformers need to generate answers, and Subquadratic’s entry forces competitors and customers to revisit assumptions about costs, speed, and energy. In BCI, the bottleneck is getting systems that can do useful work in a real person’s life, and the proof is trial growth plus at least one major regulatory milestone in China.
For executives, the strategic implications are straightforward even if the tech is not. On the AI side, a credible path to faster, cheaper, and lower-energy LLMs could reshape what “efficient enough” means for your product roadmap. If receipts hold up, evaluation processes may need to move from headline performance to compute efficiency, energy cost, and deployment feasibility. On the BCI side, accelerating trial enrollment and the shift toward medical approval means companies should treat clinical progress and regulatory pathways as core strategy, not downstream paperwork. In both cases, the boards and investment committees should pay attention now because the winners are likely to be the teams that can convert technical progress into repeatable, comparable outcomes under real constraints.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Black Forest Labs launches FLUX 3 with 20-second audio-video, but no pricing yet
Enter gated early access for video and action, while image waits, benchmarks stay preliminary, and costs are a mystery.

Lisa Su pushes disaggregated inference with Cerebras, aiming Helios at Nvidia’s Vera Rubin
AMD’s new Cerebras pact is a direct bet that AI inference should be split across specialized chips, not one dominant rack.

AMD’s Helios claims 30% per-dollar lead over Nvidia’s Vera Rubin
The 72-GPU rack looks like a head-on spec war, but real-world FLOPS and power efficiency decide the winner.

