Loading market data...

Startups Claim Faster, Cheaper Alternatives to the Transformer

Two startups are betting that the transformer — the architecture behind every major large language model — is too slow and too expensive to keep carrying the industry forward. Subquadratic, based in Miami, says its SubQ model uses a sparse attention mechanism that rivals top mainstream LLMs on search and coding tasks. Manifest AI, in San Francisco, is developing a mechanism it calls "power retention." Both are chasing the same goal: making LLMs dramatically cheaper to run.

The quadratic wall

The transformer was introduced in a 2017 Google paper, "Attention Is All You Need." It works by giving every token in a sequence attention to every other token — which gets expensive fast. A 10,000-word document might require 50 million multiplications. That quadratic scaling is why context windows are capped and why compute bills keep climbing. OpenAI's Greg Brockman said this year the company is set to spend $50 billion on computing.

📊 Market Data Snapshot

24h Change
-0.10%
7d Change
-3.30%
Fear & Greed
34 Fear
Sentiment
🔴 slightly bearish
Bitcoin (BTC): $62,980 Rank #1

Two paths to a cheaper attention

Subquadratic's sparse attention skips over irrelevant parts of the sequence, rather than processing everything. The company claims its SubQ model already matches the performance of mainstream LLMs on tasks like search and coding. Manifest AI's power retention is a different mechanism entirely, though the company has released few technical details. Both are early, and neither has published peer-reviewed results.

The AI narrative in crypto has been built on the assumption that compute demand will keep skyrocketing. Projects that sell GPU compute or tokenize AI services are priced on that scarcity. If transformer alternatives actually work, that assumption breaks. The compute needed per token could fall by an order of magnitude, and the value of hoarding GPUs would shrink accordingly.

The energy forecast could shift

The International Energy Agency predicts data center electricity consumption will double by 2030. That forecast assumes transformer-based AI stays the norm. Sparse attention or power retention could flatten that curve — which would be good for the grid and bad for anyone who bought into the "AI will eat all the power" trade.

The claims are unverified, and scaling from a research demo to production is the hard part. But if either startup delivers, the transformer's ten-year run at the center of AI could start to fade. For crypto, the question is whether the market has priced in a future that never arrives.