56x Faster AI Model
Is this Miami startup the real deal? Subquadratic just dropped independent proof that its new SubQ model is 56x faster than FlashAttention. ⚡️
What the video says
Could a Miami startup have quietly solved the mathematical bottleneck that has held back large language models for nearly a decade? When Subquadratic came out of stealth last month claiming its new SubQ model bypasses the costly scaling limits of transformers, skeptics warned it could be the AI equivalent of Theranos.
But the company has now shared independent evaluation results from Appen, bringing the first real receipts to back up its massive claims. The evaluation logged that SubQ is 56 times faster than models using flash attention in a baseline speed test, while dynamically selecting key token relationships on the fly.
In massive data retrieval trials, SubQ sustained a notable 98% score on Needle in a Haystack tests, with context windows of 6 million and 12 million tokens. The model also scored 89.7% on LiveCodeBench, matching the ballpark performance of top-tier coding models from Google, OpenAI, and Anthropic.
While the startup bootstrapped SubQ by reusing weights from the Chinese open-source model Qwen, the results have validated its sparse attention architecture. If these efficiency gains hold up, it could fundamentally reshape how future AI models are built.