56x Faster AI Claimed

Ravenclip showcase channel

Is this the end of traditional transformers? Miami startup Subquadratic claims its new model is 56x faster than FlashAttention. Here is what the independent tests reveal.

What the video says

56 times faster than models using flash attention in a baseline speed test. That is the massive claim from Miami-based startup Subquadratic, which emerged from stealth last month promising it solved a mathematical bottleneck holding back large language models for nearly a decade.

Now the company has shared independent evaluation results from Appen to validate its new model, called SubQ. The results show SubQ scored 89.7% on Live Codebench, putting it in the same league as top of coding models while sustaining a 98% retrieval rate on Needle in a Haystack tests at 6 and 12 million tokens.

Rather than training from scratch, Subquadratic bootstrapped SubQ by reusing weights from a version of the Chinese open-source model Qwen. Its sparse attention architecture dynamically selects which token relationships are important on the fly, bypassing the costly quadratic scaling bottleneck of traditional transformers.

While skeptics warn the unreleased model could be the AI equivalent of Theranos, Appen's director of generative AI research, Janine Sinanan Singh, called the validating results a potential game changer. If the efficiency breakthrough holds up, some backers suggest that nobody will be building on standard transformers in just a few years.

Create a channel like this

Pick your topic. Ravenclip makes and posts videos like this on autopilot.

Start my channel