5 Comments
User's avatar
JP's avatar

The SRAM vs HBM framing nails the core tradeoff. Cerebras went big wafer, but there's an even more radical play emerging. Taalas hardwired Llama 3.1 8B directly into silicon and claims 17,000 tokens/sec on a chip costing a fraction per million tokens. Different bet entirely: model-specific silicon vs general-purpose wafer scale. Dug into it here: https://reading.sh/what-happens-when-ai-inference-gets-10-times-faster-bf0286a34a45?sk=8dfc863d0c5e9e9d15da1b2d49737b6b

Jordamøn's avatar

holy smokes, if that's for real it's brilliant

3rdWorldBear's avatar

Thanks a lot for sharing such an interesting move in the AI industry and also explaining the relevant technicalities in great detail.

Curzon's avatar

Thanks Jordamon, I enioy these kinds of posts.

J Scott's avatar

Very interesting.

The innovation is good. PC gaming used to push this, but its been static for 15 years.