Original ·2026452488434651264· Andrej Karpathy @karpathy · · en XCitaba a ·2026351870852358492· Reiner Pope @reinerpope · · en XWe’re building an LLM chip that delivers much higher throughput than any other chip while also achieving the lowest latency. We call it the MatX One.
The MatX One chip is based on a splittable systolic array, which has the energy and area efficiency that large systolic arraysWith the coming tsunami of demand for tokens, there are significant opportunities to orchestrate the underlying memory+compute *just right* for LLMs.
The fundamental and non-obvious constraint is that due to the chip fabrication process, you get two completely distinct pools of
RT @karpathy: With the coming tsunami of demand for tokens, there are significant opportunities to orchestrate the underlying memory+comput…