Community Discussion · Policy
Hardcoding Models into Silicon: The Ultimate Form of Hardcore Hacking
Last weekend at an edge computing hackathon, I saw someone running llama.cpp on an FPGA. When we chatted, he said it's still too "soft"—every matrix multiplication has to go through the instruction pipeline, wasting at least two orders of magnitude in energy efficiency. I replied then: why not just tape out and hardwire the weights into silicon? He froze for two seconds and asked how much that would cost. Now Alphabet tells us they're already doing it.
Physix Frontier