AMD acquires Taalas to boost inference performance by etching models in silicon

Early tech demos show model-specific integrated circuits churning out up to 17,000 tokens a second