What is Speculative Decoding and How Does a Draft Model Work?
I was speaking to a colleague the other day about how LLMs process tokens and predict the next word (discussing the technical details of LLMs excites me), and when I brought up speculative decoding, which