speculative decoding in agents being effective is interesting.
speculative decoding in agents being effective is interesting.
- It indicates there is a bias towards locality, and the extreme path dependence of LLMs
- A small model can, with high probability guess a "correct" next token at each step. However, there is such incredible path dependence this does not work overall