speculative decoding in agents being effective is interesting.

speculative decoding in agents being effective is interesting.

  • It indicates there is a bias towards locality, and the extreme path dependence of LLMs
  • A small model can, with high probability guess a "correct" next token at each step. However, there is such incredible path dependence this does not work overall