speculative-decoding
from Orchestra-Research/AI-research-SKILLs
Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time ap
v1.0.0MIT
468
Lines
1,625
Words
17
Code Blocks
Languages
bashpython
19-emerging-techniques/speculative-decoding/SKILL.md