ECHO framework improves LLM inference efficiency with hierarchical speculative decoding
ECHO introduces a hierarchical dual-loop approach that leverages early LLM layers' discriminative power to reduce stale draft candidates and verification costs in speculative decoding. This method addresses key inefficiencies in draft-model-free LLM inference.