AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
AngelSpec unified training framework adaptively switches between multi-token prediction and block-parallel diffusion for speculative LLM decoding across heterogeneous workloads.