Reduced Matrix Multiplication Cuts Transformer Inference Costs
Reduced Matrix Multiplication (RMM) is a training-free, input-adaptive method that lowers Transformer inference costs by selecting informative slices in matrix multiplications without changing model weights. This approach reduces the high-dimensional matrix products common in large language model i...