E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring
TL;DR AI
2 min readKey summary
Researchers introduced E-PMQ, a post-merge quantization framework for merged neural networks.
It uses source expert weights and merged-weight anchoring to calibrate low-bit quantization more effectively.
On vision and language benchmarks, E-PMQ significantly outperformed standard GPTQ, preserving accuracy better after merging.
The approach could make combined expert models cheaper to serve by reducing memory use without major performance loss.
