Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

TL;DR AI
2 min readKey summary
Researchers introduced ESRT, an edge-cloud speech translation system that keeps a lightweight encoder on-device and sends compressed features to the cloud instead of raw audio.
The design reduces voice leakage risk and cuts bandwidth use by up to 10x, while improving multilingual translation with curriculum learning and data balancing.
On FLEURS, ESRT achieved state-of-the-art results for 45-language many-to-many translation, including strong cross-lingual performance.
The authors released code and models, including ESRT-4B and ESRT-12B.
