Switch language한국어
Back to the list

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced SplitQ, a low-bit PTQ method for vision-language models that tackles text-image activation imbalance.

  2. The paper pinpoints uneven modality-specific outlier channels as a major source of quantization error.

  3. SplitQ combines channel splitting with adaptive cross-modal calibration to reduce these errors.

  4. It improves accuracy across multiple datasets and settings, including W4A8, W4A4, W3A3, and W3A2.

  5. The method could make large VLMs more practical on resource-limited devices without losing as much accuracy.

Read the original