Switch language한국어
Back to the list

NVIDIA Introduces a 4-Bit Pretraining Methodology Using NVFP4, Validated on a 12B Hybrid Mamba-Transformer at 10T Token Horizon

TL;DR AI

Key summary

2 min read
  1. NVIDIA introduced an NVFP4-based 4-bit pretraining method for large language models.

  2. Using this approach, a 12B hybrid Mamba-Transformer was trained on 10T tokens.

  3. The recipe combines selective BF16 layers, random Hadamard transforms, and delayed quantization.

  4. The resulting model matched an FP8 baseline on MMLU-Pro while leveraging Blackwell hardware support.

Read the original