Skip to content
worth noting Open-source

Authors release the QATFactory tool and checkpoints for training LLMs with target quantization

only one source so far

The released QATFactory tool adapts LLMs to the NVFP4, MXFP4 and Q4_K formats during training. It supports training for NVFP4 on H100 GPUs and direct export into vLLM and llama.cpp. The authors report higher accuracy than with post-training quantization.

The authors have released the complete training code for the QATFactory tool and the resulting checkpoints. The open-source tool supports quantization-aware distillation (QAD) and reinforcement learning adapted for quantization (QARL), the NVFP4, MXFP4 and Q4_K formats, both dense and mixture-of-experts models, and training of all parameters as well as via LoRA.

During training, the tool simulates the quantization used at deployment, while performing matrix computations in BF16. According to the authors, this enables training for NVFP4 on H100 GPUs, which do not have FP4 Tensor Cores. According to the authors, checkpoints can be exported directly into vLLM and llama.cpp without any further lossy conversion and without additional overhead during inference.

The experiments covered models with 8 to 230 billion parameters. For the Qwen3.5-9B model, the QAD method achieved an average benchmark accuracy of 68.9% for NVFP4 and 66.0% for MXFP4; the best post-training quantization (PTQ) results were 65.4% and 56.4%. According to the authors, the appropriate strategy varies by format: for NVFP4, quantizing only the weights during training generally performed better, while for MXFP4, quantizing both weights and activations performed better. With the same number of training tokens, using a smaller number of 32K-length sequences increased average accuracy by 1.9 points compared to a larger number of 4K-length sequences.

What changed

Why it matters

Developers can adapt a model to the target quantization format already during training and, according to the authors' results, thereby limit the accuracy loss compared to PTQ. For teams preparing deployment, the practical option is to train for NVFP4 on H100 GPUs and export the results directly into vLLM and llama.cpp.

Two audiences, two different impacts

What this means

01

For individuals

When training a quantized model, the choice of target format also affects the appropriate training strategy: the authors' results support different handling of weights and activations for NVFP4 and MXFP4.

What to do For your own model, compare the accuracy after QAD with the PTQ result for the selected target format.
More practical updates →
02

For a business

Teams developing LLMs can connect training with deployment in vLLM and llama.cpp via direct export. According to the authors, for training with the NVFP4 target format, they can use H100 GPUs without native support for this format.

Development
What to decide Verify the export and quality of the resulting model in the vLLM or llama.cpp tool you are using before incorporating it into your company's deployment process.
More business impacts →
llama.cpp MXFP4 NVFP4 Q4_K QATFactory vLLM

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
arXiv cs.LG (Machine Learning) research source · first detected QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs