Hugging Face's 4-bit quantized model outperforms its full-precision original via Quantization-Aware Healing.
quantization
15 tracked signals on quantization.
Baseten raised $13B Series F as inference engineering emerges as a critical AI discipline
“How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?”
NVIDIA's Nemotron 3.5 Lightning NVFP4 delivers 4x faster throughput compressed from 66GB to 22GB
“preserves accuracy while unlocking up to 4x faster throughput”
Panel from NVIDIA, Unsloth, HuggingFace, and Ollama frames quantization as AI democratization at the edge.
“same cost more intelligence”
Hugging Face integrates Nunchaku 4-bit quantization into Diffusers for faster diffusion inference
TurboQuant compresses agent retrieval embeddings to 3-4 bits, cutting memory cost 5x without degrading search quality.
“Today, we will see how you can cut memory cost of agent retrieval five times without breaking your search.”
Hugging Face released LFM2.5 Q4_0 checkpoints via quantization-aware distillation
Unsloth fixed a gradient accumulation bug improving training accuracy by 1-3% across the entire stack.
“fixed a gradient accumulation bug fix um which increased accuracy by 1 to 3% um across the entire training stack”
NVIDIA released a Nemotron 3 Ultra checkpoint quantized with the NVFP4 4-bit floating point format using NVIDIA Model Optimizer.
“NVFP4, an innovative 4-bit floating point introduced with NVIDIA Blackwell architecture”
Quantization shrinks AI models via fewer bits per value, trading some quality for smaller size and faster inference.
“a slightly worse model that actually fits your setup can still be much more useful than the full precision version”
Apple's TensorOps library now natively accelerates quantized ML kernels on Apple Silicon, leveraging the M5 chip's new neural accelerator.
“The neural accelerator is a new hardware block in M5, located directly in each shader core.”
Hugging Face hosts advanced local AI education series on inference engines and quantization
“Everybody should know about this and they should know how to do this.”
NVIDIA explains converting FP8-quantized checkpoints into TensorRT engines for faster production inference.
“Converting a quantized checkpoint into an NVIDIA TensorRT engine bridges the gap between model optimization and production deployment, enabling faster inference, higher throughput, and more efficient GPU utilization at scale.”
DeepLearning.AI and Red Hat launch a course on efficient open-source LLM inference using vLLM.
“The techniques you learn in this course are what power efficient LM serving in production today.”
NVIDIA gave Unsloth a DGX Station to accelerate its model quantization and RL research.
“We're going to be use utilizing this to do all of the models. And just provide more and more for the community.”