The Stack
AI 101 — What is model quantization?
A large language model is a giant bag of numbers. Quantization is the trick of shrinking those numbers so the model takes up less memory and runs faster — usually with only a small, barely-noticeable drop in quality. If you have ever tried to run a powerful open-weight model