TOPIC #181Advanced 13 min read

Model Quantization: INT8, INT4, FP16, BF16

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Every number in a model costs bytes, and bytes decide what fits on a GPU and how fast tokens come out. This topic is the map of numeric formats — FP32/FP16/BF16/FP8/INT8/INT4 — plus the post-training quantization methods (GPTQ, AWQ, SmoothQuant) and the memory-bandwidth math behind them.

01.The Problem: The Model Weighs Too Many Bytes

A model is just a giant list of numbers.

That sounds trivial until you do the arithmetic. A 70-billion-parameter model:

  • stored in FP32 (4 bytes per number) → 280 GB
  • stored in BF16 (2 bytes) → 140 GB
  • stored in INT8 (1 byte) → 70 GB
  • stored in INT4 (half a byte) → 35 GB

No single GPU carries 280 GB. Some carry 80. A gamer's card carries 24.

So the question becomes

Insight

Does every number really need 32 bits of precision — or can we round them and still get great answers?

You already round numbers every day. Nobody writes "₹4999.37"; a shop posts "₹5000". The value barely changed; the storage got simpler.

Quantization is exactly this, applied to every weight in a model: replace each precise number with the nearest value from a smaller menu of allowed values. Fewer menu entries → fewer bits per number → smaller model, less memory traffic, faster inference.

The catch, of course: round too aggressively and the model starts giving worse answers. The entire art of this topic is rounding cleverly enough that the answers barely change.

This is also the foundation stone for the QLoRA topic you just met: QLoRA's NF4 is one point on the map we are about to draw.

Precision Formats and Quantization Paths 🧮

PRO Architecture Blueprint

Precision Formats and Quantization Paths 🧮

Bytes per parameter fall from 4.0 (FP32) to 0.5 (INT4); BF16 dominates training for its FP32-like dynamic range, while INT4 weight-only dominates edge/local inference.

Precision Formats and Quantization Paths 🧮
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #181: Model Quantization: INT8, INT4, FP16, BF16

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?