TOPIC #223Intermediate 12 min read

ONNX: The Interchange Format for Models

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

ONNX is the USB-C of machine learning: one self-contained model file — a computation graph plus weights, drawn from a versioned dictionary of standard operators — that PyTorch or TensorFlow exports and ONNX Runtime, TensorRT, CoreML, or a browser then executes. Covers opsets, export workflows, dynamic shapes, quantization, custom ops, and where interchange still breaks down.

01.The Problem: You Trained in PyTorch, but Production Runs Everything Else

Your team trains a ranking model in PyTorch. It works. Time to ship it.

Then reality shows up:

  • The scoring service is a plain CPU microservice. No GPU at all.
  • The marketing site wants the model inside the browser tab.
  • The mobile app wants it on the iPhone.
  • One partner runs NVIDIA H100s, another ships Intel laptops, a third builds car dashboards on automotive NPUs.

You could hand each of them the PyTorch checkpoint and tell them to install PyTorch and replay your exact training code. Nobody wants that. A checkpoint is framework-private: it only speaks Python and PyTorch.

So the question becomes

Insight

Can we describe a trained model once, in a neutral language, so that any runtime on any hardware can execute it?

That one question is the entire job of ONNX.

02.The Idea in Plain Words: A Graph Plus a Dictionary of Agreed Operations

Open Neural Network Exchange (ONNX) is simply

Insight

A neutral file format that stores a model as a computation graph — a list of standard operations plus the weights — so any compatible runtime can execute it.

The project was founded in 2017 by Microsoft and OpenAI with Facebook and Amazon, and today lives under the Linux Foundation as ONNX AI. It standardizes exactly two things:

  1. A graph format (protobuf): named nodes, typed inputs/outputs, and initializers — the weights — stored right in the file as tensor constants. The file says "apply this Convolution, then that MatMul," not "call this PyTorch function."
  2. A standard operator library: Conv, MatMul, Softmax, Attention... each with precisely defined math. The library is versioned by opsets: "opset 17" means a pinned operator contract — every operator up to 17 behaves in exactly one agreed way. Ai.onnx.ml adds classical-ML operators (random forests, scalers, SVMs).

An ONNX file is self-contained and runtime-neutral: weights plus graph, one artifact. That is why heterogeneous production stacks — PyTorch training, CPU inference microservices, automotive NPUs, browsers via onnxruntime-web — can share a single file.

How does a file get made? The export paths:

  • PyTorch via torch.onnx.export — tracing: run the model on a sample input once and record which operations fired.
  • The newer torch_onnx_dynamo_exporter / torch.export-to-ONNX path (2024+), which handles dynamic behavior much better.
  • TensorFlow/Keras via tf2onnx.

03.A Worked Example: One Ranker, Four Devices

Take a small image ranker and walk one release end to end.

Step 1 — export. Training ends; you call the exporter and get ranker.onnx, a single file holding graph + weights.

Step 2 — look inside. The graph is just named boxes and arrows:

code
  image -> Conv -> ReLU -> MatMul -> Softmax -> score
              (all weights stored here, as plain tensors)
  no Python, no framework — just the recipe and the ingredients

Step 3 — plug the same file into four sockets.

  • CPU microservice: ONNX Runtime with its built-in CPU (MLAS) kernels.
  • iPhone: CoreML execution provider.
  • Browser tab: onnxruntime-web (WASM).
  • H100 server: CUDA or TensorRT execution provider.

Nobody retrained anything. One file, four targets.

Step 4 — the parity check, with real numbers. Feed one golden input through the framework model and through the exported file:

  • PyTorch says: 0.7312
  • ONNX Runtime says: 0.7309
  • Max absolute difference: 0.0003 — under the usual 1e-3 tolerance. Ship it.

Had the difference been 0.05, something broke during export — a Python constant baked in by tracing, an unsupported operator quietly replaced with a wrong fallback — and the file never reaches production. That single subtraction is what makes interchange safe.

04.The Analogy: USB-C for Models

ONNX is best pictured as the USB-C port of machine learning.

One plug shape; every device accepts it — laptop, phone, monitor, headphones. And a standards document says exactly what each pin means. Map that onto models:

  • The connector shape = the graph format. One file fits every runtime.
  • The standards document = the opset. "USB spec revision 2.1" and "opset 17" both promise: here is exactly what every pin/operator does.
  • Adapter dongles = execution providers. The same cable drives HDMI or VGA because the adapter translates to that port's dialect; the same ONNX graph drives NVIDIA, Apple, or Intel silicon because CUDA/CoreML/OpenVINO providers translate it.
  • Voltage negotiation = quantization. The same cable delivers 5V or 20V depending on what the device asks for; the same graph runs at float32, float16, or int8 depending on your latency and memory budget.

Before USB-C, every phone needed its own cable and every drawer had five of them. Before ONNX, every framework needed its own deployment stack and every team had five of those.

Keep one caution attached to the analogy: plugging in does not make your phone charge faster — the brick does that. Likewise, exporting to ONNX makes a model pluggable; speed comes from whoever executes it. Which brings us to the runtime.

05.Running the File: ONNX Runtime Does the Actual Work

A file cannot execute itself. ONNX Runtime (ORT) is the reference execution engine — the power brick with interchangeable tips.

ORT partitions the graph across execution providers (EPs): CUDA, TensorRT, DirectML, OpenVINO, CoreML, and CPU MLAS kernels. One graph runs on an H100, an Intel laptop, or an iPhone by swapping EPs — the graph itself never changes.

On top of plain "run the graph," ORT layers a few ideas worth knowing by name:

  • Session config: graph optimizations (basic/extended/layout) fuse patterns like Conv-BatchNorm-ReLU into a single kernel at load time. Fewer kernels = fewer trips to memory, the classic GPU win.
  • Dynamic shapes: declare input axes as strings ({0: "batch", 1: "seq"}) instead of re-exporting one file per batch size; runtimes bucket kernels accordingly.
  • IO-binding and zero-copy on GPU remove host round-trips in tight serving loops.
  • Quantization — shrink floats to integers so the same file costs less to run:
    • Dynamic int8: quantize the weights, no calibration dataset needed. Easy, usually good enough.
    • Static PTQ (post-training quantization): run a small calibration dataset first to measure real activation ranges.
    • QAT outputs: results from quantization-aware training, for when PTQ hurts accuracy.
    • ORT quantizes ONNX graphs to int4/int8 for CPU and GPU, and the GenAI package (2024) serves quantized LLMs including AWQ/INT4 paths.
  • Float16 and bfloat16 conversion for small-footprint GPU inference.
python— Export, verify, and run with dynamic batch
import torch, onnxruntime as ort

# 1) Export from a model in eval mode; mark dynamic axes
dummy = torch.randn(1, 3, 224, 224)
torch.onnx.export(model, dummy, "ranker.onnx", opset_version=18,
                  input_names=["image"], dynamic_axes={"image": {0: "batch"}})

# 2) Verify numerical parity against the source framework
import numpy as np
sess = ort.InferenceSession("ranker.onnx", providers=["CUDAExecutionProvider"])
got = sess.run(None, {"image": dummy.numpy()})[0]
ref = model.eval()(dummy).detach().numpy()
print("max abs diff:", np.abs(got - ref).max())   # aim for < 1e-3

06.The Rough Edges: Where Interchange Still Breaks Down

The USB-C dream never fully eliminated friction. Four known rough edges:

  • Operator coverage: exotic ops — custom flash kernels, some grid_sampler variants, dynamic control flow captured by tracing — are not in the standard dictionary. Exporting falls back to a custom/ragged op or splits the graph. In practice, exporting forces your model into the supported subset.
  • Tracing artifacts: the old torch.onnx.export baked Python constants and control flow into the graph structure — dead branches frozen in, hidden assumptions like "batch is always 1." The torch.export-based exporter fixes this class of bugs.
  • Numerical drift: fusion and reordering can shift outputs by small epsilons. Always run parity checks (allclose against the source model) in CI, and treat fp16 ONNX as a separate validation target — half precision drifts more.
  • LLMs: a plain ONNX export of a decoder-only transformer is a non-starter for decode loops (the sampling loop lives outside the graph). The LLM world converged on safetensors weights + engine runtimes (vLLM/TRT-LLM), with ONNX mainly used for encoders, small models, and non-NVIDIA targets (optimum exports to ORT-GenAI).

The mindset: ONNX is a contract problem, not a magic converter. Every export ships with a verified parity number.

07.In Practice: ONNX in the MLOps Pipeline

Where does the file belong in a real pipeline? Practical placement: training framework -> ONNX -> execution.

  1. Export at the end of training as a release artifact — version it in the model registry alongside the checkpoint (one registry version can carry weights, ONNX export, and tokenizer; see the registry topic).
  2. Validate in CI: schema/opset check, parity test on golden inputs, latency benchmark on target hardware. Three cheap gates that catch nearly every export bug.
  3. Optionally compile further: pass ONNX through TensorRT/polygraphy to build an engine for latency-critical GPUs (see the TensorRT topic) — same graph, two execution backends: portable ONNX versus a hardware-locked engine.
  4. Serve via ORT in Triton (onnxruntime backend) or lightweight CPU fleets.

The toolbox, by name:

  • netron — open a file and look at the graph visually.
  • onnx-simplifier — graph surgery: drop dead nodes, fold constants.
  • polygraphy — debug and compare engines against the source graph.
  • onnx Python API — rewrite operators programmatically.

Teams that treat ONNX as "an extra export script" get bitten by silent drift. Teams that treat it as a versioned contract with CI checks ship one artifact across CUDA, CPU, CoreML, OpenVINO, web, and automotive NPUs — and sleep fine.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • One artifact runs across CUDA, CPU, DirectML, CoreML, OpenVINO, web (WASM), and automotive NPUs via execution providers.
  • ORT session optimizations and int8/uint4 quantization cut CPU inference cost dramatically.
  • Opset versioning plus protobuf schemas make release auditing tractable.

Trade-offs & Constraints

  • Custom or unsupported ops break the graph or force Python fallback, negating portability.
  • Tracing-based exports embed shape/control-flow assumptions; verification is mandatory per export.
  • For big autoregressive LLMs ONNX is not the serving fast path; engine runtimes are.
Production Implementation in Big Tech
Microsoft Azure AI / Copilot stack• CPU + GPU portable inference

Azure ML exports trained models once to ONNX and deploys to fleets mixing GPU and CPU nodes through ONNX Runtime execution providers; Office and Edge features reuse the same graphs via onnxruntime-web/node — ONNX as the single release format across silicon.

Staff+ Engineering Takeaways

  • ONNX = self-contained graph (protobuf) + versioned operator library (opsets); runtime-neutral by design.
  • ONNX Runtime executes one graph across CUDA/TensorRT/CPU/CoreML/OpenVINO by swapping execution providers.
  • Quantization (dynamic int8, PTQ, int4) and graph fusion are ONNX-side wins, but every export needs numerical parity validation.
  • For large LLM serving the industry uses engine runtimes (vLLM/TRT-LLM) with safetensors; ONNX shines for everything else.

Topic Knowledge Check

Exercise 1 of 2 • Test your architectural comprehension.

Exercise 1 of 20 answered
1

What does an ONNX opset version pin?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?