MIGraphX Field Guide
Practical AMD GPU inference documentation for builders using MIGraphX, ROCm, ONNX, C++, Python, Docker, and production services.
MIGraphX is AMD's graph compiler and inference runtime for optimized model execution on AMD GPUs through ROCm. Use it when you want the AMD equivalent of a compiled inference backend instead of running models through an unoptimized CPU or generic runtime path.
This guide is written as an implementation manual. It focuses on the operations you need in real projects:
- install ROCm and MIGraphX
- validate an AMD GPU runtime
- compile ONNX models for the
gputarget - save and load compiled programs
- call MIGraphX from the command line, C++, and Python
- run in Docker containers
- benchmark latency and throughput
- design an AMD backend beside TensorRT
Core Mental Model
model.onnx
-> MIGraphX parser
-> graph optimization and lowering
-> compile for target: gpu
-> saved compiled program
-> runtime execution with input tensorsFor a C++ application, the practical flow usually looks like this:
build mode:
parse ONNX -> compile for gpu -> save program
run mode:
load program -> allocate input/output arguments -> eval/runThat maps cleanly to services that already have a TensorRT-style build and run split.
Where To Start
Quick start
Install, validate, compile, and run a model.
C++ API
Embed MIGraphX in a native inference service.
Argus backend notes
Port a TensorRT video pipeline to an AMD backend.
Docker
Run MIGraphX inside ROCm containers.
What This Guide Is Not
This is not a replacement for AMD's official reference. It is a field guide that reorganizes the official material into workflows. When exact API signatures matter, prefer the upstream reference linked from the source links page.