AMD GPU inference documentation

MIGraphX, explained for people building real inference systems.

A practical field guide for installing MIGraphX, compiling ONNX models, calling it from C++ and Python, running containers, and designing an AMD backend beside TensorRT.

Input

model.onnx + fp32 tensors

Compile

parse_onnx -> compile(gpu)

Deploy

load .mxr -> run batches

Run a model

Compile ONNX once, load the saved program, and execute inference on an AMD GPU.

Embed in C++

Use parse, compile, save, load, and eval from a native service or video pipeline.

Port from TensorRT

Map CUDA plus TensorRT concepts to HIP plus MIGraphX for an AMD backend.