MIGraphX Field Guide

TensorRT To MIGraphX Backend

Design notes for adding an AMD ROCm backend beside a TensorRT backend in Argus-style C++ inference pipelines.

MIGraphX is useful for the AMD inference part of an Argus-style pipeline. It does not replace FFmpeg, decode, encode, stream output, preprocessing, or postprocessing.

Concept Mapping

NVIDIA pathAMD path
CUDA runtimeHIP/ROCm runtime
TensorRT builderMIGraphX parse and compile
TensorRT engine filesaved MIGraphX program, often .mxr
ICudaEngine / execution contextmigraphx::program
CUDA device buffersMIGraphX arguments and ROCm/HIP memory integration
enqueueV3program eval/run

Create a backend-neutral interface:

struct InferenceOutput
{
    std::vector<float> data;
    std::size_t elements_per_frame = 0;
};

class InferenceBackend
{
  public:
    virtual ~InferenceBackend() = default;
    virtual bool ready() const = 0;
    virtual std::size_t output_elements_per_frame() const = 0;
    virtual InferenceOutput run(const float* input, int batch_size) = 0;
};

Then implement:

TensorRtBackend
MIGraphXBackend

Keep TensorRT and MIGraphX headers out of shared pipeline headers. Backend-specific types should live in backend-specific .cxx files where possible.

Build Flow

TensorRT:

ONNX -> TensorRT build -> .engine

MIGraphX:

ONNX -> MIGraphX compile(gpu) -> .mxr

CLI shape:

./argus --backend migraphx --build \
  --model input/models/yolo.onnx \
  --engine output/yolo.mxr \
  --batch 16

Run shape:

./argus --backend migraphx \
  --engine output/yolo.mxr \
  --input rtsp://source \
  --output rtsp://localhost:8554/argus

CMake Shape

cmake -B build-amd -DARGUS_BACKEND=migraphx
cmake --build build-amd

Avoid silently choosing the backend only from the detected GPU. Build hosts and deployment hosts can differ. Use auto-detect only as a local convenience.

Performance Expectations

MIGraphX is the right AMD-native component to evaluate for optimized inference. Expect it to beat CPU inference on supported AMD GPUs. Do not assume it will match a TensorRT result for every model without benchmarking.

Benchmark both paths with:

same model export
same input size
same batch size
same precision
same NMS location
same postprocess

First AMD Proof

Before refactoring the whole pipeline:

migraphx-driver compile input/models/yolo.onnx --gpu --save output/yolo.mxr
migraphx-driver perf output/yolo.mxr

Then write a small standalone C++ smoke test that loads output/yolo.mxr, runs one generated or captured tensor, and confirms the output shape your postprocess expects.

On this page