TensorRT To MIGraphX Backend
Design notes for adding an AMD ROCm backend beside a TensorRT backend in Argus-style C++ inference pipelines.
MIGraphX is useful for the AMD inference part of an Argus-style pipeline. It does not replace FFmpeg, decode, encode, stream output, preprocessing, or postprocessing.
Concept Mapping
| NVIDIA path | AMD path |
|---|---|
| CUDA runtime | HIP/ROCm runtime |
| TensorRT builder | MIGraphX parse and compile |
| TensorRT engine file | saved MIGraphX program, often .mxr |
ICudaEngine / execution context | migraphx::program |
| CUDA device buffers | MIGraphX arguments and ROCm/HIP memory integration |
enqueueV3 | program eval/run |
Recommended Boundary
Create a backend-neutral interface:
struct InferenceOutput
{
std::vector<float> data;
std::size_t elements_per_frame = 0;
};
class InferenceBackend
{
public:
virtual ~InferenceBackend() = default;
virtual bool ready() const = 0;
virtual std::size_t output_elements_per_frame() const = 0;
virtual InferenceOutput run(const float* input, int batch_size) = 0;
};Then implement:
TensorRtBackend
MIGraphXBackendKeep TensorRT and MIGraphX headers out of shared pipeline headers. Backend-specific types should live in backend-specific .cxx files where possible.
Build Flow
TensorRT:
ONNX -> TensorRT build -> .engineMIGraphX:
ONNX -> MIGraphX compile(gpu) -> .mxrCLI shape:
./argus --backend migraphx --build \
--model input/models/yolo.onnx \
--engine output/yolo.mxr \
--batch 16Run shape:
./argus --backend migraphx \
--engine output/yolo.mxr \
--input rtsp://source \
--output rtsp://localhost:8554/argusCMake Shape
cmake -B build-amd -DARGUS_BACKEND=migraphx
cmake --build build-amdAvoid silently choosing the backend only from the detected GPU. Build hosts and deployment hosts can differ. Use auto-detect only as a local convenience.
Performance Expectations
MIGraphX is the right AMD-native component to evaluate for optimized inference. Expect it to beat CPU inference on supported AMD GPUs. Do not assume it will match a TensorRT result for every model without benchmarking.
Benchmark both paths with:
same model export
same input size
same batch size
same precision
same NMS location
same postprocessFirst AMD Proof
Before refactoring the whole pipeline:
migraphx-driver compile input/models/yolo.onnx --gpu --save output/yolo.mxr
migraphx-driver perf output/yolo.mxrThen write a small standalone C++ smoke test that loads output/yolo.mxr, runs one generated or captured tensor, and confirms the output shape your postprocess expects.