MIGraphX Field Guide
Performance

Benchmarking

Measure MIGraphX latency and throughput in a way that maps to production behavior.

Benchmarking should answer the same question your service asks: how many real inputs can be processed at the latency budget you care about?

Driver Benchmark

Start with:

migraphx-driver perf model.onnx --gpu

Then benchmark the saved artifact:

migraphx-driver compile model.onnx --gpu --save model.mxr
migraphx-driver perf model.mxr

The saved-artifact benchmark is closer to production startup behavior.

Application Benchmark

Measure the full path separately:

decode
preprocess
host-to-device or argument fill
MIGraphX execution
device-to-host or output read
postprocess
encode/output

If model execution is fast but the full pipeline is slow, the bottleneck is outside MIGraphX.

Batch Sweep

For video or stream inference:

batch=1
batch=2
batch=4
batch=8
batch=16

Record:

median latency
p95 latency
throughput
GPU utilization
memory usage
input queue depth

Warm-Up

Always warm up before recording numbers:

load program
run N warm-up iterations
start timer
run measured iterations
stop timer

Compare Fairly

When comparing TensorRT and MIGraphX:

same model architecture
same input size
same batch size
same precision
same NMS location
same postprocess code where possible
same stream count and frame source

Do not compare TensorRT FP16 against MIGraphX FP32 and call it a backend result.

On this page