Performance
Benchmarking
Measure MIGraphX latency and throughput in a way that maps to production behavior.
Benchmarking should answer the same question your service asks: how many real inputs can be processed at the latency budget you care about?
Driver Benchmark
Start with:
migraphx-driver perf model.onnx --gpuThen benchmark the saved artifact:
migraphx-driver compile model.onnx --gpu --save model.mxr
migraphx-driver perf model.mxrThe saved-artifact benchmark is closer to production startup behavior.
Application Benchmark
Measure the full path separately:
decode
preprocess
host-to-device or argument fill
MIGraphX execution
device-to-host or output read
postprocess
encode/outputIf model execution is fast but the full pipeline is slow, the bottleneck is outside MIGraphX.
Batch Sweep
For video or stream inference:
batch=1
batch=2
batch=4
batch=8
batch=16Record:
median latency
p95 latency
throughput
GPU utilization
memory usage
input queue depthWarm-Up
Always warm up before recording numbers:
load program
run N warm-up iterations
start timer
run measured iterations
stop timerCompare Fairly
When comparing TensorRT and MIGraphX:
same model architecture
same input size
same batch size
same precision
same NMS location
same postprocess code where possible
same stream count and frame sourceDo not compare TensorRT FP16 against MIGraphX FP32 and call it a backend result.