Performance
Production Runtime
Runtime patterns for using MIGraphX in services and high-throughput pipelines.
Production performance usually comes from keeping the inference path boring.
Startup
Do this at startup:
load saved program
inspect parameter names and shapes
allocate reusable input/output storage
run warm-up iterations
publish readiness only after warm-up succeedsHot Path
Keep the hot path short:
read next batch
fill input argument
run program
copy or read output
push result to postprocessAvoid:
- parsing ONNX in the request loop
- compiling in the request loop
- allocating large buffers per frame
- logging per inference at info level
- changing shapes every batch
Failure Modes
Treat these as startup failures:
saved program missing
wrong model version
unexpected input name
unexpected output shape
GPU target unavailable
ROCm runtime unavailableTreat these as runtime failures:
input queue timeout
bad frame shape
GPU execution error
output validation failureKeep logs separate so operational debugging is not mixed with model compatibility debugging.