Performance
Precision
Choose FP32, FP16, or quantized paths deliberately and verify numerical behavior.
Start with FP32. Move to faster precision after correctness and shape behavior are known.
FP32 Baseline
migraphx-driver verify model.onnx --gpu
migraphx-driver perf model.onnx --gpuRecord this as the correctness baseline.
FP16
FP16 can improve throughput and memory bandwidth on supported AMD GPUs. Validate it with real samples, not only generated input.
Suggested process:
1. run FP32 reference
2. compile or configure FP16 path
3. compare outputs on real validation data
4. check detection metrics or task-specific accuracy
5. benchmark latency and throughputQuantization
Quantized inference requires more model-specific validation. Do not enable quantized paths only because they are faster in a synthetic benchmark.
Track:
calibration data
calibration method
operator coverage
accuracy delta
per-model enablement
fallback policyProduction Rule
Every precision mode should have a recorded output tolerance:
model: yolo11s
precision: fp16
metric: mAP or task-specific score
max accepted delta: agreed threshold
sample set: exact dataset revision