C++ API
Embed MIGraphX in a native C++ application with parse, compile, save, load, and eval.
Use the C++ API when MIGraphX is part of a native service, stream processor, or performance-sensitive application.
Minimal Run
#include <migraphx/migraphx.hpp>
#include <iostream>
#include <map>
#include <string>
int main()
{
migraphx::program program = migraphx::parse_onnx("model.onnx");
program.compile(migraphx::target{"gpu"});
std::map<std::string, migraphx::argument> inputs;
for(const auto& name : program.get_parameter_names())
{
auto shape = program.get_parameter_shape(name);
inputs[name] = migraphx::generate_argument(shape);
}
auto outputs = program.eval(inputs);
std::cout << "outputs: " << outputs.size() << "\n";
}This is good for a smoke test. Production code should fill arguments with real preprocessed input data.
Build Once, Save Program
#include <migraphx/migraphx.hpp>
int main()
{
migraphx::program program = migraphx::parse_onnx("model.onnx");
program.compile(migraphx::target{"gpu"});
migraphx::save(program, "model.mxr");
}Load And Run
#include <migraphx/migraphx.hpp>
#include <map>
#include <string>
int main()
{
migraphx::program program = migraphx::load("model.mxr");
std::map<std::string, migraphx::argument> inputs;
for(const auto& name : program.get_parameter_names())
{
auto shape = program.get_parameter_shape(name);
inputs[name] = migraphx::generate_argument(shape);
}
auto outputs = program.eval(inputs);
(void)outputs;
}Working With Real Buffers
Most services already have a contiguous input tensor after preprocessing:
float input[batch * channels * height * width]The integration task is to create or fill a MIGraphX argument that matches the model parameter shape, then call eval or the equivalent run function. Keep these checks close to startup:
input element count == program input shape element count
input type == model input type
output shape == postprocess expectation
batch <= compiled batch capacityError Handling
Fail fast at load time:
missing model artifact
failed load
input names do not match
unexpected output rank
unexpected output element typeAvoid discovering these in the hot inference loop.
Threading Pattern
For a pipeline such as video inference:
one compiled program per inference worker
preprocess queue -> input tensor -> program eval -> output queueDo not share mutable input/output buffers across workers without explicit synchronization. If you need multiple inference threads, benchmark one program per thread against a single batching worker before assuming more threads improve throughput.