MIGraphX Field Guide
Languages

C++ API

Embed MIGraphX in a native C++ application with parse, compile, save, load, and eval.

Use the C++ API when MIGraphX is part of a native service, stream processor, or performance-sensitive application.

Minimal Run

#include <migraphx/migraphx.hpp>

#include <iostream>
#include <map>
#include <string>

int main()
{
    migraphx::program program = migraphx::parse_onnx("model.onnx");
    program.compile(migraphx::target{"gpu"});

    std::map<std::string, migraphx::argument> inputs;
    for(const auto& name : program.get_parameter_names())
    {
        auto shape = program.get_parameter_shape(name);
        inputs[name] = migraphx::generate_argument(shape);
    }

    auto outputs = program.eval(inputs);
    std::cout << "outputs: " << outputs.size() << "\n";
}

This is good for a smoke test. Production code should fill arguments with real preprocessed input data.

Build Once, Save Program

#include <migraphx/migraphx.hpp>

int main()
{
    migraphx::program program = migraphx::parse_onnx("model.onnx");
    program.compile(migraphx::target{"gpu"});
    migraphx::save(program, "model.mxr");
}

Load And Run

#include <migraphx/migraphx.hpp>

#include <map>
#include <string>

int main()
{
    migraphx::program program = migraphx::load("model.mxr");

    std::map<std::string, migraphx::argument> inputs;
    for(const auto& name : program.get_parameter_names())
    {
        auto shape = program.get_parameter_shape(name);
        inputs[name] = migraphx::generate_argument(shape);
    }

    auto outputs = program.eval(inputs);
    (void)outputs;
}

Working With Real Buffers

Most services already have a contiguous input tensor after preprocessing:

float input[batch * channels * height * width]

The integration task is to create or fill a MIGraphX argument that matches the model parameter shape, then call eval or the equivalent run function. Keep these checks close to startup:

input element count == program input shape element count
input type == model input type
output shape == postprocess expectation
batch <= compiled batch capacity

Error Handling

Fail fast at load time:

missing model artifact
failed load
input names do not match
unexpected output rank
unexpected output element type

Avoid discovering these in the hot inference loop.

Threading Pattern

For a pipeline such as video inference:

one compiled program per inference worker
preprocess queue -> input tensor -> program eval -> output queue

Do not share mutable input/output buffers across workers without explicit synchronization. If you need multiple inference threads, benchmark one program per thread against a single batching worker before assuming more threads improve throughput.

On this page