Choose the Four Components
The command’s positional order is Benchmark, Harness, Model:
sample_ids selects the task, and concurrency 1 makes the initial run easier to inspect. Once it completes, use the command builder to configure a larger evaluation.
Choose an Interface
Both interfaces use the same runtime and run configuration. Save repeated settings in a YAML file and load it with
--config; see configuration files and precedence. The model under test is supplied through the command, an orchestration request, or the SDK.
