Skip to content

Run inference with free-form payload

client.inference.create(InferenceCreateParams { model, args, inference_configuration } body, RequestOptionsoptions?): InferenceCreateResponse
POST/v5/inference

Runs a model using a free-form, vendor-native request payload rather than a fixed OpenAI schema.

Use this endpoint when the target model does not fit the OpenAI chat, responses, or text-completion contracts: the args field is an arbitrary dict passed straight through to the selected vendor gateway, and the reply is returned inside response as arbitrary JSON (object, array, string, number, or boolean). Prefer /v5/chat/completions, /v5/responses, or /v5/completions when your request matches one of those OpenAI-standard shapes, since those return typed OpenAI response objects. The model is chosen from the model field formatted as vendor/name, which selects the per-vendor inference gateway. When the request enables streaming, the response is sent as server-sent events with each chunk wrapped as a generic_inference.chunk object; otherwise a single generic_inference object is returned.

ParametersExpand Collapse
body: InferenceCreateParams { model, args, inference_configuration }
model: string

model specified as vendor/name (ex. openai/gpt-5)

args?: Record<string, unknown>

Arguments passed into model

inference_configuration?: LaunchInferenceConfiguration { num_retries, timeout_seconds }

Vendor specific configuration

num_retries?: number
timeout_seconds?: number
ReturnsExpand Collapse
InferenceCreateResponse = InferenceResponse { response, object } | InferenceResponseChunk { response, object }
One of the following:
InferenceResponse { response, object }
response: Record<string, unknown> | Array<unknown> | string | 2 more | null
One of the following:
Record<string, unknown>
Array<unknown>
string
number
boolean
object?: "generic_inference"
InferenceResponseChunk { response, object }
response: Record<string, unknown> | Array<unknown> | string | 2 more | null
One of the following:
Record<string, unknown>
Array<unknown>
string
number
boolean
object?: "generic_inference.chunk"

Run inference with free-form payload

import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const inference = await client.inference.create({ model: 'model' });

console.log(inference);
{
  "response": {
    "foo": "bar"
  },
  "object": "generic_inference"
}
Returns Examples
{
  "response": {
    "foo": "bar"
  },
  "object": "generic_inference"
}