Skip to content

Run inference with free-form payload

inference.create(InferenceCreateParams**kwargs) -> InferenceCreateResponse
POST/v5/inference

Runs a model using a free-form, vendor-native request payload rather than a fixed OpenAI schema.

Use this endpoint when the target model does not fit the OpenAI chat, responses, or text-completion contracts: the args field is an arbitrary dict passed straight through to the selected vendor gateway, and the reply is returned inside response as arbitrary JSON (object, array, string, number, or boolean). Prefer /v5/chat/completions, /v5/responses, or /v5/completions when your request matches one of those OpenAI-standard shapes, since those return typed OpenAI response objects. The model is chosen from the model field formatted as vendor/name, which selects the per-vendor inference gateway. When the request enables streaming, the response is sent as server-sent events with each chunk wrapped as a generic_inference.chunk object; otherwise a single generic_inference object is returned.

ParametersExpand Collapse
model: str

model specified as vendor/name (ex. openai/gpt-5)

args: Optional[Dict[str, object]]

Arguments passed into model

inference_configuration: Optional[LaunchInferenceConfigurationParam]

Vendor specific configuration

num_retries: Optional[int]
timeout_seconds: Optional[int]
ReturnsExpand Collapse
One of the following:
class InferenceResponse: …
response: Union[Dict[str, object], List[object], str, 3 more]
One of the following:
Dict[str, object]
List[object]
str
float
bool
object: Optional[Literal["generic_inference"]]
class InferenceResponseChunk: …
response: Union[Dict[str, object], List[object], str, 3 more]
One of the following:
Dict[str, object]
List[object]
str
float
bool
object: Optional[Literal["generic_inference.chunk"]]

Run inference with free-form payload

import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
inference = client.inference.create(
    model="model",
)
print(inference)
{
  "response": {
    "foo": "bar"
  },
  "object": "generic_inference"
}
Returns Examples
{
  "response": {
    "foo": "bar"
  },
  "object": "generic_inference"
}