# Inference

## Run inference with free-form payload

`client.inference.create(InferenceCreateParamsbody, RequestOptionsoptions?): InferenceCreateResponse`

**post** `/v5/inference`

Runs a model using a free-form, vendor-native request payload rather than a fixed OpenAI schema.

Use this endpoint when the target model does not fit the OpenAI chat, responses, or
text-completion contracts: the `args` field is an arbitrary dict passed straight through to the
selected vendor gateway, and the reply is returned inside `response` as arbitrary JSON (object,
array, string, number, or boolean). Prefer /v5/chat/completions, /v5/responses, or /v5/completions
when your request matches one of those OpenAI-standard shapes, since those return typed OpenAI
response objects. The model is chosen from the `model` field formatted as `vendor/name`, which
selects the per-vendor inference gateway. When the request enables streaming, the response is sent
as server-sent events with each chunk wrapped as a `generic_inference.chunk` object; otherwise a
single `generic_inference` object is returned.

### Parameters

- `body: InferenceCreateParams`

  - `model: string`

    model specified as `vendor/name` (ex. openai/gpt-5)

  - `args?: Record<string, unknown>`

    Arguments passed into model

  - `inference_configuration?: LaunchInferenceConfiguration`

    Vendor specific configuration

    - `num_retries?: number`

    - `timeout_seconds?: number`

### Returns

- `InferenceCreateResponse = InferenceResponse | InferenceResponseChunk`

  - `InferenceResponse`

    - `response: Record<string, unknown> | Array<unknown> | string | 2 more | null`

      - `Record<string, unknown>`

      - `Array<unknown>`

      - `string`

      - `number`

      - `boolean`

    - `object?: "generic_inference"`

      - `"generic_inference"`

  - `InferenceResponseChunk`

    - `response: Record<string, unknown> | Array<unknown> | string | 2 more | null`

      - `Record<string, unknown>`

      - `Array<unknown>`

      - `string`

      - `number`

      - `boolean`

    - `object?: "generic_inference.chunk"`

      - `"generic_inference.chunk"`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const inference = await client.inference.create({ model: 'model' });

console.log(inference);
```

#### Response

```json
{
  "response": {
    "foo": "bar"
  },
  "object": "generic_inference"
}
```

## Domain Types

### Inference Response

- `InferenceResponse`

  - `response: Record<string, unknown> | Array<unknown> | string | 2 more | null`

    - `Record<string, unknown>`

    - `Array<unknown>`

    - `string`

    - `number`

    - `boolean`

  - `object?: "generic_inference"`

    - `"generic_inference"`

### Inference Response Chunk

- `InferenceResponseChunk`

  - `response: Record<string, unknown> | Array<unknown> | string | 2 more | null`

    - `Record<string, unknown>`

    - `Array<unknown>`

    - `string`

    - `number`

    - `boolean`

  - `object?: "generic_inference.chunk"`

    - `"generic_inference.chunk"`

### Launch Inference Configuration

- `LaunchInferenceConfiguration`

  - `num_retries?: number`

  - `timeout_seconds?: number`

### Inference Create Response

- `InferenceCreateResponse = InferenceResponse | InferenceResponseChunk`

  - `InferenceResponse`

    - `response: Record<string, unknown> | Array<unknown> | string | 2 more | null`

      - `Record<string, unknown>`

      - `Array<unknown>`

      - `string`

      - `number`

      - `boolean`

    - `object?: "generic_inference"`

      - `"generic_inference"`

  - `InferenceResponseChunk`

    - `response: Record<string, unknown> | Array<unknown> | string | 2 more | null`

      - `Record<string, unknown>`

      - `Array<unknown>`

      - `string`

      - `number`

      - `boolean`

    - `object?: "generic_inference.chunk"`

      - `"generic_inference.chunk"`
