# Inference

## Run inference with free-form payload

`client.Inference.New(ctx, body) (*InferenceNewResponseUnion, error)`

**post** `/v5/inference`

Runs a model using a free-form, vendor-native request payload rather than a fixed OpenAI schema.

Use this endpoint when the target model does not fit the OpenAI chat, responses, or
text-completion contracts: the `args` field is an arbitrary dict passed straight through to the
selected vendor gateway, and the reply is returned inside `response` as arbitrary JSON (object,
array, string, number, or boolean). Prefer /v5/chat/completions, /v5/responses, or /v5/completions
when your request matches one of those OpenAI-standard shapes, since those return typed OpenAI
response objects. The model is chosen from the `model` field formatted as `vendor/name`, which
selects the per-vendor inference gateway. When the request enables streaming, the response is sent
as server-sent events with each chunk wrapped as a `generic_inference.chunk` object; otherwise a
single `generic_inference` object is returned.

### Parameters

- `body InferenceNewParams`

  - `Model param.Field[string]`

    model specified as `vendor/name` (ex. openai/gpt-5)

  - `Args param.Field[map[string, any]]`

    Arguments passed into model

  - `InferenceConfiguration param.Field[LaunchInferenceConfiguration]`

    Vendor specific configuration

### Returns

- `type InferenceNewResponseUnion interface{…}`

  - `type InferenceResponse struct{…}`

    - `Response InferenceResponseResponseUnion`

      - `type InferenceResponseResponseMap map[string, any]`

      - `type InferenceResponseResponseArray []any`

      - `string`

      - `float64`

      - `bool`

    - `Object InferenceResponseObject`

      - `const InferenceResponseObjectGenericInference InferenceResponseObject = "generic_inference"`

  - `type InferenceResponseChunk struct{…}`

    - `Response InferenceResponseChunkResponseUnion`

      - `type InferenceResponseChunkResponseMap map[string, any]`

      - `type InferenceResponseChunkResponseArray []any`

      - `string`

      - `float64`

      - `bool`

    - `Object InferenceResponseChunkObject`

      - `const InferenceResponseChunkObjectGenericInferenceChunk InferenceResponseChunkObject = "generic_inference.chunk"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  inference, err := client.Inference.New(context.TODO(), sgpdev.InferenceNewParams{
    Model: "model",
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", inference)
}
```

#### Response

```json
{
  "response": {
    "foo": "bar"
  },
  "object": "generic_inference"
}
```

## Domain Types

### Inference Response

- `type InferenceResponse struct{…}`

  - `Response InferenceResponseResponseUnion`

    - `type InferenceResponseResponseMap map[string, any]`

    - `type InferenceResponseResponseArray []any`

    - `string`

    - `float64`

    - `bool`

  - `Object InferenceResponseObject`

    - `const InferenceResponseObjectGenericInference InferenceResponseObject = "generic_inference"`

### Inference Response Chunk

- `type InferenceResponseChunk struct{…}`

  - `Response InferenceResponseChunkResponseUnion`

    - `type InferenceResponseChunkResponseMap map[string, any]`

    - `type InferenceResponseChunkResponseArray []any`

    - `string`

    - `float64`

    - `bool`

  - `Object InferenceResponseChunkObject`

    - `const InferenceResponseChunkObjectGenericInferenceChunk InferenceResponseChunkObject = "generic_inference.chunk"`

### Launch Inference Configuration

- `type LaunchInferenceConfiguration struct{…}`

  - `NumRetries int64`

  - `TimeoutSeconds int64`
