## Create a custom model

`client.Models.New(ctx, body) (*InferenceModel, error)`

**post** `/v5/models`

Create a custom model record in your account and begin deploying it through a supported serving vendor.

A model here is a record for a model you deploy and serve through Scale's own inference vendors: only the `launch` and `llmengine` vendors are accepted and any other vendor is rejected. This is distinct from `GET /v5/chat/completions/models`, which lists the models already available to call for chat completions rather than creating or managing these records. The call is asynchronous — the record is created in a deploying status, a deployment job is recorded, and a Temporal workflow is started to perform the deployment, so the model is not ready for inference when this returns. A model name must be unique per vendor within your account; if a model with the same name and vendor already exists the request fails unless `on_conflict` is set to `update`, in which case the existing model is updated instead.

### Parameters

- `body ModelNewParams`

  - `Model param.Field[ModelNewParamsModelUnion]`

    Register a model already served by an external / proxy-served vendor
    (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

    Unlike launch/llmengine, no Scale-side deployment is performed: the record is
    created READY and is immediately callable via /v5/chat/completions. Accepted only
    when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
    covers every vendor except launch/llmengine, and no vendor_configuration applies.

    - `type ModelNewParamsModelLaunch struct{…}`

      - `Name string`

        Unique name to reference your model

      - `VendorConfiguration LaunchVendorConfiguration`

        - `ModelImage LaunchVendorConfigurationModelImage`

          - `Command []string`

          - `Registry string`

          - `Repository string`

          - `Tag string`

          - `EnvVars map[string, any]`

          - `HealthcheckRoute string`

          - `PredictRoute string`

          - `ReadinessDelay int64`

          - `RequestSchema map[string, any]`

          - `ResponseSchema map[string, any]`

          - `StreamingCommand []string`

          - `StreamingPredictRoute string`

        - `ModelInfra LaunchVendorConfigurationModelInfra`

          - `CPUs LaunchVendorConfigurationModelInfraCPUsUnion`

            - `string`

            - `int64`

          - `EndpointType string`

            - `const LaunchVendorConfigurationModelInfraEndpointTypeAsync LaunchVendorConfigurationModelInfraEndpointType = "async"`

            - `const LaunchVendorConfigurationModelInfraEndpointTypeSync LaunchVendorConfigurationModelInfraEndpointType = "sync"`

            - `const LaunchVendorConfigurationModelInfraEndpointTypeStreaming LaunchVendorConfigurationModelInfraEndpointType = "streaming"`

          - `GPUType string`

            - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaTeslaT4 LaunchVendorConfigurationModelInfraGPUType = "nvidia-tesla-t4"`

            - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA10 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a10"`

            - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100"`

            - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100e LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100e"`

            - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100"`

            - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_1g20gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-1g20gb"`

            - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_3g40gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-3g40gb"`

          - `GPUs int64`

          - `HighPriority bool`

          - `Labels map[string, string]`

          - `MaxWorkers int64`

          - `Memory string`

          - `MinWorkers int64`

          - `PerWorker int64`

          - `PublicInference bool`

          - `Storage string`

      - `ModelMetadata map[string, any]`

      - `ModelType string`

        - `const ModelNewParamsModelLaunchModelTypeGeneric ModelNewParamsModelLaunchModelType = "generic"`

      - `ModelVendor string`

        - `const ModelNewParamsModelLaunchModelVendorLaunch ModelNewParamsModelLaunchModelVendor = "launch"`

      - `OnConflict string`

        - `const ModelNewParamsModelLaunchOnConflictError ModelNewParamsModelLaunchOnConflict = "error"`

        - `const ModelNewParamsModelLaunchOnConflictUpdate ModelNewParamsModelLaunchOnConflict = "update"`

    - `type ModelNewParamsModelLlmengine struct{…}`

      - `Name string`

        Unique name to reference your model

      - `VendorConfiguration LlmEngineVendorConfiguration`

        - `Model string`

        - `ChatTemplateOverride string`

        - `CheckpointPath string`

        - `CPUs int64`

        - `DefaultCallbackURL string`

        - `EndpointType string`

        - `GPUType string`

        - `GPUs int64`

        - `HighPriority bool`

        - `InferenceFramework string`

        - `InferenceFrameworkImageTag string`

        - `Labels map[string, string]`

        - `MaxWorkers int64`

        - `Memory string`

        - `MinWorkers int64`

        - `NodesPerWorker int64`

        - `NumShards int64`

        - `PerWorker int64`

        - `PostInferenceHooks []string`

        - `PublicInference bool`

        - `Quantize string`

        - `Source string`

        - `Storage string`

      - `ModelMetadata map[string, any]`

      - `ModelType string`

        - `const ModelNewParamsModelLlmengineModelTypeChatCompletion ModelNewParamsModelLlmengineModelType = "chat_completion"`

      - `ModelVendor string`

        - `const ModelNewParamsModelLlmengineModelVendorLlmengine ModelNewParamsModelLlmengineModelVendor = "llmengine"`

      - `OnConflict string`

        - `const ModelNewParamsModelLlmengineOnConflictError ModelNewParamsModelLlmengineOnConflict = "error"`

        - `const ModelNewParamsModelLlmengineOnConflictUpdate ModelNewParamsModelLlmengineOnConflict = "update"`

    - `type ModelNewParamsModelHostedModelCreateRequest struct{…}`

      Register a model already served by an external / proxy-served vendor
      (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

      Unlike launch/llmengine, no Scale-side deployment is performed: the record is
      created READY and is immediately callable via /v5/chat/completions. Accepted only
      when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
      covers every vendor except launch/llmengine, and no vendor_configuration applies.

      - `ModelType InferenceModelType`

        Type of model, for example `chat_completion`

        - `const InferenceModelTypeGeneric InferenceModelType = "generic"`

        - `const InferenceModelTypeCompletion InferenceModelType = "completion"`

        - `const InferenceModelTypeChatCompletion InferenceModelType = "chat_completion"`

      - `ModelVendor string`

        Vendor to serve/create model

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorOpenAI ModelNewParamsModelHostedModelCreateRequestModelVendor = "openai"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorCohere ModelNewParamsModelHostedModelCreateRequestModelVendor = "cohere"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorVertexAI ModelNewParamsModelHostedModelCreateRequestModelVendor = "vertex_ai"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorAnthropic ModelNewParamsModelHostedModelCreateRequestModelVendor = "anthropic"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorAzure ModelNewParamsModelHostedModelCreateRequestModelVendor = "azure"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorGemini ModelNewParamsModelHostedModelCreateRequestModelVendor = "gemini"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorModelZoo ModelNewParamsModelHostedModelCreateRequestModelVendor = "model_zoo"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorBedrock ModelNewParamsModelHostedModelCreateRequestModelVendor = "bedrock"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorXai ModelNewParamsModelHostedModelCreateRequestModelVendor = "xai"`

        - `const ModelNewParamsModelHostedModelCreateRequestModelVendorFireworksAI ModelNewParamsModelHostedModelCreateRequestModelVendor = "fireworks_ai"`

      - `Name string`

        Unique name to reference your model

      - `ModelMetadata map[string, any]`

      - `OnConflict string`

        - `const ModelNewParamsModelHostedModelCreateRequestOnConflictError ModelNewParamsModelHostedModelCreateRequestOnConflict = "error"`

        - `const ModelNewParamsModelHostedModelCreateRequestOnConflictUpdate ModelNewParamsModelHostedModelCreateRequestOnConflict = "update"`

### Returns

- `type InferenceModel struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedByIdentityType InferenceModelCreatedByIdentityType`

    The type of identity that created the entity.

    - `const InferenceModelCreatedByIdentityTypeUser InferenceModelCreatedByIdentityType = "user"`

    - `const InferenceModelCreatedByIdentityTypeServiceAccount InferenceModelCreatedByIdentityType = "service_account"`

  - `CreatedByUserID string`

    The user who originally created the entity.

  - `ModelType InferenceModelType`

    - `const InferenceModelTypeGeneric InferenceModelType = "generic"`

    - `const InferenceModelTypeCompletion InferenceModelType = "completion"`

    - `const InferenceModelTypeChatCompletion InferenceModelType = "chat_completion"`

  - `ModelVendor InferenceModelVendor`

    - `const InferenceModelVendorOpenAI InferenceModelVendor = "openai"`

    - `const InferenceModelVendorCohere InferenceModelVendor = "cohere"`

    - `const InferenceModelVendorVertexAI InferenceModelVendor = "vertex_ai"`

    - `const InferenceModelVendorAnthropic InferenceModelVendor = "anthropic"`

    - `const InferenceModelVendorAzure InferenceModelVendor = "azure"`

    - `const InferenceModelVendorGemini InferenceModelVendor = "gemini"`

    - `const InferenceModelVendorLaunch InferenceModelVendor = "launch"`

    - `const InferenceModelVendorLlmengine InferenceModelVendor = "llmengine"`

    - `const InferenceModelVendorModelZoo InferenceModelVendor = "model_zoo"`

    - `const InferenceModelVendorBedrock InferenceModelVendor = "bedrock"`

    - `const InferenceModelVendorXai InferenceModelVendor = "xai"`

    - `const InferenceModelVendorFireworksAI InferenceModelVendor = "fireworks_ai"`

  - `Name string`

  - `Status InferenceModelStatus`

    - `const InferenceModelStatusFailed InferenceModelStatus = "failed"`

    - `const InferenceModelStatusReady InferenceModelStatus = "ready"`

    - `const InferenceModelStatusDeploying InferenceModelStatus = "deploying"`

    - `const InferenceModelStatusDeploymentTimeout InferenceModelStatus = "deployment_timeout"`

  - `ModelAvailability InferenceModelAvailability`

    - `const InferenceModelAvailabilityUnknown InferenceModelAvailability = "unknown"`

    - `const InferenceModelAvailabilityAvailable InferenceModelAvailability = "available"`

    - `const InferenceModelAvailabilityUnavailable InferenceModelAvailability = "unavailable"`

  - `ModelMetadata map[string, any]`

  - `Object InferenceModelObject`

    - `const InferenceModelObjectModel InferenceModelObject = "model"`

  - `StatusReason string`

  - `VendorConfiguration InferenceModelVendorConfigurationUnion`

    - `type LaunchVendorConfiguration struct{…}`

      - `ModelImage LaunchVendorConfigurationModelImage`

        - `Command []string`

        - `Registry string`

        - `Repository string`

        - `Tag string`

        - `EnvVars map[string, any]`

        - `HealthcheckRoute string`

        - `PredictRoute string`

        - `ReadinessDelay int64`

        - `RequestSchema map[string, any]`

        - `ResponseSchema map[string, any]`

        - `StreamingCommand []string`

        - `StreamingPredictRoute string`

      - `ModelInfra LaunchVendorConfigurationModelInfra`

        - `CPUs LaunchVendorConfigurationModelInfraCPUsUnion`

          - `string`

          - `int64`

        - `EndpointType string`

          - `const LaunchVendorConfigurationModelInfraEndpointTypeAsync LaunchVendorConfigurationModelInfraEndpointType = "async"`

          - `const LaunchVendorConfigurationModelInfraEndpointTypeSync LaunchVendorConfigurationModelInfraEndpointType = "sync"`

          - `const LaunchVendorConfigurationModelInfraEndpointTypeStreaming LaunchVendorConfigurationModelInfraEndpointType = "streaming"`

        - `GPUType string`

          - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaTeslaT4 LaunchVendorConfigurationModelInfraGPUType = "nvidia-tesla-t4"`

          - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA10 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a10"`

          - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100"`

          - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100e LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100e"`

          - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100"`

          - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_1g20gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-1g20gb"`

          - `const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_3g40gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-3g40gb"`

        - `GPUs int64`

        - `HighPriority bool`

        - `Labels map[string, string]`

        - `MaxWorkers int64`

        - `Memory string`

        - `MinWorkers int64`

        - `PerWorker int64`

        - `PublicInference bool`

        - `Storage string`

    - `type LlmEngineVendorConfiguration struct{…}`

      - `Model string`

      - `ChatTemplateOverride string`

      - `CheckpointPath string`

      - `CPUs int64`

      - `DefaultCallbackURL string`

      - `EndpointType string`

      - `GPUType string`

      - `GPUs int64`

      - `HighPriority bool`

      - `InferenceFramework string`

      - `InferenceFrameworkImageTag string`

      - `Labels map[string, string]`

      - `MaxWorkers int64`

      - `Memory string`

      - `MinWorkers int64`

      - `NodesPerWorker int64`

      - `NumShards int64`

      - `PerWorker int64`

      - `PostInferenceHooks []string`

      - `PublicInference bool`

      - `Quantize string`

      - `Source string`

      - `Storage string`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  inferenceModel, err := client.Models.New(context.TODO(), sgpdev.ModelNewParams{
    OfLaunch: &sgpdev.ModelNewParamsModelLaunch{
      Name: "name",
      VendorConfiguration: sgpdev.LaunchVendorConfigurationParam{
        ModelImage: sgpdev.LaunchVendorConfigurationModelImageParam{
          Command: []string{"string"},
          Registry: "registry",
          Repository: "repository",
          Tag: "tag",
        },
        ModelInfra: sgpdev.LaunchVendorConfigurationModelInfraParam{

        },
      },
      ModelVendor: "launch",
    },
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", inferenceModel.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```
