# Models

## Create a custom model

`client.models.create(ModelCreateParamsparams, RequestOptionsoptions?): InferenceModel`

**post** `/v5/models`

Create a custom model record in your account and begin deploying it through a supported serving vendor.

A model here is a record for a model you deploy and serve through Scale's own inference vendors: only the `launch` and `llmengine` vendors are accepted and any other vendor is rejected. This is distinct from `GET /v5/chat/completions/models`, which lists the models already available to call for chat completions rather than creating or managing these records. The call is asynchronous — the record is created in a deploying status, a deployment job is recorded, and a Temporal workflow is started to perform the deployment, so the model is not ready for inference when this returns. A model name must be unique per vendor within your account; if a model with the same name and vendor already exists the request fails unless `on_conflict` is set to `update`, in which case the existing model is updated instead.

### Parameters

- `params: ModelCreateParams`

  - `model: LaunchModelCreateRequest | LlmEngineModelCreateRequest | HostedModelCreateRequest`

    Register a model already served by an external / proxy-served vendor
    (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

    Unlike launch/llmengine, no Scale-side deployment is performed: the record is
    created READY and is immediately callable via /v5/chat/completions. Accepted only
    when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
    covers every vendor except launch/llmengine, and no vendor_configuration applies.

    - `LaunchModelCreateRequest`

      - `name: string`

        Unique name to reference your model

      - `vendor_configuration: LaunchVendorConfiguration`

        - `model_image: ModelImage`

          - `command: Array<string>`

          - `registry: string`

          - `repository: string`

          - `tag: string`

          - `env_vars?: Record<string, unknown>`

          - `healthcheck_route?: string`

          - `predict_route?: string`

          - `readiness_delay?: number`

          - `request_schema?: Record<string, unknown>`

          - `response_schema?: Record<string, unknown>`

          - `streaming_command?: Array<string>`

          - `streaming_predict_route?: string`

        - `model_infra: ModelInfra`

          - `cpus?: string | number`

            - `string`

            - `number`

          - `endpoint_type?: "async" | "sync" | "streaming"`

            - `"async"`

            - `"sync"`

            - `"streaming"`

          - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

            - `"nvidia-tesla-t4"`

            - `"nvidia-ampere-a10"`

            - `"nvidia-ampere-a100"`

            - `"nvidia-ampere-a100e"`

            - `"nvidia-hopper-h100"`

            - `"nvidia-hopper-h100-1g20gb"`

            - `"nvidia-hopper-h100-3g40gb"`

          - `gpus?: number`

          - `high_priority?: boolean`

          - `labels?: Record<string, string>`

          - `max_workers?: number`

          - `memory?: string`

          - `min_workers?: number`

          - `per_worker?: number`

          - `public_inference?: boolean`

          - `storage?: string`

      - `model_metadata?: Record<string, unknown>`

      - `model_type?: "generic"`

        - `"generic"`

      - `model_vendor?: "launch"`

        - `"launch"`

      - `on_conflict?: "error" | "update"`

        - `"error"`

        - `"update"`

    - `LlmEngineModelCreateRequest`

      - `name: string`

        Unique name to reference your model

      - `vendor_configuration: LlmEngineVendorConfiguration`

        - `model: string`

        - `chat_template_override?: string`

        - `checkpoint_path?: string`

        - `cpus?: number`

        - `default_callback_url?: string`

        - `endpoint_type?: string`

        - `gpu_type?: string`

        - `gpus?: number`

        - `high_priority?: boolean`

        - `inference_framework?: string`

        - `inference_framework_image_tag?: string`

        - `labels?: Record<string, string>`

        - `max_workers?: number`

        - `memory?: string`

        - `min_workers?: number`

        - `nodes_per_worker?: number`

        - `num_shards?: number`

        - `per_worker?: number`

        - `post_inference_hooks?: Array<string>`

        - `public_inference?: boolean`

        - `quantize?: string`

        - `source?: string`

        - `storage?: string`

      - `model_metadata?: Record<string, unknown>`

      - `model_type?: "chat_completion"`

        - `"chat_completion"`

      - `model_vendor?: "llmengine"`

        - `"llmengine"`

      - `on_conflict?: "error" | "update"`

        - `"error"`

        - `"update"`

    - `HostedModelCreateRequest`

      Register a model already served by an external / proxy-served vendor
      (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

      Unlike launch/llmengine, no Scale-side deployment is performed: the record is
      created READY and is immediately callable via /v5/chat/completions. Accepted only
      when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
      covers every vendor except launch/llmengine, and no vendor_configuration applies.

      - `model_type: InferenceModelType`

        Type of model, for example `chat_completion`

        - `"generic"`

        - `"completion"`

        - `"chat_completion"`

      - `model_vendor: "openai" | "cohere" | "vertex_ai" | 7 more`

        Vendor to serve/create model

        - `"openai"`

        - `"cohere"`

        - `"vertex_ai"`

        - `"anthropic"`

        - `"azure"`

        - `"gemini"`

        - `"model_zoo"`

        - `"bedrock"`

        - `"xai"`

        - `"fireworks_ai"`

      - `name: string`

        Unique name to reference your model

      - `model_metadata?: Record<string, unknown>`

      - `on_conflict?: "error" | "update"`

        - `"error"`

        - `"update"`

### Returns

- `InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" | "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" | "ready" | "deploying" | "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability?: InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata?: Record<string, unknown>`

  - `object?: "model"`

    - `"model"`

  - `status_reason?: string`

  - `vendor_configuration?: LaunchVendorConfiguration | LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration`

      - `model_image: ModelImage`

        - `command: Array<string>`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars?: Record<string, unknown>`

        - `healthcheck_route?: string`

        - `predict_route?: string`

        - `readiness_delay?: number`

        - `request_schema?: Record<string, unknown>`

        - `response_schema?: Record<string, unknown>`

        - `streaming_command?: Array<string>`

        - `streaming_predict_route?: string`

      - `model_infra: ModelInfra`

        - `cpus?: string | number`

          - `string`

          - `number`

        - `endpoint_type?: "async" | "sync" | "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus?: number`

        - `high_priority?: boolean`

        - `labels?: Record<string, string>`

        - `max_workers?: number`

        - `memory?: string`

        - `min_workers?: number`

        - `per_worker?: number`

        - `public_inference?: boolean`

        - `storage?: string`

    - `LlmEngineVendorConfiguration`

      - `model: string`

      - `chat_template_override?: string`

      - `checkpoint_path?: string`

      - `cpus?: number`

      - `default_callback_url?: string`

      - `endpoint_type?: string`

      - `gpu_type?: string`

      - `gpus?: number`

      - `high_priority?: boolean`

      - `inference_framework?: string`

      - `inference_framework_image_tag?: string`

      - `labels?: Record<string, string>`

      - `max_workers?: number`

      - `memory?: string`

      - `min_workers?: number`

      - `nodes_per_worker?: number`

      - `num_shards?: number`

      - `per_worker?: number`

      - `post_inference_hooks?: Array<string>`

      - `public_inference?: boolean`

      - `quantize?: string`

      - `source?: string`

      - `storage?: string`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const inferenceModel = await client.models.create({
  model: {
    name: 'name',
    vendor_configuration: {
      model_image: {
        command: ['string'],
        registry: 'registry',
        repository: 'repository',
        tag: 'tag',
      },
      model_infra: {},
    },
    model_vendor: 'launch',
  },
});

console.log(inferenceModel.id);
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## List custom models

`client.models.list(ModelListParamsquery?, RequestOptionsoptions?): CursorPage<InferenceModel>`

**get** `/v5/models`

List the custom model records registered in your account.

Returns a paginated list of the model records managed through this API — models your account deploys through the `launch` or `llmengine` serving vendors — optionally filtered by name and by model vendor, and scoped to the caller's account. This is different from `GET /v5/chat/completions/models`, which lists the models available to invoke for chat completions; this endpoint returns the managed records along with their deployment status, not the catalog of callable completion models.

### Parameters

- `query: ModelListParams`

  - `ending_before?: string`

  - `limit?: number`

  - `model_vendor?: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name?: string`

  - `sort_by?: string`

  - `sort_order?: SortOrder`

    - `"asc"`

    - `"desc"`

  - `starting_after?: string`

### Returns

- `InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" | "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" | "ready" | "deploying" | "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability?: InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata?: Record<string, unknown>`

  - `object?: "model"`

    - `"model"`

  - `status_reason?: string`

  - `vendor_configuration?: LaunchVendorConfiguration | LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration`

      - `model_image: ModelImage`

        - `command: Array<string>`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars?: Record<string, unknown>`

        - `healthcheck_route?: string`

        - `predict_route?: string`

        - `readiness_delay?: number`

        - `request_schema?: Record<string, unknown>`

        - `response_schema?: Record<string, unknown>`

        - `streaming_command?: Array<string>`

        - `streaming_predict_route?: string`

      - `model_infra: ModelInfra`

        - `cpus?: string | number`

          - `string`

          - `number`

        - `endpoint_type?: "async" | "sync" | "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus?: number`

        - `high_priority?: boolean`

        - `labels?: Record<string, string>`

        - `max_workers?: number`

        - `memory?: string`

        - `min_workers?: number`

        - `per_worker?: number`

        - `public_inference?: boolean`

        - `storage?: string`

    - `LlmEngineVendorConfiguration`

      - `model: string`

      - `chat_template_override?: string`

      - `checkpoint_path?: string`

      - `cpus?: number`

      - `default_callback_url?: string`

      - `endpoint_type?: string`

      - `gpu_type?: string`

      - `gpus?: number`

      - `high_priority?: boolean`

      - `inference_framework?: string`

      - `inference_framework_image_tag?: string`

      - `labels?: Record<string, string>`

      - `max_workers?: number`

      - `memory?: string`

      - `min_workers?: number`

      - `nodes_per_worker?: number`

      - `num_shards?: number`

      - `per_worker?: number`

      - `post_inference_hooks?: Array<string>`

      - `public_inference?: boolean`

      - `quantize?: string`

      - `source?: string`

      - `storage?: string`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

// Automatically fetches more pages as needed.
for await (const inferenceModel of client.models.list()) {
  console.log(inferenceModel.id);
}
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by_identity_type": "user",
      "created_by_user_id": "created_by_user_id",
      "model_type": "generic",
      "model_vendor": "openai",
      "name": "name",
      "status": "failed",
      "model_availability": "unknown",
      "model_metadata": {
        "foo": "bar"
      },
      "object": "model",
      "status_reason": "status_reason",
      "vendor_configuration": {
        "model_image": {
          "command": [
            "string"
          ],
          "registry": "registry",
          "repository": "repository",
          "tag": "tag",
          "env_vars": {
            "foo": "bar"
          },
          "healthcheck_route": "healthcheck_route",
          "predict_route": "predict_route",
          "readiness_delay": 0,
          "request_schema": {
            "foo": "bar"
          },
          "response_schema": {
            "foo": "bar"
          },
          "streaming_command": [
            "string"
          ],
          "streaming_predict_route": "streaming_predict_route"
        },
        "model_infra": {
          "cpus": "string",
          "endpoint_type": "async",
          "gpu_type": "nvidia-tesla-t4",
          "gpus": 0,
          "high_priority": true,
          "labels": {
            "foo": "string"
          },
          "max_workers": 0,
          "memory": "memory",
          "min_workers": 0,
          "per_worker": 0,
          "public_inference": true,
          "storage": "storage"
        }
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Update a custom model

`client.models.update(stringmodelID, ModelUpdateParamsparams, RequestOptionsoptions?): InferenceModel`

**patch** `/v5/models/{model_id}`

Update a custom model record; vendor-configuration changes are applied asynchronously by redeploying the model.

This supports three kinds of update: changing model metadata only, renaming the model, and changing the vendor configuration. A vendor-configuration change is asynchronous — it puts the model back into a deploying status, records an update job, and starts a Temporal workflow to redeploy, so the new configuration is not live when this returns; metadata-only and rename changes take effect immediately. The vendor configuration supplied must match the model's own vendor (`launch` or `llmengine`), and only those two vendors are supported. A model that is currently deploying cannot be modified and the request fails until deployment finishes. When renaming with `on_conflict` set to `swap`, the name is exchanged with an existing model of the same name and vendor instead of failing on the uniqueness constraint.

### Parameters

- `modelID: string`

- `params: ModelUpdateParams`

  - `model: DefaultModelPatchRequest | ModelConfigurationPatchRequest | SwapNamesModelPatchRequest`

    - `DefaultModelPatchRequest`

      - `model_metadata?: Record<string, unknown>`

    - `ModelConfigurationPatchRequest`

      - `vendor_configuration: PartialLaunchVendorConfiguration | PartialLlmEngineVendorConfiguration`

        - `PartialLaunchVendorConfiguration`

          - `model_image?: ModelImage`

            - `command?: Array<string>`

            - `env_vars?: Record<string, unknown>`

            - `healthcheck_route?: string`

            - `predict_route?: string`

            - `readiness_delay?: number`

            - `registry?: string`

            - `repository?: string`

            - `request_schema?: Record<string, unknown>`

            - `response_schema?: Record<string, unknown>`

            - `streaming_command?: Array<string>`

            - `streaming_predict_route?: string`

            - `tag?: string`

          - `model_infra?: ModelInfra`

            - `cpus?: string | number`

              - `string`

              - `number`

            - `endpoint_type?: "async" | "sync" | "streaming"`

              - `"async"`

              - `"sync"`

              - `"streaming"`

            - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

              - `"nvidia-tesla-t4"`

              - `"nvidia-ampere-a10"`

              - `"nvidia-ampere-a100"`

              - `"nvidia-ampere-a100e"`

              - `"nvidia-hopper-h100"`

              - `"nvidia-hopper-h100-1g20gb"`

              - `"nvidia-hopper-h100-3g40gb"`

            - `gpus?: number`

            - `high_priority?: boolean`

            - `labels?: Record<string, string>`

            - `max_workers?: number`

            - `memory?: string`

            - `min_workers?: number`

            - `per_worker?: number`

            - `public_inference?: boolean`

            - `storage?: string`

        - `PartialLlmEngineVendorConfiguration`

          - `chat_template_override?: string`

          - `checkpoint_path?: string`

          - `cpus?: number`

          - `default_callback_url?: string`

          - `endpoint_type?: string`

          - `gpu_type?: string`

          - `gpus?: number`

          - `high_priority?: boolean`

          - `inference_framework?: string`

          - `inference_framework_image_tag?: string`

          - `labels?: Record<string, string>`

          - `max_workers?: number`

          - `memory?: string`

          - `min_workers?: number`

          - `model?: string`

          - `nodes_per_worker?: number`

          - `num_shards?: number`

          - `per_worker?: number`

          - `post_inference_hooks?: Array<string>`

          - `public_inference?: boolean`

          - `quantize?: string`

          - `source?: string`

          - `storage?: string`

      - `model_metadata?: Record<string, unknown>`

    - `SwapNamesModelPatchRequest`

      - `name: string`

      - `on_conflict?: "error" | "swap"`

        - `"error"`

        - `"swap"`

### Returns

- `InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" | "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" | "ready" | "deploying" | "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability?: InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata?: Record<string, unknown>`

  - `object?: "model"`

    - `"model"`

  - `status_reason?: string`

  - `vendor_configuration?: LaunchVendorConfiguration | LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration`

      - `model_image: ModelImage`

        - `command: Array<string>`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars?: Record<string, unknown>`

        - `healthcheck_route?: string`

        - `predict_route?: string`

        - `readiness_delay?: number`

        - `request_schema?: Record<string, unknown>`

        - `response_schema?: Record<string, unknown>`

        - `streaming_command?: Array<string>`

        - `streaming_predict_route?: string`

      - `model_infra: ModelInfra`

        - `cpus?: string | number`

          - `string`

          - `number`

        - `endpoint_type?: "async" | "sync" | "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus?: number`

        - `high_priority?: boolean`

        - `labels?: Record<string, string>`

        - `max_workers?: number`

        - `memory?: string`

        - `min_workers?: number`

        - `per_worker?: number`

        - `public_inference?: boolean`

        - `storage?: string`

    - `LlmEngineVendorConfiguration`

      - `model: string`

      - `chat_template_override?: string`

      - `checkpoint_path?: string`

      - `cpus?: number`

      - `default_callback_url?: string`

      - `endpoint_type?: string`

      - `gpu_type?: string`

      - `gpus?: number`

      - `high_priority?: boolean`

      - `inference_framework?: string`

      - `inference_framework_image_tag?: string`

      - `labels?: Record<string, string>`

      - `max_workers?: number`

      - `memory?: string`

      - `min_workers?: number`

      - `nodes_per_worker?: number`

      - `num_shards?: number`

      - `per_worker?: number`

      - `post_inference_hooks?: Array<string>`

      - `public_inference?: boolean`

      - `quantize?: string`

      - `source?: string`

      - `storage?: string`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const inferenceModel = await client.models.update('model_id', { model: {} });

console.log(inferenceModel.id);
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## Delete a custom model

`client.models.delete(stringmodelID, RequestOptionsoptions?): ModelDeleteResponse`

**delete** `/v5/models/{model_id}`

Permanently delete a custom model record and tear down its deployment at the serving vendor.

This is a hard delete: the model row is removed from your account entirely and cannot be restored afterward. Before the record is removed, if the model has an associated vendor deployment that deployment is torn down at its serving vendor. A model that is currently deploying cannot be deleted and the request fails until deployment finishes. This operates on the model records managed by this API, distinct from the `GET /v5/chat/completions/models` catalog of models callable for chat completions.

### Parameters

- `modelID: string`

### Returns

- `ModelDeleteResponse`

  - `id: string`

  - `deleted: boolean`

  - `object?: "model"`

    - `"model"`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const model = await client.models.delete('model_id');

console.log(model.id);
```

#### Response

```json
{
  "id": "id",
  "deleted": true,
  "object": "model"
}
```

## Get a custom model

`client.models.retrieve(stringmodelID, RequestOptionsoptions?): InferenceModel`

**get** `/v5/models/{model_id}`

Retrieve a single custom model record by its ID.

Returns the model record — including its vendor, configuration, and current deployment status — for a model managed through this API and owned by the caller's account. This is distinct from `GET /v5/chat/completions/models`, which lists the models available to call for chat completions rather than returning a single managed record.

### Parameters

- `modelID: string`

### Returns

- `InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" | "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" | "ready" | "deploying" | "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability?: InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata?: Record<string, unknown>`

  - `object?: "model"`

    - `"model"`

  - `status_reason?: string`

  - `vendor_configuration?: LaunchVendorConfiguration | LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration`

      - `model_image: ModelImage`

        - `command: Array<string>`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars?: Record<string, unknown>`

        - `healthcheck_route?: string`

        - `predict_route?: string`

        - `readiness_delay?: number`

        - `request_schema?: Record<string, unknown>`

        - `response_schema?: Record<string, unknown>`

        - `streaming_command?: Array<string>`

        - `streaming_predict_route?: string`

      - `model_infra: ModelInfra`

        - `cpus?: string | number`

          - `string`

          - `number`

        - `endpoint_type?: "async" | "sync" | "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus?: number`

        - `high_priority?: boolean`

        - `labels?: Record<string, string>`

        - `max_workers?: number`

        - `memory?: string`

        - `min_workers?: number`

        - `per_worker?: number`

        - `public_inference?: boolean`

        - `storage?: string`

    - `LlmEngineVendorConfiguration`

      - `model: string`

      - `chat_template_override?: string`

      - `checkpoint_path?: string`

      - `cpus?: number`

      - `default_callback_url?: string`

      - `endpoint_type?: string`

      - `gpu_type?: string`

      - `gpus?: number`

      - `high_priority?: boolean`

      - `inference_framework?: string`

      - `inference_framework_image_tag?: string`

      - `labels?: Record<string, string>`

      - `max_workers?: number`

      - `memory?: string`

      - `min_workers?: number`

      - `nodes_per_worker?: number`

      - `num_shards?: number`

      - `per_worker?: number`

      - `post_inference_hooks?: Array<string>`

      - `public_inference?: boolean`

      - `quantize?: string`

      - `source?: string`

      - `storage?: string`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const inferenceModel = await client.models.retrieve('model_id');

console.log(inferenceModel.id);
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## Domain Types

### Inference Model

- `InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" | "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" | "ready" | "deploying" | "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability?: InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata?: Record<string, unknown>`

  - `object?: "model"`

    - `"model"`

  - `status_reason?: string`

  - `vendor_configuration?: LaunchVendorConfiguration | LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration`

      - `model_image: ModelImage`

        - `command: Array<string>`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars?: Record<string, unknown>`

        - `healthcheck_route?: string`

        - `predict_route?: string`

        - `readiness_delay?: number`

        - `request_schema?: Record<string, unknown>`

        - `response_schema?: Record<string, unknown>`

        - `streaming_command?: Array<string>`

        - `streaming_predict_route?: string`

      - `model_infra: ModelInfra`

        - `cpus?: string | number`

          - `string`

          - `number`

        - `endpoint_type?: "async" | "sync" | "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus?: number`

        - `high_priority?: boolean`

        - `labels?: Record<string, string>`

        - `max_workers?: number`

        - `memory?: string`

        - `min_workers?: number`

        - `per_worker?: number`

        - `public_inference?: boolean`

        - `storage?: string`

    - `LlmEngineVendorConfiguration`

      - `model: string`

      - `chat_template_override?: string`

      - `checkpoint_path?: string`

      - `cpus?: number`

      - `default_callback_url?: string`

      - `endpoint_type?: string`

      - `gpu_type?: string`

      - `gpus?: number`

      - `high_priority?: boolean`

      - `inference_framework?: string`

      - `inference_framework_image_tag?: string`

      - `labels?: Record<string, string>`

      - `max_workers?: number`

      - `memory?: string`

      - `min_workers?: number`

      - `nodes_per_worker?: number`

      - `num_shards?: number`

      - `per_worker?: number`

      - `post_inference_hooks?: Array<string>`

      - `public_inference?: boolean`

      - `quantize?: string`

      - `source?: string`

      - `storage?: string`

### Inference Model Availability

- `InferenceModelAvailability = "unknown" | "available" | "unavailable"`

  - `"unknown"`

  - `"available"`

  - `"unavailable"`

### Inference Model Type

- `InferenceModelType = "generic" | "completion" | "chat_completion"`

  - `"generic"`

  - `"completion"`

  - `"chat_completion"`

### Launch Vendor Configuration

- `LaunchVendorConfiguration`

  - `model_image: ModelImage`

    - `command: Array<string>`

    - `registry: string`

    - `repository: string`

    - `tag: string`

    - `env_vars?: Record<string, unknown>`

    - `healthcheck_route?: string`

    - `predict_route?: string`

    - `readiness_delay?: number`

    - `request_schema?: Record<string, unknown>`

    - `response_schema?: Record<string, unknown>`

    - `streaming_command?: Array<string>`

    - `streaming_predict_route?: string`

  - `model_infra: ModelInfra`

    - `cpus?: string | number`

      - `string`

      - `number`

    - `endpoint_type?: "async" | "sync" | "streaming"`

      - `"async"`

      - `"sync"`

      - `"streaming"`

    - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

      - `"nvidia-tesla-t4"`

      - `"nvidia-ampere-a10"`

      - `"nvidia-ampere-a100"`

      - `"nvidia-ampere-a100e"`

      - `"nvidia-hopper-h100"`

      - `"nvidia-hopper-h100-1g20gb"`

      - `"nvidia-hopper-h100-3g40gb"`

    - `gpus?: number`

    - `high_priority?: boolean`

    - `labels?: Record<string, string>`

    - `max_workers?: number`

    - `memory?: string`

    - `min_workers?: number`

    - `per_worker?: number`

    - `public_inference?: boolean`

    - `storage?: string`

### Llm Engine Vendor Configuration

- `LlmEngineVendorConfiguration`

  - `model: string`

  - `chat_template_override?: string`

  - `checkpoint_path?: string`

  - `cpus?: number`

  - `default_callback_url?: string`

  - `endpoint_type?: string`

  - `gpu_type?: string`

  - `gpus?: number`

  - `high_priority?: boolean`

  - `inference_framework?: string`

  - `inference_framework_image_tag?: string`

  - `labels?: Record<string, string>`

  - `max_workers?: number`

  - `memory?: string`

  - `min_workers?: number`

  - `nodes_per_worker?: number`

  - `num_shards?: number`

  - `per_worker?: number`

  - `post_inference_hooks?: Array<string>`

  - `public_inference?: boolean`

  - `quantize?: string`

  - `source?: string`

  - `storage?: string`

### Model Delete Response

- `ModelDeleteResponse`

  - `id: string`

  - `deleted: boolean`

  - `object?: "model"`

    - `"model"`
