## Update a custom model

`client.models.update(stringmodelID, ModelUpdateParamsparams, RequestOptionsoptions?): InferenceModel`

**patch** `/v5/models/{model_id}`

Update a custom model record; vendor-configuration changes are applied asynchronously by redeploying the model.

This supports three kinds of update: changing model metadata only, renaming the model, and changing the vendor configuration. A vendor-configuration change is asynchronous — it puts the model back into a deploying status, records an update job, and starts a Temporal workflow to redeploy, so the new configuration is not live when this returns; metadata-only and rename changes take effect immediately. The vendor configuration supplied must match the model's own vendor (`launch` or `llmengine`), and only those two vendors are supported. A model that is currently deploying cannot be modified and the request fails until deployment finishes. When renaming with `on_conflict` set to `swap`, the name is exchanged with an existing model of the same name and vendor instead of failing on the uniqueness constraint.

### Parameters

- `modelID: string`

- `params: ModelUpdateParams`

  - `model: DefaultModelPatchRequest | ModelConfigurationPatchRequest | SwapNamesModelPatchRequest`

    - `DefaultModelPatchRequest`

      - `model_metadata?: Record<string, unknown>`

    - `ModelConfigurationPatchRequest`

      - `vendor_configuration: PartialLaunchVendorConfiguration | PartialLlmEngineVendorConfiguration`

        - `PartialLaunchVendorConfiguration`

          - `model_image?: ModelImage`

            - `command?: Array<string>`

            - `env_vars?: Record<string, unknown>`

            - `healthcheck_route?: string`

            - `predict_route?: string`

            - `readiness_delay?: number`

            - `registry?: string`

            - `repository?: string`

            - `request_schema?: Record<string, unknown>`

            - `response_schema?: Record<string, unknown>`

            - `streaming_command?: Array<string>`

            - `streaming_predict_route?: string`

            - `tag?: string`

          - `model_infra?: ModelInfra`

            - `cpus?: string | number`

              - `string`

              - `number`

            - `endpoint_type?: "async" | "sync" | "streaming"`

              - `"async"`

              - `"sync"`

              - `"streaming"`

            - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

              - `"nvidia-tesla-t4"`

              - `"nvidia-ampere-a10"`

              - `"nvidia-ampere-a100"`

              - `"nvidia-ampere-a100e"`

              - `"nvidia-hopper-h100"`

              - `"nvidia-hopper-h100-1g20gb"`

              - `"nvidia-hopper-h100-3g40gb"`

            - `gpus?: number`

            - `high_priority?: boolean`

            - `labels?: Record<string, string>`

            - `max_workers?: number`

            - `memory?: string`

            - `min_workers?: number`

            - `per_worker?: number`

            - `public_inference?: boolean`

            - `storage?: string`

        - `PartialLlmEngineVendorConfiguration`

          - `chat_template_override?: string`

          - `checkpoint_path?: string`

          - `cpus?: number`

          - `default_callback_url?: string`

          - `endpoint_type?: string`

          - `gpu_type?: string`

          - `gpus?: number`

          - `high_priority?: boolean`

          - `inference_framework?: string`

          - `inference_framework_image_tag?: string`

          - `labels?: Record<string, string>`

          - `max_workers?: number`

          - `memory?: string`

          - `min_workers?: number`

          - `model?: string`

          - `nodes_per_worker?: number`

          - `num_shards?: number`

          - `per_worker?: number`

          - `post_inference_hooks?: Array<string>`

          - `public_inference?: boolean`

          - `quantize?: string`

          - `source?: string`

          - `storage?: string`

      - `model_metadata?: Record<string, unknown>`

    - `SwapNamesModelPatchRequest`

      - `name: string`

      - `on_conflict?: "error" | "swap"`

        - `"error"`

        - `"swap"`

### Returns

- `InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" | "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" | "ready" | "deploying" | "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability?: InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata?: Record<string, unknown>`

  - `object?: "model"`

    - `"model"`

  - `status_reason?: string`

  - `vendor_configuration?: LaunchVendorConfiguration | LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration`

      - `model_image: ModelImage`

        - `command: Array<string>`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars?: Record<string, unknown>`

        - `healthcheck_route?: string`

        - `predict_route?: string`

        - `readiness_delay?: number`

        - `request_schema?: Record<string, unknown>`

        - `response_schema?: Record<string, unknown>`

        - `streaming_command?: Array<string>`

        - `streaming_predict_route?: string`

      - `model_infra: ModelInfra`

        - `cpus?: string | number`

          - `string`

          - `number`

        - `endpoint_type?: "async" | "sync" | "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus?: number`

        - `high_priority?: boolean`

        - `labels?: Record<string, string>`

        - `max_workers?: number`

        - `memory?: string`

        - `min_workers?: number`

        - `per_worker?: number`

        - `public_inference?: boolean`

        - `storage?: string`

    - `LlmEngineVendorConfiguration`

      - `model: string`

      - `chat_template_override?: string`

      - `checkpoint_path?: string`

      - `cpus?: number`

      - `default_callback_url?: string`

      - `endpoint_type?: string`

      - `gpu_type?: string`

      - `gpus?: number`

      - `high_priority?: boolean`

      - `inference_framework?: string`

      - `inference_framework_image_tag?: string`

      - `labels?: Record<string, string>`

      - `max_workers?: number`

      - `memory?: string`

      - `min_workers?: number`

      - `nodes_per_worker?: number`

      - `num_shards?: number`

      - `per_worker?: number`

      - `post_inference_hooks?: Array<string>`

      - `public_inference?: boolean`

      - `quantize?: string`

      - `source?: string`

      - `storage?: string`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const inferenceModel = await client.models.update('model_id', { model: {} });

console.log(inferenceModel.id);
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```
