# Models

## Create a custom model

`models.create(ModelCreateParams**kwargs)  -> InferenceModel`

**post** `/v5/models`

Create a custom model record in your account and begin deploying it through a supported serving vendor.

A model here is a record for a model you deploy and serve through Scale's own inference vendors: only the `launch` and `llmengine` vendors are accepted and any other vendor is rejected. This is distinct from `GET /v5/chat/completions/models`, which lists the models already available to call for chat completions rather than creating or managing these records. The call is asynchronous — the record is created in a deploying status, a deployment job is recorded, and a Temporal workflow is started to perform the deployment, so the model is not ready for inference when this returns. A model name must be unique per vendor within your account; if a model with the same name and vendor already exists the request fails unless `on_conflict` is set to `update`, in which case the existing model is updated instead.

### Parameters

- `model: Model`

  Register a model already served by an external / proxy-served vendor
  (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

  Unlike launch/llmengine, no Scale-side deployment is performed: the record is
  created READY and is immediately callable via /v5/chat/completions. Accepted only
  when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
  covers every vendor except launch/llmengine, and no vendor_configuration applies.

  - `class ModelLaunchModelCreateRequest: …`

    - `name: str`

      Unique name to reference your model

    - `vendor_configuration: LaunchVendorConfigurationParam`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `model_metadata: Optional[Dict[str, object]]`

    - `model_type: Optional[Literal["generic"]]`

      - `"generic"`

    - `model_vendor: Optional[Literal["launch"]]`

      - `"launch"`

    - `on_conflict: Optional[Literal["error", "update"]]`

      - `"error"`

      - `"update"`

  - `class ModelLlmEngineModelCreateRequest: …`

    - `name: str`

      Unique name to reference your model

    - `vendor_configuration: LlmEngineVendorConfigurationParam`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

    - `model_metadata: Optional[Dict[str, object]]`

    - `model_type: Optional[Literal["chat_completion"]]`

      - `"chat_completion"`

    - `model_vendor: Optional[Literal["llmengine"]]`

      - `"llmengine"`

    - `on_conflict: Optional[Literal["error", "update"]]`

      - `"error"`

      - `"update"`

  - `class ModelHostedModelCreateRequest: …`

    Register a model already served by an external / proxy-served vendor
    (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

    Unlike launch/llmengine, no Scale-side deployment is performed: the record is
    created READY and is immediately callable via /v5/chat/completions. Accepted only
    when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
    covers every vendor except launch/llmengine, and no vendor_configuration applies.

    - `model_type: InferenceModelType`

      Type of model, for example `chat_completion`

      - `"generic"`

      - `"completion"`

      - `"chat_completion"`

    - `model_vendor: Literal["openai", "cohere", "vertex_ai", 7 more]`

      Vendor to serve/create model

      - `"openai"`

      - `"cohere"`

      - `"vertex_ai"`

      - `"anthropic"`

      - `"azure"`

      - `"gemini"`

      - `"model_zoo"`

      - `"bedrock"`

      - `"xai"`

      - `"fireworks_ai"`

    - `name: str`

      Unique name to reference your model

    - `model_metadata: Optional[Dict[str, object]]`

    - `on_conflict: Optional[Literal["error", "update"]]`

      - `"error"`

      - `"update"`

### Returns

- `class InferenceModel: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: Literal["user", "service_account"]`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: str`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: str`

  - `status: Literal["failed", "ready", "deploying", "deployment_timeout"]`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: Optional[InferenceModelAvailability]`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: Optional[Dict[str, object]]`

  - `object: Optional[Literal["model"]]`

    - `"model"`

  - `status_reason: Optional[str]`

  - `vendor_configuration: Optional[VendorConfiguration]`

    - `class LaunchVendorConfiguration: …`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `class LlmEngineVendorConfiguration: …`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
inference_model = client.models.create(
    model={
        "name": "name",
        "vendor_configuration": {
            "model_image": {
                "command": ["string"],
                "registry": "registry",
                "repository": "repository",
                "tag": "tag",
            },
            "model_infra": {},
        },
        "model_vendor": "launch",
    },
)
print(inference_model.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## List custom models

`models.list(ModelListParams**kwargs)  -> SyncCursorPage[InferenceModel]`

**get** `/v5/models`

List the custom model records registered in your account.

Returns a paginated list of the model records managed through this API — models your account deploys through the `launch` or `llmengine` serving vendors — optionally filtered by name and by model vendor, and scoped to the caller's account. This is different from `GET /v5/chat/completions/models`, which lists the models available to invoke for chat completions; this endpoint returns the managed records along with their deployment status, not the catalog of callable completion models.

### Parameters

- `ending_before: Optional[str]`

- `limit: Optional[int]`

- `model_vendor: Optional[InferenceModelVendor]`

  - `"openai"`

  - `"cohere"`

  - `"vertex_ai"`

  - `"anthropic"`

  - `"azure"`

  - `"gemini"`

  - `"launch"`

  - `"llmengine"`

  - `"model_zoo"`

  - `"bedrock"`

  - `"xai"`

  - `"fireworks_ai"`

- `name: Optional[str]`

- `sort_by: Optional[str]`

- `sort_order: Optional[SortOrder]`

  - `"asc"`

  - `"desc"`

- `starting_after: Optional[str]`

### Returns

- `class InferenceModel: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: Literal["user", "service_account"]`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: str`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: str`

  - `status: Literal["failed", "ready", "deploying", "deployment_timeout"]`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: Optional[InferenceModelAvailability]`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: Optional[Dict[str, object]]`

  - `object: Optional[Literal["model"]]`

    - `"model"`

  - `status_reason: Optional[str]`

  - `vendor_configuration: Optional[VendorConfiguration]`

    - `class LaunchVendorConfiguration: …`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `class LlmEngineVendorConfiguration: …`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
page = client.models.list()
page = page.items[0]
print(page.id)
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by_identity_type": "user",
      "created_by_user_id": "created_by_user_id",
      "model_type": "generic",
      "model_vendor": "openai",
      "name": "name",
      "status": "failed",
      "model_availability": "unknown",
      "model_metadata": {
        "foo": "bar"
      },
      "object": "model",
      "status_reason": "status_reason",
      "vendor_configuration": {
        "model_image": {
          "command": [
            "string"
          ],
          "registry": "registry",
          "repository": "repository",
          "tag": "tag",
          "env_vars": {
            "foo": "bar"
          },
          "healthcheck_route": "healthcheck_route",
          "predict_route": "predict_route",
          "readiness_delay": 0,
          "request_schema": {
            "foo": "bar"
          },
          "response_schema": {
            "foo": "bar"
          },
          "streaming_command": [
            "string"
          ],
          "streaming_predict_route": "streaming_predict_route"
        },
        "model_infra": {
          "cpus": "string",
          "endpoint_type": "async",
          "gpu_type": "nvidia-tesla-t4",
          "gpus": 0,
          "high_priority": true,
          "labels": {
            "foo": "string"
          },
          "max_workers": 0,
          "memory": "memory",
          "min_workers": 0,
          "per_worker": 0,
          "public_inference": true,
          "storage": "storage"
        }
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Update a custom model

`models.update(strmodel_id, ModelUpdateParams**kwargs)  -> InferenceModel`

**patch** `/v5/models/{model_id}`

Update a custom model record; vendor-configuration changes are applied asynchronously by redeploying the model.

This supports three kinds of update: changing model metadata only, renaming the model, and changing the vendor configuration. A vendor-configuration change is asynchronous — it puts the model back into a deploying status, records an update job, and starts a Temporal workflow to redeploy, so the new configuration is not live when this returns; metadata-only and rename changes take effect immediately. The vendor configuration supplied must match the model's own vendor (`launch` or `llmengine`), and only those two vendors are supported. A model that is currently deploying cannot be modified and the request fails until deployment finishes. When renaming with `on_conflict` set to `swap`, the name is exchanged with an existing model of the same name and vendor instead of failing on the uniqueness constraint.

### Parameters

- `model_id: str`

- `model: Model`

  - `class ModelDefaultModelPatchRequest: …`

    - `model_metadata: Optional[Dict[str, object]]`

  - `class ModelModelConfigurationPatchRequest: …`

    - `vendor_configuration: ModelModelConfigurationPatchRequestVendorConfiguration`

      - `class ModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfiguration: …`

        - `model_image: Optional[ModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelImage]`

          - `command: Optional[Sequence[str]]`

          - `env_vars: Optional[Dict[str, object]]`

          - `healthcheck_route: Optional[str]`

          - `predict_route: Optional[str]`

          - `readiness_delay: Optional[int]`

          - `registry: Optional[str]`

          - `repository: Optional[str]`

          - `request_schema: Optional[Dict[str, object]]`

          - `response_schema: Optional[Dict[str, object]]`

          - `streaming_command: Optional[Sequence[str]]`

          - `streaming_predict_route: Optional[str]`

          - `tag: Optional[str]`

        - `model_infra: Optional[ModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfra]`

          - `cpus: Optional[Union[str, int]]`

            - `str`

            - `int`

          - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

            - `"async"`

            - `"sync"`

            - `"streaming"`

          - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

            - `"nvidia-tesla-t4"`

            - `"nvidia-ampere-a10"`

            - `"nvidia-ampere-a100"`

            - `"nvidia-ampere-a100e"`

            - `"nvidia-hopper-h100"`

            - `"nvidia-hopper-h100-1g20gb"`

            - `"nvidia-hopper-h100-3g40gb"`

          - `gpus: Optional[int]`

          - `high_priority: Optional[bool]`

          - `labels: Optional[Dict[str, str]]`

          - `max_workers: Optional[int]`

          - `memory: Optional[str]`

          - `min_workers: Optional[int]`

          - `per_worker: Optional[int]`

          - `public_inference: Optional[bool]`

          - `storage: Optional[str]`

      - `class ModelModelConfigurationPatchRequestVendorConfigurationPartialLlmEngineVendorConfiguration: …`

        - `chat_template_override: Optional[str]`

        - `checkpoint_path: Optional[str]`

        - `cpus: Optional[int]`

        - `default_callback_url: Optional[str]`

        - `endpoint_type: Optional[str]`

        - `gpu_type: Optional[str]`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `inference_framework: Optional[str]`

        - `inference_framework_image_tag: Optional[str]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `model: Optional[str]`

        - `nodes_per_worker: Optional[int]`

        - `num_shards: Optional[int]`

        - `per_worker: Optional[int]`

        - `post_inference_hooks: Optional[Sequence[str]]`

        - `public_inference: Optional[bool]`

        - `quantize: Optional[str]`

        - `source: Optional[str]`

        - `storage: Optional[str]`

    - `model_metadata: Optional[Dict[str, object]]`

  - `class ModelSwapNamesModelPatchRequest: …`

    - `name: str`

    - `on_conflict: Optional[Literal["error", "swap"]]`

      - `"error"`

      - `"swap"`

### Returns

- `class InferenceModel: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: Literal["user", "service_account"]`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: str`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: str`

  - `status: Literal["failed", "ready", "deploying", "deployment_timeout"]`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: Optional[InferenceModelAvailability]`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: Optional[Dict[str, object]]`

  - `object: Optional[Literal["model"]]`

    - `"model"`

  - `status_reason: Optional[str]`

  - `vendor_configuration: Optional[VendorConfiguration]`

    - `class LaunchVendorConfiguration: …`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `class LlmEngineVendorConfiguration: …`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
inference_model = client.models.update(
    model_id="model_id",
    model={},
)
print(inference_model.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## Delete a custom model

`models.delete(strmodel_id)  -> ModelDeleteResponse`

**delete** `/v5/models/{model_id}`

Permanently delete a custom model record and tear down its deployment at the serving vendor.

This is a hard delete: the model row is removed from your account entirely and cannot be restored afterward. Before the record is removed, if the model has an associated vendor deployment that deployment is torn down at its serving vendor. A model that is currently deploying cannot be deleted and the request fails until deployment finishes. This operates on the model records managed by this API, distinct from the `GET /v5/chat/completions/models` catalog of models callable for chat completions.

### Parameters

- `model_id: str`

### Returns

- `class ModelDeleteResponse: …`

  - `id: str`

  - `deleted: bool`

  - `object: Optional[Literal["model"]]`

    - `"model"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
model = client.models.delete(
    "model_id",
)
print(model.id)
```

#### Response

```json
{
  "id": "id",
  "deleted": true,
  "object": "model"
}
```

## Get a custom model

`models.retrieve(strmodel_id)  -> InferenceModel`

**get** `/v5/models/{model_id}`

Retrieve a single custom model record by its ID.

Returns the model record — including its vendor, configuration, and current deployment status — for a model managed through this API and owned by the caller's account. This is distinct from `GET /v5/chat/completions/models`, which lists the models available to call for chat completions rather than returning a single managed record.

### Parameters

- `model_id: str`

### Returns

- `class InferenceModel: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: Literal["user", "service_account"]`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: str`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: str`

  - `status: Literal["failed", "ready", "deploying", "deployment_timeout"]`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: Optional[InferenceModelAvailability]`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: Optional[Dict[str, object]]`

  - `object: Optional[Literal["model"]]`

    - `"model"`

  - `status_reason: Optional[str]`

  - `vendor_configuration: Optional[VendorConfiguration]`

    - `class LaunchVendorConfiguration: …`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `class LlmEngineVendorConfiguration: …`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
inference_model = client.models.retrieve(
    "model_id",
)
print(inference_model.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## Domain Types

### Inference Model

- `class InferenceModel: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: Literal["user", "service_account"]`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: str`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: str`

  - `status: Literal["failed", "ready", "deploying", "deployment_timeout"]`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: Optional[InferenceModelAvailability]`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: Optional[Dict[str, object]]`

  - `object: Optional[Literal["model"]]`

    - `"model"`

  - `status_reason: Optional[str]`

  - `vendor_configuration: Optional[VendorConfiguration]`

    - `class LaunchVendorConfiguration: …`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `class LlmEngineVendorConfiguration: …`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

### Inference Model Availability

- `Literal["unknown", "available", "unavailable"]`

  - `"unknown"`

  - `"available"`

  - `"unavailable"`

### Inference Model Type

- `Literal["generic", "completion", "chat_completion"]`

  - `"generic"`

  - `"completion"`

  - `"chat_completion"`

### Launch Vendor Configuration

- `class LaunchVendorConfiguration: …`

  - `model_image: ModelImage`

    - `command: List[str]`

    - `registry: str`

    - `repository: str`

    - `tag: str`

    - `env_vars: Optional[Dict[str, object]]`

    - `healthcheck_route: Optional[str]`

    - `predict_route: Optional[str]`

    - `readiness_delay: Optional[int]`

    - `request_schema: Optional[Dict[str, object]]`

    - `response_schema: Optional[Dict[str, object]]`

    - `streaming_command: Optional[List[str]]`

    - `streaming_predict_route: Optional[str]`

  - `model_infra: ModelInfra`

    - `cpus: Optional[Union[str, int, null]]`

      - `str`

      - `int`

    - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

      - `"async"`

      - `"sync"`

      - `"streaming"`

    - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

      - `"nvidia-tesla-t4"`

      - `"nvidia-ampere-a10"`

      - `"nvidia-ampere-a100"`

      - `"nvidia-ampere-a100e"`

      - `"nvidia-hopper-h100"`

      - `"nvidia-hopper-h100-1g20gb"`

      - `"nvidia-hopper-h100-3g40gb"`

    - `gpus: Optional[int]`

    - `high_priority: Optional[bool]`

    - `labels: Optional[Dict[str, str]]`

    - `max_workers: Optional[int]`

    - `memory: Optional[str]`

    - `min_workers: Optional[int]`

    - `per_worker: Optional[int]`

    - `public_inference: Optional[bool]`

    - `storage: Optional[str]`

### Llm Engine Vendor Configuration

- `class LlmEngineVendorConfiguration: …`

  - `model: str`

  - `chat_template_override: Optional[str]`

  - `checkpoint_path: Optional[str]`

  - `cpus: Optional[int]`

  - `default_callback_url: Optional[str]`

  - `endpoint_type: Optional[str]`

  - `gpu_type: Optional[str]`

  - `gpus: Optional[int]`

  - `high_priority: Optional[bool]`

  - `inference_framework: Optional[str]`

  - `inference_framework_image_tag: Optional[str]`

  - `labels: Optional[Dict[str, str]]`

  - `max_workers: Optional[int]`

  - `memory: Optional[str]`

  - `min_workers: Optional[int]`

  - `nodes_per_worker: Optional[int]`

  - `num_shards: Optional[int]`

  - `per_worker: Optional[int]`

  - `post_inference_hooks: Optional[List[str]]`

  - `public_inference: Optional[bool]`

  - `quantize: Optional[str]`

  - `source: Optional[str]`

  - `storage: Optional[str]`

### Model Delete Response

- `class ModelDeleteResponse: …`

  - `id: str`

  - `deleted: bool`

  - `object: Optional[Literal["model"]]`

    - `"model"`
