# Models

## Create a custom model

**post** `/v5/models`

Create a custom model record in your account and begin deploying it through a supported serving vendor.

A model here is a record for a model you deploy and serve through Scale's own inference vendors: only the `launch` and `llmengine` vendors are accepted and any other vendor is rejected. This is distinct from `GET /v5/chat/completions/models`, which lists the models already available to call for chat completions rather than creating or managing these records. The call is asynchronous — the record is created in a deploying status, a deployment job is recorded, and a Temporal workflow is started to perform the deployment, so the model is not ready for inference when this returns. A model name must be unique per vendor within your account; if a model with the same name and vendor already exists the request fails unless `on_conflict` is set to `update`, in which case the existing model is updated instead.

### Body Parameters

- `model: object { name, vendor_configuration, model_metadata, 3 more }  or object { name, vendor_configuration, model_metadata, 3 more }  or object { model_type, model_vendor, name, 2 more }`

  Register a model already served by an external / proxy-served vendor
  (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

  Unlike launch/llmengine, no Scale-side deployment is performed: the record is
  created READY and is immediately callable via /v5/chat/completions. Accepted only
  when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
  covers every vendor except launch/llmengine, and no vendor_configuration applies.

  - `Launch object { name, vendor_configuration, model_metadata, 3 more }`

    - `name: string`

      Unique name to reference your model

    - `vendor_configuration: LaunchVendorConfiguration`

      - `model_image: object { command, registry, repository, 9 more }`

        - `command: array of string`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars: optional map[unknown]`

        - `healthcheck_route: optional string`

        - `predict_route: optional string`

        - `readiness_delay: optional number`

        - `request_schema: optional map[unknown]`

        - `response_schema: optional map[unknown]`

        - `streaming_command: optional array of string`

        - `streaming_predict_route: optional string`

      - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

        - `cpus: optional string or number`

          - `string`

          - `number`

        - `endpoint_type: optional "async" or "sync" or "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `per_worker: optional number`

        - `public_inference: optional boolean`

        - `storage: optional string`

    - `model_metadata: optional map[unknown]`

    - `model_type: optional "generic"`

      - `"generic"`

    - `model_vendor: optional "launch"`

      - `"launch"`

    - `on_conflict: optional "error" or "update"`

      - `"error"`

      - `"update"`

  - `Llmengine object { name, vendor_configuration, model_metadata, 3 more }`

    - `name: string`

      Unique name to reference your model

    - `vendor_configuration: LlmEngineVendorConfiguration`

      - `model: string`

      - `chat_template_override: optional string`

      - `checkpoint_path: optional string`

      - `cpus: optional number`

      - `default_callback_url: optional string`

      - `endpoint_type: optional string`

      - `gpu_type: optional string`

      - `gpus: optional number`

      - `high_priority: optional boolean`

      - `inference_framework: optional string`

      - `inference_framework_image_tag: optional string`

      - `labels: optional map[string]`

      - `max_workers: optional number`

      - `memory: optional string`

      - `min_workers: optional number`

      - `nodes_per_worker: optional number`

      - `num_shards: optional number`

      - `per_worker: optional number`

      - `post_inference_hooks: optional array of string`

      - `public_inference: optional boolean`

      - `quantize: optional string`

      - `source: optional string`

      - `storage: optional string`

    - `model_metadata: optional map[unknown]`

    - `model_type: optional "chat_completion"`

      - `"chat_completion"`

    - `model_vendor: optional "llmengine"`

      - `"llmengine"`

    - `on_conflict: optional "error" or "update"`

      - `"error"`

      - `"update"`

  - `HostedModelCreateRequest object { model_type, model_vendor, name, 2 more }`

    Register a model already served by an external / proxy-served vendor
    (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

    Unlike launch/llmengine, no Scale-side deployment is performed: the record is
    created READY and is immediately callable via /v5/chat/completions. Accepted only
    when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
    covers every vendor except launch/llmengine, and no vendor_configuration applies.

    - `model_type: InferenceModelType`

      Type of model, for example `chat_completion`

      - `"generic"`

      - `"completion"`

      - `"chat_completion"`

    - `model_vendor: "openai" or "cohere" or "vertex_ai" or 7 more`

      Vendor to serve/create model

      - `"openai"`

      - `"cohere"`

      - `"vertex_ai"`

      - `"anthropic"`

      - `"azure"`

      - `"gemini"`

      - `"model_zoo"`

      - `"bedrock"`

      - `"xai"`

      - `"fireworks_ai"`

    - `name: string`

      Unique name to reference your model

    - `model_metadata: optional map[unknown]`

    - `on_conflict: optional "error" or "update"`

      - `"error"`

      - `"update"`

### Returns

- `InferenceModel object { id, created_at, created_by_identity_type, 10 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" or "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" or "ready" or "deploying" or "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: optional InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: optional map[unknown]`

  - `object: optional "model"`

    - `"model"`

  - `status_reason: optional string`

  - `vendor_configuration: optional LaunchVendorConfiguration or LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration object { model_image, model_infra }`

      - `model_image: object { command, registry, repository, 9 more }`

        - `command: array of string`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars: optional map[unknown]`

        - `healthcheck_route: optional string`

        - `predict_route: optional string`

        - `readiness_delay: optional number`

        - `request_schema: optional map[unknown]`

        - `response_schema: optional map[unknown]`

        - `streaming_command: optional array of string`

        - `streaming_predict_route: optional string`

      - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

        - `cpus: optional string or number`

          - `string`

          - `number`

        - `endpoint_type: optional "async" or "sync" or "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `per_worker: optional number`

        - `public_inference: optional boolean`

        - `storage: optional string`

    - `LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }`

      - `model: string`

      - `chat_template_override: optional string`

      - `checkpoint_path: optional string`

      - `cpus: optional number`

      - `default_callback_url: optional string`

      - `endpoint_type: optional string`

      - `gpu_type: optional string`

      - `gpus: optional number`

      - `high_priority: optional boolean`

      - `inference_framework: optional string`

      - `inference_framework_image_tag: optional string`

      - `labels: optional map[string]`

      - `max_workers: optional number`

      - `memory: optional string`

      - `min_workers: optional number`

      - `nodes_per_worker: optional number`

      - `num_shards: optional number`

      - `per_worker: optional number`

      - `post_inference_hooks: optional array of string`

      - `public_inference: optional boolean`

      - `quantize: optional string`

      - `source: optional string`

      - `storage: optional string`

### Example

```http
curl https://api.egp.scale.com/v5/models \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "name": "name",
          "vendor_configuration": {
            "model_image": {
              "command": [
                "string"
              ],
              "registry": "registry",
              "repository": "repository",
              "tag": "tag",
              "env_vars": {
                "foo": "bar"
              },
              "healthcheck_route": "healthcheck_route",
              "predict_route": "predict_route",
              "readiness_delay": 0,
              "request_schema": {
                "foo": "bar"
              },
              "response_schema": {
                "foo": "bar"
              },
              "streaming_command": [
                "string"
              ],
              "streaming_predict_route": "streaming_predict_route"
            },
            "model_infra": {
              "cpus": "string",
              "endpoint_type": "async",
              "gpu_type": "nvidia-tesla-t4",
              "gpus": 0,
              "high_priority": true,
              "labels": {
                "foo": "string"
              },
              "max_workers": 0,
              "memory": "memory",
              "min_workers": 0,
              "per_worker": 0,
              "public_inference": true,
              "storage": "storage"
            }
          },
          "model_metadata": {
            "foo": "bar"
          },
          "model_type": "generic",
          "model_vendor": "launch",
          "on_conflict": "error"
        }'
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## List custom models

**get** `/v5/models`

List the custom model records registered in your account.

Returns a paginated list of the model records managed through this API — models your account deploys through the `launch` or `llmengine` serving vendors — optionally filtered by name and by model vendor, and scoped to the caller's account. This is different from `GET /v5/chat/completions/models`, which lists the models available to invoke for chat completions; this endpoint returns the managed records along with their deployment status, not the catalog of callable completion models.

### Query Parameters

- `ending_before: optional string`

- `limit: optional number`

- `model_vendor: optional InferenceModelVendor`

  - `"openai"`

  - `"cohere"`

  - `"vertex_ai"`

  - `"anthropic"`

  - `"azure"`

  - `"gemini"`

  - `"launch"`

  - `"llmengine"`

  - `"model_zoo"`

  - `"bedrock"`

  - `"xai"`

  - `"fireworks_ai"`

- `name: optional string`

- `sort_by: optional string`

- `sort_order: optional SortOrder`

  - `"asc"`

  - `"desc"`

- `starting_after: optional string`

### Returns

- `has_more: boolean`

  Whether there are more items left to be fetched.

- `items: array of InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" or "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" or "ready" or "deploying" or "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: optional InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: optional map[unknown]`

  - `object: optional "model"`

    - `"model"`

  - `status_reason: optional string`

  - `vendor_configuration: optional LaunchVendorConfiguration or LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration object { model_image, model_infra }`

      - `model_image: object { command, registry, repository, 9 more }`

        - `command: array of string`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars: optional map[unknown]`

        - `healthcheck_route: optional string`

        - `predict_route: optional string`

        - `readiness_delay: optional number`

        - `request_schema: optional map[unknown]`

        - `response_schema: optional map[unknown]`

        - `streaming_command: optional array of string`

        - `streaming_predict_route: optional string`

      - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

        - `cpus: optional string or number`

          - `string`

          - `number`

        - `endpoint_type: optional "async" or "sync" or "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `per_worker: optional number`

        - `public_inference: optional boolean`

        - `storage: optional string`

    - `LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }`

      - `model: string`

      - `chat_template_override: optional string`

      - `checkpoint_path: optional string`

      - `cpus: optional number`

      - `default_callback_url: optional string`

      - `endpoint_type: optional string`

      - `gpu_type: optional string`

      - `gpus: optional number`

      - `high_priority: optional boolean`

      - `inference_framework: optional string`

      - `inference_framework_image_tag: optional string`

      - `labels: optional map[string]`

      - `max_workers: optional number`

      - `memory: optional string`

      - `min_workers: optional number`

      - `nodes_per_worker: optional number`

      - `num_shards: optional number`

      - `per_worker: optional number`

      - `post_inference_hooks: optional array of string`

      - `public_inference: optional boolean`

      - `quantize: optional string`

      - `source: optional string`

      - `storage: optional string`

- `total: number`

  The total of items that match the query. This is greater than or equal to the number of items returned.

- `limit: optional number`

  The maximum number of items to return.

- `object: optional "list"`

  - `"list"`

### Example

```http
curl https://api.egp.scale.com/v5/models \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by_identity_type": "user",
      "created_by_user_id": "created_by_user_id",
      "model_type": "generic",
      "model_vendor": "openai",
      "name": "name",
      "status": "failed",
      "model_availability": "unknown",
      "model_metadata": {
        "foo": "bar"
      },
      "object": "model",
      "status_reason": "status_reason",
      "vendor_configuration": {
        "model_image": {
          "command": [
            "string"
          ],
          "registry": "registry",
          "repository": "repository",
          "tag": "tag",
          "env_vars": {
            "foo": "bar"
          },
          "healthcheck_route": "healthcheck_route",
          "predict_route": "predict_route",
          "readiness_delay": 0,
          "request_schema": {
            "foo": "bar"
          },
          "response_schema": {
            "foo": "bar"
          },
          "streaming_command": [
            "string"
          ],
          "streaming_predict_route": "streaming_predict_route"
        },
        "model_infra": {
          "cpus": "string",
          "endpoint_type": "async",
          "gpu_type": "nvidia-tesla-t4",
          "gpus": 0,
          "high_priority": true,
          "labels": {
            "foo": "string"
          },
          "max_workers": 0,
          "memory": "memory",
          "min_workers": 0,
          "per_worker": 0,
          "public_inference": true,
          "storage": "storage"
        }
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Update a custom model

**patch** `/v5/models/{model_id}`

Update a custom model record; vendor-configuration changes are applied asynchronously by redeploying the model.

This supports three kinds of update: changing model metadata only, renaming the model, and changing the vendor configuration. A vendor-configuration change is asynchronous — it puts the model back into a deploying status, records an update job, and starts a Temporal workflow to redeploy, so the new configuration is not live when this returns; metadata-only and rename changes take effect immediately. The vendor configuration supplied must match the model's own vendor (`launch` or `llmengine`), and only those two vendors are supported. A model that is currently deploying cannot be modified and the request fails until deployment finishes. When renaming with `on_conflict` set to `swap`, the name is exchanged with an existing model of the same name and vendor instead of failing on the uniqueness constraint.

### Path Parameters

- `model_id: string`

### Body Parameters

- `model: object { model_metadata }  or object { vendor_configuration, model_metadata }  or object { name, on_conflict }`

  - `DefaultModelPatchRequest object { model_metadata }`

    - `model_metadata: optional map[unknown]`

  - `ModelConfigurationPatchRequest object { vendor_configuration, model_metadata }`

    - `vendor_configuration: object { model_image, model_infra }  or object { chat_template_override, checkpoint_path, cpus, 20 more }`

      - `PartialLaunchVendorConfiguration object { model_image, model_infra }`

        - `model_image: optional object { command, env_vars, healthcheck_route, 9 more }`

          - `command: optional array of string`

          - `env_vars: optional map[unknown]`

          - `healthcheck_route: optional string`

          - `predict_route: optional string`

          - `readiness_delay: optional number`

          - `registry: optional string`

          - `repository: optional string`

          - `request_schema: optional map[unknown]`

          - `response_schema: optional map[unknown]`

          - `streaming_command: optional array of string`

          - `streaming_predict_route: optional string`

          - `tag: optional string`

        - `model_infra: optional object { cpus, endpoint_type, gpu_type, 9 more }`

          - `cpus: optional string or number`

            - `string`

            - `number`

          - `endpoint_type: optional "async" or "sync" or "streaming"`

            - `"async"`

            - `"sync"`

            - `"streaming"`

          - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

            - `"nvidia-tesla-t4"`

            - `"nvidia-ampere-a10"`

            - `"nvidia-ampere-a100"`

            - `"nvidia-ampere-a100e"`

            - `"nvidia-hopper-h100"`

            - `"nvidia-hopper-h100-1g20gb"`

            - `"nvidia-hopper-h100-3g40gb"`

          - `gpus: optional number`

          - `high_priority: optional boolean`

          - `labels: optional map[string]`

          - `max_workers: optional number`

          - `memory: optional string`

          - `min_workers: optional number`

          - `per_worker: optional number`

          - `public_inference: optional boolean`

          - `storage: optional string`

      - `PartialLlmEngineVendorConfiguration object { chat_template_override, checkpoint_path, cpus, 20 more }`

        - `chat_template_override: optional string`

        - `checkpoint_path: optional string`

        - `cpus: optional number`

        - `default_callback_url: optional string`

        - `endpoint_type: optional string`

        - `gpu_type: optional string`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `inference_framework: optional string`

        - `inference_framework_image_tag: optional string`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `model: optional string`

        - `nodes_per_worker: optional number`

        - `num_shards: optional number`

        - `per_worker: optional number`

        - `post_inference_hooks: optional array of string`

        - `public_inference: optional boolean`

        - `quantize: optional string`

        - `source: optional string`

        - `storage: optional string`

    - `model_metadata: optional map[unknown]`

  - `SwapNamesModelPatchRequest object { name, on_conflict }`

    - `name: string`

    - `on_conflict: optional "error" or "swap"`

      - `"error"`

      - `"swap"`

### Returns

- `InferenceModel object { id, created_at, created_by_identity_type, 10 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" or "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" or "ready" or "deploying" or "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: optional InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: optional map[unknown]`

  - `object: optional "model"`

    - `"model"`

  - `status_reason: optional string`

  - `vendor_configuration: optional LaunchVendorConfiguration or LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration object { model_image, model_infra }`

      - `model_image: object { command, registry, repository, 9 more }`

        - `command: array of string`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars: optional map[unknown]`

        - `healthcheck_route: optional string`

        - `predict_route: optional string`

        - `readiness_delay: optional number`

        - `request_schema: optional map[unknown]`

        - `response_schema: optional map[unknown]`

        - `streaming_command: optional array of string`

        - `streaming_predict_route: optional string`

      - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

        - `cpus: optional string or number`

          - `string`

          - `number`

        - `endpoint_type: optional "async" or "sync" or "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `per_worker: optional number`

        - `public_inference: optional boolean`

        - `storage: optional string`

    - `LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }`

      - `model: string`

      - `chat_template_override: optional string`

      - `checkpoint_path: optional string`

      - `cpus: optional number`

      - `default_callback_url: optional string`

      - `endpoint_type: optional string`

      - `gpu_type: optional string`

      - `gpus: optional number`

      - `high_priority: optional boolean`

      - `inference_framework: optional string`

      - `inference_framework_image_tag: optional string`

      - `labels: optional map[string]`

      - `max_workers: optional number`

      - `memory: optional string`

      - `min_workers: optional number`

      - `nodes_per_worker: optional number`

      - `num_shards: optional number`

      - `per_worker: optional number`

      - `post_inference_hooks: optional array of string`

      - `public_inference: optional boolean`

      - `quantize: optional string`

      - `source: optional string`

      - `storage: optional string`

### Example

```http
curl https://api.egp.scale.com/v5/models/$MODEL_ID \
    -X PATCH \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "model_metadata": {
            "foo": "bar"
          }
        }'
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## Delete a custom model

**delete** `/v5/models/{model_id}`

Permanently delete a custom model record and tear down its deployment at the serving vendor.

This is a hard delete: the model row is removed from your account entirely and cannot be restored afterward. Before the record is removed, if the model has an associated vendor deployment that deployment is torn down at its serving vendor. A model that is currently deploying cannot be deleted and the request fails until deployment finishes. This operates on the model records managed by this API, distinct from the `GET /v5/chat/completions/models` catalog of models callable for chat completions.

### Path Parameters

- `model_id: string`

### Returns

- `id: string`

- `deleted: boolean`

- `object: optional "model"`

  - `"model"`

### Example

```http
curl https://api.egp.scale.com/v5/models/$MODEL_ID \
    -X DELETE \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "id": "id",
  "deleted": true,
  "object": "model"
}
```

## Get a custom model

**get** `/v5/models/{model_id}`

Retrieve a single custom model record by its ID.

Returns the model record — including its vendor, configuration, and current deployment status — for a model managed through this API and owned by the caller's account. This is distinct from `GET /v5/chat/completions/models`, which lists the models available to call for chat completions rather than returning a single managed record.

### Path Parameters

- `model_id: string`

### Returns

- `InferenceModel object { id, created_at, created_by_identity_type, 10 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" or "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" or "ready" or "deploying" or "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: optional InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: optional map[unknown]`

  - `object: optional "model"`

    - `"model"`

  - `status_reason: optional string`

  - `vendor_configuration: optional LaunchVendorConfiguration or LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration object { model_image, model_infra }`

      - `model_image: object { command, registry, repository, 9 more }`

        - `command: array of string`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars: optional map[unknown]`

        - `healthcheck_route: optional string`

        - `predict_route: optional string`

        - `readiness_delay: optional number`

        - `request_schema: optional map[unknown]`

        - `response_schema: optional map[unknown]`

        - `streaming_command: optional array of string`

        - `streaming_predict_route: optional string`

      - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

        - `cpus: optional string or number`

          - `string`

          - `number`

        - `endpoint_type: optional "async" or "sync" or "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `per_worker: optional number`

        - `public_inference: optional boolean`

        - `storage: optional string`

    - `LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }`

      - `model: string`

      - `chat_template_override: optional string`

      - `checkpoint_path: optional string`

      - `cpus: optional number`

      - `default_callback_url: optional string`

      - `endpoint_type: optional string`

      - `gpu_type: optional string`

      - `gpus: optional number`

      - `high_priority: optional boolean`

      - `inference_framework: optional string`

      - `inference_framework_image_tag: optional string`

      - `labels: optional map[string]`

      - `max_workers: optional number`

      - `memory: optional string`

      - `min_workers: optional number`

      - `nodes_per_worker: optional number`

      - `num_shards: optional number`

      - `per_worker: optional number`

      - `post_inference_hooks: optional array of string`

      - `public_inference: optional boolean`

      - `quantize: optional string`

      - `source: optional string`

      - `storage: optional string`

### Example

```http
curl https://api.egp.scale.com/v5/models/$MODEL_ID \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```

## Domain Types

### Inference Model

- `InferenceModel object { id, created_at, created_by_identity_type, 10 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" or "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" or "ready" or "deploying" or "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: optional InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: optional map[unknown]`

  - `object: optional "model"`

    - `"model"`

  - `status_reason: optional string`

  - `vendor_configuration: optional LaunchVendorConfiguration or LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration object { model_image, model_infra }`

      - `model_image: object { command, registry, repository, 9 more }`

        - `command: array of string`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars: optional map[unknown]`

        - `healthcheck_route: optional string`

        - `predict_route: optional string`

        - `readiness_delay: optional number`

        - `request_schema: optional map[unknown]`

        - `response_schema: optional map[unknown]`

        - `streaming_command: optional array of string`

        - `streaming_predict_route: optional string`

      - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

        - `cpus: optional string or number`

          - `string`

          - `number`

        - `endpoint_type: optional "async" or "sync" or "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `per_worker: optional number`

        - `public_inference: optional boolean`

        - `storage: optional string`

    - `LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }`

      - `model: string`

      - `chat_template_override: optional string`

      - `checkpoint_path: optional string`

      - `cpus: optional number`

      - `default_callback_url: optional string`

      - `endpoint_type: optional string`

      - `gpu_type: optional string`

      - `gpus: optional number`

      - `high_priority: optional boolean`

      - `inference_framework: optional string`

      - `inference_framework_image_tag: optional string`

      - `labels: optional map[string]`

      - `max_workers: optional number`

      - `memory: optional string`

      - `min_workers: optional number`

      - `nodes_per_worker: optional number`

      - `num_shards: optional number`

      - `per_worker: optional number`

      - `post_inference_hooks: optional array of string`

      - `public_inference: optional boolean`

      - `quantize: optional string`

      - `source: optional string`

      - `storage: optional string`

### Inference Model Availability

- `InferenceModelAvailability = "unknown" or "available" or "unavailable"`

  - `"unknown"`

  - `"available"`

  - `"unavailable"`

### Inference Model Type

- `InferenceModelType = "generic" or "completion" or "chat_completion"`

  - `"generic"`

  - `"completion"`

  - `"chat_completion"`

### Launch Vendor Configuration

- `LaunchVendorConfiguration object { model_image, model_infra }`

  - `model_image: object { command, registry, repository, 9 more }`

    - `command: array of string`

    - `registry: string`

    - `repository: string`

    - `tag: string`

    - `env_vars: optional map[unknown]`

    - `healthcheck_route: optional string`

    - `predict_route: optional string`

    - `readiness_delay: optional number`

    - `request_schema: optional map[unknown]`

    - `response_schema: optional map[unknown]`

    - `streaming_command: optional array of string`

    - `streaming_predict_route: optional string`

  - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

    - `cpus: optional string or number`

      - `string`

      - `number`

    - `endpoint_type: optional "async" or "sync" or "streaming"`

      - `"async"`

      - `"sync"`

      - `"streaming"`

    - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

      - `"nvidia-tesla-t4"`

      - `"nvidia-ampere-a10"`

      - `"nvidia-ampere-a100"`

      - `"nvidia-ampere-a100e"`

      - `"nvidia-hopper-h100"`

      - `"nvidia-hopper-h100-1g20gb"`

      - `"nvidia-hopper-h100-3g40gb"`

    - `gpus: optional number`

    - `high_priority: optional boolean`

    - `labels: optional map[string]`

    - `max_workers: optional number`

    - `memory: optional string`

    - `min_workers: optional number`

    - `per_worker: optional number`

    - `public_inference: optional boolean`

    - `storage: optional string`

### Llm Engine Vendor Configuration

- `LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }`

  - `model: string`

  - `chat_template_override: optional string`

  - `checkpoint_path: optional string`

  - `cpus: optional number`

  - `default_callback_url: optional string`

  - `endpoint_type: optional string`

  - `gpu_type: optional string`

  - `gpus: optional number`

  - `high_priority: optional boolean`

  - `inference_framework: optional string`

  - `inference_framework_image_tag: optional string`

  - `labels: optional map[string]`

  - `max_workers: optional number`

  - `memory: optional string`

  - `min_workers: optional number`

  - `nodes_per_worker: optional number`

  - `num_shards: optional number`

  - `per_worker: optional number`

  - `post_inference_hooks: optional array of string`

  - `public_inference: optional boolean`

  - `quantize: optional string`

  - `source: optional string`

  - `storage: optional string`

### Model Delete Response

- `ModelDeleteResponse object { id, deleted, object }`

  - `id: string`

  - `deleted: boolean`

  - `object: optional "model"`

    - `"model"`
