Skip to content

Update a custom model

models.update(strmodel_id, ModelUpdateParams**kwargs) -> InferenceModel
PATCH/v5/models/{model_id}

Update a custom model record; vendor-configuration changes are applied asynchronously by redeploying the model.

This supports three kinds of update: changing model metadata only, renaming the model, and changing the vendor configuration. A vendor-configuration change is asynchronous — it puts the model back into a deploying status, records an update job, and starts a Temporal workflow to redeploy, so the new configuration is not live when this returns; metadata-only and rename changes take effect immediately. The vendor configuration supplied must match the model’s own vendor (launch or llmengine), and only those two vendors are supported. A model that is currently deploying cannot be modified and the request fails until deployment finishes. When renaming with on_conflict set to swap, the name is exchanged with an existing model of the same name and vendor instead of failing on the uniqueness constraint.

ParametersExpand Collapse
model_id: str
model: Model
One of the following:
class ModelDefaultModelPatchRequest: …
model_metadata: Optional[Dict[str, object]]
class ModelModelConfigurationPatchRequest: …
vendor_configuration: ModelModelConfigurationPatchRequestVendorConfiguration
One of the following:
class ModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfiguration: …
model_image: Optional[ModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelImage]
command: Optional[Sequence[str]]
env_vars: Optional[Dict[str, object]]
healthcheck_route: Optional[str]
predict_route: Optional[str]
readiness_delay: Optional[int]
registry: Optional[str]
repository: Optional[str]
request_schema: Optional[Dict[str, object]]
response_schema: Optional[Dict[str, object]]
streaming_command: Optional[Sequence[str]]
streaming_predict_route: Optional[str]
tag: Optional[str]
model_infra: Optional[ModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfra]
cpus: Optional[Union[str, int]]
One of the following:
str
int
endpoint_type: Optional[Literal["async", "sync", "streaming"]]
One of the following:
"async"
"sync"
"streaming"
gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus: Optional[int]
high_priority: Optional[bool]
labels: Optional[Dict[str, str]]
max_workers: Optional[int]
memory: Optional[str]
min_workers: Optional[int]
per_worker: Optional[int]
public_inference: Optional[bool]
storage: Optional[str]
class ModelModelConfigurationPatchRequestVendorConfigurationPartialLlmEngineVendorConfiguration: …
chat_template_override: Optional[str]
checkpoint_path: Optional[str]
cpus: Optional[int]
default_callback_url: Optional[str]
endpoint_type: Optional[str]
gpu_type: Optional[str]
gpus: Optional[int]
high_priority: Optional[bool]
inference_framework: Optional[str]
inference_framework_image_tag: Optional[str]
labels: Optional[Dict[str, str]]
max_workers: Optional[int]
memory: Optional[str]
min_workers: Optional[int]
model: Optional[str]
nodes_per_worker: Optional[int]
num_shards: Optional[int]
per_worker: Optional[int]
post_inference_hooks: Optional[Sequence[str]]
public_inference: Optional[bool]
quantize: Optional[str]
source: Optional[str]
storage: Optional[str]
model_metadata: Optional[Dict[str, object]]
class ModelSwapNamesModelPatchRequest: …
name: str
on_conflict: Optional[Literal["error", "swap"]]
One of the following:
"error"
"swap"
ReturnsExpand Collapse
class InferenceModel: …
id: str

The unique identifier of the entity.

created_at: datetime

The date and time when the entity was created in ISO format.

formatdate-time
created_by_identity_type: Literal["user", "service_account"]

The type of identity that created the entity.

One of the following:
"user"
"service_account"
created_by_user_id: str

The user who originally created the entity.

model_type: InferenceModelType
One of the following:
"generic"
"completion"
"chat_completion"
model_vendor: InferenceModelVendor
One of the following:
"openai"
"cohere"
"vertex_ai"
"anthropic"
"azure"
"gemini"
"launch"
"llmengine"
"model_zoo"
"bedrock"
"xai"
"fireworks_ai"
name: str
status: Literal["failed", "ready", "deploying", "deployment_timeout"]
One of the following:
"failed"
"ready"
"deploying"
"deployment_timeout"
model_availability: Optional[InferenceModelAvailability]
One of the following:
"unknown"
"available"
"unavailable"
model_metadata: Optional[Dict[str, object]]
object: Optional[Literal["model"]]
status_reason: Optional[str]
vendor_configuration: Optional[VendorConfiguration]
One of the following:
class LaunchVendorConfiguration: …
model_image: ModelImage
command: List[str]
registry: str
repository: str
tag: str
env_vars: Optional[Dict[str, object]]
healthcheck_route: Optional[str]
predict_route: Optional[str]
readiness_delay: Optional[int]
request_schema: Optional[Dict[str, object]]
response_schema: Optional[Dict[str, object]]
streaming_command: Optional[List[str]]
streaming_predict_route: Optional[str]
model_infra: ModelInfra
cpus: Optional[Union[str, int, null]]
One of the following:
str
int
endpoint_type: Optional[Literal["async", "sync", "streaming"]]
One of the following:
"async"
"sync"
"streaming"
gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus: Optional[int]
high_priority: Optional[bool]
labels: Optional[Dict[str, str]]
max_workers: Optional[int]
memory: Optional[str]
min_workers: Optional[int]
per_worker: Optional[int]
public_inference: Optional[bool]
storage: Optional[str]
class LlmEngineVendorConfiguration: …
model: str
chat_template_override: Optional[str]
checkpoint_path: Optional[str]
cpus: Optional[int]
default_callback_url: Optional[str]
endpoint_type: Optional[str]
gpu_type: Optional[str]
gpus: Optional[int]
high_priority: Optional[bool]
inference_framework: Optional[str]
inference_framework_image_tag: Optional[str]
labels: Optional[Dict[str, str]]
max_workers: Optional[int]
memory: Optional[str]
min_workers: Optional[int]
nodes_per_worker: Optional[int]
num_shards: Optional[int]
per_worker: Optional[int]
post_inference_hooks: Optional[List[str]]
public_inference: Optional[bool]
quantize: Optional[str]
source: Optional[str]
storage: Optional[str]

Update a custom model

import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
inference_model = client.models.update(
    model_id="model_id",
    model={},
)
print(inference_model.id)
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
Returns Examples
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}