Skip to content

Update a custom model

client.Models.Update(ctx, modelID, body) (*InferenceModel, error)
PATCH/v5/models/{model_id}

Update a custom model record; vendor-configuration changes are applied asynchronously by redeploying the model.

This supports three kinds of update: changing model metadata only, renaming the model, and changing the vendor configuration. A vendor-configuration change is asynchronous — it puts the model back into a deploying status, records an update job, and starts a Temporal workflow to redeploy, so the new configuration is not live when this returns; metadata-only and rename changes take effect immediately. The vendor configuration supplied must match the model’s own vendor (launch or llmengine), and only those two vendors are supported. A model that is currently deploying cannot be modified and the request fails until deployment finishes. When renaming with on_conflict set to swap, the name is exchanged with an existing model of the same name and vendor instead of failing on the uniqueness constraint.

ParametersExpand Collapse
modelID string
body ModelUpdateParams
Model param.Field[ModelUpdateParamsModelUnion]
type ModelUpdateParamsModelDefaultModelPatchRequest struct{…}
ModelMetadata map[string, any]Optional
type ModelUpdateParamsModelModelConfigurationPatchRequest struct{…}
VendorConfiguration ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationUnion
One of the following:
type ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfiguration struct{…}
ModelImage ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelImageOptional
Command []stringOptional
EnvVars map[string, any]Optional
HealthcheckRoute stringOptional
PredictRoute stringOptional
ReadinessDelay int64Optional
Registry stringOptional
Repository stringOptional
RequestSchema map[string, any]Optional
ResponseSchema map[string, any]Optional
StreamingCommand []stringOptional
StreamingPredictRoute stringOptional
Tag stringOptional
ModelInfra ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraOptional
CPUs ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraCPUsUnionOptional
One of the following:
string
int64
EndpointType stringOptional
One of the following:
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraEndpointTypeAsync ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraEndpointType = "async"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraEndpointTypeSync ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraEndpointType = "sync"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraEndpointTypeStreaming ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraEndpointType = "streaming"
GPUType stringOptional
One of the following:
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUTypeNvidiaTeslaT4 ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUType = "nvidia-tesla-t4"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA10 ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a10"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100 ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100e ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100e"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100 ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_1g20gb ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-1g20gb"
const ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_3g40gb ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-3g40gb"
GPUs int64Optional
HighPriority boolOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
PerWorker int64Optional
PublicInference boolOptional
Storage stringOptional
type ModelUpdateParamsModelModelConfigurationPatchRequestVendorConfigurationPartialLlmEngineVendorConfiguration struct{…}
ChatTemplateOverride stringOptional
CheckpointPath stringOptional
CPUs int64Optional
DefaultCallbackURL stringOptional
EndpointType stringOptional
GPUType stringOptional
GPUs int64Optional
HighPriority boolOptional
InferenceFramework stringOptional
InferenceFrameworkImageTag stringOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
Model stringOptional
NodesPerWorker int64Optional
NumShards int64Optional
PerWorker int64Optional
PostInferenceHooks []stringOptional
PublicInference boolOptional
Quantize stringOptional
Source stringOptional
Storage stringOptional
ModelMetadata map[string, any]Optional
type ModelUpdateParamsModelSwapNamesModelPatchRequest struct{…}
Name string
OnConflict stringOptional
One of the following:
const ModelUpdateParamsModelSwapNamesModelPatchRequestOnConflictError ModelUpdateParamsModelSwapNamesModelPatchRequestOnConflict = "error"
const ModelUpdateParamsModelSwapNamesModelPatchRequestOnConflictSwap ModelUpdateParamsModelSwapNamesModelPatchRequestOnConflict = "swap"
ReturnsExpand Collapse
type InferenceModel struct{…}
ID string

The unique identifier of the entity.

CreatedAt Time

The date and time when the entity was created in ISO format.

formatdate-time
CreatedByIdentityType InferenceModelCreatedByIdentityType

The type of identity that created the entity.

One of the following:
const InferenceModelCreatedByIdentityTypeUser InferenceModelCreatedByIdentityType = "user"
const InferenceModelCreatedByIdentityTypeServiceAccount InferenceModelCreatedByIdentityType = "service_account"
CreatedByUserID string

The user who originally created the entity.

One of the following:
const InferenceModelTypeGeneric InferenceModelType = "generic"
const InferenceModelTypeCompletion InferenceModelType = "completion"
const InferenceModelTypeChatCompletion InferenceModelType = "chat_completion"
One of the following:
const InferenceModelVendorOpenAI InferenceModelVendor = "openai"
const InferenceModelVendorCohere InferenceModelVendor = "cohere"
const InferenceModelVendorVertexAI InferenceModelVendor = "vertex_ai"
const InferenceModelVendorAnthropic InferenceModelVendor = "anthropic"
const InferenceModelVendorAzure InferenceModelVendor = "azure"
const InferenceModelVendorGemini InferenceModelVendor = "gemini"
const InferenceModelVendorLaunch InferenceModelVendor = "launch"
const InferenceModelVendorLlmengine InferenceModelVendor = "llmengine"
const InferenceModelVendorModelZoo InferenceModelVendor = "model_zoo"
const InferenceModelVendorBedrock InferenceModelVendor = "bedrock"
const InferenceModelVendorXai InferenceModelVendor = "xai"
const InferenceModelVendorFireworksAI InferenceModelVendor = "fireworks_ai"
Name string
Status InferenceModelStatus
One of the following:
const InferenceModelStatusFailed InferenceModelStatus = "failed"
const InferenceModelStatusReady InferenceModelStatus = "ready"
const InferenceModelStatusDeploying InferenceModelStatus = "deploying"
const InferenceModelStatusDeploymentTimeout InferenceModelStatus = "deployment_timeout"
ModelAvailability InferenceModelAvailabilityOptional
One of the following:
const InferenceModelAvailabilityUnknown InferenceModelAvailability = "unknown"
const InferenceModelAvailabilityAvailable InferenceModelAvailability = "available"
const InferenceModelAvailabilityUnavailable InferenceModelAvailability = "unavailable"
ModelMetadata map[string, any]Optional
Object InferenceModelObjectOptional
StatusReason stringOptional
VendorConfiguration InferenceModelVendorConfigurationUnionOptional
One of the following:
type LaunchVendorConfiguration struct{…}
ModelImage LaunchVendorConfigurationModelImage
Command []string
Registry string
Repository string
Tag string
EnvVars map[string, any]Optional
HealthcheckRoute stringOptional
PredictRoute stringOptional
ReadinessDelay int64Optional
RequestSchema map[string, any]Optional
ResponseSchema map[string, any]Optional
StreamingCommand []stringOptional
StreamingPredictRoute stringOptional
ModelInfra LaunchVendorConfigurationModelInfra
CPUs LaunchVendorConfigurationModelInfraCPUsUnionOptional
One of the following:
string
int64
EndpointType stringOptional
One of the following:
const LaunchVendorConfigurationModelInfraEndpointTypeAsync LaunchVendorConfigurationModelInfraEndpointType = "async"
const LaunchVendorConfigurationModelInfraEndpointTypeSync LaunchVendorConfigurationModelInfraEndpointType = "sync"
const LaunchVendorConfigurationModelInfraEndpointTypeStreaming LaunchVendorConfigurationModelInfraEndpointType = "streaming"
GPUType stringOptional
One of the following:
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaTeslaT4 LaunchVendorConfigurationModelInfraGPUType = "nvidia-tesla-t4"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA10 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a10"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100e LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100e"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_1g20gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-1g20gb"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_3g40gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-3g40gb"
GPUs int64Optional
HighPriority boolOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
PerWorker int64Optional
PublicInference boolOptional
Storage stringOptional
type LlmEngineVendorConfiguration struct{…}
Model string
ChatTemplateOverride stringOptional
CheckpointPath stringOptional
CPUs int64Optional
DefaultCallbackURL stringOptional
EndpointType stringOptional
GPUType stringOptional
GPUs int64Optional
HighPriority boolOptional
InferenceFramework stringOptional
InferenceFrameworkImageTag stringOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
NodesPerWorker int64Optional
NumShards int64Optional
PerWorker int64Optional
PostInferenceHooks []stringOptional
PublicInference boolOptional
Quantize stringOptional
Source stringOptional
Storage stringOptional

Update a custom model

package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  inferenceModel, err := client.Models.Update(
    context.TODO(),
    "model_id",
    sgpdev.ModelUpdateParams{
      OfDefaultModelPatchRequest: &sgpdev.ModelUpdateParamsModelDefaultModelPatchRequest{

      },
    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", inferenceModel.ID)
}
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
Returns Examples
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}