Skip to content

Create a custom model

client.Models.New(ctx, body) (*InferenceModel, error)
POST/v5/models

Create a custom model record in your account and begin deploying it through a supported serving vendor.

A model here is a record for a model you deploy and serve through Scale’s own inference vendors: only the launch and llmengine vendors are accepted and any other vendor is rejected. This is distinct from GET /v5/chat/completions/models, which lists the models already available to call for chat completions rather than creating or managing these records. The call is asynchronous — the record is created in a deploying status, a deployment job is recorded, and a Temporal workflow is started to perform the deployment, so the model is not ready for inference when this returns. A model name must be unique per vendor within your account; if a model with the same name and vendor already exists the request fails unless on_conflict is set to update, in which case the existing model is updated instead.

ParametersExpand Collapse
body ModelNewParams
Model param.Field[ModelNewParamsModelUnion]

Register a model already served by an external / proxy-served vendor (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

Unlike launch/llmengine, no Scale-side deployment is performed: the record is created READY and is immediately callable via /v5/chat/completions. Accepted only when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor) covers every vendor except launch/llmengine, and no vendor_configuration applies.

type ModelNewParamsModelLaunch struct{…}
Name string

Unique name to reference your model

VendorConfiguration LaunchVendorConfiguration
ModelImage LaunchVendorConfigurationModelImage
Command []string
Registry string
Repository string
Tag string
EnvVars map[string, any]Optional
HealthcheckRoute stringOptional
PredictRoute stringOptional
ReadinessDelay int64Optional
RequestSchema map[string, any]Optional
ResponseSchema map[string, any]Optional
StreamingCommand []stringOptional
StreamingPredictRoute stringOptional
ModelInfra LaunchVendorConfigurationModelInfra
CPUs LaunchVendorConfigurationModelInfraCPUsUnionOptional
One of the following:
string
int64
EndpointType stringOptional
One of the following:
const LaunchVendorConfigurationModelInfraEndpointTypeAsync LaunchVendorConfigurationModelInfraEndpointType = "async"
const LaunchVendorConfigurationModelInfraEndpointTypeSync LaunchVendorConfigurationModelInfraEndpointType = "sync"
const LaunchVendorConfigurationModelInfraEndpointTypeStreaming LaunchVendorConfigurationModelInfraEndpointType = "streaming"
GPUType stringOptional
One of the following:
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaTeslaT4 LaunchVendorConfigurationModelInfraGPUType = "nvidia-tesla-t4"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA10 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a10"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100e LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100e"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_1g20gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-1g20gb"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_3g40gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-3g40gb"
GPUs int64Optional
HighPriority boolOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
PerWorker int64Optional
PublicInference boolOptional
Storage stringOptional
ModelMetadata map[string, any]Optional
ModelType stringOptional
ModelVendor stringOptional
OnConflict stringOptional
One of the following:
const ModelNewParamsModelLaunchOnConflictError ModelNewParamsModelLaunchOnConflict = "error"
const ModelNewParamsModelLaunchOnConflictUpdate ModelNewParamsModelLaunchOnConflict = "update"
type ModelNewParamsModelLlmengine struct{…}
Name string

Unique name to reference your model

VendorConfiguration LlmEngineVendorConfiguration
Model string
ChatTemplateOverride stringOptional
CheckpointPath stringOptional
CPUs int64Optional
DefaultCallbackURL stringOptional
EndpointType stringOptional
GPUType stringOptional
GPUs int64Optional
HighPriority boolOptional
InferenceFramework stringOptional
InferenceFrameworkImageTag stringOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
NodesPerWorker int64Optional
NumShards int64Optional
PerWorker int64Optional
PostInferenceHooks []stringOptional
PublicInference boolOptional
Quantize stringOptional
Source stringOptional
Storage stringOptional
ModelMetadata map[string, any]Optional
ModelType stringOptional
ModelVendor stringOptional
OnConflict stringOptional
One of the following:
const ModelNewParamsModelLlmengineOnConflictError ModelNewParamsModelLlmengineOnConflict = "error"
const ModelNewParamsModelLlmengineOnConflictUpdate ModelNewParamsModelLlmengineOnConflict = "update"
type ModelNewParamsModelHostedModelCreateRequest struct{…}

Register a model already served by an external / proxy-served vendor (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

Unlike launch/llmengine, no Scale-side deployment is performed: the record is created READY and is immediately callable via /v5/chat/completions. Accepted only when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor) covers every vendor except launch/llmengine, and no vendor_configuration applies.

Type of model, for example chat_completion

One of the following:
const InferenceModelTypeGeneric InferenceModelType = "generic"
const InferenceModelTypeCompletion InferenceModelType = "completion"
const InferenceModelTypeChatCompletion InferenceModelType = "chat_completion"
ModelVendor string

Vendor to serve/create model

One of the following:
const ModelNewParamsModelHostedModelCreateRequestModelVendorOpenAI ModelNewParamsModelHostedModelCreateRequestModelVendor = "openai"
const ModelNewParamsModelHostedModelCreateRequestModelVendorCohere ModelNewParamsModelHostedModelCreateRequestModelVendor = "cohere"
const ModelNewParamsModelHostedModelCreateRequestModelVendorVertexAI ModelNewParamsModelHostedModelCreateRequestModelVendor = "vertex_ai"
const ModelNewParamsModelHostedModelCreateRequestModelVendorAnthropic ModelNewParamsModelHostedModelCreateRequestModelVendor = "anthropic"
const ModelNewParamsModelHostedModelCreateRequestModelVendorAzure ModelNewParamsModelHostedModelCreateRequestModelVendor = "azure"
const ModelNewParamsModelHostedModelCreateRequestModelVendorGemini ModelNewParamsModelHostedModelCreateRequestModelVendor = "gemini"
const ModelNewParamsModelHostedModelCreateRequestModelVendorModelZoo ModelNewParamsModelHostedModelCreateRequestModelVendor = "model_zoo"
const ModelNewParamsModelHostedModelCreateRequestModelVendorBedrock ModelNewParamsModelHostedModelCreateRequestModelVendor = "bedrock"
const ModelNewParamsModelHostedModelCreateRequestModelVendorXai ModelNewParamsModelHostedModelCreateRequestModelVendor = "xai"
const ModelNewParamsModelHostedModelCreateRequestModelVendorFireworksAI ModelNewParamsModelHostedModelCreateRequestModelVendor = "fireworks_ai"
Name string

Unique name to reference your model

ModelMetadata map[string, any]Optional
OnConflict stringOptional
One of the following:
const ModelNewParamsModelHostedModelCreateRequestOnConflictError ModelNewParamsModelHostedModelCreateRequestOnConflict = "error"
const ModelNewParamsModelHostedModelCreateRequestOnConflictUpdate ModelNewParamsModelHostedModelCreateRequestOnConflict = "update"
ReturnsExpand Collapse
type InferenceModel struct{…}
ID string

The unique identifier of the entity.

CreatedAt Time

The date and time when the entity was created in ISO format.

formatdate-time
CreatedByIdentityType InferenceModelCreatedByIdentityType

The type of identity that created the entity.

One of the following:
const InferenceModelCreatedByIdentityTypeUser InferenceModelCreatedByIdentityType = "user"
const InferenceModelCreatedByIdentityTypeServiceAccount InferenceModelCreatedByIdentityType = "service_account"
CreatedByUserID string

The user who originally created the entity.

One of the following:
const InferenceModelTypeGeneric InferenceModelType = "generic"
const InferenceModelTypeCompletion InferenceModelType = "completion"
const InferenceModelTypeChatCompletion InferenceModelType = "chat_completion"
One of the following:
const InferenceModelVendorOpenAI InferenceModelVendor = "openai"
const InferenceModelVendorCohere InferenceModelVendor = "cohere"
const InferenceModelVendorVertexAI InferenceModelVendor = "vertex_ai"
const InferenceModelVendorAnthropic InferenceModelVendor = "anthropic"
const InferenceModelVendorAzure InferenceModelVendor = "azure"
const InferenceModelVendorGemini InferenceModelVendor = "gemini"
const InferenceModelVendorLaunch InferenceModelVendor = "launch"
const InferenceModelVendorLlmengine InferenceModelVendor = "llmengine"
const InferenceModelVendorModelZoo InferenceModelVendor = "model_zoo"
const InferenceModelVendorBedrock InferenceModelVendor = "bedrock"
const InferenceModelVendorXai InferenceModelVendor = "xai"
const InferenceModelVendorFireworksAI InferenceModelVendor = "fireworks_ai"
Name string
Status InferenceModelStatus
One of the following:
const InferenceModelStatusFailed InferenceModelStatus = "failed"
const InferenceModelStatusReady InferenceModelStatus = "ready"
const InferenceModelStatusDeploying InferenceModelStatus = "deploying"
const InferenceModelStatusDeploymentTimeout InferenceModelStatus = "deployment_timeout"
ModelAvailability InferenceModelAvailabilityOptional
One of the following:
const InferenceModelAvailabilityUnknown InferenceModelAvailability = "unknown"
const InferenceModelAvailabilityAvailable InferenceModelAvailability = "available"
const InferenceModelAvailabilityUnavailable InferenceModelAvailability = "unavailable"
ModelMetadata map[string, any]Optional
Object InferenceModelObjectOptional
StatusReason stringOptional
VendorConfiguration InferenceModelVendorConfigurationUnionOptional
One of the following:
type LaunchVendorConfiguration struct{…}
ModelImage LaunchVendorConfigurationModelImage
Command []string
Registry string
Repository string
Tag string
EnvVars map[string, any]Optional
HealthcheckRoute stringOptional
PredictRoute stringOptional
ReadinessDelay int64Optional
RequestSchema map[string, any]Optional
ResponseSchema map[string, any]Optional
StreamingCommand []stringOptional
StreamingPredictRoute stringOptional
ModelInfra LaunchVendorConfigurationModelInfra
CPUs LaunchVendorConfigurationModelInfraCPUsUnionOptional
One of the following:
string
int64
EndpointType stringOptional
One of the following:
const LaunchVendorConfigurationModelInfraEndpointTypeAsync LaunchVendorConfigurationModelInfraEndpointType = "async"
const LaunchVendorConfigurationModelInfraEndpointTypeSync LaunchVendorConfigurationModelInfraEndpointType = "sync"
const LaunchVendorConfigurationModelInfraEndpointTypeStreaming LaunchVendorConfigurationModelInfraEndpointType = "streaming"
GPUType stringOptional
One of the following:
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaTeslaT4 LaunchVendorConfigurationModelInfraGPUType = "nvidia-tesla-t4"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA10 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a10"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaAmpereA100e LaunchVendorConfigurationModelInfraGPUType = "nvidia-ampere-a100e"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100 LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_1g20gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-1g20gb"
const LaunchVendorConfigurationModelInfraGPUTypeNvidiaHopperH100_3g40gb LaunchVendorConfigurationModelInfraGPUType = "nvidia-hopper-h100-3g40gb"
GPUs int64Optional
HighPriority boolOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
PerWorker int64Optional
PublicInference boolOptional
Storage stringOptional
type LlmEngineVendorConfiguration struct{…}
Model string
ChatTemplateOverride stringOptional
CheckpointPath stringOptional
CPUs int64Optional
DefaultCallbackURL stringOptional
EndpointType stringOptional
GPUType stringOptional
GPUs int64Optional
HighPriority boolOptional
InferenceFramework stringOptional
InferenceFrameworkImageTag stringOptional
Labels map[string, string]Optional
MaxWorkers int64Optional
Memory stringOptional
MinWorkers int64Optional
NodesPerWorker int64Optional
NumShards int64Optional
PerWorker int64Optional
PostInferenceHooks []stringOptional
PublicInference boolOptional
Quantize stringOptional
Source stringOptional
Storage stringOptional

Create a custom model

package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  inferenceModel, err := client.Models.New(context.TODO(), sgpdev.ModelNewParams{
    OfLaunch: &sgpdev.ModelNewParamsModelLaunch{
      Name: "name",
      VendorConfiguration: sgpdev.LaunchVendorConfigurationParam{
        ModelImage: sgpdev.LaunchVendorConfigurationModelImageParam{
          Command: []string{"string"},
          Registry: "registry",
          Repository: "repository",
          Tag: "tag",
        },
        ModelInfra: sgpdev.LaunchVendorConfigurationModelInfraParam{

        },
      },
      ModelVendor: "launch",
    },
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", inferenceModel.ID)
}
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
Returns Examples
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}