Skip to content

Create a custom model

client.models.create(ModelCreateParams { model } params, RequestOptionsoptions?): InferenceModel { id, created_at, created_by_identity_type, 10 more }
POST/v5/models

Create a custom model record in your account and begin deploying it through a supported serving vendor.

A model here is a record for a model you deploy and serve through Scale’s own inference vendors: only the launch and llmengine vendors are accepted and any other vendor is rejected. This is distinct from GET /v5/chat/completions/models, which lists the models already available to call for chat completions rather than creating or managing these records. The call is asynchronous — the record is created in a deploying status, a deployment job is recorded, and a Temporal workflow is started to perform the deployment, so the model is not ready for inference when this returns. A model name must be unique per vendor within your account; if a model with the same name and vendor already exists the request fails unless on_conflict is set to update, in which case the existing model is updated instead.

ParametersExpand Collapse
params: ModelCreateParams { model }
model: LaunchModelCreateRequest { name, vendor_configuration, model_metadata, 3 more } | LlmEngineModelCreateRequest { name, vendor_configuration, model_metadata, 3 more } | HostedModelCreateRequest { model_type, model_vendor, name, 2 more }

Register a model already served by an external / proxy-served vendor (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

Unlike launch/llmengine, no Scale-side deployment is performed: the record is created READY and is immediately callable via /v5/chat/completions. Accepted only when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor) covers every vendor except launch/llmengine, and no vendor_configuration applies.

One of the following:
LaunchModelCreateRequest { name, vendor_configuration, model_metadata, 3 more }
name: string

Unique name to reference your model

vendor_configuration: LaunchVendorConfiguration { model_image, model_infra }
model_image: ModelImage { command, registry, repository, 9 more }
command: Array<string>
registry: string
repository: string
tag: string
env_vars?: Record<string, unknown>
healthcheck_route?: string
predict_route?: string
readiness_delay?: number
request_schema?: Record<string, unknown>
response_schema?: Record<string, unknown>
streaming_command?: Array<string>
streaming_predict_route?: string
model_infra: ModelInfra { cpus, endpoint_type, gpu_type, 9 more }
cpus?: string | number
One of the following:
string
number
endpoint_type?: "async" | "sync" | "streaming"
One of the following:
"async"
"sync"
"streaming"
gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus?: number
high_priority?: boolean
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
per_worker?: number
public_inference?: boolean
storage?: string
model_metadata?: Record<string, unknown>
model_type?: "generic"
model_vendor?: "launch"
on_conflict?: "error" | "update"
One of the following:
"error"
"update"
LlmEngineModelCreateRequest { name, vendor_configuration, model_metadata, 3 more }
name: string

Unique name to reference your model

vendor_configuration: LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
model: string
chat_template_override?: string
checkpoint_path?: string
cpus?: number
default_callback_url?: string
endpoint_type?: string
gpu_type?: string
gpus?: number
high_priority?: boolean
inference_framework?: string
inference_framework_image_tag?: string
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
nodes_per_worker?: number
num_shards?: number
per_worker?: number
post_inference_hooks?: Array<string>
public_inference?: boolean
quantize?: string
source?: string
storage?: string
model_metadata?: Record<string, unknown>
model_type?: "chat_completion"
model_vendor?: "llmengine"
on_conflict?: "error" | "update"
One of the following:
"error"
"update"
HostedModelCreateRequest { model_type, model_vendor, name, 2 more }

Register a model already served by an external / proxy-served vendor (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

Unlike launch/llmengine, no Scale-side deployment is performed: the record is created READY and is immediately callable via /v5/chat/completions. Accepted only when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor) covers every vendor except launch/llmengine, and no vendor_configuration applies.

model_type: InferenceModelType

Type of model, for example chat_completion

One of the following:
"generic"
"completion"
"chat_completion"
model_vendor: "openai" | "cohere" | "vertex_ai" | 7 more

Vendor to serve/create model

One of the following:
"openai"
"cohere"
"vertex_ai"
"anthropic"
"azure"
"gemini"
"model_zoo"
"bedrock"
"xai"
"fireworks_ai"
name: string

Unique name to reference your model

model_metadata?: Record<string, unknown>
on_conflict?: "error" | "update"
One of the following:
"error"
"update"
ReturnsExpand Collapse
InferenceModel { id, created_at, created_by_identity_type, 10 more }
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by_identity_type: "user" | "service_account"

The type of identity that created the entity.

One of the following:
"user"
"service_account"
created_by_user_id: string

The user who originally created the entity.

model_type: InferenceModelType
One of the following:
"generic"
"completion"
"chat_completion"
model_vendor: InferenceModelVendor
One of the following:
"openai"
"cohere"
"vertex_ai"
"anthropic"
"azure"
"gemini"
"launch"
"llmengine"
"model_zoo"
"bedrock"
"xai"
"fireworks_ai"
name: string
status: "failed" | "ready" | "deploying" | "deployment_timeout"
One of the following:
"failed"
"ready"
"deploying"
"deployment_timeout"
model_availability?: InferenceModelAvailability
One of the following:
"unknown"
"available"
"unavailable"
model_metadata?: Record<string, unknown>
object?: "model"
status_reason?: string
vendor_configuration?: LaunchVendorConfiguration { model_image, model_infra } | LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
One of the following:
LaunchVendorConfiguration { model_image, model_infra }
model_image: ModelImage { command, registry, repository, 9 more }
command: Array<string>
registry: string
repository: string
tag: string
env_vars?: Record<string, unknown>
healthcheck_route?: string
predict_route?: string
readiness_delay?: number
request_schema?: Record<string, unknown>
response_schema?: Record<string, unknown>
streaming_command?: Array<string>
streaming_predict_route?: string
model_infra: ModelInfra { cpus, endpoint_type, gpu_type, 9 more }
cpus?: string | number
One of the following:
string
number
endpoint_type?: "async" | "sync" | "streaming"
One of the following:
"async"
"sync"
"streaming"
gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus?: number
high_priority?: boolean
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
per_worker?: number
public_inference?: boolean
storage?: string
LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
model: string
chat_template_override?: string
checkpoint_path?: string
cpus?: number
default_callback_url?: string
endpoint_type?: string
gpu_type?: string
gpus?: number
high_priority?: boolean
inference_framework?: string
inference_framework_image_tag?: string
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
nodes_per_worker?: number
num_shards?: number
per_worker?: number
post_inference_hooks?: Array<string>
public_inference?: boolean
quantize?: string
source?: string
storage?: string

Create a custom model

import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const inferenceModel = await client.models.create({
  model: {
    name: 'name',
    vendor_configuration: {
      model_image: {
        command: ['string'],
        registry: 'registry',
        repository: 'repository',
        tag: 'tag',
      },
      model_infra: {},
    },
    model_vendor: 'launch',
  },
});

console.log(inferenceModel.id);
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
Returns Examples
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}