Skip to content

Get a custom model

GET/v5/models/{model_id}

Retrieve a single custom model record by its ID.

Returns the model record — including its vendor, configuration, and current deployment status — for a model managed through this API and owned by the caller’s account. This is distinct from GET /v5/chat/completions/models, which lists the models available to call for chat completions rather than returning a single managed record.

Path ParametersExpand Collapse
model_id: string
ReturnsExpand Collapse
InferenceModel object { id, created_at, created_by_identity_type, 10 more }
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by_identity_type: "user" or "service_account"

The type of identity that created the entity.

One of the following:
"user"
"service_account"
created_by_user_id: string

The user who originally created the entity.

model_type: InferenceModelType
One of the following:
"generic"
"completion"
"chat_completion"
model_vendor: InferenceModelVendor
One of the following:
"openai"
"cohere"
"vertex_ai"
"anthropic"
"azure"
"gemini"
"launch"
"llmengine"
"model_zoo"
"bedrock"
"xai"
"fireworks_ai"
name: string
status: "failed" or "ready" or "deploying" or "deployment_timeout"
One of the following:
"failed"
"ready"
"deploying"
"deployment_timeout"
model_availability: optional InferenceModelAvailability
One of the following:
"unknown"
"available"
"unavailable"
model_metadata: optional map[unknown]
object: optional "model"
status_reason: optional string
vendor_configuration: optional LaunchVendorConfiguration { model_image, model_infra } or LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
One of the following:
LaunchVendorConfiguration object { model_image, model_infra }
model_image: object { command, registry, repository, 9 more }
command: array of string
registry: string
repository: string
tag: string
env_vars: optional map[unknown]
healthcheck_route: optional string
predict_route: optional string
readiness_delay: optional number
request_schema: optional map[unknown]
response_schema: optional map[unknown]
streaming_command: optional array of string
streaming_predict_route: optional string
model_infra: object { cpus, endpoint_type, gpu_type, 9 more }
cpus: optional string or number
One of the following:
string
number
endpoint_type: optional "async" or "sync" or "streaming"
One of the following:
"async"
"sync"
"streaming"
gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus: optional number
high_priority: optional boolean
labels: optional map[string]
max_workers: optional number
memory: optional string
min_workers: optional number
per_worker: optional number
public_inference: optional boolean
storage: optional string
LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }
model: string
chat_template_override: optional string
checkpoint_path: optional string
cpus: optional number
default_callback_url: optional string
endpoint_type: optional string
gpu_type: optional string
gpus: optional number
high_priority: optional boolean
inference_framework: optional string
inference_framework_image_tag: optional string
labels: optional map[string]
max_workers: optional number
memory: optional string
min_workers: optional number
nodes_per_worker: optional number
num_shards: optional number
per_worker: optional number
post_inference_hooks: optional array of string
public_inference: optional boolean
quantize: optional string
source: optional string
storage: optional string

Get a custom model

curl https://api.egp.scale.com/v5/models/$MODEL_ID \
    -H "x-api-key: $SGP_API_KEY"
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
Returns Examples
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}