Skip to content

List custom models

GET/v5/models

List the custom model records registered in your account.

Returns a paginated list of the model records managed through this API — models your account deploys through the launch or llmengine serving vendors — optionally filtered by name and by model vendor, and scoped to the caller’s account. This is different from GET /v5/chat/completions/models, which lists the models available to invoke for chat completions; this endpoint returns the managed records along with their deployment status, not the catalog of callable completion models.

Query ParametersExpand Collapse
ending_before: optional string
limit: optional number
maximum10000
minimum1
model_vendor: optional InferenceModelVendor
One of the following:
"openai"
"cohere"
"vertex_ai"
"anthropic"
"azure"
"gemini"
"launch"
"llmengine"
"model_zoo"
"bedrock"
"xai"
"fireworks_ai"
name: optional string
sort_by: optional string
sort_order: optional SortOrder
One of the following:
"asc"
"desc"
starting_after: optional string
ReturnsExpand Collapse
has_more: boolean

Whether there are more items left to be fetched.

items: array of InferenceModel { id, created_at, created_by_identity_type, 10 more }
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by_identity_type: "user" or "service_account"

The type of identity that created the entity.

One of the following:
"user"
"service_account"
created_by_user_id: string

The user who originally created the entity.

model_type: InferenceModelType
One of the following:
"generic"
"completion"
"chat_completion"
model_vendor: InferenceModelVendor
One of the following:
"openai"
"cohere"
"vertex_ai"
"anthropic"
"azure"
"gemini"
"launch"
"llmengine"
"model_zoo"
"bedrock"
"xai"
"fireworks_ai"
name: string
status: "failed" or "ready" or "deploying" or "deployment_timeout"
One of the following:
"failed"
"ready"
"deploying"
"deployment_timeout"
model_availability: optional InferenceModelAvailability
One of the following:
"unknown"
"available"
"unavailable"
model_metadata: optional map[unknown]
object: optional "model"
status_reason: optional string
vendor_configuration: optional LaunchVendorConfiguration { model_image, model_infra } or LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
One of the following:
LaunchVendorConfiguration object { model_image, model_infra }
model_image: object { command, registry, repository, 9 more }
command: array of string
registry: string
repository: string
tag: string
env_vars: optional map[unknown]
healthcheck_route: optional string
predict_route: optional string
readiness_delay: optional number
request_schema: optional map[unknown]
response_schema: optional map[unknown]
streaming_command: optional array of string
streaming_predict_route: optional string
model_infra: object { cpus, endpoint_type, gpu_type, 9 more }
cpus: optional string or number
One of the following:
string
number
endpoint_type: optional "async" or "sync" or "streaming"
One of the following:
"async"
"sync"
"streaming"
gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus: optional number
high_priority: optional boolean
labels: optional map[string]
max_workers: optional number
memory: optional string
min_workers: optional number
per_worker: optional number
public_inference: optional boolean
storage: optional string
LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }
model: string
chat_template_override: optional string
checkpoint_path: optional string
cpus: optional number
default_callback_url: optional string
endpoint_type: optional string
gpu_type: optional string
gpus: optional number
high_priority: optional boolean
inference_framework: optional string
inference_framework_image_tag: optional string
labels: optional map[string]
max_workers: optional number
memory: optional string
min_workers: optional number
nodes_per_worker: optional number
num_shards: optional number
per_worker: optional number
post_inference_hooks: optional array of string
public_inference: optional boolean
quantize: optional string
source: optional string
storage: optional string
total: number

The total of items that match the query. This is greater than or equal to the number of items returned.

limit: optional number

The maximum number of items to return.

object: optional "list"

List custom models

curl https://api.egp.scale.com/v5/models \
    -H "x-api-key: $SGP_API_KEY"
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by_identity_type": "user",
      "created_by_user_id": "created_by_user_id",
      "model_type": "generic",
      "model_vendor": "openai",
      "name": "name",
      "status": "failed",
      "model_availability": "unknown",
      "model_metadata": {
        "foo": "bar"
      },
      "object": "model",
      "status_reason": "status_reason",
      "vendor_configuration": {
        "model_image": {
          "command": [
            "string"
          ],
          "registry": "registry",
          "repository": "repository",
          "tag": "tag",
          "env_vars": {
            "foo": "bar"
          },
          "healthcheck_route": "healthcheck_route",
          "predict_route": "predict_route",
          "readiness_delay": 0,
          "request_schema": {
            "foo": "bar"
          },
          "response_schema": {
            "foo": "bar"
          },
          "streaming_command": [
            "string"
          ],
          "streaming_predict_route": "streaming_predict_route"
        },
        "model_infra": {
          "cpus": "string",
          "endpoint_type": "async",
          "gpu_type": "nvidia-tesla-t4",
          "gpus": 0,
          "high_priority": true,
          "labels": {
            "foo": "string"
          },
          "max_workers": 0,
          "memory": "memory",
          "min_workers": 0,
          "per_worker": 0,
          "public_inference": true,
          "storage": "storage"
        }
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
Returns Examples
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by_identity_type": "user",
      "created_by_user_id": "created_by_user_id",
      "model_type": "generic",
      "model_vendor": "openai",
      "name": "name",
      "status": "failed",
      "model_availability": "unknown",
      "model_metadata": {
        "foo": "bar"
      },
      "object": "model",
      "status_reason": "status_reason",
      "vendor_configuration": {
        "model_image": {
          "command": [
            "string"
          ],
          "registry": "registry",
          "repository": "repository",
          "tag": "tag",
          "env_vars": {
            "foo": "bar"
          },
          "healthcheck_route": "healthcheck_route",
          "predict_route": "predict_route",
          "readiness_delay": 0,
          "request_schema": {
            "foo": "bar"
          },
          "response_schema": {
            "foo": "bar"
          },
          "streaming_command": [
            "string"
          ],
          "streaming_predict_route": "streaming_predict_route"
        },
        "model_infra": {
          "cpus": "string",
          "endpoint_type": "async",
          "gpu_type": "nvidia-tesla-t4",
          "gpus": 0,
          "high_priority": true,
          "labels": {
            "foo": "string"
          },
          "max_workers": 0,
          "memory": "memory",
          "min_workers": 0,
          "per_worker": 0,
          "public_inference": true,
          "storage": "storage"
        }
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}