## List custom models

**get** `/v5/models`

List the custom model records registered in your account.

Returns a paginated list of the model records managed through this API — models your account deploys through the `launch` or `llmengine` serving vendors — optionally filtered by name and by model vendor, and scoped to the caller's account. This is different from `GET /v5/chat/completions/models`, which lists the models available to invoke for chat completions; this endpoint returns the managed records along with their deployment status, not the catalog of callable completion models.

### Query Parameters

- `ending_before: optional string`

- `limit: optional number`

- `model_vendor: optional InferenceModelVendor`

  - `"openai"`

  - `"cohere"`

  - `"vertex_ai"`

  - `"anthropic"`

  - `"azure"`

  - `"gemini"`

  - `"launch"`

  - `"llmengine"`

  - `"model_zoo"`

  - `"bedrock"`

  - `"xai"`

  - `"fireworks_ai"`

- `name: optional string`

- `sort_by: optional string`

- `sort_order: optional SortOrder`

  - `"asc"`

  - `"desc"`

- `starting_after: optional string`

### Returns

- `has_more: boolean`

  Whether there are more items left to be fetched.

- `items: array of InferenceModel`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: "user" or "service_account"`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: string`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: string`

  - `status: "failed" or "ready" or "deploying" or "deployment_timeout"`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: optional InferenceModelAvailability`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: optional map[unknown]`

  - `object: optional "model"`

    - `"model"`

  - `status_reason: optional string`

  - `vendor_configuration: optional LaunchVendorConfiguration or LlmEngineVendorConfiguration`

    - `LaunchVendorConfiguration object { model_image, model_infra }`

      - `model_image: object { command, registry, repository, 9 more }`

        - `command: array of string`

        - `registry: string`

        - `repository: string`

        - `tag: string`

        - `env_vars: optional map[unknown]`

        - `healthcheck_route: optional string`

        - `predict_route: optional string`

        - `readiness_delay: optional number`

        - `request_schema: optional map[unknown]`

        - `response_schema: optional map[unknown]`

        - `streaming_command: optional array of string`

        - `streaming_predict_route: optional string`

      - `model_infra: object { cpus, endpoint_type, gpu_type, 9 more }`

        - `cpus: optional string or number`

          - `string`

          - `number`

        - `endpoint_type: optional "async" or "sync" or "streaming"`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: optional "nvidia-tesla-t4" or "nvidia-ampere-a10" or "nvidia-ampere-a100" or 4 more`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: optional number`

        - `high_priority: optional boolean`

        - `labels: optional map[string]`

        - `max_workers: optional number`

        - `memory: optional string`

        - `min_workers: optional number`

        - `per_worker: optional number`

        - `public_inference: optional boolean`

        - `storage: optional string`

    - `LlmEngineVendorConfiguration object { model, chat_template_override, checkpoint_path, 20 more }`

      - `model: string`

      - `chat_template_override: optional string`

      - `checkpoint_path: optional string`

      - `cpus: optional number`

      - `default_callback_url: optional string`

      - `endpoint_type: optional string`

      - `gpu_type: optional string`

      - `gpus: optional number`

      - `high_priority: optional boolean`

      - `inference_framework: optional string`

      - `inference_framework_image_tag: optional string`

      - `labels: optional map[string]`

      - `max_workers: optional number`

      - `memory: optional string`

      - `min_workers: optional number`

      - `nodes_per_worker: optional number`

      - `num_shards: optional number`

      - `per_worker: optional number`

      - `post_inference_hooks: optional array of string`

      - `public_inference: optional boolean`

      - `quantize: optional string`

      - `source: optional string`

      - `storage: optional string`

- `total: number`

  The total of items that match the query. This is greater than or equal to the number of items returned.

- `limit: optional number`

  The maximum number of items to return.

- `object: optional "list"`

  - `"list"`

### Example

```http
curl https://api.egp.scale.com/v5/models \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by_identity_type": "user",
      "created_by_user_id": "created_by_user_id",
      "model_type": "generic",
      "model_vendor": "openai",
      "name": "name",
      "status": "failed",
      "model_availability": "unknown",
      "model_metadata": {
        "foo": "bar"
      },
      "object": "model",
      "status_reason": "status_reason",
      "vendor_configuration": {
        "model_image": {
          "command": [
            "string"
          ],
          "registry": "registry",
          "repository": "repository",
          "tag": "tag",
          "env_vars": {
            "foo": "bar"
          },
          "healthcheck_route": "healthcheck_route",
          "predict_route": "predict_route",
          "readiness_delay": 0,
          "request_schema": {
            "foo": "bar"
          },
          "response_schema": {
            "foo": "bar"
          },
          "streaming_command": [
            "string"
          ],
          "streaming_predict_route": "streaming_predict_route"
        },
        "model_infra": {
          "cpus": "string",
          "endpoint_type": "async",
          "gpu_type": "nvidia-tesla-t4",
          "gpus": 0,
          "high_priority": true,
          "labels": {
            "foo": "string"
          },
          "max_workers": 0,
          "memory": "memory",
          "min_workers": 0,
          "per_worker": 0,
          "public_inference": true,
          "storage": "storage"
        }
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```
