## Create a custom model

`models.create(ModelCreateParams**kwargs)  -> InferenceModel`

**post** `/v5/models`

Create a custom model record in your account and begin deploying it through a supported serving vendor.

A model here is a record for a model you deploy and serve through Scale's own inference vendors: only the `launch` and `llmengine` vendors are accepted and any other vendor is rejected. This is distinct from `GET /v5/chat/completions/models`, which lists the models already available to call for chat completions rather than creating or managing these records. The call is asynchronous — the record is created in a deploying status, a deployment job is recorded, and a Temporal workflow is started to perform the deployment, so the model is not ready for inference when this returns. A model name must be unique per vendor within your account; if a model with the same name and vendor already exists the request fails unless `on_conflict` is set to `update`, in which case the existing model is updated instead.

### Parameters

- `model: Model`

  Register a model already served by an external / proxy-served vendor
  (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

  Unlike launch/llmengine, no Scale-side deployment is performed: the record is
  created READY and is immediately callable via /v5/chat/completions. Accepted only
  when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
  covers every vendor except launch/llmengine, and no vendor_configuration applies.

  - `class ModelLaunchModelCreateRequest: …`

    - `name: str`

      Unique name to reference your model

    - `vendor_configuration: LaunchVendorConfigurationParam`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `model_metadata: Optional[Dict[str, object]]`

    - `model_type: Optional[Literal["generic"]]`

      - `"generic"`

    - `model_vendor: Optional[Literal["launch"]]`

      - `"launch"`

    - `on_conflict: Optional[Literal["error", "update"]]`

      - `"error"`

      - `"update"`

  - `class ModelLlmEngineModelCreateRequest: …`

    - `name: str`

      Unique name to reference your model

    - `vendor_configuration: LlmEngineVendorConfigurationParam`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

    - `model_metadata: Optional[Dict[str, object]]`

    - `model_type: Optional[Literal["chat_completion"]]`

      - `"chat_completion"`

    - `model_vendor: Optional[Literal["llmengine"]]`

      - `"llmengine"`

    - `on_conflict: Optional[Literal["error", "update"]]`

      - `"error"`

      - `"update"`

  - `class ModelHostedModelCreateRequest: …`

    Register a model already served by an external / proxy-served vendor
    (e.g. an OpenAI-compatible self-hosted model behind the inference proxy).

    Unlike launch/llmengine, no Scale-side deployment is performed: the record is
    created READY and is immediately callable via /v5/chat/completions. Accepted only
    when NATIVE_OPENAI_INFERENCE_GATEWAY is enabled. The discriminator (model_vendor)
    covers every vendor except launch/llmengine, and no vendor_configuration applies.

    - `model_type: InferenceModelType`

      Type of model, for example `chat_completion`

      - `"generic"`

      - `"completion"`

      - `"chat_completion"`

    - `model_vendor: Literal["openai", "cohere", "vertex_ai", 7 more]`

      Vendor to serve/create model

      - `"openai"`

      - `"cohere"`

      - `"vertex_ai"`

      - `"anthropic"`

      - `"azure"`

      - `"gemini"`

      - `"model_zoo"`

      - `"bedrock"`

      - `"xai"`

      - `"fireworks_ai"`

    - `name: str`

      Unique name to reference your model

    - `model_metadata: Optional[Dict[str, object]]`

    - `on_conflict: Optional[Literal["error", "update"]]`

      - `"error"`

      - `"update"`

### Returns

- `class InferenceModel: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by_identity_type: Literal["user", "service_account"]`

    The type of identity that created the entity.

    - `"user"`

    - `"service_account"`

  - `created_by_user_id: str`

    The user who originally created the entity.

  - `model_type: InferenceModelType`

    - `"generic"`

    - `"completion"`

    - `"chat_completion"`

  - `model_vendor: InferenceModelVendor`

    - `"openai"`

    - `"cohere"`

    - `"vertex_ai"`

    - `"anthropic"`

    - `"azure"`

    - `"gemini"`

    - `"launch"`

    - `"llmengine"`

    - `"model_zoo"`

    - `"bedrock"`

    - `"xai"`

    - `"fireworks_ai"`

  - `name: str`

  - `status: Literal["failed", "ready", "deploying", "deployment_timeout"]`

    - `"failed"`

    - `"ready"`

    - `"deploying"`

    - `"deployment_timeout"`

  - `model_availability: Optional[InferenceModelAvailability]`

    - `"unknown"`

    - `"available"`

    - `"unavailable"`

  - `model_metadata: Optional[Dict[str, object]]`

  - `object: Optional[Literal["model"]]`

    - `"model"`

  - `status_reason: Optional[str]`

  - `vendor_configuration: Optional[VendorConfiguration]`

    - `class LaunchVendorConfiguration: …`

      - `model_image: ModelImage`

        - `command: List[str]`

        - `registry: str`

        - `repository: str`

        - `tag: str`

        - `env_vars: Optional[Dict[str, object]]`

        - `healthcheck_route: Optional[str]`

        - `predict_route: Optional[str]`

        - `readiness_delay: Optional[int]`

        - `request_schema: Optional[Dict[str, object]]`

        - `response_schema: Optional[Dict[str, object]]`

        - `streaming_command: Optional[List[str]]`

        - `streaming_predict_route: Optional[str]`

      - `model_infra: ModelInfra`

        - `cpus: Optional[Union[str, int, null]]`

          - `str`

          - `int`

        - `endpoint_type: Optional[Literal["async", "sync", "streaming"]]`

          - `"async"`

          - `"sync"`

          - `"streaming"`

        - `gpu_type: Optional[Literal["nvidia-tesla-t4", "nvidia-ampere-a10", "nvidia-ampere-a100", 4 more]]`

          - `"nvidia-tesla-t4"`

          - `"nvidia-ampere-a10"`

          - `"nvidia-ampere-a100"`

          - `"nvidia-ampere-a100e"`

          - `"nvidia-hopper-h100"`

          - `"nvidia-hopper-h100-1g20gb"`

          - `"nvidia-hopper-h100-3g40gb"`

        - `gpus: Optional[int]`

        - `high_priority: Optional[bool]`

        - `labels: Optional[Dict[str, str]]`

        - `max_workers: Optional[int]`

        - `memory: Optional[str]`

        - `min_workers: Optional[int]`

        - `per_worker: Optional[int]`

        - `public_inference: Optional[bool]`

        - `storage: Optional[str]`

    - `class LlmEngineVendorConfiguration: …`

      - `model: str`

      - `chat_template_override: Optional[str]`

      - `checkpoint_path: Optional[str]`

      - `cpus: Optional[int]`

      - `default_callback_url: Optional[str]`

      - `endpoint_type: Optional[str]`

      - `gpu_type: Optional[str]`

      - `gpus: Optional[int]`

      - `high_priority: Optional[bool]`

      - `inference_framework: Optional[str]`

      - `inference_framework_image_tag: Optional[str]`

      - `labels: Optional[Dict[str, str]]`

      - `max_workers: Optional[int]`

      - `memory: Optional[str]`

      - `min_workers: Optional[int]`

      - `nodes_per_worker: Optional[int]`

      - `num_shards: Optional[int]`

      - `per_worker: Optional[int]`

      - `post_inference_hooks: Optional[List[str]]`

      - `public_inference: Optional[bool]`

      - `quantize: Optional[str]`

      - `source: Optional[str]`

      - `storage: Optional[str]`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
inference_model = client.models.create(
    model={
        "name": "name",
        "vendor_configuration": {
            "model_image": {
                "command": ["string"],
                "registry": "registry",
                "repository": "repository",
                "tag": "tag",
            },
            "model_infra": {},
        },
        "model_vendor": "launch",
    },
)
print(inference_model.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by_identity_type": "user",
  "created_by_user_id": "created_by_user_id",
  "model_type": "generic",
  "model_vendor": "openai",
  "name": "name",
  "status": "failed",
  "model_availability": "unknown",
  "model_metadata": {
    "foo": "bar"
  },
  "object": "model",
  "status_reason": "status_reason",
  "vendor_configuration": {
    "model_image": {
      "command": [
        "string"
      ],
      "registry": "registry",
      "repository": "repository",
      "tag": "tag",
      "env_vars": {
        "foo": "bar"
      },
      "healthcheck_route": "healthcheck_route",
      "predict_route": "predict_route",
      "readiness_delay": 0,
      "request_schema": {
        "foo": "bar"
      },
      "response_schema": {
        "foo": "bar"
      },
      "streaming_command": [
        "string"
      ],
      "streaming_predict_route": "streaming_predict_route"
    },
    "model_infra": {
      "cpus": "string",
      "endpoint_type": "async",
      "gpu_type": "nvidia-tesla-t4",
      "gpus": 0,
      "high_priority": true,
      "labels": {
        "foo": "string"
      },
      "max_workers": 0,
      "memory": "memory",
      "min_workers": 0,
      "per_worker": 0,
      "public_inference": true,
      "storage": "storage"
    }
  }
}
```
