Skip to content

Models

Create a custom model
client.models.create(ModelCreateParams { model } params, RequestOptionsoptions?): InferenceModel { id, created_at, created_by_identity_type, 10 more }
POST/v5/models
List custom models
client.models.list(ModelListParams { ending_before, limit, model_vendor, 4 more } query?, RequestOptionsoptions?): CursorPage<InferenceModel { id, created_at, created_by_identity_type, 10 more } >
GET/v5/models
Update a custom model
client.models.update(stringmodelID, ModelUpdateParams { model } params, RequestOptionsoptions?): InferenceModel { id, created_at, created_by_identity_type, 10 more }
PATCH/v5/models/{model_id}
Delete a custom model
client.models.delete(stringmodelID, RequestOptionsoptions?): ModelDeleteResponse { id, deleted, object }
DELETE/v5/models/{model_id}
Get a custom model
client.models.retrieve(stringmodelID, RequestOptionsoptions?): InferenceModel { id, created_at, created_by_identity_type, 10 more }
GET/v5/models/{model_id}
ModelsExpand Collapse
InferenceModel { id, created_at, created_by_identity_type, 10 more }
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by_identity_type: "user" | "service_account"

The type of identity that created the entity.

One of the following:
"user"
"service_account"
created_by_user_id: string

The user who originally created the entity.

model_type: InferenceModelType
One of the following:
"generic"
"completion"
"chat_completion"
model_vendor: InferenceModelVendor
One of the following:
"openai"
"cohere"
"vertex_ai"
"anthropic"
"azure"
"gemini"
"launch"
"llmengine"
"model_zoo"
"bedrock"
"xai"
"fireworks_ai"
name: string
status: "failed" | "ready" | "deploying" | "deployment_timeout"
One of the following:
"failed"
"ready"
"deploying"
"deployment_timeout"
model_availability?: InferenceModelAvailability
One of the following:
"unknown"
"available"
"unavailable"
model_metadata?: Record<string, unknown>
object?: "model"
status_reason?: string
vendor_configuration?: LaunchVendorConfiguration { model_image, model_infra } | LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
One of the following:
LaunchVendorConfiguration { model_image, model_infra }
model_image: ModelImage { command, registry, repository, 9 more }
command: Array<string>
registry: string
repository: string
tag: string
env_vars?: Record<string, unknown>
healthcheck_route?: string
predict_route?: string
readiness_delay?: number
request_schema?: Record<string, unknown>
response_schema?: Record<string, unknown>
streaming_command?: Array<string>
streaming_predict_route?: string
model_infra: ModelInfra { cpus, endpoint_type, gpu_type, 9 more }
cpus?: string | number
One of the following:
string
number
endpoint_type?: "async" | "sync" | "streaming"
One of the following:
"async"
"sync"
"streaming"
gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus?: number
high_priority?: boolean
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
per_worker?: number
public_inference?: boolean
storage?: string
LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
model: string
chat_template_override?: string
checkpoint_path?: string
cpus?: number
default_callback_url?: string
endpoint_type?: string
gpu_type?: string
gpus?: number
high_priority?: boolean
inference_framework?: string
inference_framework_image_tag?: string
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
nodes_per_worker?: number
num_shards?: number
per_worker?: number
post_inference_hooks?: Array<string>
public_inference?: boolean
quantize?: string
source?: string
storage?: string
InferenceModelAvailability = "unknown" | "available" | "unavailable"
One of the following:
"unknown"
"available"
"unavailable"
InferenceModelType = "generic" | "completion" | "chat_completion"
One of the following:
"generic"
"completion"
"chat_completion"
LaunchVendorConfiguration { model_image, model_infra }
model_image: ModelImage { command, registry, repository, 9 more }
command: Array<string>
registry: string
repository: string
tag: string
env_vars?: Record<string, unknown>
healthcheck_route?: string
predict_route?: string
readiness_delay?: number
request_schema?: Record<string, unknown>
response_schema?: Record<string, unknown>
streaming_command?: Array<string>
streaming_predict_route?: string
model_infra: ModelInfra { cpus, endpoint_type, gpu_type, 9 more }
cpus?: string | number
One of the following:
string
number
endpoint_type?: "async" | "sync" | "streaming"
One of the following:
"async"
"sync"
"streaming"
gpu_type?: "nvidia-tesla-t4" | "nvidia-ampere-a10" | "nvidia-ampere-a100" | 4 more
One of the following:
"nvidia-tesla-t4"
"nvidia-ampere-a10"
"nvidia-ampere-a100"
"nvidia-ampere-a100e"
"nvidia-hopper-h100"
"nvidia-hopper-h100-1g20gb"
"nvidia-hopper-h100-3g40gb"
gpus?: number
high_priority?: boolean
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
per_worker?: number
public_inference?: boolean
storage?: string
LlmEngineVendorConfiguration { model, chat_template_override, checkpoint_path, 20 more }
model: string
chat_template_override?: string
checkpoint_path?: string
cpus?: number
default_callback_url?: string
endpoint_type?: string
gpu_type?: string
gpus?: number
high_priority?: boolean
inference_framework?: string
inference_framework_image_tag?: string
labels?: Record<string, string>
max_workers?: number
memory?: string
min_workers?: number
nodes_per_worker?: number
num_shards?: number
per_worker?: number
post_inference_hooks?: Array<string>
public_inference?: boolean
quantize?: string
source?: string
storage?: string
ModelDeleteResponse { id, deleted, object }
id: string
deleted: boolean
object?: "model"