Skip to content

Generate OpenAI chat completion from messages

chat.completions.create(CompletionCreateParams**kwargs) -> CompletionCreateResponse
POST/v5/chat/completions

Generates a chat completion from an OpenAI-style messages array.

Use this endpoint for standard messages-based chat inference; use /v5/completions instead when you have a raw text prompt rather than messages, /v5/responses for the OpenAI Responses API contract, and /v5/inference when the payload does not follow any OpenAI schema. The request accepts the OpenAI Chat Completions parameters (extra fields are allowed and forwarded), and the model is selected from model given as vendor/name; most vendors are served through the litellm proxy gateway, while OpenAI may use a native gateway. When stream is set the response is delivered as server-sent events of chat.completion.chunk; otherwise a single chat.completion object is returned. Token usage is recorded for the account, and for streaming responses it is read from the final chunk.

ParametersExpand Collapse
messages: Iterable[Dict[str, object]]

openai standard message format

model: str

model specified as model_vendor/model, for example openai/gpt-4o

audio: Optional[Dict[str, object]]

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

frequency_penalty: Optional[float]

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

maximum2
minimum-2
function_call: Optional[Dict[str, object]]

Deprecated in favor of tool_choice. Controls which function is called by the model.

functions: Optional[Iterable[Dict[str, object]]]

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

logit_bias: Optional[Dict[str, int]]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

logprobs: Optional[bool]

Whether to return log probabilities of the output tokens or not.

max_completion_tokens: Optional[int]

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

max_tokens: Optional[int]

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

metadata: Optional[Dict[str, str]]

Developer-defined tags and values used for filtering completions in the dashboard.

modalities: Optional[Sequence[str]]

Output types that you would like the model to generate for this request.

n: Optional[int]

How many chat completion choices to generate for each input message.

parallel_tool_calls: Optional[bool]

Whether to enable parallel function calling during tool use.

prediction: Optional[Dict[str, object]]

Static predicted output content, such as the content of a text file being regenerated.

presence_penalty: Optional[float]

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

maximum2
minimum-2
reasoning_effort: Optional[str]

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format: Optional[Dict[str, object]]

An object specifying the format that the model must output.

seed: Optional[int]

If specified, system will attempt to sample deterministically for repeated requests with same seed.

stop: Optional[Union[str, Sequence[str]]]

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
str
Sequence[str]
store: Optional[bool]

Whether to store the output for use in model distillation or evals products.

stream: Optional[Literal[false]]

If true, partial message deltas will be sent as server-sent events.

stream_options: Optional[Dict[str, object]]

Options for streaming response. Only set this when stream is true.

temperature: Optional[float]

What sampling temperature to use. Higher values make output more random, lower more focused.

maximum2
minimum0
tool_choice: Optional[Union[str, Dict[str, object]]]

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
str
Dict[str, object]
tools: Optional[Iterable[Dict[str, object]]]

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

top_k: Optional[int]

Only sample from the top K options for each subsequent token

top_logprobs: Optional[int]

Number of most likely tokens to return at each position, with associated log probability.

maximum20
minimum0
top_p: Optional[float]

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

maximum1
minimum0
x_openai_api_key: Optional[str]
ReturnsExpand Collapse
One of the following:
class ChatCompletion: …
id: str
choices: List[Choice]
finish_reason: Literal["stop", "length", "tool_calls", 2 more]
One of the following:
"stop"
"length"
"tool_calls"
"content_filter"
"function_call"
index: int
message: ChoiceMessage

A chat completion message generated by the model.

role: Literal["assistant"]
annotations: Optional[List[ChoiceMessageAnnotation]]
type: Literal["url_citation"]
url_citation: ChoiceMessageAnnotationURLCitation

A URL citation when using web search.

end_index: int
start_index: int
title: str
url: str
audio: Optional[ChoiceMessageAudio]

If the audio output modality is requested, this object contains data about the audio response from the model. Learn more.

id: str
data: str
expires_at: int
transcript: str
content: Optional[str]
function_call: Optional[ChoiceMessageFunctionCall]

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

arguments: str
name: str
refusal: Optional[str]
tool_calls: Optional[List[ChoiceMessageToolCall]]
One of the following:
class ChoiceMessageToolCallChatCompletionMessageFunctionToolCall: …

A call to a function tool created by the model.

id: str
function: ChoiceMessageToolCallChatCompletionMessageFunctionToolCallFunction

The function that the model called.

arguments: str
name: str
type: Literal["function"]
class ChoiceMessageToolCallChatCompletionMessageCustomToolCall: …

A call to a custom tool created by the model.

id: str
custom: ChoiceMessageToolCallChatCompletionMessageCustomToolCallCustom

The custom tool that the model called.

input: str
name: str
type: Literal["custom"]
logprobs: Optional[ChoiceLogprobs]

Log probability information for the choice.

content: Optional[List[ChatCompletionTokenLogprob]]
token: str
logprob: float
top_logprobs: List[TopLogprob]
token: str
logprob: float
bytes: Optional[List[int]]
bytes: Optional[List[int]]
refusal: Optional[List[ChatCompletionTokenLogprob]]
token: str
logprob: float
top_logprobs: List[TopLogprob]
token: str
logprob: float
bytes: Optional[List[int]]
bytes: Optional[List[int]]
created: int
model: str
object: Optional[Literal["chat.completion"]]
service_tier: Optional[Literal["auto", "default", "flex", 2 more]]
One of the following:
"auto"
"default"
"flex"
"scale"
"priority"
system_fingerprint: Optional[str]
usage: Optional[CompletionUsage]

Usage statistics for the completion request.

completion_tokens: int
prompt_tokens: int
total_tokens: int
completion_tokens_details: Optional[CompletionTokensDetails]

Breakdown of tokens used in a completion.

accepted_prediction_tokens: Optional[int]
audio_tokens: Optional[int]
reasoning_tokens: Optional[int]
rejected_prediction_tokens: Optional[int]
prompt_tokens_details: Optional[PromptTokensDetails]

Breakdown of tokens used in the prompt.

audio_tokens: Optional[int]
cached_tokens: Optional[int]
class ChatCompletionChunk: …
id: str
choices: List[Choice]
delta: ChoiceDelta

A chat completion delta generated by streamed model responses.

content: Optional[str]
function_call: Optional[ChoiceDeltaFunctionCall]

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

arguments: Optional[str]
name: Optional[str]
refusal: Optional[str]
role: Optional[Literal["developer", "system", "user", 2 more]]
One of the following:
"developer"
"system"
"user"
"assistant"
"tool"
tool_calls: Optional[List[ChoiceDeltaToolCall]]
index: int
id: Optional[str]
function: Optional[ChoiceDeltaToolCallFunction]
arguments: Optional[str]
name: Optional[str]
type: Optional[Literal["function"]]
index: int
finish_reason: Optional[Literal["stop", "length", "tool_calls", 2 more]]
One of the following:
"stop"
"length"
"tool_calls"
"content_filter"
"function_call"
logprobs: Optional[ChoiceLogprobs]

Log probability information for the choice.

content: Optional[List[ChatCompletionTokenLogprob]]
token: str
logprob: float
top_logprobs: List[TopLogprob]
token: str
logprob: float
bytes: Optional[List[int]]
bytes: Optional[List[int]]
refusal: Optional[List[ChatCompletionTokenLogprob]]
token: str
logprob: float
top_logprobs: List[TopLogprob]
token: str
logprob: float
bytes: Optional[List[int]]
bytes: Optional[List[int]]
created: int
model: str
object: Optional[Literal["chat.completion.chunk"]]
service_tier: Optional[Literal["auto", "default", "flex", 2 more]]
One of the following:
"auto"
"default"
"flex"
"scale"
"priority"
system_fingerprint: Optional[str]
usage: Optional[CompletionUsage]

Usage statistics for the completion request.

completion_tokens: int
prompt_tokens: int
total_tokens: int
completion_tokens_details: Optional[CompletionTokensDetails]

Breakdown of tokens used in a completion.

accepted_prediction_tokens: Optional[int]
audio_tokens: Optional[int]
reasoning_tokens: Optional[int]
rejected_prediction_tokens: Optional[int]
prompt_tokens_details: Optional[PromptTokensDetails]

Breakdown of tokens used in the prompt.

audio_tokens: Optional[int]
cached_tokens: Optional[int]
class ChatCompletionChunk: …
id: str
choices: List[Choice]
delta: ChoiceDelta

A chat completion delta generated by streamed model responses.

content: Optional[str]
function_call: Optional[ChoiceDeltaFunctionCall]

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

arguments: Optional[str]
name: Optional[str]
refusal: Optional[str]
role: Optional[Literal["developer", "system", "user", 2 more]]
One of the following:
"developer"
"system"
"user"
"assistant"
"tool"
tool_calls: Optional[List[ChoiceDeltaToolCall]]
index: int
id: Optional[str]
function: Optional[ChoiceDeltaToolCallFunction]
arguments: Optional[str]
name: Optional[str]
type: Optional[Literal["function"]]
index: int
finish_reason: Optional[Literal["stop", "length", "tool_calls", 2 more]]
One of the following:
"stop"
"length"
"tool_calls"
"content_filter"
"function_call"
logprobs: Optional[ChoiceLogprobs]

Log probability information for the choice.

content: Optional[List[ChatCompletionTokenLogprob]]
token: str
logprob: float
top_logprobs: List[TopLogprob]
token: str
logprob: float
bytes: Optional[List[int]]
bytes: Optional[List[int]]
refusal: Optional[List[ChatCompletionTokenLogprob]]
token: str
logprob: float
top_logprobs: List[TopLogprob]
token: str
logprob: float
bytes: Optional[List[int]]
bytes: Optional[List[int]]
created: int
model: str
object: Optional[Literal["chat.completion.chunk"]]
service_tier: Optional[Literal["auto", "default", "flex", 2 more]]
One of the following:
"auto"
"default"
"flex"
"scale"
"priority"
system_fingerprint: Optional[str]
usage: Optional[CompletionUsage]

Usage statistics for the completion request.

completion_tokens: int
prompt_tokens: int
total_tokens: int
completion_tokens_details: Optional[CompletionTokensDetails]

Breakdown of tokens used in a completion.

accepted_prediction_tokens: Optional[int]
audio_tokens: Optional[int]
reasoning_tokens: Optional[int]
rejected_prediction_tokens: Optional[int]
prompt_tokens_details: Optional[PromptTokensDetails]

Breakdown of tokens used in the prompt.

audio_tokens: Optional[int]
cached_tokens: Optional[int]

Generate OpenAI chat completion from messages

import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
for completion in client.chat.completions.create(
    messages=[{
        "foo": "bar"
    }],
    model="model",
):
  print(completion)
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "role": "assistant",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": {
              "end_index": 0,
              "start_index": 0,
              "title": "title",
              "url": "url"
            }
          }
        ],
        "audio": {
          "id": "id",
          "data": "data",
          "expires_at": 0,
          "transcript": "transcript"
        },
        "content": "content",
        "function_call": {
          "arguments": "arguments",
          "name": "name"
        },
        "refusal": "refusal",
        "tool_calls": [
          {
            "id": "id",
            "function": {
              "arguments": "arguments",
              "name": "name"
            },
            "type": "function"
          }
        ]
      },
      "logprobs": {
        "content": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ],
        "refusal": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "chat.completion",
  "service_tier": "auto",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}
Returns Examples
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "role": "assistant",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": {
              "end_index": 0,
              "start_index": 0,
              "title": "title",
              "url": "url"
            }
          }
        ],
        "audio": {
          "id": "id",
          "data": "data",
          "expires_at": 0,
          "transcript": "transcript"
        },
        "content": "content",
        "function_call": {
          "arguments": "arguments",
          "name": "name"
        },
        "refusal": "refusal",
        "tool_calls": [
          {
            "id": "id",
            "function": {
              "arguments": "arguments",
              "name": "name"
            },
            "type": "function"
          }
        ]
      },
      "logprobs": {
        "content": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ],
        "refusal": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "chat.completion",
  "service_tier": "auto",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}