Generate OpenAI chat completion from messages
Generates a chat completion from an OpenAI-style messages array.
Use this endpoint for standard messages-based chat inference; use /v5/completions instead when you
have a raw text prompt rather than messages, /v5/responses for the OpenAI Responses API contract,
and /v5/inference when the payload does not follow any OpenAI schema. The request accepts the OpenAI
Chat Completions parameters (extra fields are allowed and forwarded), and the model is selected from
model given as vendor/name; most vendors are served through the litellm proxy gateway, while
OpenAI may use a native gateway. When stream is set the response is delivered as server-sent events
of chat.completion.chunk; otherwise a single chat.completion object is returned. Token usage is
recorded for the account, and for streaming responses it is read from the final chunk.
Parameters
Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Deprecated in favor of tool_choice. Controls which function is called by the model.
Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Developer-defined tags and values used for filtering completions in the dashboard.
Output types that you would like the model to generate for this request.
Static predicted output content, such as the content of a text file being regenerated.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
For o1 models only. Constrains effort on reasoning. Values: low, medium, high.
An object specifying the format that the model must output.
If specified, system will attempt to sample deterministically for repeated requests with same seed.
Options for streaming response. Only set this when stream is true.
What sampling temperature to use. Higher values make output more random, lower more focused.
A list of tools the model may call. Currently, only functions are supported. Max 128 functions.
Number of most likely tokens to return at each position, with associated log probability.
Generate OpenAI chat completion from messages
import os
from scale_gp_beta import SGPClient
client = SGPClient(
api_key=os.environ.get("SGP_API_KEY"), # This is the default and can be omitted
)
for completion in client.chat.completions.create(
messages=[{
"foo": "bar"
}],
model="model",
):
print(completion){
"id": "id",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"role": "assistant",
"annotations": [
{
"type": "url_citation",
"url_citation": {
"end_index": 0,
"start_index": 0,
"title": "title",
"url": "url"
}
}
],
"audio": {
"id": "id",
"data": "data",
"expires_at": 0,
"transcript": "transcript"
},
"content": "content",
"function_call": {
"arguments": "arguments",
"name": "name"
},
"refusal": "refusal",
"tool_calls": [
{
"id": "id",
"function": {
"arguments": "arguments",
"name": "name"
},
"type": "function"
}
]
},
"logprobs": {
"content": [
{
"token": "token",
"logprob": 0,
"top_logprobs": [
{
"token": "token",
"logprob": 0,
"bytes": [
0
]
}
],
"bytes": [
0
]
}
],
"refusal": [
{
"token": "token",
"logprob": 0,
"top_logprobs": [
{
"token": "token",
"logprob": 0,
"bytes": [
0
]
}
],
"bytes": [
0
]
}
]
}
}
],
"created": 0,
"model": "model",
"object": "chat.completion",
"service_tier": "auto",
"system_fingerprint": "system_fingerprint",
"usage": {
"completion_tokens": 0,
"prompt_tokens": 0,
"total_tokens": 0,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 0,
"rejected_prediction_tokens": 0
},
"prompt_tokens_details": {
"audio_tokens": 0,
"cached_tokens": 0
}
}
}Returns Examples
{
"id": "id",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"role": "assistant",
"annotations": [
{
"type": "url_citation",
"url_citation": {
"end_index": 0,
"start_index": 0,
"title": "title",
"url": "url"
}
}
],
"audio": {
"id": "id",
"data": "data",
"expires_at": 0,
"transcript": "transcript"
},
"content": "content",
"function_call": {
"arguments": "arguments",
"name": "name"
},
"refusal": "refusal",
"tool_calls": [
{
"id": "id",
"function": {
"arguments": "arguments",
"name": "name"
},
"type": "function"
}
]
},
"logprobs": {
"content": [
{
"token": "token",
"logprob": 0,
"top_logprobs": [
{
"token": "token",
"logprob": 0,
"bytes": [
0
]
}
],
"bytes": [
0
]
}
],
"refusal": [
{
"token": "token",
"logprob": 0,
"top_logprobs": [
{
"token": "token",
"logprob": 0,
"bytes": [
0
]
}
],
"bytes": [
0
]
}
]
}
}
],
"created": 0,
"model": "model",
"object": "chat.completion",
"service_tier": "auto",
"system_fingerprint": "system_fingerprint",
"usage": {
"completion_tokens": 0,
"prompt_tokens": 0,
"total_tokens": 0,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 0,
"rejected_prediction_tokens": 0
},
"prompt_tokens_details": {
"audio_tokens": 0,
"cached_tokens": 0
}
}
}