Skip to content

Generate OpenAI chat completion from messages

POST/v5/chat/completions

Generates a chat completion from an OpenAI-style messages array.

Use this endpoint for standard messages-based chat inference; use /v5/completions instead when you have a raw text prompt rather than messages, /v5/responses for the OpenAI Responses API contract, and /v5/inference when the payload does not follow any OpenAI schema. The request accepts the OpenAI Chat Completions parameters (extra fields are allowed and forwarded), and the model is selected from model given as vendor/name; most vendors are served through the litellm proxy gateway, while OpenAI may use a native gateway. When stream is set the response is delivered as server-sent events of chat.completion.chunk; otherwise a single chat.completion object is returned. Token usage is recorded for the account, and for streaming responses it is read from the final chunk.

Header ParametersExpand Collapse
"x-openai-api-key": optional string
Body ParametersJSONExpand Collapse
messages: array of map[unknown]

openai standard message format

model: string

model specified as model_vendor/model, for example openai/gpt-4o

audio: optional map[unknown]

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

frequency_penalty: optional number

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

maximum2
minimum-2
function_call: optional map[unknown]

Deprecated in favor of tool_choice. Controls which function is called by the model.

functions: optional array of map[unknown]

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

logit_bias: optional map[number]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

logprobs: optional boolean

Whether to return log probabilities of the output tokens or not.

max_completion_tokens: optional number

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

max_tokens: optional number

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

metadata: optional map[string]

Developer-defined tags and values used for filtering completions in the dashboard.

modalities: optional array of string

Output types that you would like the model to generate for this request.

n: optional number

How many chat completion choices to generate for each input message.

parallel_tool_calls: optional boolean

Whether to enable parallel function calling during tool use.

prediction: optional map[unknown]

Static predicted output content, such as the content of a text file being regenerated.

presence_penalty: optional number

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

maximum2
minimum-2
reasoning_effort: optional string

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format: optional map[unknown]

An object specifying the format that the model must output.

seed: optional number

If specified, system will attempt to sample deterministically for repeated requests with same seed.

stop: optional string or array of string

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
string
array of string
store: optional boolean

Whether to store the output for use in model distillation or evals products.

stream: optional boolean

If true, partial message deltas will be sent as server-sent events.

stream_options: optional map[unknown]

Options for streaming response. Only set this when stream is true.

temperature: optional number

What sampling temperature to use. Higher values make output more random, lower more focused.

maximum2
minimum0
tool_choice: optional string or map[unknown]

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
string
map[unknown]
tools: optional array of map[unknown]

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

top_k: optional number

Only sample from the top K options for each subsequent token

top_logprobs: optional number

Number of most likely tokens to return at each position, with associated log probability.

maximum20
minimum0
top_p: optional number

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

maximum1
minimum0
ReturnsExpand Collapse
ChatCompletion object { id, choices, created, 5 more }
id: string
choices: array of object { finish_reason, index, message, logprobs }
finish_reason: "stop" or "length" or "tool_calls" or 2 more
One of the following:
"stop"
"length"
"tool_calls"
"content_filter"
"function_call"
index: number
message: object { role, annotations, audio, 4 more }

A chat completion message generated by the model.

role: "assistant"
annotations: optional array of object { type, url_citation }
type: "url_citation"
url_citation: object { end_index, start_index, title, url }

A URL citation when using web search.

end_index: number
start_index: number
title: string
url: string
audio: optional object { id, data, expires_at, transcript }

If the audio output modality is requested, this object contains data about the audio response from the model. Learn more.

id: string
data: string
expires_at: number
transcript: string
content: optional string
function_call: optional object { arguments, name }

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

arguments: string
name: string
refusal: optional string
tool_calls: optional array of object { id, function, type } or object { id, custom, type }
One of the following:
ChatCompletionMessageFunctionToolCall object { id, function, type }

A call to a function tool created by the model.

id: string
function: object { arguments, name }

The function that the model called.

arguments: string
name: string
type: "function"
ChatCompletionMessageCustomToolCall object { id, custom, type }

A call to a custom tool created by the model.

id: string
custom: object { input, name }

The custom tool that the model called.

input: string
name: string
type: "custom"
logprobs: optional ChoiceLogprobs { content, refusal }

Log probability information for the choice.

content: optional array of ChatCompletionTokenLogprob { token, logprob, top_logprobs, bytes }
token: string
logprob: number
top_logprobs: array of object { token, logprob, bytes }
token: string
logprob: number
bytes: optional array of number
bytes: optional array of number
refusal: optional array of ChatCompletionTokenLogprob { token, logprob, top_logprobs, bytes }
token: string
logprob: number
top_logprobs: array of object { token, logprob, bytes }
token: string
logprob: number
bytes: optional array of number
bytes: optional array of number
created: number
model: string
object: optional "chat.completion"
service_tier: optional "auto" or "default" or "flex" or 2 more
One of the following:
"auto"
"default"
"flex"
"scale"
"priority"
system_fingerprint: optional string
usage: optional CompletionUsage { completion_tokens, prompt_tokens, total_tokens, 2 more }

Usage statistics for the completion request.

completion_tokens: number
prompt_tokens: number
total_tokens: number
completion_tokens_details: optional object { accepted_prediction_tokens, audio_tokens, reasoning_tokens, rejected_prediction_tokens }

Breakdown of tokens used in a completion.

accepted_prediction_tokens: optional number
audio_tokens: optional number
reasoning_tokens: optional number
rejected_prediction_tokens: optional number
prompt_tokens_details: optional object { audio_tokens, cached_tokens }

Breakdown of tokens used in the prompt.

audio_tokens: optional number
cached_tokens: optional number
ChatCompletionChunk object { id, choices, created, 5 more }
id: string
choices: array of object { delta, index, finish_reason, logprobs }
delta: object { content, function_call, refusal, 2 more }

A chat completion delta generated by streamed model responses.

content: optional string
function_call: optional object { arguments, name }

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

arguments: optional string
name: optional string
refusal: optional string
role: optional "developer" or "system" or "user" or 2 more
One of the following:
"developer"
"system"
"user"
"assistant"
"tool"
tool_calls: optional array of object { index, id, function, type }
index: number
id: optional string
function: optional object { arguments, name }
arguments: optional string
name: optional string
type: optional "function"
index: number
finish_reason: optional "stop" or "length" or "tool_calls" or 2 more
One of the following:
"stop"
"length"
"tool_calls"
"content_filter"
"function_call"
logprobs: optional ChoiceLogprobs { content, refusal }

Log probability information for the choice.

content: optional array of ChatCompletionTokenLogprob { token, logprob, top_logprobs, bytes }
token: string
logprob: number
top_logprobs: array of object { token, logprob, bytes }
token: string
logprob: number
bytes: optional array of number
bytes: optional array of number
refusal: optional array of ChatCompletionTokenLogprob { token, logprob, top_logprobs, bytes }
token: string
logprob: number
top_logprobs: array of object { token, logprob, bytes }
token: string
logprob: number
bytes: optional array of number
bytes: optional array of number
created: number
model: string
object: optional "chat.completion.chunk"
service_tier: optional "auto" or "default" or "flex" or 2 more
One of the following:
"auto"
"default"
"flex"
"scale"
"priority"
system_fingerprint: optional string
usage: optional CompletionUsage { completion_tokens, prompt_tokens, total_tokens, 2 more }

Usage statistics for the completion request.

completion_tokens: number
prompt_tokens: number
total_tokens: number
completion_tokens_details: optional object { accepted_prediction_tokens, audio_tokens, reasoning_tokens, rejected_prediction_tokens }

Breakdown of tokens used in a completion.

accepted_prediction_tokens: optional number
audio_tokens: optional number
reasoning_tokens: optional number
rejected_prediction_tokens: optional number
prompt_tokens_details: optional object { audio_tokens, cached_tokens }

Breakdown of tokens used in the prompt.

audio_tokens: optional number
cached_tokens: optional number
ChatCompletionChunk object { id, choices, created, 5 more }
id: string
choices: array of object { delta, index, finish_reason, logprobs }
delta: object { content, function_call, refusal, 2 more }

A chat completion delta generated by streamed model responses.

content: optional string
function_call: optional object { arguments, name }

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

arguments: optional string
name: optional string
refusal: optional string
role: optional "developer" or "system" or "user" or 2 more
One of the following:
"developer"
"system"
"user"
"assistant"
"tool"
tool_calls: optional array of object { index, id, function, type }
index: number
id: optional string
function: optional object { arguments, name }
arguments: optional string
name: optional string
type: optional "function"
index: number
finish_reason: optional "stop" or "length" or "tool_calls" or 2 more
One of the following:
"stop"
"length"
"tool_calls"
"content_filter"
"function_call"
logprobs: optional ChoiceLogprobs { content, refusal }

Log probability information for the choice.

content: optional array of ChatCompletionTokenLogprob { token, logprob, top_logprobs, bytes }
token: string
logprob: number
top_logprobs: array of object { token, logprob, bytes }
token: string
logprob: number
bytes: optional array of number
bytes: optional array of number
refusal: optional array of ChatCompletionTokenLogprob { token, logprob, top_logprobs, bytes }
token: string
logprob: number
top_logprobs: array of object { token, logprob, bytes }
token: string
logprob: number
bytes: optional array of number
bytes: optional array of number
created: number
model: string
object: optional "chat.completion.chunk"
service_tier: optional "auto" or "default" or "flex" or 2 more
One of the following:
"auto"
"default"
"flex"
"scale"
"priority"
system_fingerprint: optional string
usage: optional CompletionUsage { completion_tokens, prompt_tokens, total_tokens, 2 more }

Usage statistics for the completion request.

completion_tokens: number
prompt_tokens: number
total_tokens: number
completion_tokens_details: optional object { accepted_prediction_tokens, audio_tokens, reasoning_tokens, rejected_prediction_tokens }

Breakdown of tokens used in a completion.

accepted_prediction_tokens: optional number
audio_tokens: optional number
reasoning_tokens: optional number
rejected_prediction_tokens: optional number
prompt_tokens_details: optional object { audio_tokens, cached_tokens }

Breakdown of tokens used in the prompt.

audio_tokens: optional number
cached_tokens: optional number

Generate OpenAI chat completion from messages

curl https://api.egp.scale.com/v5/chat/completions \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "messages": [
            {
              "foo": "bar"
            }
          ],
          "model": "model"
        }'
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "role": "assistant",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": {
              "end_index": 0,
              "start_index": 0,
              "title": "title",
              "url": "url"
            }
          }
        ],
        "audio": {
          "id": "id",
          "data": "data",
          "expires_at": 0,
          "transcript": "transcript"
        },
        "content": "content",
        "function_call": {
          "arguments": "arguments",
          "name": "name"
        },
        "refusal": "refusal",
        "tool_calls": [
          {
            "id": "id",
            "function": {
              "arguments": "arguments",
              "name": "name"
            },
            "type": "function"
          }
        ]
      },
      "logprobs": {
        "content": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ],
        "refusal": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "chat.completion",
  "service_tier": "auto",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}
Returns Examples
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "role": "assistant",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": {
              "end_index": 0,
              "start_index": 0,
              "title": "title",
              "url": "url"
            }
          }
        ],
        "audio": {
          "id": "id",
          "data": "data",
          "expires_at": 0,
          "transcript": "transcript"
        },
        "content": "content",
        "function_call": {
          "arguments": "arguments",
          "name": "name"
        },
        "refusal": "refusal",
        "tool_calls": [
          {
            "id": "id",
            "function": {
              "arguments": "arguments",
              "name": "name"
            },
            "type": "function"
          }
        ]
      },
      "logprobs": {
        "content": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ],
        "refusal": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "chat.completion",
  "service_tier": "auto",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}