## Generate OpenAI chat completion from messages

**post** `/v5/chat/completions`

Generates a chat completion from an OpenAI-style `messages` array.

Use this endpoint for standard messages-based chat inference; use /v5/completions instead when you
have a raw text `prompt` rather than messages, /v5/responses for the OpenAI Responses API contract,
and /v5/inference when the payload does not follow any OpenAI schema. The request accepts the OpenAI
Chat Completions parameters (extra fields are allowed and forwarded), and the model is selected from
`model` given as `vendor/name`; most vendors are served through the litellm proxy gateway, while
OpenAI may use a native gateway. When `stream` is set the response is delivered as server-sent events
of `chat.completion.chunk`; otherwise a single `chat.completion` object is returned. Token usage is
recorded for the account, and for streaming responses it is read from the final chunk.

### Header Parameters

- `"x-openai-api-key": optional string`

### Body Parameters

- `messages: array of map[unknown]`

  openai standard message format

- `model: string`

  model specified as `model_vendor/model`, for example `openai/gpt-4o`

- `audio: optional map[unknown]`

  Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

- `frequency_penalty: optional number`

  Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

- `function_call: optional map[unknown]`

  Deprecated in favor of tool_choice. Controls which function is called by the model.

- `functions: optional array of map[unknown]`

  Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

- `logit_bias: optional map[number]`

  Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

- `logprobs: optional boolean`

  Whether to return log probabilities of the output tokens or not.

- `max_completion_tokens: optional number`

  An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

- `max_tokens: optional number`

  Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

- `metadata: optional map[string]`

  Developer-defined tags and values used for filtering completions in the dashboard.

- `modalities: optional array of string`

  Output types that you would like the model to generate for this request.

- `n: optional number`

  How many chat completion choices to generate for each input message.

- `parallel_tool_calls: optional boolean`

  Whether to enable parallel function calling during tool use.

- `prediction: optional map[unknown]`

  Static predicted output content, such as the content of a text file being regenerated.

- `presence_penalty: optional number`

  Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

- `reasoning_effort: optional string`

  For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

- `response_format: optional map[unknown]`

  An object specifying the format that the model must output.

- `seed: optional number`

  If specified, system will attempt to sample deterministically for repeated requests with same seed.

- `stop: optional string or array of string`

  Up to 4 sequences where the API will stop generating further tokens.

  - `string`

  - `array of string`

- `store: optional boolean`

  Whether to store the output for use in model distillation or evals products.

- `stream: optional boolean`

  If true, partial message deltas will be sent as server-sent events.

- `stream_options: optional map[unknown]`

  Options for streaming response. Only set this when stream is true.

- `temperature: optional number`

  What sampling temperature to use. Higher values make output more random, lower more focused.

- `tool_choice: optional string or map[unknown]`

  Controls which tool is called by the model. Values: none, auto, required, or specific tool.

  - `string`

  - `map[unknown]`

- `tools: optional array of map[unknown]`

  A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

- `top_k: optional number`

  Only sample from the top K options for each subsequent token

- `top_logprobs: optional number`

  Number of most likely tokens to return at each position, with associated log probability.

- `top_p: optional number`

  Alternative to temperature. Only tokens comprising top_p probability mass are considered.

### Returns

- `ChatCompletion object { id, choices, created, 5 more }`

  - `id: string`

  - `choices: array of object { finish_reason, index, message, logprobs }`

    - `finish_reason: "stop" or "length" or "tool_calls" or 2 more`

      - `"stop"`

      - `"length"`

      - `"tool_calls"`

      - `"content_filter"`

      - `"function_call"`

    - `index: number`

    - `message: object { role, annotations, audio, 4 more }`

      A chat completion message generated by the model.

      - `role: "assistant"`

        - `"assistant"`

      - `annotations: optional array of object { type, url_citation }`

        - `type: "url_citation"`

          - `"url_citation"`

        - `url_citation: object { end_index, start_index, title, url }`

          A URL citation when using web search.

          - `end_index: number`

          - `start_index: number`

          - `title: string`

          - `url: string`

      - `audio: optional object { id, data, expires_at, transcript }`

        If the audio output modality is requested, this object contains data
        about the audio response from the model. [Learn more](https://platform.openai.com/docs/guides/audio).

        - `id: string`

        - `data: string`

        - `expires_at: number`

        - `transcript: string`

      - `content: optional string`

      - `function_call: optional object { arguments, name }`

        Deprecated and replaced by `tool_calls`.

        The name and arguments of a function that should be called, as generated by the model.

        - `arguments: string`

        - `name: string`

      - `refusal: optional string`

      - `tool_calls: optional array of object { id, function, type }  or object { id, custom, type }`

        - `ChatCompletionMessageFunctionToolCall object { id, function, type }`

          A call to a function tool created by the model.

          - `id: string`

          - `function: object { arguments, name }`

            The function that the model called.

            - `arguments: string`

            - `name: string`

          - `type: "function"`

            - `"function"`

        - `ChatCompletionMessageCustomToolCall object { id, custom, type }`

          A call to a custom tool created by the model.

          - `id: string`

          - `custom: object { input, name }`

            The custom tool that the model called.

            - `input: string`

            - `name: string`

          - `type: "custom"`

            - `"custom"`

    - `logprobs: optional ChoiceLogprobs`

      Log probability information for the choice.

      - `content: optional array of ChatCompletionTokenLogprob`

        - `token: string`

        - `logprob: number`

        - `top_logprobs: array of object { token, logprob, bytes }`

          - `token: string`

          - `logprob: number`

          - `bytes: optional array of number`

        - `bytes: optional array of number`

      - `refusal: optional array of ChatCompletionTokenLogprob`

        - `token: string`

        - `logprob: number`

        - `top_logprobs: array of object { token, logprob, bytes }`

        - `bytes: optional array of number`

  - `created: number`

  - `model: string`

  - `object: optional "chat.completion"`

    - `"chat.completion"`

  - `service_tier: optional "auto" or "default" or "flex" or 2 more`

    - `"auto"`

    - `"default"`

    - `"flex"`

    - `"scale"`

    - `"priority"`

  - `system_fingerprint: optional string`

  - `usage: optional CompletionUsage`

    Usage statistics for the completion request.

    - `completion_tokens: number`

    - `prompt_tokens: number`

    - `total_tokens: number`

    - `completion_tokens_details: optional object { accepted_prediction_tokens, audio_tokens, reasoning_tokens, rejected_prediction_tokens }`

      Breakdown of tokens used in a completion.

      - `accepted_prediction_tokens: optional number`

      - `audio_tokens: optional number`

      - `reasoning_tokens: optional number`

      - `rejected_prediction_tokens: optional number`

    - `prompt_tokens_details: optional object { audio_tokens, cached_tokens }`

      Breakdown of tokens used in the prompt.

      - `audio_tokens: optional number`

      - `cached_tokens: optional number`

- `ChatCompletionChunk object { id, choices, created, 5 more }`

  - `id: string`

  - `choices: array of object { delta, index, finish_reason, logprobs }`

    - `delta: object { content, function_call, refusal, 2 more }`

      A chat completion delta generated by streamed model responses.

      - `content: optional string`

      - `function_call: optional object { arguments, name }`

        Deprecated and replaced by `tool_calls`.

        The name and arguments of a function that should be called, as generated by the model.

        - `arguments: optional string`

        - `name: optional string`

      - `refusal: optional string`

      - `role: optional "developer" or "system" or "user" or 2 more`

        - `"developer"`

        - `"system"`

        - `"user"`

        - `"assistant"`

        - `"tool"`

      - `tool_calls: optional array of object { index, id, function, type }`

        - `index: number`

        - `id: optional string`

        - `function: optional object { arguments, name }`

          - `arguments: optional string`

          - `name: optional string`

        - `type: optional "function"`

          - `"function"`

    - `index: number`

    - `finish_reason: optional "stop" or "length" or "tool_calls" or 2 more`

      - `"stop"`

      - `"length"`

      - `"tool_calls"`

      - `"content_filter"`

      - `"function_call"`

    - `logprobs: optional ChoiceLogprobs`

      Log probability information for the choice.

  - `created: number`

  - `model: string`

  - `object: optional "chat.completion.chunk"`

    - `"chat.completion.chunk"`

  - `service_tier: optional "auto" or "default" or "flex" or 2 more`

    - `"auto"`

    - `"default"`

    - `"flex"`

    - `"scale"`

    - `"priority"`

  - `system_fingerprint: optional string`

  - `usage: optional CompletionUsage`

    Usage statistics for the completion request.

### Example

```http
curl https://api.egp.scale.com/v5/chat/completions \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "messages": [
            {
              "foo": "bar"
            }
          ],
          "model": "model"
        }'
```

#### Response

```json
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "role": "assistant",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": {
              "end_index": 0,
              "start_index": 0,
              "title": "title",
              "url": "url"
            }
          }
        ],
        "audio": {
          "id": "id",
          "data": "data",
          "expires_at": 0,
          "transcript": "transcript"
        },
        "content": "content",
        "function_call": {
          "arguments": "arguments",
          "name": "name"
        },
        "refusal": "refusal",
        "tool_calls": [
          {
            "id": "id",
            "function": {
              "arguments": "arguments",
              "name": "name"
            },
            "type": "function"
          }
        ]
      },
      "logprobs": {
        "content": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ],
        "refusal": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "chat.completion",
  "service_tier": "auto",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}
```
