Skip to content

Generate OpenAI chat completion from messages

client.Chat.Completions.New(ctx, params) (*ChatCompletionNewResponseUnion, error)
POST/v5/chat/completions

Generates a chat completion from an OpenAI-style messages array.

Use this endpoint for standard messages-based chat inference; use /v5/completions instead when you have a raw text prompt rather than messages, /v5/responses for the OpenAI Responses API contract, and /v5/inference when the payload does not follow any OpenAI schema. The request accepts the OpenAI Chat Completions parameters (extra fields are allowed and forwarded), and the model is selected from model given as vendor/name; most vendors are served through the litellm proxy gateway, while OpenAI may use a native gateway. When stream is set the response is delivered as server-sent events of chat.completion.chunk; otherwise a single chat.completion object is returned. Token usage is recorded for the account, and for streaming responses it is read from the final chunk.

ParametersExpand Collapse
params ChatCompletionNewParams
Messages param.Field[[]map[string, any]]

Body param: openai standard message format

Model param.Field[string]

Body param: model specified as model_vendor/model, for example openai/gpt-4o

Audio param.Field[map[string, any]]Optional

Body param: Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

FrequencyPenalty param.Field[float64]Optional

Body param: Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

maximum2
minimum-2
FunctionCall param.Field[map[string, any]]Optional

Body param: Deprecated in favor of tool_choice. Controls which function is called by the model.

Functions param.Field[[]map[string, any]]Optional

Body param: Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

LogitBias param.Field[map[string, int64]]Optional

Body param: Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

Logprobs param.Field[bool]Optional

Body param: Whether to return log probabilities of the output tokens or not.

MaxCompletionTokens param.Field[int64]Optional

Body param: An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

MaxTokens param.Field[int64]Optional

Body param: Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

Metadata param.Field[map[string, string]]Optional

Body param: Developer-defined tags and values used for filtering completions in the dashboard.

Modalities param.Field[[]string]Optional

Body param: Output types that you would like the model to generate for this request.

N param.Field[int64]Optional

Body param: How many chat completion choices to generate for each input message.

ParallelToolCalls param.Field[bool]Optional

Body param: Whether to enable parallel function calling during tool use.

Prediction param.Field[map[string, any]]Optional

Body param: Static predicted output content, such as the content of a text file being regenerated.

PresencePenalty param.Field[float64]Optional

Body param: Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

maximum2
minimum-2
ReasoningEffort param.Field[string]Optional

Body param: For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

ResponseFormat param.Field[map[string, any]]Optional

Body param: An object specifying the format that the model must output.

Seed param.Field[int64]Optional

Body param: If specified, system will attempt to sample deterministically for repeated requests with same seed.

Stop param.Field[ChatCompletionNewParamsStopUnion]Optional

Body param: Up to 4 sequences where the API will stop generating further tokens.

string
[]string
Store param.Field[bool]Optional

Body param: Whether to store the output for use in model distillation or evals products.

StreamOptions param.Field[map[string, any]]Optional

Body param: Options for streaming response. Only set this when stream is true.

Temperature param.Field[float64]Optional

Body param: What sampling temperature to use. Higher values make output more random, lower more focused.

maximum2
minimum0
ToolChoice param.Field[ChatCompletionNewParamsToolChoiceUnion]Optional

Body param: Controls which tool is called by the model. Values: none, auto, required, or specific tool.

string
map[string, any]
Tools param.Field[[]map[string, any]]Optional

Body param: A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

TopK param.Field[int64]Optional

Body param: Only sample from the top K options for each subsequent token

TopLogprobs param.Field[int64]Optional

Body param: Number of most likely tokens to return at each position, with associated log probability.

maximum20
minimum0
TopP param.Field[float64]Optional

Body param: Alternative to temperature. Only tokens comprising top_p probability mass are considered.

maximum1
minimum0
XOpenAIAPIKey param.Field[string]Optional

Header param

ReturnsExpand Collapse
type ChatCompletionNewResponseUnion interface{…}
One of the following:
type ChatCompletion struct{…}
ID string
Choices []ChatCompletionChoice
FinishReason string
One of the following:
const ChatCompletionChoiceFinishReasonStop ChatCompletionChoiceFinishReason = "stop"
const ChatCompletionChoiceFinishReasonLength ChatCompletionChoiceFinishReason = "length"
const ChatCompletionChoiceFinishReasonToolCalls ChatCompletionChoiceFinishReason = "tool_calls"
const ChatCompletionChoiceFinishReasonContentFilter ChatCompletionChoiceFinishReason = "content_filter"
const ChatCompletionChoiceFinishReasonFunctionCall ChatCompletionChoiceFinishReason = "function_call"
Index int64
Message ChatCompletionChoiceMessage

A chat completion message generated by the model.

Role Assistant
Annotations []ChatCompletionChoiceMessageAnnotationOptional
Type URLCitation
URLCitation ChatCompletionChoiceMessageAnnotationURLCitation

A URL citation when using web search.

EndIndex int64
StartIndex int64
Title string
URL string
Audio ChatCompletionChoiceMessageAudioOptional

If the audio output modality is requested, this object contains data about the audio response from the model. Learn more.

ID string
Data string
ExpiresAt int64
Transcript string
Content stringOptional
FunctionCall ChatCompletionChoiceMessageFunctionCallOptional

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

Arguments string
Name string
Refusal stringOptional
ToolCalls []ChatCompletionChoiceMessageToolCallUnionOptional
One of the following:
type ChatCompletionChoiceMessageToolCallChatCompletionMessageFunctionToolCall struct{…}

A call to a function tool created by the model.

ID string
Function ChatCompletionChoiceMessageToolCallChatCompletionMessageFunctionToolCallFunction

The function that the model called.

Arguments string
Name string
Type Function
type ChatCompletionChoiceMessageToolCallChatCompletionMessageCustomToolCall struct{…}

A call to a custom tool created by the model.

ID string
Custom ChatCompletionChoiceMessageToolCallChatCompletionMessageCustomToolCallCustom

The custom tool that the model called.

Input string
Name string
Type Custom
Logprobs ChoiceLogprobsOptional

Log probability information for the choice.

Content []ChatCompletionTokenLogprobOptional
Token string
Logprob float64
TopLogprobs []ChatCompletionTokenLogprobTopLogprob
Token string
Logprob float64
Bytes []int64Optional
Bytes []int64Optional
Refusal []ChatCompletionTokenLogprobOptional
Token string
Logprob float64
TopLogprobs []ChatCompletionTokenLogprobTopLogprob
Token string
Logprob float64
Bytes []int64Optional
Bytes []int64Optional
Created int64
Model string
Object ChatCompletionObjectOptional
ServiceTier ChatCompletionServiceTierOptional
One of the following:
const ChatCompletionServiceTierAuto ChatCompletionServiceTier = "auto"
const ChatCompletionServiceTierDefault ChatCompletionServiceTier = "default"
const ChatCompletionServiceTierFlex ChatCompletionServiceTier = "flex"
const ChatCompletionServiceTierScale ChatCompletionServiceTier = "scale"
const ChatCompletionServiceTierPriority ChatCompletionServiceTier = "priority"
SystemFingerprint stringOptional
Usage CompletionUsageOptional

Usage statistics for the completion request.

CompletionTokens int64
PromptTokens int64
TotalTokens int64
CompletionTokensDetails CompletionUsageCompletionTokensDetailsOptional

Breakdown of tokens used in a completion.

AcceptedPredictionTokens int64Optional
AudioTokens int64Optional
ReasoningTokens int64Optional
RejectedPredictionTokens int64Optional
PromptTokensDetails CompletionUsagePromptTokensDetailsOptional

Breakdown of tokens used in the prompt.

AudioTokens int64Optional
CachedTokens int64Optional
type ChatCompletionChunk struct{…}
ID string
Choices []ChatCompletionChunkChoice
Delta ChatCompletionChunkChoiceDelta

A chat completion delta generated by streamed model responses.

Content stringOptional
FunctionCall ChatCompletionChunkChoiceDeltaFunctionCallOptional

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

Arguments stringOptional
Name stringOptional
Refusal stringOptional
Role stringOptional
One of the following:
const ChatCompletionChunkChoiceDeltaRoleDeveloper ChatCompletionChunkChoiceDeltaRole = "developer"
const ChatCompletionChunkChoiceDeltaRoleSystem ChatCompletionChunkChoiceDeltaRole = "system"
const ChatCompletionChunkChoiceDeltaRoleUser ChatCompletionChunkChoiceDeltaRole = "user"
const ChatCompletionChunkChoiceDeltaRoleAssistant ChatCompletionChunkChoiceDeltaRole = "assistant"
const ChatCompletionChunkChoiceDeltaRoleTool ChatCompletionChunkChoiceDeltaRole = "tool"
ToolCalls []ChatCompletionChunkChoiceDeltaToolCallOptional
Index int64
ID stringOptional
Function ChatCompletionChunkChoiceDeltaToolCallFunctionOptional
Arguments stringOptional
Name stringOptional
Type stringOptional
Index int64
FinishReason stringOptional
One of the following:
const ChatCompletionChunkChoiceFinishReasonStop ChatCompletionChunkChoiceFinishReason = "stop"
const ChatCompletionChunkChoiceFinishReasonLength ChatCompletionChunkChoiceFinishReason = "length"
const ChatCompletionChunkChoiceFinishReasonToolCalls ChatCompletionChunkChoiceFinishReason = "tool_calls"
const ChatCompletionChunkChoiceFinishReasonContentFilter ChatCompletionChunkChoiceFinishReason = "content_filter"
const ChatCompletionChunkChoiceFinishReasonFunctionCall ChatCompletionChunkChoiceFinishReason = "function_call"
Logprobs ChoiceLogprobsOptional

Log probability information for the choice.

Content []ChatCompletionTokenLogprobOptional
Token string
Logprob float64
TopLogprobs []ChatCompletionTokenLogprobTopLogprob
Token string
Logprob float64
Bytes []int64Optional
Bytes []int64Optional
Refusal []ChatCompletionTokenLogprobOptional
Token string
Logprob float64
TopLogprobs []ChatCompletionTokenLogprobTopLogprob
Token string
Logprob float64
Bytes []int64Optional
Bytes []int64Optional
Created int64
Model string
Object ChatCompletionChunkObjectOptional
ServiceTier ChatCompletionChunkServiceTierOptional
One of the following:
const ChatCompletionChunkServiceTierAuto ChatCompletionChunkServiceTier = "auto"
const ChatCompletionChunkServiceTierDefault ChatCompletionChunkServiceTier = "default"
const ChatCompletionChunkServiceTierFlex ChatCompletionChunkServiceTier = "flex"
const ChatCompletionChunkServiceTierScale ChatCompletionChunkServiceTier = "scale"
const ChatCompletionChunkServiceTierPriority ChatCompletionChunkServiceTier = "priority"
SystemFingerprint stringOptional
Usage CompletionUsageOptional

Usage statistics for the completion request.

CompletionTokens int64
PromptTokens int64
TotalTokens int64
CompletionTokensDetails CompletionUsageCompletionTokensDetailsOptional

Breakdown of tokens used in a completion.

AcceptedPredictionTokens int64Optional
AudioTokens int64Optional
ReasoningTokens int64Optional
RejectedPredictionTokens int64Optional
PromptTokensDetails CompletionUsagePromptTokensDetailsOptional

Breakdown of tokens used in the prompt.

AudioTokens int64Optional
CachedTokens int64Optional
type ChatCompletionChunk struct{…}
ID string
Choices []ChatCompletionChunkChoice
Delta ChatCompletionChunkChoiceDelta

A chat completion delta generated by streamed model responses.

Content stringOptional
FunctionCall ChatCompletionChunkChoiceDeltaFunctionCallOptional

Deprecated and replaced by tool_calls.

The name and arguments of a function that should be called, as generated by the model.

Arguments stringOptional
Name stringOptional
Refusal stringOptional
Role stringOptional
One of the following:
const ChatCompletionChunkChoiceDeltaRoleDeveloper ChatCompletionChunkChoiceDeltaRole = "developer"
const ChatCompletionChunkChoiceDeltaRoleSystem ChatCompletionChunkChoiceDeltaRole = "system"
const ChatCompletionChunkChoiceDeltaRoleUser ChatCompletionChunkChoiceDeltaRole = "user"
const ChatCompletionChunkChoiceDeltaRoleAssistant ChatCompletionChunkChoiceDeltaRole = "assistant"
const ChatCompletionChunkChoiceDeltaRoleTool ChatCompletionChunkChoiceDeltaRole = "tool"
ToolCalls []ChatCompletionChunkChoiceDeltaToolCallOptional
Index int64
ID stringOptional
Function ChatCompletionChunkChoiceDeltaToolCallFunctionOptional
Arguments stringOptional
Name stringOptional
Type stringOptional
Index int64
FinishReason stringOptional
One of the following:
const ChatCompletionChunkChoiceFinishReasonStop ChatCompletionChunkChoiceFinishReason = "stop"
const ChatCompletionChunkChoiceFinishReasonLength ChatCompletionChunkChoiceFinishReason = "length"
const ChatCompletionChunkChoiceFinishReasonToolCalls ChatCompletionChunkChoiceFinishReason = "tool_calls"
const ChatCompletionChunkChoiceFinishReasonContentFilter ChatCompletionChunkChoiceFinishReason = "content_filter"
const ChatCompletionChunkChoiceFinishReasonFunctionCall ChatCompletionChunkChoiceFinishReason = "function_call"
Logprobs ChoiceLogprobsOptional

Log probability information for the choice.

Content []ChatCompletionTokenLogprobOptional
Token string
Logprob float64
TopLogprobs []ChatCompletionTokenLogprobTopLogprob
Token string
Logprob float64
Bytes []int64Optional
Bytes []int64Optional
Refusal []ChatCompletionTokenLogprobOptional
Token string
Logprob float64
TopLogprobs []ChatCompletionTokenLogprobTopLogprob
Token string
Logprob float64
Bytes []int64Optional
Bytes []int64Optional
Created int64
Model string
Object ChatCompletionChunkObjectOptional
ServiceTier ChatCompletionChunkServiceTierOptional
One of the following:
const ChatCompletionChunkServiceTierAuto ChatCompletionChunkServiceTier = "auto"
const ChatCompletionChunkServiceTierDefault ChatCompletionChunkServiceTier = "default"
const ChatCompletionChunkServiceTierFlex ChatCompletionChunkServiceTier = "flex"
const ChatCompletionChunkServiceTierScale ChatCompletionChunkServiceTier = "scale"
const ChatCompletionChunkServiceTierPriority ChatCompletionChunkServiceTier = "priority"
SystemFingerprint stringOptional
Usage CompletionUsageOptional

Usage statistics for the completion request.

CompletionTokens int64
PromptTokens int64
TotalTokens int64
CompletionTokensDetails CompletionUsageCompletionTokensDetailsOptional

Breakdown of tokens used in a completion.

AcceptedPredictionTokens int64Optional
AudioTokens int64Optional
ReasoningTokens int64Optional
RejectedPredictionTokens int64Optional
PromptTokensDetails CompletionUsagePromptTokensDetailsOptional

Breakdown of tokens used in the prompt.

AudioTokens int64Optional
CachedTokens int64Optional

Generate OpenAI chat completion from messages

package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  completion, err := client.Chat.Completions.New(context.TODO(), sgpdev.ChatCompletionNewParams{
    Messages: []map[string]any{map[string]any{
    "foo": "bar",
    }},
    Model: "model",
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", completion)
}
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "role": "assistant",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": {
              "end_index": 0,
              "start_index": 0,
              "title": "title",
              "url": "url"
            }
          }
        ],
        "audio": {
          "id": "id",
          "data": "data",
          "expires_at": 0,
          "transcript": "transcript"
        },
        "content": "content",
        "function_call": {
          "arguments": "arguments",
          "name": "name"
        },
        "refusal": "refusal",
        "tool_calls": [
          {
            "id": "id",
            "function": {
              "arguments": "arguments",
              "name": "name"
            },
            "type": "function"
          }
        ]
      },
      "logprobs": {
        "content": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ],
        "refusal": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "chat.completion",
  "service_tier": "auto",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}
Returns Examples
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "role": "assistant",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": {
              "end_index": 0,
              "start_index": 0,
              "title": "title",
              "url": "url"
            }
          }
        ],
        "audio": {
          "id": "id",
          "data": "data",
          "expires_at": 0,
          "transcript": "transcript"
        },
        "content": "content",
        "function_call": {
          "arguments": "arguments",
          "name": "name"
        },
        "refusal": "refusal",
        "tool_calls": [
          {
            "id": "id",
            "function": {
              "arguments": "arguments",
              "name": "name"
            },
            "type": "function"
          }
        ]
      },
      "logprobs": {
        "content": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ],
        "refusal": [
          {
            "token": "token",
            "logprob": 0,
            "top_logprobs": [
              {
                "token": "token",
                "logprob": 0,
                "bytes": [
                  0
                ]
              }
            ],
            "bytes": [
              0
            ]
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "chat.completion",
  "service_tier": "auto",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}