Skip to content

Generate legacy text completion from prompt

completions.create(CompletionCreateParams**kwargs) -> Completion
POST/v5/completions

Generates a legacy text completion from a raw prompt (a string or list of strings).

Use this endpoint for non-chat, prompt-in/text-out inference using the OpenAI text-completion contract; use /v5/chat/completions when you have a structured messages array, /v5/responses for the OpenAI Responses API, and /v5/inference for payloads that follow no OpenAI schema. The model is selected from model given as vendor/name and routed to the matching per-vendor gateway. When stream is set the response is delivered as server-sent events; otherwise a single text_completion object is returned. Token usage is recorded for the account, read from the final chunk on streaming responses.

ParametersExpand Collapse
model: str

model specified as model_vendor/model, for example openai/gpt-4o

prompt: Union[str, Sequence[str]]

The prompt to generate completions for, encoded as a string

One of the following:
str
Sequence[str]
best_of: Optional[int]

Generates best_of completions server-side and returns the best one. Must be greater than n when used together.

echo: Optional[bool]

Echo back the prompt in addition to the completion

frequency_penalty: Optional[float]

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text.

logit_bias: Optional[Dict[str, int]]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

logprobs: Optional[int]

Include log probabilities of the most likely tokens. Maximum value is 5.

max_tokens: Optional[int]

The maximum number of tokens that can be generated in the completion.

n: Optional[int]

How many completions to generate for each prompt.

presence_penalty: Optional[float]

Number between -2.0 and 2.0. Positive values penalize new tokens based on their presence in the text so far.

seed: Optional[int]

If specified, attempts to generate deterministic samples. Determinism is not guaranteed.

stop: Optional[Union[str, Sequence[str]]]

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
str
Sequence[str]
stream: Optional[Literal[false]]

Whether to stream back partial progress. If set, tokens will be sent as data-only server-sent events.

stream_options: Optional[Dict[str, object]]

Options for streaming response. Only set this when stream is True.

suffix: Optional[str]

The suffix that comes after a completion of inserted text. Only supported for gpt-3.5-turbo-instruct.

temperature: Optional[float]

Sampling temperature between 0 and 2. Higher values make output more random, lower more focused.

top_p: Optional[float]

Alternative to temperature. Consider only tokens with top_p probability mass. Range 0-1.

user: Optional[str]

A unique identifier representing your end-user, which can help OpenAI monitor and detect abuse.

ReturnsExpand Collapse
class Completion: …
id: str
choices: List[Choice]
finish_reason: Literal["stop", "length", "content_filter"]
One of the following:
"stop"
"length"
"content_filter"
index: int
text: str
logprobs: Optional[ChoiceLogprobs]
text_offset: Optional[List[int]]
token_logprobs: Optional[List[float]]
tokens: Optional[List[str]]
top_logprobs: Optional[List[Dict[str, float]]]
created: int
model: str
object: Optional[Literal["text_completion"]]
system_fingerprint: Optional[str]
usage: Optional[CompletionUsage]

Usage statistics for the completion request.

completion_tokens: int
prompt_tokens: int
total_tokens: int
completion_tokens_details: Optional[CompletionTokensDetails]

Breakdown of tokens used in a completion.

accepted_prediction_tokens: Optional[int]
audio_tokens: Optional[int]
reasoning_tokens: Optional[int]
rejected_prediction_tokens: Optional[int]
prompt_tokens_details: Optional[PromptTokensDetails]

Breakdown of tokens used in the prompt.

audio_tokens: Optional[int]
cached_tokens: Optional[int]
class Completion: …
id: str
choices: List[Choice]
finish_reason: Literal["stop", "length", "content_filter"]
One of the following:
"stop"
"length"
"content_filter"
index: int
text: str
logprobs: Optional[ChoiceLogprobs]
text_offset: Optional[List[int]]
token_logprobs: Optional[List[float]]
tokens: Optional[List[str]]
top_logprobs: Optional[List[Dict[str, float]]]
created: int
model: str
object: Optional[Literal["text_completion"]]
system_fingerprint: Optional[str]
usage: Optional[CompletionUsage]

Usage statistics for the completion request.

completion_tokens: int
prompt_tokens: int
total_tokens: int
completion_tokens_details: Optional[CompletionTokensDetails]

Breakdown of tokens used in a completion.

accepted_prediction_tokens: Optional[int]
audio_tokens: Optional[int]
reasoning_tokens: Optional[int]
rejected_prediction_tokens: Optional[int]
prompt_tokens_details: Optional[PromptTokensDetails]

Breakdown of tokens used in the prompt.

audio_tokens: Optional[int]
cached_tokens: Optional[int]

Generate legacy text completion from prompt

import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
for completion in client.completions.create(
    model="model",
    prompt="string",
):
  print(completion)
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "text": "text",
      "logprobs": {
        "text_offset": [
          0
        ],
        "token_logprobs": [
          0
        ],
        "tokens": [
          "string"
        ],
        "top_logprobs": [
          {
            "foo": 0
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "text_completion",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}
Returns Examples
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "text": "text",
      "logprobs": {
        "text_offset": [
          0
        ],
        "token_logprobs": [
          0
        ],
        "tokens": [
          "string"
        ],
        "top_logprobs": [
          {
            "foo": 0
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "text_completion",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}