Skip to content

Generate legacy text completion from prompt

POST/v5/completions

Generates a legacy text completion from a raw prompt (a string or list of strings).

Use this endpoint for non-chat, prompt-in/text-out inference using the OpenAI text-completion contract; use /v5/chat/completions when you have a structured messages array, /v5/responses for the OpenAI Responses API, and /v5/inference for payloads that follow no OpenAI schema. The model is selected from model given as vendor/name and routed to the matching per-vendor gateway. When stream is set the response is delivered as server-sent events; otherwise a single text_completion object is returned. Token usage is recorded for the account, read from the final chunk on streaming responses.

Body ParametersJSONExpand Collapse
model: string

model specified as model_vendor/model, for example openai/gpt-4o

prompt: string or array of string

The prompt to generate completions for, encoded as a string

One of the following:
string
array of string
best_of: optional number

Generates best_of completions server-side and returns the best one. Must be greater than n when used together.

echo: optional boolean

Echo back the prompt in addition to the completion

frequency_penalty: optional number

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text.

logit_bias: optional map[number]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

logprobs: optional number

Include log probabilities of the most likely tokens. Maximum value is 5.

max_tokens: optional number

The maximum number of tokens that can be generated in the completion.

n: optional number

How many completions to generate for each prompt.

presence_penalty: optional number

Number between -2.0 and 2.0. Positive values penalize new tokens based on their presence in the text so far.

seed: optional number

If specified, attempts to generate deterministic samples. Determinism is not guaranteed.

stop: optional string or array of string

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
string
array of string
stream: optional boolean

Whether to stream back partial progress. If set, tokens will be sent as data-only server-sent events.

stream_options: optional map[unknown]

Options for streaming response. Only set this when stream is True.

suffix: optional string

The suffix that comes after a completion of inserted text. Only supported for gpt-3.5-turbo-instruct.

temperature: optional number

Sampling temperature between 0 and 2. Higher values make output more random, lower more focused.

top_p: optional number

Alternative to temperature. Consider only tokens with top_p probability mass. Range 0-1.

user: optional string

A unique identifier representing your end-user, which can help OpenAI monitor and detect abuse.

ReturnsExpand Collapse
Completion object { id, choices, created, 4 more }
id: string
choices: array of object { finish_reason, index, text, logprobs }
finish_reason: "stop" or "length" or "content_filter"
One of the following:
"stop"
"length"
"content_filter"
index: number
text: string
logprobs: optional object { text_offset, token_logprobs, tokens, top_logprobs }
text_offset: optional array of number
token_logprobs: optional array of number
tokens: optional array of string
top_logprobs: optional array of map[number]
created: number
model: string
object: optional "text_completion"
system_fingerprint: optional string
usage: optional CompletionUsage { completion_tokens, prompt_tokens, total_tokens, 2 more }

Usage statistics for the completion request.

completion_tokens: number
prompt_tokens: number
total_tokens: number
completion_tokens_details: optional object { accepted_prediction_tokens, audio_tokens, reasoning_tokens, rejected_prediction_tokens }

Breakdown of tokens used in a completion.

accepted_prediction_tokens: optional number
audio_tokens: optional number
reasoning_tokens: optional number
rejected_prediction_tokens: optional number
prompt_tokens_details: optional object { audio_tokens, cached_tokens }

Breakdown of tokens used in the prompt.

audio_tokens: optional number
cached_tokens: optional number
Completion object { id, choices, created, 4 more }
id: string
choices: array of object { finish_reason, index, text, logprobs }
finish_reason: "stop" or "length" or "content_filter"
One of the following:
"stop"
"length"
"content_filter"
index: number
text: string
logprobs: optional object { text_offset, token_logprobs, tokens, top_logprobs }
text_offset: optional array of number
token_logprobs: optional array of number
tokens: optional array of string
top_logprobs: optional array of map[number]
created: number
model: string
object: optional "text_completion"
system_fingerprint: optional string
usage: optional CompletionUsage { completion_tokens, prompt_tokens, total_tokens, 2 more }

Usage statistics for the completion request.

completion_tokens: number
prompt_tokens: number
total_tokens: number
completion_tokens_details: optional object { accepted_prediction_tokens, audio_tokens, reasoning_tokens, rejected_prediction_tokens }

Breakdown of tokens used in a completion.

accepted_prediction_tokens: optional number
audio_tokens: optional number
reasoning_tokens: optional number
rejected_prediction_tokens: optional number
prompt_tokens_details: optional object { audio_tokens, cached_tokens }

Breakdown of tokens used in the prompt.

audio_tokens: optional number
cached_tokens: optional number

Generate legacy text completion from prompt

curl https://api.egp.scale.com/v5/completions \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "model": "model",
          "prompt": "string"
        }'
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "text": "text",
      "logprobs": {
        "text_offset": [
          0
        ],
        "token_logprobs": [
          0
        ],
        "tokens": [
          "string"
        ],
        "top_logprobs": [
          {
            "foo": 0
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "text_completion",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}
Returns Examples
{
  "id": "id",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "text": "text",
      "logprobs": {
        "text_offset": [
          0
        ],
        "token_logprobs": [
          0
        ],
        "tokens": [
          "string"
        ],
        "top_logprobs": [
          {
            "foo": 0
          }
        ]
      }
    }
  ],
  "created": 0,
  "model": "model",
  "object": "text_completion",
  "system_fingerprint": "system_fingerprint",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    }
  }
}