Skip to content

Evaluations

Create Evaluation
client.evaluations.create(EvaluationCreateParams { evaluation } params, RequestOptionsoptions?): Evaluation { id, created_at, created_by, 12 more }
POST/v5/evaluations
List Evaluations
client.evaluations.list(EvaluationListParams { ending_before, include_archived, limit, 6 more } query?, RequestOptionsoptions?): CursorPage<Evaluation { id, created_at, created_by, 12 more } >
GET/v5/evaluations
Get Evaluation
client.evaluations.retrieve(stringevaluationID, EvaluationRetrieveParams { include_archived, views } query?, RequestOptionsoptions?): Evaluation { id, created_at, created_by, 12 more }
GET/v5/evaluations/{evaluation_id}
Archive Evaluation
client.evaluations.archive(stringevaluationID, RequestOptionsoptions?): Evaluation { id, created_at, created_by, 12 more }
DELETE/v5/evaluations/{evaluation_id}
Update or Restore Evaluation
client.evaluations.update(stringevaluationID, EvaluationUpdateParams { evaluation } params, RequestOptionsoptions?): Evaluation { id, created_at, created_by, 12 more }
PATCH/v5/evaluations/{evaluation_id}
Get Evaluation Data Schema
client.evaluations.retrieveSchema(stringevaluationID, EvaluationRetrieveSchemaParams { include_archived } query?, RequestOptionsoptions?): EvaluationSchemaResponse { evaluation_id, fields, total_items, 3 more }
GET/v5/evaluations/{evaluation_id}/schema
Filter Evaluations
client.evaluations.filter(EvaluationFilterParams { filters, ending_before, include_archived, 5 more } params, RequestOptionsoptions?): CursorPage<Evaluation { id, created_at, created_by, 12 more } >
POST/v5/evaluations/filter
Get Evaluation Taxonomy
client.evaluations.retrieveTaxonomy(stringevaluationID, RequestOptionsoptions?): EvaluationRetrieveTaxonomyResponse
GET/v5/evaluations/{evaluation_id}/taxonomy
ModelsExpand Collapse
AndEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "and"
AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: Array<string>
data_fields?: Array<string>
designated_to?: ApeAgent { config, agent_name } | IfAgent { config, agent_name } | TruthfulnessAgent { config, agent_name } | BaseAgent { config, agent_name }
One of the following:
ApeAgent { config, agent_name }
config: Config { model, temperature }
model?: string
temperature?: number
agent_name?: "APEAgent"
IfAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "IFAgent"
TruthfulnessAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "TruthfulnessAgent"
BaseAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "BaseAgent"
output_type?: "text" | "integer" | "float" | "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values?: Array<string | number | boolean>
One of the following:
string
number
boolean
rubric_id?: string
rubric_version?: number
EqEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "eq"
Evaluation { id, created_at, created_by, 12 more }
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity { id, type, object }

The identity that created the entity.

id: string
type: "user" | "service_account"
One of the following:
"user"
"service_account"
object?: "identity"
datasets: Array<Dataset { id, created_at, created_by, 6 more } > | null
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity { id, type, object }

The identity that created the entity.

id: string
type: "user" | "service_account"
One of the following:
"user"
"service_account"
object?: "identity"
current_version_num: number
name: string
tags: Array<string> | null

The tags associated with the entity

archived_at?: string

The date and time when the entity was archived in ISO format.

formatdate-time
description?: string
object?: "dataset"
name: string
status: "failed" | "completed" | "running"
One of the following:
"failed"
"completed"
"running"
tags: Array<string> | null

The tags associated with the entity

archived_at?: string

The date and time when the entity was archived in ISO format.

formatdate-time
description?: string
error_count?: number

Number of task errors across all items in this evaluation.

metadata?: Record<string, unknown>

Metadata key-value pairs for the evaluation

object?: "evaluation"
progress?: EvaluationTasksProgressSchema { items, workflows }

Progress of the evaluation’s underlying async job

items?: Items { failed, pending, successful, 2 more }
failed: number
pending: number
successful: number
total: number
failed_items?: Array<FailedItem>
item_id: string
error?: string
error_type?: string
workflows?: Workflows { completed, failed, pending, total }
completed: number
failed: number
pending: number
total: number
status_reason?: string

Reason for evaluation status

tasks?: Array<EvaluationTask>

Tasks executed during evaluation. Populated with optional task view.

One of the following:
ChatCompletionEvaluationTask { configuration, alias, task_type }
configuration: Configuration { messages, model, audio, 24 more }
messages: Array<Record<string, unknown>> | ItemLocator

openai standard message format

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
model: string

model specified as model_vendor/model, for example openai/gpt-4o

audio?: Record<string, unknown> | ItemLocator

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
Record<string, unknown>
ItemLocator = string
frequency_penalty?: number | ItemLocator

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
number
ItemLocator = string
function_call?: Record<string, unknown> | ItemLocator

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
Record<string, unknown>
ItemLocator = string
functions?: Array<Record<string, unknown>> | ItemLocator

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
logit_bias?: Record<string, number> | ItemLocator

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
Record<string, number>
ItemLocator = string
logprobs?: boolean | ItemLocator

Whether to return log probabilities of the output tokens or not.

One of the following:
boolean
ItemLocator = string
max_completion_tokens?: number | ItemLocator

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
number
ItemLocator = string
max_tokens?: number | ItemLocator

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
number
ItemLocator = string
metadata?: Record<string, string> | ItemLocator

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
Record<string, string>
ItemLocator = string
modalities?: Array<string> | ItemLocator

Output types that you would like the model to generate for this request.

One of the following:
Array<string>
ItemLocator = string
n?: number | ItemLocator

How many chat completion choices to generate for each input message.

One of the following:
number
ItemLocator = string
parallel_tool_calls?: boolean | ItemLocator

Whether to enable parallel function calling during tool use.

One of the following:
boolean
ItemLocator = string
prediction?: Record<string, unknown> | ItemLocator

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
Record<string, unknown>
ItemLocator = string
presence_penalty?: number | ItemLocator

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
number
ItemLocator = string
reasoning_effort?: string

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format?: Record<string, unknown> | ItemLocator

An object specifying the format that the model must output.

One of the following:
Record<string, unknown>
ItemLocator = string
seed?: number | ItemLocator

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
number
ItemLocator = string
stop?: string | Array<string>

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
string
Array<string>
store?: boolean | ItemLocator

Whether to store the output for use in model distillation or evals products.

One of the following:
boolean
ItemLocator = string
temperature?: number | ItemLocator

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
number
ItemLocator = string
tool_choice?: string | Record<string, unknown>

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
string
Record<string, unknown>
tools?: Array<Record<string, unknown>> | ItemLocator

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
top_k?: number | ItemLocator

Only sample from the top K options for each subsequent token

One of the following:
number
ItemLocator = string
top_logprobs?: number | ItemLocator

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
number
ItemLocator = string
top_p?: number | ItemLocator

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the chat_completion

task_type?: "chat_completion"
GenericInferenceEvaluationTask { configuration, alias, task_type }
configuration: Configuration { model, args, inference_configuration }
model: string

model specified as vendor/name (ex. openai/gpt-5)

args?: Record<string, unknown> | ItemLocator

Arguments passed into model

One of the following:
Record<string, unknown>
ItemLocator = string
inference_configuration?: LaunchInferenceConfiguration { num_retries, timeout_seconds } | ItemLocator

Vendor specific configuration

One of the following:
LaunchInferenceConfiguration { num_retries, timeout_seconds }
num_retries?: number
timeout_seconds?: number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the inference

task_type?: "inference"
ApplicationVariantV1EvaluationTask { configuration, alias, task_type }
configuration: Configuration { application_variant_id, inputs, history, 2 more }
application_variant_id: string
inputs: Record<string, unknown> | ItemLocator

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
Record<string, unknown>
ItemLocator = string
history?: Array<ApplicationRequestResponsePairArray> | ItemLocator

History of the application

One of the following:
Array<ApplicationRequestResponsePairArray>
request: string

Request inputs

response: string

Response outputs

session_data?: Record<string, unknown>

Session data corresponding to the request response pair

ItemLocator = string
operation_metadata?: Record<string, unknown> | ItemLocator

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
Record<string, unknown>
ItemLocator = string
overrides?: AgenticApplicationOverrides { concurrent, initial_state, partial_trace, 2 more } | Record<string, KnowledgeBaseNodeOverride> | ItemLocator

Optional overrides for the application

One of the following:
AgenticApplicationOverrides { concurrent, initial_state, partial_trace, 2 more }

Execution override options for agentic applications

concurrent?: boolean
initial_state?: InitialState { current_node, state }
current_node: string
state: Record<string, unknown>
partial_trace?: Array<PartialTrace>
duration_ms: number
node_id: string
operation_input: string
operation_output: string
operation_type: string
start_timestamp: string
workflow_id: string
operation_metadata?: Record<string, unknown>
return_span?: boolean
use_channels?: boolean
Record<string, KnowledgeBaseNodeOverride>
artifact_ids_filter?: Array<string>
artifact_name_regex?: Array<string>
type?: "knowledge_base_schema"
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the application_variant

task_type?: "application_variant"
AgentexOutputEvaluationTask { configuration, alias, task_type }
configuration: Configuration { agentex_agent_id, input_column, agent_task_params, 6 more }
agentex_agent_id: string

The ID of the Agentex agent to use

input_column: string | Record<string, unknown> | Array<unknown>

The dataset column to use as input for the agent

One of the following:
string
Record<string, unknown>
Array<unknown>
agent_task_params?: Record<string, unknown> | ItemLocator

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
Record<string, unknown>
ItemLocator = string
completion_mode?: "first_message" | "turn_quiescence"

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id?: string

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces?: boolean | ItemLocator

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
boolean
ItemLocator = string
input_mode?: "text" | "data"

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds?: number | ItemLocator

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
number
ItemLocator = string
timeout_seconds?: number | ItemLocator

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the agentex_output

task_type?: "agentex_output"
MetricEvaluationTask { configuration, alias, task_type }
configuration: BleuScorerConfigWithItemLocator { candidate, reference, type } | MeteorScorerConfigWithItemLocator { candidate, reference, type } | CosineSimilarityScorerConfigWithItemLocator { candidate, reference, type } | 4 more
One of the following:
BleuScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "bleu"
MeteorScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "meteor"
CosineSimilarityScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "cosine_similarity"
F1ScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "f1"
RougeScorer1ConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rouge1"
RougeScorer2ConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rouge2"
RougeScorerLConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rougeL"
alias?: string

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type?: "metric"
AutoEvaluationQuestionTask { configuration, alias, task_type }
configuration: Configuration { model, prompt, question_id }
model: string

model specified as model_vendor/model_name

prompt: string
question_id: string

question to be evaluated

alias?: string

Alias to title the results column. Defaults to the auto_evaluation_question

task_type?: "auto_evaluation.question"
AutoEvaluationGuidedDecodingEvaluationTask { configuration, alias, task_type }
configuration: AutoEvaluationStructuredOutputTaskRequestWithItemLocator { model, prompt, response_format, 3 more } | AutoEvaluationGuidedDecodingTaskRequestWithItemLocator { choices, model, prompt, 3 more } | AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
One of the following:
AutoEvaluationStructuredOutputTaskRequestWithItemLocator { model, prompt, response_format, 3 more }
model: string

model specified as model_vendor/model_name

prompt: string
response_format: Record<string, unknown>

JSON schema used for structuring the model response

inference_args?: Record<string, unknown>

Additional arguments to pass to the inference request

run_condition?: ConstEvaluationRunCondition { op, value } | VarEvaluationRunCondition { path, op } | EqEvaluationRunCondition { left, right, op } | 12 more
One of the following:
ConstEvaluationRunCondition { op, value }
op?: "const"
value?: string | number | boolean
One of the following:
string
number
boolean
VarEvaluationRunCondition { path, op }
path: string
op?: "var"
EqEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "eq"
NeEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "ne"
LtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lt"
LteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lte"
GtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gt"
GteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gte"
AndEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "and"
OrEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "or"
InEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "in"
NotInEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "not_in"
NotEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "not"
IsNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_null"
IsNotNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_not_null"
system_prompt?: string
AutoEvaluationGuidedDecodingTaskRequestWithItemLocator { choices, model, prompt, 3 more }
choices: Array<string>

Choices array cannot be empty

model: string

model specified as model_vendor/model_name

prompt: string
inference_args?: Record<string, unknown>

Additional arguments to pass to the inference request

run_condition?: ConstEvaluationRunCondition { op, value } | VarEvaluationRunCondition { path, op } | EqEvaluationRunCondition { left, right, op } | 12 more
One of the following:
ConstEvaluationRunCondition { op, value }
op?: "const"
value?: string | number | boolean
One of the following:
string
number
boolean
VarEvaluationRunCondition { path, op }
path: string
op?: "var"
EqEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "eq"
NeEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "ne"
LtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lt"
LteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lte"
GtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gt"
GteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gte"
AndEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "and"
OrEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "or"
InEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "in"
NotInEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "not_in"
NotEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "not"
IsNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_null"
IsNotNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_not_null"
system_prompt?: string
AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: Array<string>
data_fields?: Array<string>
designated_to?: ApeAgent { config, agent_name } | IfAgent { config, agent_name } | TruthfulnessAgent { config, agent_name } | BaseAgent { config, agent_name }
One of the following:
ApeAgent { config, agent_name }
config: Config { model, temperature }
model?: string
temperature?: number
agent_name?: "APEAgent"
IfAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "IFAgent"
TruthfulnessAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "TruthfulnessAgent"
BaseAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "BaseAgent"
output_type?: "text" | "integer" | "float" | "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values?: Array<string | number | boolean>
One of the following:
string
number
boolean
rubric_id?: string
rubric_version?: number
alias?: string

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type?: "auto_evaluation.guided_decoding"
AutoEvaluationAgentEvaluationTask { configuration, alias, task_type }
configuration: AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: Array<string>
data_fields?: Array<string>
designated_to?: ApeAgent { config, agent_name } | IfAgent { config, agent_name } | TruthfulnessAgent { config, agent_name } | BaseAgent { config, agent_name }
One of the following:
ApeAgent { config, agent_name }
config: Config { model, temperature }
model?: string
temperature?: number
agent_name?: "APEAgent"
IfAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "IFAgent"
TruthfulnessAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "TruthfulnessAgent"
BaseAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "BaseAgent"
output_type?: "text" | "integer" | "float" | "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values?: Array<string | number | boolean>
One of the following:
string
number
boolean
rubric_id?: string
rubric_version?: number
alias?: string

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type?: "auto_evaluation.agent"
ContributorEvaluationQuestionTask { configuration, alias, task_type }
configuration: Configuration { layout, question_id, prefill_from, 3 more }
layout: Container { children, direction }
children: Array<Container { children, direction } | Component { data, label } >

The children to be displayed within the container

One of the following:
Container = Container { children, direction }
Component { data, label }

A pointer to the data in each evaluation item to be displayed within the component

label?: string
direction?: "row" | "column"

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: string
prefill_from?: string

Dataset column to prefill contributor question task result

minLength1
queue_id?: string

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required?: boolean

Whether the question is required to be answered

rubric_id?: string

ID of the rubric to use for scoring this evaluation question

alias?: string

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type?: "contributor_evaluation.question"
CustomFunctionEvaluationTask { configuration, alias, task_type }
configuration: Configuration { function_source, arg_mapping, config_args, outputs }

Configuration for a custom Python function evaluation task.

function_source: string

Python function source code

maxLength10000
arg_mapping?: Record<string, string>

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args?: Record<string, unknown>

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs?: Array<Output>

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: string

Dot path in the custom function return value to materialize.

minLength1
alias?: string

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias?: string

Alias to title the results column. Defaults to the function name.

task_type?: "custom_function"
EvaluationSchemaResponse { evaluation_id, fields, total_items, 3 more }

Schema information for an evaluation’s item data structure

evaluation_id: string

The ID of the evaluation

fields: Array<Field>

List of all discovered fields, ordered alphabetically by field_name

data_type: string

JSON type: ‘string’, ‘number’, ‘boolean’, ‘object’, ‘array’, or ‘null’

field_name: string

The flattened JSON key path (e.g., ‘metadata.category’)

item_count: number

Number of evaluation items containing this field

minimum0
source: "data" | "task_result_cache"

The source of the field: ‘data’ or ‘task_result_cache’

One of the following:
"data"
"task_result_cache"
object?: "field_schema"
total_items: number

Total number of evaluation items

minimum0
is_sampled?: boolean

Whether schema was computed from a sample of items (for large evaluations)

object?: "evaluation_schema"
sample_size?: number

Number of items sampled for schema inference, if applicable

minimum0
EvaluationTask = ChatCompletionEvaluationTask { configuration, alias, task_type } | GenericInferenceEvaluationTask { configuration, alias, task_type } | ApplicationVariantV1EvaluationTask { configuration, alias, task_type } | 7 more
One of the following:
ChatCompletionEvaluationTask { configuration, alias, task_type }
configuration: Configuration { messages, model, audio, 24 more }
messages: Array<Record<string, unknown>> | ItemLocator

openai standard message format

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
model: string

model specified as model_vendor/model, for example openai/gpt-4o

audio?: Record<string, unknown> | ItemLocator

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
Record<string, unknown>
ItemLocator = string
frequency_penalty?: number | ItemLocator

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
number
ItemLocator = string
function_call?: Record<string, unknown> | ItemLocator

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
Record<string, unknown>
ItemLocator = string
functions?: Array<Record<string, unknown>> | ItemLocator

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
logit_bias?: Record<string, number> | ItemLocator

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
Record<string, number>
ItemLocator = string
logprobs?: boolean | ItemLocator

Whether to return log probabilities of the output tokens or not.

One of the following:
boolean
ItemLocator = string
max_completion_tokens?: number | ItemLocator

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
number
ItemLocator = string
max_tokens?: number | ItemLocator

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
number
ItemLocator = string
metadata?: Record<string, string> | ItemLocator

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
Record<string, string>
ItemLocator = string
modalities?: Array<string> | ItemLocator

Output types that you would like the model to generate for this request.

One of the following:
Array<string>
ItemLocator = string
n?: number | ItemLocator

How many chat completion choices to generate for each input message.

One of the following:
number
ItemLocator = string
parallel_tool_calls?: boolean | ItemLocator

Whether to enable parallel function calling during tool use.

One of the following:
boolean
ItemLocator = string
prediction?: Record<string, unknown> | ItemLocator

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
Record<string, unknown>
ItemLocator = string
presence_penalty?: number | ItemLocator

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
number
ItemLocator = string
reasoning_effort?: string

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format?: Record<string, unknown> | ItemLocator

An object specifying the format that the model must output.

One of the following:
Record<string, unknown>
ItemLocator = string
seed?: number | ItemLocator

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
number
ItemLocator = string
stop?: string | Array<string>

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
string
Array<string>
store?: boolean | ItemLocator

Whether to store the output for use in model distillation or evals products.

One of the following:
boolean
ItemLocator = string
temperature?: number | ItemLocator

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
number
ItemLocator = string
tool_choice?: string | Record<string, unknown>

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
string
Record<string, unknown>
tools?: Array<Record<string, unknown>> | ItemLocator

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
top_k?: number | ItemLocator

Only sample from the top K options for each subsequent token

One of the following:
number
ItemLocator = string
top_logprobs?: number | ItemLocator

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
number
ItemLocator = string
top_p?: number | ItemLocator

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the chat_completion

task_type?: "chat_completion"
GenericInferenceEvaluationTask { configuration, alias, task_type }
configuration: Configuration { model, args, inference_configuration }
model: string

model specified as vendor/name (ex. openai/gpt-5)

args?: Record<string, unknown> | ItemLocator

Arguments passed into model

One of the following:
Record<string, unknown>
ItemLocator = string
inference_configuration?: LaunchInferenceConfiguration { num_retries, timeout_seconds } | ItemLocator

Vendor specific configuration

One of the following:
LaunchInferenceConfiguration { num_retries, timeout_seconds }
num_retries?: number
timeout_seconds?: number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the inference

task_type?: "inference"
ApplicationVariantV1EvaluationTask { configuration, alias, task_type }
configuration: Configuration { application_variant_id, inputs, history, 2 more }
application_variant_id: string
inputs: Record<string, unknown> | ItemLocator

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
Record<string, unknown>
ItemLocator = string
history?: Array<ApplicationRequestResponsePairArray> | ItemLocator

History of the application

One of the following:
Array<ApplicationRequestResponsePairArray>
request: string

Request inputs

response: string

Response outputs

session_data?: Record<string, unknown>

Session data corresponding to the request response pair

ItemLocator = string
operation_metadata?: Record<string, unknown> | ItemLocator

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
Record<string, unknown>
ItemLocator = string
overrides?: AgenticApplicationOverrides { concurrent, initial_state, partial_trace, 2 more } | Record<string, KnowledgeBaseNodeOverride> | ItemLocator

Optional overrides for the application

One of the following:
AgenticApplicationOverrides { concurrent, initial_state, partial_trace, 2 more }

Execution override options for agentic applications

concurrent?: boolean
initial_state?: InitialState { current_node, state }
current_node: string
state: Record<string, unknown>
partial_trace?: Array<PartialTrace>
duration_ms: number
node_id: string
operation_input: string
operation_output: string
operation_type: string
start_timestamp: string
workflow_id: string
operation_metadata?: Record<string, unknown>
return_span?: boolean
use_channels?: boolean
Record<string, KnowledgeBaseNodeOverride>
artifact_ids_filter?: Array<string>
artifact_name_regex?: Array<string>
type?: "knowledge_base_schema"
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the application_variant

task_type?: "application_variant"
AgentexOutputEvaluationTask { configuration, alias, task_type }
configuration: Configuration { agentex_agent_id, input_column, agent_task_params, 6 more }
agentex_agent_id: string

The ID of the Agentex agent to use

input_column: string | Record<string, unknown> | Array<unknown>

The dataset column to use as input for the agent

One of the following:
string
Record<string, unknown>
Array<unknown>
agent_task_params?: Record<string, unknown> | ItemLocator

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
Record<string, unknown>
ItemLocator = string
completion_mode?: "first_message" | "turn_quiescence"

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id?: string

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces?: boolean | ItemLocator

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
boolean
ItemLocator = string
input_mode?: "text" | "data"

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds?: number | ItemLocator

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
number
ItemLocator = string
timeout_seconds?: number | ItemLocator

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the agentex_output

task_type?: "agentex_output"
MetricEvaluationTask { configuration, alias, task_type }
configuration: BleuScorerConfigWithItemLocator { candidate, reference, type } | MeteorScorerConfigWithItemLocator { candidate, reference, type } | CosineSimilarityScorerConfigWithItemLocator { candidate, reference, type } | 4 more
One of the following:
BleuScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "bleu"
MeteorScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "meteor"
CosineSimilarityScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "cosine_similarity"
F1ScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "f1"
RougeScorer1ConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rouge1"
RougeScorer2ConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rouge2"
RougeScorerLConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rougeL"
alias?: string

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type?: "metric"
AutoEvaluationQuestionTask { configuration, alias, task_type }
configuration: Configuration { model, prompt, question_id }
model: string

model specified as model_vendor/model_name

prompt: string
question_id: string

question to be evaluated

alias?: string

Alias to title the results column. Defaults to the auto_evaluation_question

task_type?: "auto_evaluation.question"
AutoEvaluationGuidedDecodingEvaluationTask { configuration, alias, task_type }
configuration: AutoEvaluationStructuredOutputTaskRequestWithItemLocator { model, prompt, response_format, 3 more } | AutoEvaluationGuidedDecodingTaskRequestWithItemLocator { choices, model, prompt, 3 more } | AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
One of the following:
AutoEvaluationStructuredOutputTaskRequestWithItemLocator { model, prompt, response_format, 3 more }
model: string

model specified as model_vendor/model_name

prompt: string
response_format: Record<string, unknown>

JSON schema used for structuring the model response

inference_args?: Record<string, unknown>

Additional arguments to pass to the inference request

run_condition?: ConstEvaluationRunCondition { op, value } | VarEvaluationRunCondition { path, op } | EqEvaluationRunCondition { left, right, op } | 12 more
One of the following:
ConstEvaluationRunCondition { op, value }
op?: "const"
value?: string | number | boolean
One of the following:
string
number
boolean
VarEvaluationRunCondition { path, op }
path: string
op?: "var"
EqEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "eq"
NeEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "ne"
LtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lt"
LteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lte"
GtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gt"
GteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gte"
AndEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "and"
OrEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "or"
InEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "in"
NotInEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "not_in"
NotEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "not"
IsNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_null"
IsNotNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_not_null"
system_prompt?: string
AutoEvaluationGuidedDecodingTaskRequestWithItemLocator { choices, model, prompt, 3 more }
choices: Array<string>

Choices array cannot be empty

model: string

model specified as model_vendor/model_name

prompt: string
inference_args?: Record<string, unknown>

Additional arguments to pass to the inference request

run_condition?: ConstEvaluationRunCondition { op, value } | VarEvaluationRunCondition { path, op } | EqEvaluationRunCondition { left, right, op } | 12 more
One of the following:
ConstEvaluationRunCondition { op, value }
op?: "const"
value?: string | number | boolean
One of the following:
string
number
boolean
VarEvaluationRunCondition { path, op }
path: string
op?: "var"
EqEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "eq"
NeEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "ne"
LtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lt"
LteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lte"
GtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gt"
GteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gte"
AndEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "and"
OrEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "or"
InEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "in"
NotInEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "not_in"
NotEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "not"
IsNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_null"
IsNotNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_not_null"
system_prompt?: string
AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: Array<string>
data_fields?: Array<string>
designated_to?: ApeAgent { config, agent_name } | IfAgent { config, agent_name } | TruthfulnessAgent { config, agent_name } | BaseAgent { config, agent_name }
One of the following:
ApeAgent { config, agent_name }
config: Config { model, temperature }
model?: string
temperature?: number
agent_name?: "APEAgent"
IfAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "IFAgent"
TruthfulnessAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "TruthfulnessAgent"
BaseAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "BaseAgent"
output_type?: "text" | "integer" | "float" | "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values?: Array<string | number | boolean>
One of the following:
string
number
boolean
rubric_id?: string
rubric_version?: number
alias?: string

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type?: "auto_evaluation.guided_decoding"
AutoEvaluationAgentEvaluationTask { configuration, alias, task_type }
configuration: AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: Array<string>
data_fields?: Array<string>
designated_to?: ApeAgent { config, agent_name } | IfAgent { config, agent_name } | TruthfulnessAgent { config, agent_name } | BaseAgent { config, agent_name }
One of the following:
ApeAgent { config, agent_name }
config: Config { model, temperature }
model?: string
temperature?: number
agent_name?: "APEAgent"
IfAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "IFAgent"
TruthfulnessAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "TruthfulnessAgent"
BaseAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "BaseAgent"
output_type?: "text" | "integer" | "float" | "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values?: Array<string | number | boolean>
One of the following:
string
number
boolean
rubric_id?: string
rubric_version?: number
alias?: string

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type?: "auto_evaluation.agent"
ContributorEvaluationQuestionTask { configuration, alias, task_type }
configuration: Configuration { layout, question_id, prefill_from, 3 more }
layout: Container { children, direction }
children: Array<Container { children, direction } | Component { data, label } >

The children to be displayed within the container

One of the following:
Container = Container { children, direction }
Component { data, label }

A pointer to the data in each evaluation item to be displayed within the component

label?: string
direction?: "row" | "column"

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: string
prefill_from?: string

Dataset column to prefill contributor question task result

minLength1
queue_id?: string

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required?: boolean

Whether the question is required to be answered

rubric_id?: string

ID of the rubric to use for scoring this evaluation question

alias?: string

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type?: "contributor_evaluation.question"
CustomFunctionEvaluationTask { configuration, alias, task_type }
configuration: Configuration { function_source, arg_mapping, config_args, outputs }

Configuration for a custom Python function evaluation task.

function_source: string

Python function source code

maxLength10000
arg_mapping?: Record<string, string>

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args?: Record<string, unknown>

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs?: Array<Output>

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: string

Dot path in the custom function return value to materialize.

minLength1
alias?: string

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias?: string

Alias to title the results column. Defaults to the function name.

task_type?: "custom_function"
EvaluationTasksProgressSchema { items, workflows }
items?: Items { failed, pending, successful, 2 more }
failed: number
pending: number
successful: number
total: number
failed_items?: Array<FailedItem>
item_id: string
error?: string
error_type?: string
workflows?: Workflows { completed, failed, pending, total }
completed: number
failed: number
pending: number
total: number
EvaluationViews = "tasks"
GtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gt"
GteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gte"
InEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "in"
IsNotNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_not_null"
IsNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_null"
ItemLocator = string
ItemLocatorTemplate = string
LtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lt"
LteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lte"
NeEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "ne"
NotEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "not"
NotInEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "not_in"
OrEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "or"
PaginatedListEvaluation { has_more, items, total, 2 more }
has_more: boolean

Whether there are more items left to be fetched.

items: Array<Evaluation { id, created_at, created_by, 12 more } >
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity { id, type, object }

The identity that created the entity.

id: string
type: "user" | "service_account"
One of the following:
"user"
"service_account"
object?: "identity"
datasets: Array<Dataset { id, created_at, created_by, 6 more } > | null
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity { id, type, object }

The identity that created the entity.

id: string
type: "user" | "service_account"
One of the following:
"user"
"service_account"
object?: "identity"
current_version_num: number
name: string
tags: Array<string> | null

The tags associated with the entity

archived_at?: string

The date and time when the entity was archived in ISO format.

formatdate-time
description?: string
object?: "dataset"
name: string
status: "failed" | "completed" | "running"
One of the following:
"failed"
"completed"
"running"
tags: Array<string> | null

The tags associated with the entity

archived_at?: string

The date and time when the entity was archived in ISO format.

formatdate-time
description?: string
error_count?: number

Number of task errors across all items in this evaluation.

metadata?: Record<string, unknown>

Metadata key-value pairs for the evaluation

object?: "evaluation"
progress?: EvaluationTasksProgressSchema { items, workflows }

Progress of the evaluation’s underlying async job

items?: Items { failed, pending, successful, 2 more }
failed: number
pending: number
successful: number
total: number
failed_items?: Array<FailedItem>
item_id: string
error?: string
error_type?: string
workflows?: Workflows { completed, failed, pending, total }
completed: number
failed: number
pending: number
total: number
status_reason?: string

Reason for evaluation status

tasks?: Array<EvaluationTask>

Tasks executed during evaluation. Populated with optional task view.

One of the following:
ChatCompletionEvaluationTask { configuration, alias, task_type }
configuration: Configuration { messages, model, audio, 24 more }
messages: Array<Record<string, unknown>> | ItemLocator

openai standard message format

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
model: string

model specified as model_vendor/model, for example openai/gpt-4o

audio?: Record<string, unknown> | ItemLocator

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
Record<string, unknown>
ItemLocator = string
frequency_penalty?: number | ItemLocator

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
number
ItemLocator = string
function_call?: Record<string, unknown> | ItemLocator

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
Record<string, unknown>
ItemLocator = string
functions?: Array<Record<string, unknown>> | ItemLocator

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
logit_bias?: Record<string, number> | ItemLocator

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
Record<string, number>
ItemLocator = string
logprobs?: boolean | ItemLocator

Whether to return log probabilities of the output tokens or not.

One of the following:
boolean
ItemLocator = string
max_completion_tokens?: number | ItemLocator

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
number
ItemLocator = string
max_tokens?: number | ItemLocator

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
number
ItemLocator = string
metadata?: Record<string, string> | ItemLocator

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
Record<string, string>
ItemLocator = string
modalities?: Array<string> | ItemLocator

Output types that you would like the model to generate for this request.

One of the following:
Array<string>
ItemLocator = string
n?: number | ItemLocator

How many chat completion choices to generate for each input message.

One of the following:
number
ItemLocator = string
parallel_tool_calls?: boolean | ItemLocator

Whether to enable parallel function calling during tool use.

One of the following:
boolean
ItemLocator = string
prediction?: Record<string, unknown> | ItemLocator

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
Record<string, unknown>
ItemLocator = string
presence_penalty?: number | ItemLocator

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
number
ItemLocator = string
reasoning_effort?: string

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format?: Record<string, unknown> | ItemLocator

An object specifying the format that the model must output.

One of the following:
Record<string, unknown>
ItemLocator = string
seed?: number | ItemLocator

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
number
ItemLocator = string
stop?: string | Array<string>

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
string
Array<string>
store?: boolean | ItemLocator

Whether to store the output for use in model distillation or evals products.

One of the following:
boolean
ItemLocator = string
temperature?: number | ItemLocator

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
number
ItemLocator = string
tool_choice?: string | Record<string, unknown>

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
string
Record<string, unknown>
tools?: Array<Record<string, unknown>> | ItemLocator

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
Array<Record<string, unknown>>
ItemLocator = string
top_k?: number | ItemLocator

Only sample from the top K options for each subsequent token

One of the following:
number
ItemLocator = string
top_logprobs?: number | ItemLocator

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
number
ItemLocator = string
top_p?: number | ItemLocator

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the chat_completion

task_type?: "chat_completion"
GenericInferenceEvaluationTask { configuration, alias, task_type }
configuration: Configuration { model, args, inference_configuration }
model: string

model specified as vendor/name (ex. openai/gpt-5)

args?: Record<string, unknown> | ItemLocator

Arguments passed into model

One of the following:
Record<string, unknown>
ItemLocator = string
inference_configuration?: LaunchInferenceConfiguration { num_retries, timeout_seconds } | ItemLocator

Vendor specific configuration

One of the following:
LaunchInferenceConfiguration { num_retries, timeout_seconds }
num_retries?: number
timeout_seconds?: number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the inference

task_type?: "inference"
ApplicationVariantV1EvaluationTask { configuration, alias, task_type }
configuration: Configuration { application_variant_id, inputs, history, 2 more }
application_variant_id: string
inputs: Record<string, unknown> | ItemLocator

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
Record<string, unknown>
ItemLocator = string
history?: Array<ApplicationRequestResponsePairArray> | ItemLocator

History of the application

One of the following:
Array<ApplicationRequestResponsePairArray>
request: string

Request inputs

response: string

Response outputs

session_data?: Record<string, unknown>

Session data corresponding to the request response pair

ItemLocator = string
operation_metadata?: Record<string, unknown> | ItemLocator

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
Record<string, unknown>
ItemLocator = string
overrides?: AgenticApplicationOverrides { concurrent, initial_state, partial_trace, 2 more } | Record<string, KnowledgeBaseNodeOverride> | ItemLocator

Optional overrides for the application

One of the following:
AgenticApplicationOverrides { concurrent, initial_state, partial_trace, 2 more }

Execution override options for agentic applications

concurrent?: boolean
initial_state?: InitialState { current_node, state }
current_node: string
state: Record<string, unknown>
partial_trace?: Array<PartialTrace>
duration_ms: number
node_id: string
operation_input: string
operation_output: string
operation_type: string
start_timestamp: string
workflow_id: string
operation_metadata?: Record<string, unknown>
return_span?: boolean
use_channels?: boolean
Record<string, KnowledgeBaseNodeOverride>
artifact_ids_filter?: Array<string>
artifact_name_regex?: Array<string>
type?: "knowledge_base_schema"
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the application_variant

task_type?: "application_variant"
AgentexOutputEvaluationTask { configuration, alias, task_type }
configuration: Configuration { agentex_agent_id, input_column, agent_task_params, 6 more }
agentex_agent_id: string

The ID of the Agentex agent to use

input_column: string | Record<string, unknown> | Array<unknown>

The dataset column to use as input for the agent

One of the following:
string
Record<string, unknown>
Array<unknown>
agent_task_params?: Record<string, unknown> | ItemLocator

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
Record<string, unknown>
ItemLocator = string
completion_mode?: "first_message" | "turn_quiescence"

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id?: string

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces?: boolean | ItemLocator

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
boolean
ItemLocator = string
input_mode?: "text" | "data"

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds?: number | ItemLocator

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
number
ItemLocator = string
timeout_seconds?: number | ItemLocator

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
number
ItemLocator = string
alias?: string

Alias to title the results column. Defaults to the agentex_output

task_type?: "agentex_output"
MetricEvaluationTask { configuration, alias, task_type }
configuration: BleuScorerConfigWithItemLocator { candidate, reference, type } | MeteorScorerConfigWithItemLocator { candidate, reference, type } | CosineSimilarityScorerConfigWithItemLocator { candidate, reference, type } | 4 more
One of the following:
BleuScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "bleu"
MeteorScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "meteor"
CosineSimilarityScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "cosine_similarity"
F1ScorerConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "f1"
RougeScorer1ConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rouge1"
RougeScorer2ConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rouge2"
RougeScorerLConfigWithItemLocator { candidate, reference, type }
candidate: string
reference: string
type: "rougeL"
alias?: string

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type?: "metric"
AutoEvaluationQuestionTask { configuration, alias, task_type }
configuration: Configuration { model, prompt, question_id }
model: string

model specified as model_vendor/model_name

prompt: string
question_id: string

question to be evaluated

alias?: string

Alias to title the results column. Defaults to the auto_evaluation_question

task_type?: "auto_evaluation.question"
AutoEvaluationGuidedDecodingEvaluationTask { configuration, alias, task_type }
configuration: AutoEvaluationStructuredOutputTaskRequestWithItemLocator { model, prompt, response_format, 3 more } | AutoEvaluationGuidedDecodingTaskRequestWithItemLocator { choices, model, prompt, 3 more } | AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
One of the following:
AutoEvaluationStructuredOutputTaskRequestWithItemLocator { model, prompt, response_format, 3 more }
model: string

model specified as model_vendor/model_name

prompt: string
response_format: Record<string, unknown>

JSON schema used for structuring the model response

inference_args?: Record<string, unknown>

Additional arguments to pass to the inference request

run_condition?: ConstEvaluationRunCondition { op, value } | VarEvaluationRunCondition { path, op } | EqEvaluationRunCondition { left, right, op } | 12 more
One of the following:
ConstEvaluationRunCondition { op, value }
op?: "const"
value?: string | number | boolean
One of the following:
string
number
boolean
VarEvaluationRunCondition { path, op }
path: string
op?: "var"
EqEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "eq"
NeEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "ne"
LtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lt"
LteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lte"
GtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gt"
GteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gte"
AndEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "and"
OrEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "or"
InEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "in"
NotInEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "not_in"
NotEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "not"
IsNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_null"
IsNotNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_not_null"
system_prompt?: string
AutoEvaluationGuidedDecodingTaskRequestWithItemLocator { choices, model, prompt, 3 more }
choices: Array<string>

Choices array cannot be empty

model: string

model specified as model_vendor/model_name

prompt: string
inference_args?: Record<string, unknown>

Additional arguments to pass to the inference request

run_condition?: ConstEvaluationRunCondition { op, value } | VarEvaluationRunCondition { path, op } | EqEvaluationRunCondition { left, right, op } | 12 more
One of the following:
ConstEvaluationRunCondition { op, value }
op?: "const"
value?: string | number | boolean
One of the following:
string
number
boolean
VarEvaluationRunCondition { path, op }
path: string
op?: "var"
EqEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "eq"
NeEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "ne"
LtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lt"
LteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "lte"
GtEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gt"
GteEvaluationRunCondition { left, right, op }
left: unknown
right: unknown
op?: "gte"
AndEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "and"
OrEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "or"
InEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "in"
NotInEvaluationRunCondition { left, operands, op }
left: unknown
operands: Array<unknown>
op?: "not_in"
NotEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "not"
IsNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_null"
IsNotNullEvaluationRunCondition { operands, op }
operands: Array<unknown>
op?: "is_not_null"
system_prompt?: string
AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: Array<string>
data_fields?: Array<string>
designated_to?: ApeAgent { config, agent_name } | IfAgent { config, agent_name } | TruthfulnessAgent { config, agent_name } | BaseAgent { config, agent_name }
One of the following:
ApeAgent { config, agent_name }
config: Config { model, temperature }
model?: string
temperature?: number
agent_name?: "APEAgent"
IfAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "IFAgent"
TruthfulnessAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "TruthfulnessAgent"
BaseAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "BaseAgent"
output_type?: "text" | "integer" | "float" | "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values?: Array<string | number | boolean>
One of the following:
string
number
boolean
rubric_id?: string
rubric_version?: number
alias?: string

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type?: "auto_evaluation.guided_decoding"
AutoEvaluationAgentEvaluationTask { configuration, alias, task_type }
configuration: AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: Array<string>
data_fields?: Array<string>
designated_to?: ApeAgent { config, agent_name } | IfAgent { config, agent_name } | TruthfulnessAgent { config, agent_name } | BaseAgent { config, agent_name }
One of the following:
ApeAgent { config, agent_name }
config: Config { model, temperature }
model?: string
temperature?: number
agent_name?: "APEAgent"
IfAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "IFAgent"
TruthfulnessAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "TruthfulnessAgent"
BaseAgent { config, agent_name }
config: Config { model }
model?: string
agent_name?: "BaseAgent"
output_type?: "text" | "integer" | "float" | "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values?: Array<string | number | boolean>
One of the following:
string
number
boolean
rubric_id?: string
rubric_version?: number
alias?: string

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type?: "auto_evaluation.agent"
ContributorEvaluationQuestionTask { configuration, alias, task_type }
configuration: Configuration { layout, question_id, prefill_from, 3 more }
layout: Container { children, direction }
children: Array<Container { children, direction } | Component { data, label } >

The children to be displayed within the container

One of the following:
Container = Container { children, direction }
Component { data, label }

A pointer to the data in each evaluation item to be displayed within the component

label?: string
direction?: "row" | "column"

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: string
prefill_from?: string

Dataset column to prefill contributor question task result

minLength1
queue_id?: string

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required?: boolean

Whether the question is required to be answered

rubric_id?: string

ID of the rubric to use for scoring this evaluation question

alias?: string

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type?: "contributor_evaluation.question"
CustomFunctionEvaluationTask { configuration, alias, task_type }
configuration: Configuration { function_source, arg_mapping, config_args, outputs }

Configuration for a custom Python function evaluation task.

function_source: string

Python function source code

maxLength10000
arg_mapping?: Record<string, string>

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args?: Record<string, unknown>

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs?: Array<Output>

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: string

Dot path in the custom function return value to materialize.

minLength1
alias?: string

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias?: string

Alias to title the results column. Defaults to the function name.

task_type?: "custom_function"
total: number

The total of items that match the query. This is greater than or equal to the number of items returned.

limit?: number

The maximum number of items to return.

object?: "list"
EvaluationRetrieveTaxonomyResponse = Record<string, unknown>

EvaluationsTasks

Add Test Criteria to Evaluation
client.evaluations.tasks.add(stringevaluationID, TaskAddParams { task } body, RequestOptionsoptions?): Evaluation { id, created_at, created_by, 12 more }
POST/v5/evaluations/{evaluation_id}/tasks
Update Test Criteria Configuration
client.evaluations.tasks.update(stringalias, TaskUpdateParams { evaluation_id, configuration } params, RequestOptionsoptions?): Evaluation { id, created_at, created_by, 12 more }
PATCH/v5/evaluations/{evaluation_id}/tasks/{alias}