Skip to content

Create Evaluation

evaluations.create(EvaluationCreateParams**kwargs) -> Evaluation
POST/v5/evaluations

Create an evaluation together with its items, optionally running test criteria against them.

Accepts three request shapes: standalone (inline data), from an existing dataset (dataset_id with optional per-item references), or with a new reusable dataset created inline from data. When the evaluation includes tasks that require execution (for example an LLM judge or custom function), an async job and a Temporal workflow are started and the evaluation is returned immediately with status running; task results and error_count populate asynchronously. When it includes only contributor tasks, taxonomy-only input, or no tasks, no workflow runs and it is returned with status completed. Optional tasks, metadata, tags, and taxonomy_params are persisted alongside the evaluation and its items.

ParametersExpand Collapse
evaluation: Evaluation
One of the following:
class EvaluationEvaluationStandaloneCreateRequest: …
data: Iterable[Dict[str, object]]

Items to be evaluated

name: str
maxLength128
minLength1
description: Optional[str]
maxLength280
files: Optional[Iterable[Dict[str, str]]]

Files to be associated to the evaluation

metadata: Optional[Dict[str, object]]

Optional metadata key-value pairs for the evaluation

skip_prefilled_rows: Optional[bool]

Do not queue a contributor task for prefilled questions

tags: Optional[Sequence[str]]

The tags associated with the evaluation

tasks: Optional[Iterable[EvaluationTaskParam]]

Tasks allow you to augment and evaluate your data

One of the following:
class ChatCompletionEvaluationTask: …
configuration: ChatCompletionEvaluationTaskConfiguration
messages: Union[List[Dict[str, object]], ItemLocator]

openai standard message format

One of the following:
List[Dict[str, object]]
str
model: str

model specified as model_vendor/model, for example openai/gpt-4o

audio: Optional[Union[Dict[str, object], ItemLocator, null]]

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
Dict[str, object]
str
frequency_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
float
str
function_call: Optional[Union[Dict[str, object], ItemLocator, null]]

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
Dict[str, object]
str
functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
List[Dict[str, object]]
str
logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
Dict[str, int]
str
logprobs: Optional[Union[bool, ItemLocator, null]]

Whether to return log probabilities of the output tokens or not.

One of the following:
bool
str
max_completion_tokens: Optional[Union[int, ItemLocator, null]]

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
int
str
max_tokens: Optional[Union[int, ItemLocator, null]]

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
int
str
metadata: Optional[Union[Dict[str, str], ItemLocator, null]]

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
Dict[str, str]
str
modalities: Optional[Union[List[str], ItemLocator, null]]

Output types that you would like the model to generate for this request.

One of the following:
List[str]
str
n: Optional[Union[int, ItemLocator, null]]

How many chat completion choices to generate for each input message.

One of the following:
int
str
parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]

Whether to enable parallel function calling during tool use.

One of the following:
bool
str
prediction: Optional[Union[Dict[str, object], ItemLocator, null]]

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
Dict[str, object]
str
presence_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
float
str
reasoning_effort: Optional[str]

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format: Optional[Union[Dict[str, object], ItemLocator, null]]

An object specifying the format that the model must output.

One of the following:
Dict[str, object]
str
seed: Optional[Union[int, ItemLocator, null]]

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
int
str
stop: Optional[Union[str, List[str], null]]

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
str
List[str]
store: Optional[Union[bool, ItemLocator, null]]

Whether to store the output for use in model distillation or evals products.

One of the following:
bool
str
temperature: Optional[Union[float, ItemLocator, null]]

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
float
str
tool_choice: Optional[Union[str, Dict[str, object], null]]

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
str
Dict[str, object]
tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
List[Dict[str, object]]
str
top_k: Optional[Union[int, ItemLocator, null]]

Only sample from the top K options for each subsequent token

One of the following:
int
str
top_logprobs: Optional[Union[int, ItemLocator, null]]

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
int
str
top_p: Optional[Union[float, ItemLocator, null]]

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
float
str
alias: Optional[str]

Alias to title the results column. Defaults to the chat_completion

task_type: Optional[Literal["chat_completion"]]
class GenericInferenceEvaluationTask: …
configuration: GenericInferenceEvaluationTaskConfiguration
model: str

model specified as vendor/name (ex. openai/gpt-5)

args: Optional[Union[Dict[str, object], ItemLocator, null]]

Arguments passed into model

One of the following:
Dict[str, object]
str
inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]

Vendor specific configuration

One of the following:
class LaunchInferenceConfiguration: …
num_retries: Optional[int]
timeout_seconds: Optional[int]
str
alias: Optional[str]

Alias to title the results column. Defaults to the inference

task_type: Optional[Literal["inference"]]
class ApplicationVariantV1EvaluationTask: …
configuration: ApplicationVariantV1EvaluationTaskConfiguration
application_variant_id: str
inputs: Union[Dict[str, object], ItemLocator]

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
Dict[str, object]
str
history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]

History of the application

One of the following:
List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]
request: str

Request inputs

response: str

Response outputs

session_data: Optional[Dict[str, object]]

Session data corresponding to the request response pair

str
operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
Dict[str, object]
str
overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]

Optional overrides for the application

One of the following:
class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …

Execution override options for agentic applications

concurrent: Optional[bool]
initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]
current_node: str
state: Dict[str, object]
partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]
duration_ms: int
node_id: str
operation_input: str
operation_output: str
operation_type: str
start_timestamp: str
workflow_id: str
operation_metadata: Optional[Dict[str, object]]
return_span: Optional[bool]
use_channels: Optional[bool]
Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]
artifact_ids_filter: Optional[List[str]]
artifact_name_regex: Optional[List[str]]
type: Optional[Literal["knowledge_base_schema"]]
str
alias: Optional[str]

Alias to title the results column. Defaults to the application_variant

task_type: Optional[Literal["application_variant"]]
class AgentexOutputEvaluationTask: …
configuration: AgentexOutputEvaluationTaskConfiguration
agentex_agent_id: str

The ID of the Agentex agent to use

input_column: Union[str, Dict[str, object], List[object]]

The dataset column to use as input for the agent

One of the following:
str
Dict[str, object]
List[object]
agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
Dict[str, object]
str
completion_mode: Optional[Literal["first_message", "turn_quiescence"]]

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id: Optional[str]

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces: Optional[Union[bool, ItemLocator, null]]

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
bool
str
input_mode: Optional[Literal["text", "data"]]

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds: Optional[Union[int, ItemLocator, null]]

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
int
str
timeout_seconds: Optional[Union[int, ItemLocator, null]]

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
int
str
alias: Optional[str]

Alias to title the results column. Defaults to the agentex_output

task_type: Optional[Literal["agentex_output"]]
class MetricEvaluationTask: …
configuration: MetricEvaluationTaskConfiguration
One of the following:
class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["bleu"]
class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["meteor"]
class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["cosine_similarity"]
class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["f1"]
class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge1"]
class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge2"]
class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rougeL"]
alias: Optional[str]

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type: Optional[Literal["metric"]]
class AutoEvaluationQuestionTask: …
configuration: AutoEvaluationQuestionTaskConfiguration
model: str

model specified as model_vendor/model_name

prompt: str
question_id: str

question to be evaluated

alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_question

task_type: Optional[Literal["auto_evaluation.question"]]
class AutoEvaluationGuidedDecodingEvaluationTask: …
configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …
model: str

model specified as model_vendor/model_name

prompt: str
response_format: Dict[str, object]

JSON schema used for structuring the model response

inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …
choices: List[str]

Choices array cannot be empty

model: str

model specified as model_vendor/model_name

prompt: str
inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationAgentTaskRequestWithItemLocator: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type: Optional[Literal["auto_evaluation.guided_decoding"]]
class AutoEvaluationAgentEvaluationTask: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type: Optional[Literal["auto_evaluation.agent"]]
class ContributorEvaluationQuestionTask: …
configuration: ContributorEvaluationQuestionTaskConfiguration
layout: Container
children: List[Child]

The children to be displayed within the container

One of the following:
class Component: …

A pointer to the data in each evaluation item to be displayed within the component

label: Optional[str]
direction: Optional[Literal["row", "column"]]

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: str
prefill_from: Optional[str]

Dataset column to prefill contributor question task result

minLength1
queue_id: Optional[str]

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required: Optional[bool]

Whether the question is required to be answered

rubric_id: Optional[str]

ID of the rubric to use for scoring this evaluation question

alias: Optional[str]

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type: Optional[Literal["contributor_evaluation.question"]]
class CustomFunctionEvaluationTask: …
configuration: CustomFunctionEvaluationTaskConfiguration

Configuration for a custom Python function evaluation task.

function_source: str

Python function source code

maxLength10000
arg_mapping: Optional[Dict[str, str]]

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args: Optional[Dict[str, object]]

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: str

Dot path in the custom function return value to materialize.

minLength1
alias: Optional[str]

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias: Optional[str]

Alias to title the results column. Defaults to the function name.

task_type: Optional[Literal["custom_function"]]
taxonomy_params: Optional[Dict[str, object]]

Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

class EvaluationEvaluationFromDatasetCreateRequest: …
dataset_id: str

The ID of the dataset containing the items referenced by the data field

name: str
maxLength128
minLength1
data: Optional[Iterable[EvaluationEvaluationFromDatasetCreateRequestData]]

Items to be evaluated, including references to the input dataset

dataset_item_id: str
description: Optional[str]
maxLength280
metadata: Optional[Dict[str, object]]

Optional metadata key-value pairs for the evaluation

skip_prefilled_rows: Optional[bool]

Do not queue a contributor task for prefilled questions

tags: Optional[Sequence[str]]

The tags associated with the evaluation

tasks: Optional[Iterable[EvaluationTaskParam]]

Tasks allow you to augment and evaluate your data

One of the following:
class ChatCompletionEvaluationTask: …
configuration: ChatCompletionEvaluationTaskConfiguration
messages: Union[List[Dict[str, object]], ItemLocator]

openai standard message format

One of the following:
List[Dict[str, object]]
str
model: str

model specified as model_vendor/model, for example openai/gpt-4o

audio: Optional[Union[Dict[str, object], ItemLocator, null]]

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
Dict[str, object]
str
frequency_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
float
str
function_call: Optional[Union[Dict[str, object], ItemLocator, null]]

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
Dict[str, object]
str
functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
List[Dict[str, object]]
str
logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
Dict[str, int]
str
logprobs: Optional[Union[bool, ItemLocator, null]]

Whether to return log probabilities of the output tokens or not.

One of the following:
bool
str
max_completion_tokens: Optional[Union[int, ItemLocator, null]]

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
int
str
max_tokens: Optional[Union[int, ItemLocator, null]]

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
int
str
metadata: Optional[Union[Dict[str, str], ItemLocator, null]]

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
Dict[str, str]
str
modalities: Optional[Union[List[str], ItemLocator, null]]

Output types that you would like the model to generate for this request.

One of the following:
List[str]
str
n: Optional[Union[int, ItemLocator, null]]

How many chat completion choices to generate for each input message.

One of the following:
int
str
parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]

Whether to enable parallel function calling during tool use.

One of the following:
bool
str
prediction: Optional[Union[Dict[str, object], ItemLocator, null]]

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
Dict[str, object]
str
presence_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
float
str
reasoning_effort: Optional[str]

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format: Optional[Union[Dict[str, object], ItemLocator, null]]

An object specifying the format that the model must output.

One of the following:
Dict[str, object]
str
seed: Optional[Union[int, ItemLocator, null]]

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
int
str
stop: Optional[Union[str, List[str], null]]

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
str
List[str]
store: Optional[Union[bool, ItemLocator, null]]

Whether to store the output for use in model distillation or evals products.

One of the following:
bool
str
temperature: Optional[Union[float, ItemLocator, null]]

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
float
str
tool_choice: Optional[Union[str, Dict[str, object], null]]

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
str
Dict[str, object]
tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
List[Dict[str, object]]
str
top_k: Optional[Union[int, ItemLocator, null]]

Only sample from the top K options for each subsequent token

One of the following:
int
str
top_logprobs: Optional[Union[int, ItemLocator, null]]

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
int
str
top_p: Optional[Union[float, ItemLocator, null]]

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
float
str
alias: Optional[str]

Alias to title the results column. Defaults to the chat_completion

task_type: Optional[Literal["chat_completion"]]
class GenericInferenceEvaluationTask: …
configuration: GenericInferenceEvaluationTaskConfiguration
model: str

model specified as vendor/name (ex. openai/gpt-5)

args: Optional[Union[Dict[str, object], ItemLocator, null]]

Arguments passed into model

One of the following:
Dict[str, object]
str
inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]

Vendor specific configuration

One of the following:
class LaunchInferenceConfiguration: …
num_retries: Optional[int]
timeout_seconds: Optional[int]
str
alias: Optional[str]

Alias to title the results column. Defaults to the inference

task_type: Optional[Literal["inference"]]
class ApplicationVariantV1EvaluationTask: …
configuration: ApplicationVariantV1EvaluationTaskConfiguration
application_variant_id: str
inputs: Union[Dict[str, object], ItemLocator]

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
Dict[str, object]
str
history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]

History of the application

One of the following:
List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]
request: str

Request inputs

response: str

Response outputs

session_data: Optional[Dict[str, object]]

Session data corresponding to the request response pair

str
operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
Dict[str, object]
str
overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]

Optional overrides for the application

One of the following:
class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …

Execution override options for agentic applications

concurrent: Optional[bool]
initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]
current_node: str
state: Dict[str, object]
partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]
duration_ms: int
node_id: str
operation_input: str
operation_output: str
operation_type: str
start_timestamp: str
workflow_id: str
operation_metadata: Optional[Dict[str, object]]
return_span: Optional[bool]
use_channels: Optional[bool]
Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]
artifact_ids_filter: Optional[List[str]]
artifact_name_regex: Optional[List[str]]
type: Optional[Literal["knowledge_base_schema"]]
str
alias: Optional[str]

Alias to title the results column. Defaults to the application_variant

task_type: Optional[Literal["application_variant"]]
class AgentexOutputEvaluationTask: …
configuration: AgentexOutputEvaluationTaskConfiguration
agentex_agent_id: str

The ID of the Agentex agent to use

input_column: Union[str, Dict[str, object], List[object]]

The dataset column to use as input for the agent

One of the following:
str
Dict[str, object]
List[object]
agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
Dict[str, object]
str
completion_mode: Optional[Literal["first_message", "turn_quiescence"]]

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id: Optional[str]

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces: Optional[Union[bool, ItemLocator, null]]

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
bool
str
input_mode: Optional[Literal["text", "data"]]

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds: Optional[Union[int, ItemLocator, null]]

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
int
str
timeout_seconds: Optional[Union[int, ItemLocator, null]]

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
int
str
alias: Optional[str]

Alias to title the results column. Defaults to the agentex_output

task_type: Optional[Literal["agentex_output"]]
class MetricEvaluationTask: …
configuration: MetricEvaluationTaskConfiguration
One of the following:
class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["bleu"]
class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["meteor"]
class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["cosine_similarity"]
class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["f1"]
class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge1"]
class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge2"]
class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rougeL"]
alias: Optional[str]

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type: Optional[Literal["metric"]]
class AutoEvaluationQuestionTask: …
configuration: AutoEvaluationQuestionTaskConfiguration
model: str

model specified as model_vendor/model_name

prompt: str
question_id: str

question to be evaluated

alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_question

task_type: Optional[Literal["auto_evaluation.question"]]
class AutoEvaluationGuidedDecodingEvaluationTask: …
configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …
model: str

model specified as model_vendor/model_name

prompt: str
response_format: Dict[str, object]

JSON schema used for structuring the model response

inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …
choices: List[str]

Choices array cannot be empty

model: str

model specified as model_vendor/model_name

prompt: str
inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationAgentTaskRequestWithItemLocator: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type: Optional[Literal["auto_evaluation.guided_decoding"]]
class AutoEvaluationAgentEvaluationTask: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type: Optional[Literal["auto_evaluation.agent"]]
class ContributorEvaluationQuestionTask: …
configuration: ContributorEvaluationQuestionTaskConfiguration
layout: Container
children: List[Child]

The children to be displayed within the container

One of the following:
class Component: …

A pointer to the data in each evaluation item to be displayed within the component

label: Optional[str]
direction: Optional[Literal["row", "column"]]

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: str
prefill_from: Optional[str]

Dataset column to prefill contributor question task result

minLength1
queue_id: Optional[str]

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required: Optional[bool]

Whether the question is required to be answered

rubric_id: Optional[str]

ID of the rubric to use for scoring this evaluation question

alias: Optional[str]

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type: Optional[Literal["contributor_evaluation.question"]]
class CustomFunctionEvaluationTask: …
configuration: CustomFunctionEvaluationTaskConfiguration

Configuration for a custom Python function evaluation task.

function_source: str

Python function source code

maxLength10000
arg_mapping: Optional[Dict[str, str]]

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args: Optional[Dict[str, object]]

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: str

Dot path in the custom function return value to materialize.

minLength1
alias: Optional[str]

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias: Optional[str]

Alias to title the results column. Defaults to the function name.

task_type: Optional[Literal["custom_function"]]
taxonomy_params: Optional[Dict[str, object]]

Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

class EvaluationEvaluationWithDatasetCreateRequest: …
data: Iterable[Dict[str, object]]

Items to be evaluated

dataset: EvaluationEvaluationWithDatasetCreateRequestDataset

Create a reusable dataset from items in the data field

name: str
description: Optional[str]
keys: Optional[Sequence[str]]

Keys from items in the data field that should be included in the dataset. If not provided, all keys will be included.

tags: Optional[Sequence[str]]

The tags associated with the entity

name: str
maxLength128
minLength1
description: Optional[str]
maxLength280
files: Optional[Iterable[Dict[str, str]]]

Files to be associated to the evaluation

metadata: Optional[Dict[str, object]]

Optional metadata key-value pairs for the evaluation

skip_prefilled_rows: Optional[bool]

Do not queue a contributor task for prefilled questions

tags: Optional[Sequence[str]]

The tags associated with the evaluation

tasks: Optional[Iterable[EvaluationTaskParam]]

Tasks allow you to augment and evaluate your data

One of the following:
class ChatCompletionEvaluationTask: …
configuration: ChatCompletionEvaluationTaskConfiguration
messages: Union[List[Dict[str, object]], ItemLocator]

openai standard message format

One of the following:
List[Dict[str, object]]
str
model: str

model specified as model_vendor/model, for example openai/gpt-4o

audio: Optional[Union[Dict[str, object], ItemLocator, null]]

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
Dict[str, object]
str
frequency_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
float
str
function_call: Optional[Union[Dict[str, object], ItemLocator, null]]

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
Dict[str, object]
str
functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
List[Dict[str, object]]
str
logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
Dict[str, int]
str
logprobs: Optional[Union[bool, ItemLocator, null]]

Whether to return log probabilities of the output tokens or not.

One of the following:
bool
str
max_completion_tokens: Optional[Union[int, ItemLocator, null]]

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
int
str
max_tokens: Optional[Union[int, ItemLocator, null]]

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
int
str
metadata: Optional[Union[Dict[str, str], ItemLocator, null]]

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
Dict[str, str]
str
modalities: Optional[Union[List[str], ItemLocator, null]]

Output types that you would like the model to generate for this request.

One of the following:
List[str]
str
n: Optional[Union[int, ItemLocator, null]]

How many chat completion choices to generate for each input message.

One of the following:
int
str
parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]

Whether to enable parallel function calling during tool use.

One of the following:
bool
str
prediction: Optional[Union[Dict[str, object], ItemLocator, null]]

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
Dict[str, object]
str
presence_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
float
str
reasoning_effort: Optional[str]

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format: Optional[Union[Dict[str, object], ItemLocator, null]]

An object specifying the format that the model must output.

One of the following:
Dict[str, object]
str
seed: Optional[Union[int, ItemLocator, null]]

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
int
str
stop: Optional[Union[str, List[str], null]]

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
str
List[str]
store: Optional[Union[bool, ItemLocator, null]]

Whether to store the output for use in model distillation or evals products.

One of the following:
bool
str
temperature: Optional[Union[float, ItemLocator, null]]

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
float
str
tool_choice: Optional[Union[str, Dict[str, object], null]]

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
str
Dict[str, object]
tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
List[Dict[str, object]]
str
top_k: Optional[Union[int, ItemLocator, null]]

Only sample from the top K options for each subsequent token

One of the following:
int
str
top_logprobs: Optional[Union[int, ItemLocator, null]]

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
int
str
top_p: Optional[Union[float, ItemLocator, null]]

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
float
str
alias: Optional[str]

Alias to title the results column. Defaults to the chat_completion

task_type: Optional[Literal["chat_completion"]]
class GenericInferenceEvaluationTask: …
configuration: GenericInferenceEvaluationTaskConfiguration
model: str

model specified as vendor/name (ex. openai/gpt-5)

args: Optional[Union[Dict[str, object], ItemLocator, null]]

Arguments passed into model

One of the following:
Dict[str, object]
str
inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]

Vendor specific configuration

One of the following:
class LaunchInferenceConfiguration: …
num_retries: Optional[int]
timeout_seconds: Optional[int]
str
alias: Optional[str]

Alias to title the results column. Defaults to the inference

task_type: Optional[Literal["inference"]]
class ApplicationVariantV1EvaluationTask: …
configuration: ApplicationVariantV1EvaluationTaskConfiguration
application_variant_id: str
inputs: Union[Dict[str, object], ItemLocator]

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
Dict[str, object]
str
history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]

History of the application

One of the following:
List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]
request: str

Request inputs

response: str

Response outputs

session_data: Optional[Dict[str, object]]

Session data corresponding to the request response pair

str
operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
Dict[str, object]
str
overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]

Optional overrides for the application

One of the following:
class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …

Execution override options for agentic applications

concurrent: Optional[bool]
initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]
current_node: str
state: Dict[str, object]
partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]
duration_ms: int
node_id: str
operation_input: str
operation_output: str
operation_type: str
start_timestamp: str
workflow_id: str
operation_metadata: Optional[Dict[str, object]]
return_span: Optional[bool]
use_channels: Optional[bool]
Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]
artifact_ids_filter: Optional[List[str]]
artifact_name_regex: Optional[List[str]]
type: Optional[Literal["knowledge_base_schema"]]
str
alias: Optional[str]

Alias to title the results column. Defaults to the application_variant

task_type: Optional[Literal["application_variant"]]
class AgentexOutputEvaluationTask: …
configuration: AgentexOutputEvaluationTaskConfiguration
agentex_agent_id: str

The ID of the Agentex agent to use

input_column: Union[str, Dict[str, object], List[object]]

The dataset column to use as input for the agent

One of the following:
str
Dict[str, object]
List[object]
agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
Dict[str, object]
str
completion_mode: Optional[Literal["first_message", "turn_quiescence"]]

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id: Optional[str]

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces: Optional[Union[bool, ItemLocator, null]]

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
bool
str
input_mode: Optional[Literal["text", "data"]]

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds: Optional[Union[int, ItemLocator, null]]

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
int
str
timeout_seconds: Optional[Union[int, ItemLocator, null]]

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
int
str
alias: Optional[str]

Alias to title the results column. Defaults to the agentex_output

task_type: Optional[Literal["agentex_output"]]
class MetricEvaluationTask: …
configuration: MetricEvaluationTaskConfiguration
One of the following:
class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["bleu"]
class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["meteor"]
class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["cosine_similarity"]
class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["f1"]
class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge1"]
class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge2"]
class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rougeL"]
alias: Optional[str]

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type: Optional[Literal["metric"]]
class AutoEvaluationQuestionTask: …
configuration: AutoEvaluationQuestionTaskConfiguration
model: str

model specified as model_vendor/model_name

prompt: str
question_id: str

question to be evaluated

alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_question

task_type: Optional[Literal["auto_evaluation.question"]]
class AutoEvaluationGuidedDecodingEvaluationTask: …
configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …
model: str

model specified as model_vendor/model_name

prompt: str
response_format: Dict[str, object]

JSON schema used for structuring the model response

inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …
choices: List[str]

Choices array cannot be empty

model: str

model specified as model_vendor/model_name

prompt: str
inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationAgentTaskRequestWithItemLocator: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type: Optional[Literal["auto_evaluation.guided_decoding"]]
class AutoEvaluationAgentEvaluationTask: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type: Optional[Literal["auto_evaluation.agent"]]
class ContributorEvaluationQuestionTask: …
configuration: ContributorEvaluationQuestionTaskConfiguration
layout: Container
children: List[Child]

The children to be displayed within the container

One of the following:
class Component: …

A pointer to the data in each evaluation item to be displayed within the component

label: Optional[str]
direction: Optional[Literal["row", "column"]]

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: str
prefill_from: Optional[str]

Dataset column to prefill contributor question task result

minLength1
queue_id: Optional[str]

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required: Optional[bool]

Whether the question is required to be answered

rubric_id: Optional[str]

ID of the rubric to use for scoring this evaluation question

alias: Optional[str]

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type: Optional[Literal["contributor_evaluation.question"]]
class CustomFunctionEvaluationTask: …
configuration: CustomFunctionEvaluationTaskConfiguration

Configuration for a custom Python function evaluation task.

function_source: str

Python function source code

maxLength10000
arg_mapping: Optional[Dict[str, str]]

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args: Optional[Dict[str, object]]

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: str

Dot path in the custom function return value to materialize.

minLength1
alias: Optional[str]

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias: Optional[str]

Alias to title the results column. Defaults to the function name.

task_type: Optional[Literal["custom_function"]]
taxonomy_params: Optional[Dict[str, object]]

Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

ReturnsExpand Collapse
class Evaluation: …
id: str

The unique identifier of the entity.

created_at: datetime

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity

The identity that created the entity.

id: str
type: Literal["user", "service_account"]
One of the following:
"user"
"service_account"
object: Optional[Literal["identity"]]
datasets: Optional[List[Dataset]]
id: str

The unique identifier of the entity.

created_at: datetime

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity

The identity that created the entity.

id: str
type: Literal["user", "service_account"]
One of the following:
"user"
"service_account"
object: Optional[Literal["identity"]]
current_version_num: int
name: str
tags: Optional[List[str]]

The tags associated with the entity

archived_at: Optional[datetime]

The date and time when the entity was archived in ISO format.

formatdate-time
description: Optional[str]
object: Optional[Literal["dataset"]]
name: str
status: Literal["failed", "completed", "running"]
One of the following:
"failed"
"completed"
"running"
tags: Optional[List[str]]

The tags associated with the entity

archived_at: Optional[datetime]

The date and time when the entity was archived in ISO format.

formatdate-time
description: Optional[str]
error_count: Optional[int]

Number of task errors across all items in this evaluation.

metadata: Optional[Dict[str, object]]

Metadata key-value pairs for the evaluation

object: Optional[Literal["evaluation"]]
progress: Optional[EvaluationTasksProgressSchema]

Progress of the evaluation’s underlying async job

items: Optional[Items]
failed: int
pending: int
successful: int
total: int
failed_items: Optional[List[ItemsFailedItem]]
item_id: str
error: Optional[str]
error_type: Optional[str]
workflows: Optional[Workflows]
completed: int
failed: int
pending: int
total: int
status_reason: Optional[str]

Reason for evaluation status

tasks: Optional[List[EvaluationTask]]

Tasks executed during evaluation. Populated with optional task view.

One of the following:
class ChatCompletionEvaluationTask: …
configuration: ChatCompletionEvaluationTaskConfiguration
messages: Union[List[Dict[str, object]], ItemLocator]

openai standard message format

One of the following:
List[Dict[str, object]]
str
model: str

model specified as model_vendor/model, for example openai/gpt-4o

audio: Optional[Union[Dict[str, object], ItemLocator, null]]

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
Dict[str, object]
str
frequency_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
float
str
function_call: Optional[Union[Dict[str, object], ItemLocator, null]]

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
Dict[str, object]
str
functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
List[Dict[str, object]]
str
logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
Dict[str, int]
str
logprobs: Optional[Union[bool, ItemLocator, null]]

Whether to return log probabilities of the output tokens or not.

One of the following:
bool
str
max_completion_tokens: Optional[Union[int, ItemLocator, null]]

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
int
str
max_tokens: Optional[Union[int, ItemLocator, null]]

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
int
str
metadata: Optional[Union[Dict[str, str], ItemLocator, null]]

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
Dict[str, str]
str
modalities: Optional[Union[List[str], ItemLocator, null]]

Output types that you would like the model to generate for this request.

One of the following:
List[str]
str
n: Optional[Union[int, ItemLocator, null]]

How many chat completion choices to generate for each input message.

One of the following:
int
str
parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]

Whether to enable parallel function calling during tool use.

One of the following:
bool
str
prediction: Optional[Union[Dict[str, object], ItemLocator, null]]

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
Dict[str, object]
str
presence_penalty: Optional[Union[float, ItemLocator, null]]

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
float
str
reasoning_effort: Optional[str]

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format: Optional[Union[Dict[str, object], ItemLocator, null]]

An object specifying the format that the model must output.

One of the following:
Dict[str, object]
str
seed: Optional[Union[int, ItemLocator, null]]

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
int
str
stop: Optional[Union[str, List[str], null]]

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
str
List[str]
store: Optional[Union[bool, ItemLocator, null]]

Whether to store the output for use in model distillation or evals products.

One of the following:
bool
str
temperature: Optional[Union[float, ItemLocator, null]]

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
float
str
tool_choice: Optional[Union[str, Dict[str, object], null]]

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
str
Dict[str, object]
tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
List[Dict[str, object]]
str
top_k: Optional[Union[int, ItemLocator, null]]

Only sample from the top K options for each subsequent token

One of the following:
int
str
top_logprobs: Optional[Union[int, ItemLocator, null]]

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
int
str
top_p: Optional[Union[float, ItemLocator, null]]

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
float
str
alias: Optional[str]

Alias to title the results column. Defaults to the chat_completion

task_type: Optional[Literal["chat_completion"]]
class GenericInferenceEvaluationTask: …
configuration: GenericInferenceEvaluationTaskConfiguration
model: str

model specified as vendor/name (ex. openai/gpt-5)

args: Optional[Union[Dict[str, object], ItemLocator, null]]

Arguments passed into model

One of the following:
Dict[str, object]
str
inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]

Vendor specific configuration

One of the following:
class LaunchInferenceConfiguration: …
num_retries: Optional[int]
timeout_seconds: Optional[int]
str
alias: Optional[str]

Alias to title the results column. Defaults to the inference

task_type: Optional[Literal["inference"]]
class ApplicationVariantV1EvaluationTask: …
configuration: ApplicationVariantV1EvaluationTaskConfiguration
application_variant_id: str
inputs: Union[Dict[str, object], ItemLocator]

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
Dict[str, object]
str
history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]

History of the application

One of the following:
List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]
request: str

Request inputs

response: str

Response outputs

session_data: Optional[Dict[str, object]]

Session data corresponding to the request response pair

str
operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
Dict[str, object]
str
overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]

Optional overrides for the application

One of the following:
class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …

Execution override options for agentic applications

concurrent: Optional[bool]
initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]
current_node: str
state: Dict[str, object]
partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]
duration_ms: int
node_id: str
operation_input: str
operation_output: str
operation_type: str
start_timestamp: str
workflow_id: str
operation_metadata: Optional[Dict[str, object]]
return_span: Optional[bool]
use_channels: Optional[bool]
Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]
artifact_ids_filter: Optional[List[str]]
artifact_name_regex: Optional[List[str]]
type: Optional[Literal["knowledge_base_schema"]]
str
alias: Optional[str]

Alias to title the results column. Defaults to the application_variant

task_type: Optional[Literal["application_variant"]]
class AgentexOutputEvaluationTask: …
configuration: AgentexOutputEvaluationTaskConfiguration
agentex_agent_id: str

The ID of the Agentex agent to use

input_column: Union[str, Dict[str, object], List[object]]

The dataset column to use as input for the agent

One of the following:
str
Dict[str, object]
List[object]
agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
Dict[str, object]
str
completion_mode: Optional[Literal["first_message", "turn_quiescence"]]

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id: Optional[str]

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces: Optional[Union[bool, ItemLocator, null]]

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
bool
str
input_mode: Optional[Literal["text", "data"]]

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds: Optional[Union[int, ItemLocator, null]]

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
int
str
timeout_seconds: Optional[Union[int, ItemLocator, null]]

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
int
str
alias: Optional[str]

Alias to title the results column. Defaults to the agentex_output

task_type: Optional[Literal["agentex_output"]]
class MetricEvaluationTask: …
configuration: MetricEvaluationTaskConfiguration
One of the following:
class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["bleu"]
class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["meteor"]
class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["cosine_similarity"]
class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["f1"]
class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge1"]
class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rouge2"]
class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …
candidate: str
reference: str
type: Literal["rougeL"]
alias: Optional[str]

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type: Optional[Literal["metric"]]
class AutoEvaluationQuestionTask: …
configuration: AutoEvaluationQuestionTaskConfiguration
model: str

model specified as model_vendor/model_name

prompt: str
question_id: str

question to be evaluated

alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_question

task_type: Optional[Literal["auto_evaluation.question"]]
class AutoEvaluationGuidedDecodingEvaluationTask: …
configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …
model: str

model specified as model_vendor/model_name

prompt: str
response_format: Dict[str, object]

JSON schema used for structuring the model response

inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …
choices: List[str]

Choices array cannot be empty

model: str

model specified as model_vendor/model_name

prompt: str
inference_args: Optional[Dict[str, object]]

Additional arguments to pass to the inference request

run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]
One of the following:
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
op: Optional[Literal["const"]]
value: Optional[Union[str, float, bool, null]]
One of the following:
str
float
bool
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
path: str
op: Optional[Literal["var"]]
class EqEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["eq"]]
class NeEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["ne"]]
class LtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lt"]]
class LteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["lte"]]
class GtEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gt"]]
class GteEvaluationRunCondition: …
left: object
right: object
op: Optional[Literal["gte"]]
class AndEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["and"]]
class OrEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["or"]]
class InEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["in"]]
class NotInEvaluationRunCondition: …
left: object
operands: List[object]
op: Optional[Literal["not_in"]]
class NotEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["not"]]
class IsNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_null"]]
class IsNotNullEvaluationRunCondition: …
operands: List[object]
op: Optional[Literal["is_not_null"]]
system_prompt: Optional[str]
class AutoEvaluationAgentTaskRequestWithItemLocator: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type: Optional[Literal["auto_evaluation.guided_decoding"]]
class AutoEvaluationAgentEvaluationTask: …
definition: str
name: str
output_rules: List[str]
data_fields: Optional[List[str]]
designated_to: Optional[DesignatedTo]
One of the following:
class DesignatedToApeAgent: …
config: DesignatedToApeAgentConfig
model: Optional[str]
temperature: Optional[float]
agent_name: Optional[Literal["APEAgent"]]
class DesignatedToIfAgent: …
config: DesignatedToIfAgentConfig
model: Optional[str]
agent_name: Optional[Literal["IFAgent"]]
class DesignatedToTruthfulnessAgent: …
config: DesignatedToTruthfulnessAgentConfig
model: Optional[str]
agent_name: Optional[Literal["TruthfulnessAgent"]]
class DesignatedToBaseAgent: …
config: DesignatedToBaseAgentConfig
model: Optional[str]
agent_name: Optional[Literal["BaseAgent"]]
output_type: Optional[Literal["text", "integer", "float", "boolean"]]
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: Optional[List[Union[str, float, bool]]]
One of the following:
str
float
bool
rubric_id: Optional[str]
rubric_version: Optional[int]
alias: Optional[str]

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type: Optional[Literal["auto_evaluation.agent"]]
class ContributorEvaluationQuestionTask: …
configuration: ContributorEvaluationQuestionTaskConfiguration
layout: Container
children: List[Child]

The children to be displayed within the container

One of the following:
class Component: …

A pointer to the data in each evaluation item to be displayed within the component

label: Optional[str]
direction: Optional[Literal["row", "column"]]

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: str
prefill_from: Optional[str]

Dataset column to prefill contributor question task result

minLength1
queue_id: Optional[str]

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required: Optional[bool]

Whether the question is required to be answered

rubric_id: Optional[str]

ID of the rubric to use for scoring this evaluation question

alias: Optional[str]

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type: Optional[Literal["contributor_evaluation.question"]]
class CustomFunctionEvaluationTask: …
configuration: CustomFunctionEvaluationTaskConfiguration

Configuration for a custom Python function evaluation task.

function_source: str

Python function source code

maxLength10000
arg_mapping: Optional[Dict[str, str]]

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args: Optional[Dict[str, object]]

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: str

Dot path in the custom function return value to materialize.

minLength1
alias: Optional[str]

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias: Optional[str]

Alias to title the results column. Defaults to the function name.

task_type: Optional[Literal["custom_function"]]

Create Evaluation

import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation = client.evaluations.create(
    evaluation={
        "data": [{
            "foo": "bar"
        }],
        "name": "x",
    },
)
print(evaluation.id)
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
Returns Examples
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}