Create an evaluation together with its items, optionally running test criteria against them.
Accepts three request shapes: standalone (inline data), from an existing dataset
(dataset_id with optional per-item references), or with a new reusable dataset created inline
from data. When the evaluation includes tasks that require execution (for example an LLM judge
or custom function), an async job and a Temporal workflow are started and the evaluation is
returned immediately with status running; task results and error_count populate
asynchronously. When it includes only contributor tasks, taxonomy-only input, or no tasks, no
workflow runs and it is returned with status completed. Optional tasks, metadata, tags,
and taxonomy_params are persisted alongside the evaluation and its items.
ParametersExpand Collapse
body EvaluationNewParams
Evaluation param.Field[EvaluationNewParamsEvaluationUnion]
type EvaluationNewParamsEvaluationEvaluationStandaloneCreateRequest struct{…}
Tasks allow you to augment and evaluate your data
Tasks allow you to augment and evaluate your data
type EvaluationTaskChatCompletion struct{…}
Configuration EvaluationTaskChatCompletionConfiguration
Audio EvaluationTaskChatCompletionConfigurationAudioUnionOptionalParameters for audio output. Required when audio output is requested with modalities: [‘audio’].
Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].
FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnionOptionalDeprecated in favor of tool_choice. Controls which function is called by the model.
Deprecated in favor of tool_choice. Controls which function is called by the model.
Functions EvaluationTaskChatCompletionConfigurationFunctionsUnionOptionalDeprecated in favor of tools. A list of functions the model may generate JSON inputs for.
Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.
LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnionOptionalModify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnionOptionalWhether to return log probabilities of the output tokens or not.
Whether to return log probabilities of the output tokens or not.
MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnionOptionalAn upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnionOptionalDeprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Metadata EvaluationTaskChatCompletionConfigurationMetadataUnionOptionalDeveloper-defined tags and values used for filtering completions in the dashboard.
Developer-defined tags and values used for filtering completions in the dashboard.
Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnionOptionalOutput types that you would like the model to generate for this request.
Output types that you would like the model to generate for this request.
N EvaluationTaskChatCompletionConfigurationNUnionOptionalHow many chat completion choices to generate for each input message.
How many chat completion choices to generate for each input message.
ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnionOptionalWhether to enable parallel function calling during tool use.
Whether to enable parallel function calling during tool use.
Prediction EvaluationTaskChatCompletionConfigurationPredictionUnionOptionalStatic predicted output content, such as the content of a text file being regenerated.
Static predicted output content, such as the content of a text file being regenerated.
PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
For o1 models only. Constrains effort on reasoning. Values: low, medium, high.
ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnionOptionalAn object specifying the format that the model must output.
An object specifying the format that the model must output.
Seed EvaluationTaskChatCompletionConfigurationSeedUnionOptionalIf specified, system will attempt to sample deterministically for repeated requests with same seed.
If specified, system will attempt to sample deterministically for repeated requests with same seed.
Stop EvaluationTaskChatCompletionConfigurationStopUnionOptionalUp to 4 sequences where the API will stop generating further tokens.
Up to 4 sequences where the API will stop generating further tokens.
Store EvaluationTaskChatCompletionConfigurationStoreUnionOptionalWhether to store the output for use in model distillation or evals products.
Whether to store the output for use in model distillation or evals products.
Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnionOptionalWhat sampling temperature to use. Higher values make output more random, lower more focused.
What sampling temperature to use. Higher values make output more random, lower more focused.
ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnionOptionalControls which tool is called by the model. Values: none, auto, required, or specific tool.
Controls which tool is called by the model. Values: none, auto, required, or specific tool.
Tools EvaluationTaskChatCompletionConfigurationToolsUnionOptionalA list of tools the model may call. Currently, only functions are supported. Max 128 functions.
A list of tools the model may call. Currently, only functions are supported. Max 128 functions.
TopK EvaluationTaskChatCompletionConfigurationTopKUnionOptionalOnly sample from the top K options for each subsequent token
Only sample from the top K options for each subsequent token
type EvaluationTaskInference struct{…}
type EvaluationTaskApplicationVariant struct{…}
Configuration EvaluationTaskApplicationVariantConfiguration
Inputs EvaluationTaskApplicationVariantConfigurationInputsUnionInput data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
History EvaluationTaskApplicationVariantConfigurationHistoryUnionOptionalHistory of the application
History of the application
OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnionOptionalArbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnionOptionalOptional overrides for the application
Optional overrides for the application
type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}Execution override options for agentic applications
Execution override options for agentic applications
type EvaluationTaskAgentexOutput struct{…}
Configuration EvaluationTaskAgentexOutputConfiguration
InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnionThe dataset column to use as input for the agent
The dataset column to use as input for the agent
AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnionOptionalExtra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
CompletionMode stringOptionalHow the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.
IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnionOptionalWhether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
InputMode stringOptionalHow the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnionOptionalSeconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
type EvaluationTaskMetric struct{…}
type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}
Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}
type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskAutoEvaluationAgent struct{…}
Configuration AutoEvaluationAgentTaskRequestWithItemLocator
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskContributorEvaluationQuestion struct{…}
Configuration EvaluationTaskContributorEvaluationQuestionConfiguration
Layout Container
Children []ContainerChildUnionThe children to be displayed within the container
The children to be displayed within the container
type Component struct{…}
A pointer to the data in each evaluation item to be displayed within the component
Direction ContainerDirectionOptionalThe axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
type EvaluationTaskCustomFunction struct{…}
Configuration EvaluationTaskCustomFunctionConfigurationConfiguration for a custom Python function evaluation task.
Configuration for a custom Python function evaluation task.
Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.
Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.
type EvaluationNewParamsEvaluationEvaluationFromDatasetCreateRequest struct{…}
Data []EvaluationNewParamsEvaluationEvaluationFromDatasetCreateRequestDataOptionalItems to be evaluated, including references to the input dataset
Items to be evaluated, including references to the input dataset
Tasks allow you to augment and evaluate your data
Tasks allow you to augment and evaluate your data
type EvaluationTaskChatCompletion struct{…}
Configuration EvaluationTaskChatCompletionConfiguration
Audio EvaluationTaskChatCompletionConfigurationAudioUnionOptionalParameters for audio output. Required when audio output is requested with modalities: [‘audio’].
Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].
FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnionOptionalDeprecated in favor of tool_choice. Controls which function is called by the model.
Deprecated in favor of tool_choice. Controls which function is called by the model.
Functions EvaluationTaskChatCompletionConfigurationFunctionsUnionOptionalDeprecated in favor of tools. A list of functions the model may generate JSON inputs for.
Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.
LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnionOptionalModify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnionOptionalWhether to return log probabilities of the output tokens or not.
Whether to return log probabilities of the output tokens or not.
MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnionOptionalAn upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnionOptionalDeprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Metadata EvaluationTaskChatCompletionConfigurationMetadataUnionOptionalDeveloper-defined tags and values used for filtering completions in the dashboard.
Developer-defined tags and values used for filtering completions in the dashboard.
Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnionOptionalOutput types that you would like the model to generate for this request.
Output types that you would like the model to generate for this request.
N EvaluationTaskChatCompletionConfigurationNUnionOptionalHow many chat completion choices to generate for each input message.
How many chat completion choices to generate for each input message.
ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnionOptionalWhether to enable parallel function calling during tool use.
Whether to enable parallel function calling during tool use.
Prediction EvaluationTaskChatCompletionConfigurationPredictionUnionOptionalStatic predicted output content, such as the content of a text file being regenerated.
Static predicted output content, such as the content of a text file being regenerated.
PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
For o1 models only. Constrains effort on reasoning. Values: low, medium, high.
ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnionOptionalAn object specifying the format that the model must output.
An object specifying the format that the model must output.
Seed EvaluationTaskChatCompletionConfigurationSeedUnionOptionalIf specified, system will attempt to sample deterministically for repeated requests with same seed.
If specified, system will attempt to sample deterministically for repeated requests with same seed.
Stop EvaluationTaskChatCompletionConfigurationStopUnionOptionalUp to 4 sequences where the API will stop generating further tokens.
Up to 4 sequences where the API will stop generating further tokens.
Store EvaluationTaskChatCompletionConfigurationStoreUnionOptionalWhether to store the output for use in model distillation or evals products.
Whether to store the output for use in model distillation or evals products.
Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnionOptionalWhat sampling temperature to use. Higher values make output more random, lower more focused.
What sampling temperature to use. Higher values make output more random, lower more focused.
ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnionOptionalControls which tool is called by the model. Values: none, auto, required, or specific tool.
Controls which tool is called by the model. Values: none, auto, required, or specific tool.
Tools EvaluationTaskChatCompletionConfigurationToolsUnionOptionalA list of tools the model may call. Currently, only functions are supported. Max 128 functions.
A list of tools the model may call. Currently, only functions are supported. Max 128 functions.
TopK EvaluationTaskChatCompletionConfigurationTopKUnionOptionalOnly sample from the top K options for each subsequent token
Only sample from the top K options for each subsequent token
type EvaluationTaskInference struct{…}
type EvaluationTaskApplicationVariant struct{…}
Configuration EvaluationTaskApplicationVariantConfiguration
Inputs EvaluationTaskApplicationVariantConfigurationInputsUnionInput data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
History EvaluationTaskApplicationVariantConfigurationHistoryUnionOptionalHistory of the application
History of the application
OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnionOptionalArbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnionOptionalOptional overrides for the application
Optional overrides for the application
type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}Execution override options for agentic applications
Execution override options for agentic applications
type EvaluationTaskAgentexOutput struct{…}
Configuration EvaluationTaskAgentexOutputConfiguration
InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnionThe dataset column to use as input for the agent
The dataset column to use as input for the agent
AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnionOptionalExtra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
CompletionMode stringOptionalHow the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.
IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnionOptionalWhether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
InputMode stringOptionalHow the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnionOptionalSeconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
type EvaluationTaskMetric struct{…}
type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}
Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}
type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskAutoEvaluationAgent struct{…}
Configuration AutoEvaluationAgentTaskRequestWithItemLocator
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskContributorEvaluationQuestion struct{…}
Configuration EvaluationTaskContributorEvaluationQuestionConfiguration
Layout Container
Children []ContainerChildUnionThe children to be displayed within the container
The children to be displayed within the container
type Component struct{…}
A pointer to the data in each evaluation item to be displayed within the component
Direction ContainerDirectionOptionalThe axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
type EvaluationTaskCustomFunction struct{…}
Configuration EvaluationTaskCustomFunctionConfigurationConfiguration for a custom Python function evaluation task.
Configuration for a custom Python function evaluation task.
Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.
Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.
type EvaluationNewParamsEvaluationEvaluationWithDatasetCreateRequest struct{…}
Dataset EvaluationNewParamsEvaluationEvaluationWithDatasetCreateRequestDatasetCreate a reusable dataset from items in the data field
Create a reusable dataset from items in the data field
Tasks allow you to augment and evaluate your data
Tasks allow you to augment and evaluate your data
type EvaluationTaskChatCompletion struct{…}
Configuration EvaluationTaskChatCompletionConfiguration
Audio EvaluationTaskChatCompletionConfigurationAudioUnionOptionalParameters for audio output. Required when audio output is requested with modalities: [‘audio’].
Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].
FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnionOptionalDeprecated in favor of tool_choice. Controls which function is called by the model.
Deprecated in favor of tool_choice. Controls which function is called by the model.
Functions EvaluationTaskChatCompletionConfigurationFunctionsUnionOptionalDeprecated in favor of tools. A list of functions the model may generate JSON inputs for.
Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.
LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnionOptionalModify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnionOptionalWhether to return log probabilities of the output tokens or not.
Whether to return log probabilities of the output tokens or not.
MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnionOptionalAn upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnionOptionalDeprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Metadata EvaluationTaskChatCompletionConfigurationMetadataUnionOptionalDeveloper-defined tags and values used for filtering completions in the dashboard.
Developer-defined tags and values used for filtering completions in the dashboard.
Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnionOptionalOutput types that you would like the model to generate for this request.
Output types that you would like the model to generate for this request.
N EvaluationTaskChatCompletionConfigurationNUnionOptionalHow many chat completion choices to generate for each input message.
How many chat completion choices to generate for each input message.
ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnionOptionalWhether to enable parallel function calling during tool use.
Whether to enable parallel function calling during tool use.
Prediction EvaluationTaskChatCompletionConfigurationPredictionUnionOptionalStatic predicted output content, such as the content of a text file being regenerated.
Static predicted output content, such as the content of a text file being regenerated.
PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
For o1 models only. Constrains effort on reasoning. Values: low, medium, high.
ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnionOptionalAn object specifying the format that the model must output.
An object specifying the format that the model must output.
Seed EvaluationTaskChatCompletionConfigurationSeedUnionOptionalIf specified, system will attempt to sample deterministically for repeated requests with same seed.
If specified, system will attempt to sample deterministically for repeated requests with same seed.
Stop EvaluationTaskChatCompletionConfigurationStopUnionOptionalUp to 4 sequences where the API will stop generating further tokens.
Up to 4 sequences where the API will stop generating further tokens.
Store EvaluationTaskChatCompletionConfigurationStoreUnionOptionalWhether to store the output for use in model distillation or evals products.
Whether to store the output for use in model distillation or evals products.
Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnionOptionalWhat sampling temperature to use. Higher values make output more random, lower more focused.
What sampling temperature to use. Higher values make output more random, lower more focused.
ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnionOptionalControls which tool is called by the model. Values: none, auto, required, or specific tool.
Controls which tool is called by the model. Values: none, auto, required, or specific tool.
Tools EvaluationTaskChatCompletionConfigurationToolsUnionOptionalA list of tools the model may call. Currently, only functions are supported. Max 128 functions.
A list of tools the model may call. Currently, only functions are supported. Max 128 functions.
TopK EvaluationTaskChatCompletionConfigurationTopKUnionOptionalOnly sample from the top K options for each subsequent token
Only sample from the top K options for each subsequent token
type EvaluationTaskInference struct{…}
type EvaluationTaskApplicationVariant struct{…}
Configuration EvaluationTaskApplicationVariantConfiguration
Inputs EvaluationTaskApplicationVariantConfigurationInputsUnionInput data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
History EvaluationTaskApplicationVariantConfigurationHistoryUnionOptionalHistory of the application
History of the application
OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnionOptionalArbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnionOptionalOptional overrides for the application
Optional overrides for the application
type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}Execution override options for agentic applications
Execution override options for agentic applications
type EvaluationTaskAgentexOutput struct{…}
Configuration EvaluationTaskAgentexOutputConfiguration
InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnionThe dataset column to use as input for the agent
The dataset column to use as input for the agent
AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnionOptionalExtra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
CompletionMode stringOptionalHow the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.
IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnionOptionalWhether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
InputMode stringOptionalHow the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnionOptionalSeconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
type EvaluationTaskMetric struct{…}
type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}
Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}
type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskAutoEvaluationAgent struct{…}
Configuration AutoEvaluationAgentTaskRequestWithItemLocator
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskContributorEvaluationQuestion struct{…}
Configuration EvaluationTaskContributorEvaluationQuestionConfiguration
Layout Container
Children []ContainerChildUnionThe children to be displayed within the container
The children to be displayed within the container
type Component struct{…}
A pointer to the data in each evaluation item to be displayed within the component
Direction ContainerDirectionOptionalThe axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
type EvaluationTaskCustomFunction struct{…}
Configuration EvaluationTaskCustomFunctionConfigurationConfiguration for a custom Python function evaluation task.
Configuration for a custom Python function evaluation task.
Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.
Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.
ReturnsExpand Collapse
type Evaluation struct{…}
CreatedBy IdentityThe identity that created the entity.
The identity that created the entity.
Tasks executed during evaluation. Populated with optional task view.
Tasks executed during evaluation. Populated with optional task view.
type EvaluationTaskChatCompletion struct{…}
Configuration EvaluationTaskChatCompletionConfiguration
Audio EvaluationTaskChatCompletionConfigurationAudioUnionOptionalParameters for audio output. Required when audio output is requested with modalities: [‘audio’].
Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].
FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnionOptionalDeprecated in favor of tool_choice. Controls which function is called by the model.
Deprecated in favor of tool_choice. Controls which function is called by the model.
Functions EvaluationTaskChatCompletionConfigurationFunctionsUnionOptionalDeprecated in favor of tools. A list of functions the model may generate JSON inputs for.
Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.
LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnionOptionalModify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnionOptionalWhether to return log probabilities of the output tokens or not.
Whether to return log probabilities of the output tokens or not.
MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnionOptionalAn upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnionOptionalDeprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
Metadata EvaluationTaskChatCompletionConfigurationMetadataUnionOptionalDeveloper-defined tags and values used for filtering completions in the dashboard.
Developer-defined tags and values used for filtering completions in the dashboard.
Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnionOptionalOutput types that you would like the model to generate for this request.
Output types that you would like the model to generate for this request.
N EvaluationTaskChatCompletionConfigurationNUnionOptionalHow many chat completion choices to generate for each input message.
How many chat completion choices to generate for each input message.
ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnionOptionalWhether to enable parallel function calling during tool use.
Whether to enable parallel function calling during tool use.
Prediction EvaluationTaskChatCompletionConfigurationPredictionUnionOptionalStatic predicted output content, such as the content of a text file being regenerated.
Static predicted output content, such as the content of a text file being regenerated.
PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnionOptionalNumber between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
For o1 models only. Constrains effort on reasoning. Values: low, medium, high.
ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnionOptionalAn object specifying the format that the model must output.
An object specifying the format that the model must output.
Seed EvaluationTaskChatCompletionConfigurationSeedUnionOptionalIf specified, system will attempt to sample deterministically for repeated requests with same seed.
If specified, system will attempt to sample deterministically for repeated requests with same seed.
Stop EvaluationTaskChatCompletionConfigurationStopUnionOptionalUp to 4 sequences where the API will stop generating further tokens.
Up to 4 sequences where the API will stop generating further tokens.
Store EvaluationTaskChatCompletionConfigurationStoreUnionOptionalWhether to store the output for use in model distillation or evals products.
Whether to store the output for use in model distillation or evals products.
Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnionOptionalWhat sampling temperature to use. Higher values make output more random, lower more focused.
What sampling temperature to use. Higher values make output more random, lower more focused.
ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnionOptionalControls which tool is called by the model. Values: none, auto, required, or specific tool.
Controls which tool is called by the model. Values: none, auto, required, or specific tool.
Tools EvaluationTaskChatCompletionConfigurationToolsUnionOptionalA list of tools the model may call. Currently, only functions are supported. Max 128 functions.
A list of tools the model may call. Currently, only functions are supported. Max 128 functions.
TopK EvaluationTaskChatCompletionConfigurationTopKUnionOptionalOnly sample from the top K options for each subsequent token
Only sample from the top K options for each subsequent token
type EvaluationTaskInference struct{…}
type EvaluationTaskApplicationVariant struct{…}
Configuration EvaluationTaskApplicationVariantConfiguration
Inputs EvaluationTaskApplicationVariantConfigurationInputsUnionInput data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
History EvaluationTaskApplicationVariantConfigurationHistoryUnionOptionalHistory of the application
History of the application
OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnionOptionalArbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnionOptionalOptional overrides for the application
Optional overrides for the application
type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}Execution override options for agentic applications
Execution override options for agentic applications
type EvaluationTaskAgentexOutput struct{…}
Configuration EvaluationTaskAgentexOutputConfiguration
InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnionThe dataset column to use as input for the agent
The dataset column to use as input for the agent
AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnionOptionalExtra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
CompletionMode stringOptionalHow the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.
IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnionOptionalWhether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
InputMode stringOptionalHow the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnionOptionalSeconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
type EvaluationTaskMetric struct{…}
type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}
Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}
RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnionOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}
type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskAutoEvaluationAgent struct{…}
Configuration AutoEvaluationAgentTaskRequestWithItemLocator
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
type EvaluationTaskContributorEvaluationQuestion struct{…}
Configuration EvaluationTaskContributorEvaluationQuestionConfiguration
Layout Container
Children []ContainerChildUnionThe children to be displayed within the container
The children to be displayed within the container
type Component struct{…}
A pointer to the data in each evaluation item to be displayed within the component
Direction ContainerDirectionOptionalThe axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
type EvaluationTaskCustomFunction struct{…}
Configuration EvaluationTaskCustomFunctionConfigurationConfiguration for a custom Python function evaluation task.
Configuration for a custom Python function evaluation task.
Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.
Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.
Create Evaluation
package main
import (
"context"
"fmt"
"github.com/scaleapi/sgp-dev-go"
"github.com/scaleapi/sgp-dev-go/option"
)
func main() {
client := sgpdev.NewClient(
option.WithAPIKey("My API Key"),
option.WithAccountID("My Account ID"),
)
evaluation, err := client.Evaluations.New(context.TODO(), sgpdev.EvaluationNewParams{
OfEvaluationStandaloneCreateRequest: &sgpdev.EvaluationNewParamsEvaluationEvaluationStandaloneCreateRequest{
Data: []map[string]any{map[string]any{
"foo": "bar",
}},
Name: "x",
},
})
if err != nil {
panic(err.Error())
}
fmt.Printf("%+v\n", evaluation.ID)
}
{
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"datasets": [
{
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"current_version_num": 0,
"name": "name",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"object": "dataset"
}
],
"name": "name",
"status": "failed",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"error_count": 0,
"metadata": {
"foo": "bar"
},
"object": "evaluation",
"progress": {
"items": {
"failed": 0,
"pending": 0,
"successful": 0,
"total": 0,
"failed_items": [
{
"item_id": "item_id",
"error": "error",
"error_type": "error_type"
}
]
},
"workflows": {
"completed": 0,
"failed": 0,
"pending": 0,
"total": 0
}
},
"status_reason": "status_reason",
"tasks": [
{
"configuration": {
"messages": [
{
"foo": "bar"
}
],
"model": "model",
"audio": {
"foo": "bar"
},
"frequency_penalty": -2,
"function_call": {
"foo": "bar"
},
"functions": [
{
"foo": "bar"
}
],
"logit_bias": {
"foo": 0
},
"logprobs": true,
"max_completion_tokens": 0,
"max_tokens": 0,
"metadata": {
"foo": "string"
},
"modalities": [
"string"
],
"n": 0,
"parallel_tool_calls": true,
"prediction": {
"foo": "bar"
},
"presence_penalty": -2,
"reasoning_effort": "reasoning_effort",
"response_format": {
"foo": "bar"
},
"seed": 0,
"stop": "string",
"store": true,
"temperature": 0,
"tool_choice": "string",
"tools": [
{
"foo": "bar"
}
],
"top_k": 0,
"top_logprobs": 0,
"top_p": 0
},
"alias": "alias",
"task_type": "chat_completion"
}
]
}Returns Examples
{
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"datasets": [
{
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"current_version_num": 0,
"name": "name",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"object": "dataset"
}
],
"name": "name",
"status": "failed",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"error_count": 0,
"metadata": {
"foo": "bar"
},
"object": "evaluation",
"progress": {
"items": {
"failed": 0,
"pending": 0,
"successful": 0,
"total": 0,
"failed_items": [
{
"item_id": "item_id",
"error": "error",
"error_type": "error_type"
}
]
},
"workflows": {
"completed": 0,
"failed": 0,
"pending": 0,
"total": 0
}
},
"status_reason": "status_reason",
"tasks": [
{
"configuration": {
"messages": [
{
"foo": "bar"
}
],
"model": "model",
"audio": {
"foo": "bar"
},
"frequency_penalty": -2,
"function_call": {
"foo": "bar"
},
"functions": [
{
"foo": "bar"
}
],
"logit_bias": {
"foo": 0
},
"logprobs": true,
"max_completion_tokens": 0,
"max_tokens": 0,
"metadata": {
"foo": "string"
},
"modalities": [
"string"
],
"n": 0,
"parallel_tool_calls": true,
"prediction": {
"foo": "bar"
},
"presence_penalty": -2,
"reasoning_effort": "reasoning_effort",
"response_format": {
"foo": "bar"
},
"seed": 0,
"stop": "string",
"store": true,
"temperature": 0,
"tool_choice": "string",
"tools": [
{
"foo": "bar"
}
],
"top_k": 0,
"top_logprobs": 0,
"top_p": 0
},
"alias": "alias",
"task_type": "chat_completion"
}
]
}