# Evaluations

## Create Evaluation

`client.Evaluations.New(ctx, body) (*Evaluation, error)`

**post** `/v5/evaluations`

Create an evaluation together with its items, optionally running test criteria against them.

Accepts three request shapes: standalone (inline `data`), from an existing dataset
(`dataset_id` with optional per-item references), or with a new reusable dataset created inline
from `data`. When the evaluation includes tasks that require execution (for example an LLM judge
or custom function), an async job and a Temporal workflow are started and the evaluation is
returned immediately with status `running`; task results and `error_count` populate
asynchronously. When it includes only contributor tasks, taxonomy-only input, or no tasks, no
workflow runs and it is returned with status `completed`. Optional `tasks`, `metadata`, `tags`,
and `taxonomy_params` are persisted alongside the evaluation and its items.

### Parameters

- `body EvaluationNewParams`

  - `Evaluation param.Field[EvaluationNewParamsEvaluationUnion]`

    - `type EvaluationNewParamsEvaluationEvaluationStandaloneCreateRequest struct{…}`

      - `Data []map[string, any]`

        Items to be evaluated

      - `Name string`

      - `Description string`

      - `Files []map[string, string]`

        Files to be associated to the evaluation

      - `Metadata map[string, any]`

        Optional metadata key-value pairs for the evaluation

      - `SkipPrefilledRows bool`

        Do not queue a contributor task for prefilled questions

      - `Tags []string`

        The tags associated with the evaluation

      - `Tasks []EvaluationTaskUnion`

        Tasks allow you to augment and evaluate your data

        - `type EvaluationTaskChatCompletion struct{…}`

          - `Configuration EvaluationTaskChatCompletionConfiguration`

            - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

              openai standard message format

              - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

              - `type ItemLocator string`

            - `Model string`

              model specified as `model_vendor/model`, for example `openai/gpt-4o`

            - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

              Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

              - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

              - `type ItemLocator string`

            - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

              Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

              - `float64`

              - `type ItemLocator string`

            - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

              Deprecated in favor of tool_choice. Controls which function is called by the model.

              - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

              - `type ItemLocator string`

            - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

              Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

              - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

              - `type ItemLocator string`

            - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

              Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

              - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

              - `type ItemLocator string`

            - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

              Whether to return log probabilities of the output tokens or not.

              - `bool`

              - `type ItemLocator string`

            - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

              An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

              - `int64`

              - `type ItemLocator string`

            - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

              Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

              - `int64`

              - `type ItemLocator string`

            - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

              Developer-defined tags and values used for filtering completions in the dashboard.

              - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

              - `type ItemLocator string`

            - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

              Output types that you would like the model to generate for this request.

              - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

              - `type ItemLocator string`

            - `N EvaluationTaskChatCompletionConfigurationNUnion`

              How many chat completion choices to generate for each input message.

              - `int64`

              - `type ItemLocator string`

            - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

              Whether to enable parallel function calling during tool use.

              - `bool`

              - `type ItemLocator string`

            - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

              Static predicted output content, such as the content of a text file being regenerated.

              - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

              - `type ItemLocator string`

            - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

              Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

              - `float64`

              - `type ItemLocator string`

            - `ReasoningEffort string`

              For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

            - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

              An object specifying the format that the model must output.

              - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

              - `type ItemLocator string`

            - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

              If specified, system will attempt to sample deterministically for repeated requests with same seed.

              - `int64`

              - `type ItemLocator string`

            - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

              Up to 4 sequences where the API will stop generating further tokens.

              - `string`

              - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

            - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

              Whether to store the output for use in model distillation or evals products.

              - `bool`

              - `type ItemLocator string`

            - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

              What sampling temperature to use. Higher values make output more random, lower more focused.

              - `float64`

              - `type ItemLocator string`

            - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

              Controls which tool is called by the model. Values: none, auto, required, or specific tool.

              - `string`

              - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

            - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

              A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

              - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

              - `type ItemLocator string`

            - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

              Only sample from the top K options for each subsequent token

              - `int64`

              - `type ItemLocator string`

            - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

              Number of most likely tokens to return at each position, with associated log probability.

              - `int64`

              - `type ItemLocator string`

            - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

              Alternative to temperature. Only tokens comprising top_p probability mass are considered.

              - `float64`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `chat_completion`

          - `TaskType string`

            - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

        - `type EvaluationTaskInference struct{…}`

          - `Configuration EvaluationTaskInferenceConfiguration`

            - `Model string`

              model specified as `vendor/name` (ex. openai/gpt-5)

            - `Args EvaluationTaskInferenceConfigurationArgsUnion`

              Arguments passed into model

              - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

              - `type ItemLocator string`

            - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

              Vendor specific configuration

              - `type LaunchInferenceConfiguration struct{…}`

                - `NumRetries int64`

                - `TimeoutSeconds int64`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `inference`

          - `TaskType string`

            - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

        - `type EvaluationTaskApplicationVariant struct{…}`

          - `Configuration EvaluationTaskApplicationVariantConfiguration`

            - `ApplicationVariantID string`

            - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

              Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

              - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

              - `type ItemLocator string`

            - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

              History of the application

              - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

                - `Request string`

                  Request inputs

                - `Response string`

                  Response outputs

                - `SessionData map[string, any]`

                  Session data corresponding to the request response pair

              - `type ItemLocator string`

            - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

              Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

              - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

              - `type ItemLocator string`

            - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

              Optional overrides for the application

              - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

                Execution override options for agentic applications

                - `Concurrent bool`

                - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

                  - `CurrentNode string`

                  - `State map[string, any]`

                - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

                  - `DurationMs int64`

                  - `NodeID string`

                  - `OperationInput string`

                  - `OperationOutput string`

                  - `OperationType string`

                  - `StartTimestamp string`

                  - `WorkflowID string`

                  - `OperationMetadata map[string, any]`

                - `ReturnSpan bool`

                - `UseChannels bool`

              - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

                - `ArtifactIDsFilter []string`

                - `ArtifactNameRegex []string`

                - `Type string`

                  - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `application_variant`

          - `TaskType string`

            - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

        - `type EvaluationTaskAgentexOutput struct{…}`

          - `Configuration EvaluationTaskAgentexOutputConfiguration`

            - `AgentexAgentID string`

              The ID of the Agentex agent to use

            - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

              The dataset column to use as input for the agent

              - `string`

              - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

              - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

            - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

              Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

              - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

              - `type ItemLocator string`

            - `CompletionMode string`

              How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

              - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

              - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

            - `DeploymentID string`

              Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

            - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

              Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

              - `bool`

              - `type ItemLocator string`

            - `InputMode string`

              How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

              - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

              - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

            - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

              Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

              - `int64`

              - `type ItemLocator string`

            - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

              Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

              - `int64`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `agentex_output`

          - `TaskType string`

            - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

        - `type EvaluationTaskMetric struct{…}`

          - `Configuration EvaluationTaskMetricConfigurationUnion`

            - `type EvaluationTaskMetricConfigurationBleu struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Bleu`

                - `const BleuBleu Bleu = "bleu"`

            - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Meteor`

                - `const MeteorMeteor Meteor = "meteor"`

            - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type CosineSimilarity`

                - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

            - `type EvaluationTaskMetricConfigurationF1 struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type F1`

                - `const F1F1 F1 = "f1"`

            - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Rouge1`

                - `const Rouge1Rouge1 Rouge1 = "rouge1"`

            - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Rouge2`

                - `const Rouge2Rouge2 Rouge2 = "rouge2"`

            - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type RougeL`

                - `const RougeLRougeL RougeL = "rougeL"`

          - `Alias string`

            Alias to title the results column. Defaults to the metric type specified in the configuration

          - `TaskType string`

            - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

        - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

          - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

            - `Model string`

              model specified as `model_vendor/model_name`

            - `Prompt string`

            - `QuestionID string`

              question to be evaluated

          - `Alias string`

            Alias to title the results column. Defaults to the `auto_evaluation_question`

          - `TaskType string`

            - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

        - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

          - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

              - `Model string`

                model specified as `model_vendor/model_name`

              - `Prompt string`

              - `ResponseFormat map[string, any]`

                JSON schema used for structuring the model response

              - `InferenceArgs map[string, any]`

                Additional arguments to pass to the inference request

              - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

                  - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                    - `string`

                    - `float64`

                    - `bool`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

                  - `Path string`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

                - `type EqEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op EqEvaluationRunConditionOp`

                    - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

                - `type NeEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op NeEvaluationRunConditionOp`

                    - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

                - `type LtEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op LtEvaluationRunConditionOp`

                    - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

                - `type LteEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op LteEvaluationRunConditionOp`

                    - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

                - `type GtEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op GtEvaluationRunConditionOp`

                    - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

                - `type GteEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op GteEvaluationRunConditionOp`

                    - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

                - `type AndEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op AndEvaluationRunConditionOp`

                    - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

                - `type OrEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op OrEvaluationRunConditionOp`

                    - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

                - `type InEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Operands []any`

                  - `Op InEvaluationRunConditionOp`

                    - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

                - `type NotInEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Operands []any`

                  - `Op NotInEvaluationRunConditionOp`

                    - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

                - `type NotEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op NotEvaluationRunConditionOp`

                    - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

                - `type IsNullEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op IsNullEvaluationRunConditionOp`

                    - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

                - `type IsNotNullEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op IsNotNullEvaluationRunConditionOp`

                    - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

              - `SystemPrompt string`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

              - `Choices []string`

                Choices array cannot be empty

              - `Model string`

                model specified as `model_vendor/model_name`

              - `Prompt string`

              - `InferenceArgs map[string, any]`

                Additional arguments to pass to the inference request

              - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

                  - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                    - `string`

                    - `float64`

                    - `bool`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

                  - `Path string`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

                - `type EqEvaluationRunCondition struct{…}`

                - `type NeEvaluationRunCondition struct{…}`

                - `type LtEvaluationRunCondition struct{…}`

                - `type LteEvaluationRunCondition struct{…}`

                - `type GtEvaluationRunCondition struct{…}`

                - `type GteEvaluationRunCondition struct{…}`

                - `type AndEvaluationRunCondition struct{…}`

                - `type OrEvaluationRunCondition struct{…}`

                - `type InEvaluationRunCondition struct{…}`

                - `type NotInEvaluationRunCondition struct{…}`

                - `type NotEvaluationRunCondition struct{…}`

                - `type IsNullEvaluationRunCondition struct{…}`

                - `type IsNotNullEvaluationRunCondition struct{…}`

              - `SystemPrompt string`

            - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

              - `Definition string`

              - `Name string`

              - `OutputRules []string`

              - `DataFields []string`

              - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                    - `Model string`

                    - `Temperature float64`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                    - `Model string`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                    - `Model string`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                    - `Model string`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

              - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

              - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

                - `string`

                - `float64`

                - `bool`

              - `RubricID string`

              - `RubricVersion int64`

          - `Alias string`

            Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

          - `TaskType string`

            - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

        - `type EvaluationTaskAutoEvaluationAgent struct{…}`

          - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

          - `Alias string`

            Alias to title the results column. Defaults to the `auto_evaluation_agent`

          - `TaskType string`

            - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

        - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

          - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

            - `Layout Container`

              - `Children []ContainerChildUnion`

                The children to be displayed within the container

                - `type Container struct{…}`

                - `type Component struct{…}`

                  - `Data ItemLocator`

                    A pointer to the data in each evaluation item to be displayed within the component

                  - `Label string`

              - `Direction ContainerDirection`

                The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

                - `const ContainerDirectionRow ContainerDirection = "row"`

                - `const ContainerDirectionColumn ContainerDirection = "column"`

            - `QuestionID string`

            - `PrefillFrom string`

              Dataset column to prefill contributor question task result

            - `QueueID string`

              The contributor annotation queue to include this task in. Defaults to `default`

            - `Required bool`

              Whether the question is required to be answered

            - `RubricID string`

              ID of the rubric to use for scoring this evaluation question

          - `Alias string`

            Alias to title the results column. Defaults to the `contributor_evaluation_question`

          - `TaskType string`

            - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

        - `type EvaluationTaskCustomFunction struct{…}`

          - `Configuration EvaluationTaskCustomFunctionConfiguration`

            Configuration for a custom Python function evaluation task.

            - `FunctionSource string`

              Python function source code

            - `ArgMapping map[string, string]`

              Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

            - `ConfigArgs map[string, any]`

              Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

            - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

              Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

              - `Path string`

                Dot path in the custom function return value to materialize.

              - `Alias string`

                Result column alias. Defaults to path with dots replaced by underscores.

          - `Alias string`

            Alias to title the results column. Defaults to the function name.

          - `TaskType string`

            - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

      - `TaxonomyParams map[string, any]`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

    - `type EvaluationNewParamsEvaluationEvaluationFromDatasetCreateRequest struct{…}`

      - `DatasetID string`

        The ID of the dataset containing the items referenced by the `data` field

      - `Name string`

      - `Data []EvaluationNewParamsEvaluationEvaluationFromDatasetCreateRequestData`

        Items to be evaluated, including references to the input dataset

        - `DatasetItemID string`

      - `Description string`

      - `Metadata map[string, any]`

        Optional metadata key-value pairs for the evaluation

      - `SkipPrefilledRows bool`

        Do not queue a contributor task for prefilled questions

      - `Tags []string`

        The tags associated with the evaluation

      - `Tasks []EvaluationTaskUnion`

        Tasks allow you to augment and evaluate your data

        - `type EvaluationTaskChatCompletion struct{…}`

        - `type EvaluationTaskInference struct{…}`

        - `type EvaluationTaskApplicationVariant struct{…}`

        - `type EvaluationTaskAgentexOutput struct{…}`

        - `type EvaluationTaskMetric struct{…}`

        - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

        - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

        - `type EvaluationTaskAutoEvaluationAgent struct{…}`

        - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

        - `type EvaluationTaskCustomFunction struct{…}`

      - `TaxonomyParams map[string, any]`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

    - `type EvaluationNewParamsEvaluationEvaluationWithDatasetCreateRequest struct{…}`

      - `Data []map[string, any]`

        Items to be evaluated

      - `Dataset EvaluationNewParamsEvaluationEvaluationWithDatasetCreateRequestDataset`

        Create a reusable dataset from items in the `data` field

        - `Name string`

        - `Description string`

        - `Keys []string`

          Keys from items in the `data` field that should be included in the dataset. If not provided, all keys will be included.

        - `Tags []string`

          The tags associated with the entity

      - `Name string`

      - `Description string`

      - `Files []map[string, string]`

        Files to be associated to the evaluation

      - `Metadata map[string, any]`

        Optional metadata key-value pairs for the evaluation

      - `SkipPrefilledRows bool`

        Do not queue a contributor task for prefilled questions

      - `Tags []string`

        The tags associated with the evaluation

      - `Tasks []EvaluationTaskUnion`

        Tasks allow you to augment and evaluate your data

        - `type EvaluationTaskChatCompletion struct{…}`

        - `type EvaluationTaskInference struct{…}`

        - `type EvaluationTaskApplicationVariant struct{…}`

        - `type EvaluationTaskAgentexOutput struct{…}`

        - `type EvaluationTaskMetric struct{…}`

        - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

        - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

        - `type EvaluationTaskAutoEvaluationAgent struct{…}`

        - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

        - `type EvaluationTaskCustomFunction struct{…}`

      - `TaxonomyParams map[string, any]`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluation, err := client.Evaluations.New(context.TODO(), sgpdev.EvaluationNewParams{
    OfEvaluationStandaloneCreateRequest: &sgpdev.EvaluationNewParamsEvaluationEvaluationStandaloneCreateRequest{
      Data: []map[string]any{map[string]any{
      "foo": "bar",
      }},
      Name: "x",
    },
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluation.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## List Evaluations

`client.Evaluations.List(ctx, query) (*CursorPage[Evaluation], error)`

**get** `/v5/evaluations`

List evaluations for the account, with pagination.

Supports filtering by case-insensitive name substring and by tags; archived
evaluations are excluded unless `include_archived` is set. Pass the `tasks` view to include each
evaluation's task configurations in the response. Use this for simple name or tag lookups;
to filter on metadata key-value pairs or status, use the filter endpoint instead.

### Parameters

- `query EvaluationListParams`

  - `EndingBefore param.Field[string]`

  - `IncludeArchived param.Field[bool]`

  - `Limit param.Field[int64]`

  - `Name param.Field[string]`

  - `SortBy param.Field[string]`

  - `SortOrder param.Field[SortOrder]`

  - `StartingAfter param.Field[string]`

  - `Tags param.Field[[]string]`

  - `Views param.Field[[]EvaluationViews]`

    - `const EvaluationViewsTasks EvaluationViews = "tasks"`

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  page, err := client.Evaluations.List(context.TODO(), sgpdev.EvaluationListParams{

  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", page)
}
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Get Evaluation

`client.Evaluations.Get(ctx, evaluationID, query) (*Evaluation, error)`

**get** `/v5/evaluations/{evaluation_id}`

Retrieve a single evaluation by ID.

Returns the evaluation with its datasets, async-job progress, metadata, and task-error count.
Archived evaluations are excluded unless `include_archived` is set. Pass the `tasks` view to
include the evaluation's task configurations in the response.

### Parameters

- `evaluationID string`

- `query EvaluationGetParams`

  - `IncludeArchived param.Field[bool]`

  - `Views param.Field[[]EvaluationViews]`

    - `const EvaluationViewsTasks EvaluationViews = "tasks"`

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluation, err := client.Evaluations.Get(
    context.TODO(),
    "evaluation_id",
    sgpdev.EvaluationGetParams{

    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluation.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Archive Evaluation

`client.Evaluations.Archive(ctx, evaluationID) (*Evaluation, error)`

**delete** `/v5/evaluations/{evaluation_id}`

Archive (soft-delete) an evaluation.

Sets the evaluation's archived timestamp rather than permanently deleting it, and cascades the
archive to the evaluation's items and dashboards while removing it from any evaluation groups.
The evaluation can later be brought back with a restore request to the update endpoint.

### Parameters

- `evaluationID string`

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluation, err := client.Evaluations.Archive(context.TODO(), "evaluation_id")
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluation.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Update or Restore Evaluation

`client.Evaluations.Update(ctx, evaluationID, body) (*Evaluation, error)`

**patch** `/v5/evaluations/{evaluation_id}`

Update an evaluation's mutable fields, or restore it from the archive.

The action is selected by the request body: a restore request un-archives the evaluation and
cascades the restore to its items and dashboards, while any other body applies a partial update
to fields such as name, description, tags, and metadata (metadata is applied as an RFC 7396
merge patch). Updating an already-archived evaluation is rejected — restore it first. The
evaluation row is locked for the duration of the write to avoid concurrent-update races.

### Parameters

- `evaluationID string`

- `body EvaluationUpdateParams`

  - `Evaluation param.Field[EvaluationUpdateParamsEvaluationUnion]`

    - `type EvaluationUpdateParamsEvaluationPartialEvaluationUpdateRequest struct{…}`

      - `Description string`

      - `Metadata map[string, any]`

        Optional metadata key-value pairs for the evaluation

      - `Name string`

      - `Tags []string`

        The tags associated with the evaluation

    - `type RestoreRequestParam struct{…}`

      - `Restore bool`

        Set to true to restore the entity from the database.

        - `const RestoreRequestRestoreTrue RestoreRequestRestore = true`

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluation, err := client.Evaluations.Update(
    context.TODO(),
    "evaluation_id",
    sgpdev.EvaluationUpdateParams{
      OfPartialEvaluationUpdateRequest: &sgpdev.EvaluationUpdateParamsEvaluationPartialEvaluationUpdateRequest{

      },
    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluation.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Get Evaluation Data Schema

`client.Evaluations.GetSchema(ctx, evaluationID, query) (*EvaluationSchemaResponse, error)`

**get** `/v5/evaluations/{evaluation_id}/schema`

Describe the data schema of an evaluation's items.

Inspects the item `data` and task-result fields and returns each discovered field with its
flattened key path, JSON type, source, and the number of items containing it, ordered
alphabetically by field name. For large evaluations the schema may be inferred from a sample of
items, in which case `is_sampled` is set and `sample_size` reports how many were analyzed. Set
`include_archived` to include archived items in the analysis.

### Parameters

- `evaluationID string`

- `query EvaluationGetSchemaParams`

  - `IncludeArchived param.Field[bool]`

    Include archived items in schema analysis

### Returns

- `type EvaluationSchemaResponse struct{…}`

  Schema information for an evaluation's item data structure

  - `EvaluationID string`

    The ID of the evaluation

  - `Fields []EvaluationSchemaResponseField`

    List of all discovered fields, ordered alphabetically by field_name

    - `DataType string`

      JSON type: 'string', 'number', 'boolean', 'object', 'array', or 'null'

    - `FieldName string`

      The flattened JSON key path (e.g., 'metadata.category')

    - `ItemCount int64`

      Number of evaluation items containing this field

    - `Source string`

      The source of the field: 'data' or 'task_result_cache'

      - `const EvaluationSchemaResponseFieldSourceData EvaluationSchemaResponseFieldSource = "data"`

      - `const EvaluationSchemaResponseFieldSourceTaskResultCache EvaluationSchemaResponseFieldSource = "task_result_cache"`

    - `Object string`

      - `const EvaluationSchemaResponseFieldObjectFieldSchema EvaluationSchemaResponseFieldObject = "field_schema"`

  - `TotalItems int64`

    Total number of evaluation items

  - `IsSampled bool`

    Whether schema was computed from a sample of items (for large evaluations)

  - `Object EvaluationSchemaResponseObject`

    - `const EvaluationSchemaResponseObjectEvaluationSchema EvaluationSchemaResponseObject = "evaluation_schema"`

  - `SampleSize int64`

    Number of items sampled for schema inference, if applicable

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluationSchemaResponse, err := client.Evaluations.GetSchema(
    context.TODO(),
    "evaluation_id",
    sgpdev.EvaluationGetSchemaParams{

    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluationSchemaResponse.EvaluationID)
}
```

#### Response

```json
{
  "evaluation_id": "evaluation_id",
  "fields": [
    {
      "data_type": "data_type",
      "field_name": "field_name",
      "item_count": 0,
      "source": "data",
      "object": "field_schema"
    }
  ],
  "total_items": 0,
  "is_sampled": true,
  "object": "evaluation_schema",
  "sample_size": 0
}
```

## Filter Evaluations

`client.Evaluations.Filter(ctx, params) (*CursorPage[Evaluation], error)`

**post** `/v5/evaluations/filter`

Filter evaluations by metadata, status, and tags.

Accepts up to 10 filters combined with AND logic, each comparing a key against a value with an
operator (`==`, `!=`, `>=`, `<=`, `IN`, `NOT_IN`). Filter on metadata keys returned by the
metadata-keys endpoint, plus the built-in `status` and `tag` keys. Archived evaluations are
excluded unless `include_archived` is set, and the `tasks` view includes task configurations in
each result. Use this for metadata or status filtering; for simple name or tag lookups the list
endpoint is sufficient.

### Parameters

- `params EvaluationFilterParams`

  - `Filters param.Field[[]EvaluationFilterParamsFilter]`

    Body param: List of metadata filters to apply (maximum 10)

    - `Key string`

      The metadata key to filter on

    - `Operator string`

      The comparison operator to use

      - `const EvaluationFilterParamsFilterOperatorEquals EvaluationFilterParamsFilterOperator = "=="`

      - `const EvaluationFilterParamsFilterOperatorNotEquals EvaluationFilterParamsFilterOperator = "!="`

      - `const EvaluationFilterParamsFilterOperatorGreaterOrEquals EvaluationFilterParamsFilterOperator = ">="`

      - `const EvaluationFilterParamsFilterOperatorLessOrEquals EvaluationFilterParamsFilterOperator = "<="`

      - `const EvaluationFilterParamsFilterOperatorIn EvaluationFilterParamsFilterOperator = "IN"`

      - `const EvaluationFilterParamsFilterOperatorNotIn EvaluationFilterParamsFilterOperator = "NOT_IN"`

    - `Value string`

      The value to compare against (string for all types)

    - `Object string`

      - `const EvaluationFilterParamsFilterObjectMetadataFilter EvaluationFilterParamsFilterObject = "metadata_filter"`

  - `EndingBefore param.Field[string]`

    Query param

  - `IncludeArchived param.Field[bool]`

    Query param

  - `Limit param.Field[int64]`

    Query param

  - `SortBy param.Field[string]`

    Query param

  - `SortOrder param.Field[SortOrder]`

    Query param

  - `StartingAfter param.Field[string]`

    Query param

  - `Views param.Field[[]EvaluationViews]`

    Query param

    - `const EvaluationViewsTasks EvaluationViews = "tasks"`

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  page, err := client.Evaluations.Filter(context.TODO(), sgpdev.EvaluationFilterParams{
    Filters: []sgpdev.EvaluationFilterParamsFilter{sgpdev.EvaluationFilterParamsFilter{
      Key: "key",
      Operator: "==",
      Value: "value",
    }},
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", page)
}
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Get Evaluation Taxonomy

`client.Evaluations.GetTaxonomy(ctx, evaluationID) (*EvaluationGetTaxonomyResponse, error)`

**get** `/v5/evaluations/{evaluation_id}/taxonomy`

Get the taxonomy JSON for an evaluation's contributor question tasks.

Returns the raw taxonomy document stored for the evaluation. Responds with a not-found error if
the evaluation has no taxonomy.

### Parameters

- `evaluationID string`

### Returns

- `type EvaluationGetTaxonomyResponse map[string, any]`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  response, err := client.Evaluations.GetTaxonomy(context.TODO(), "evaluation_id")
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", response)
}
```

#### Response

```json
{
  "foo": "bar"
}
```

## Domain Types

### And Evaluation Run Condition

- `type AndEvaluationRunCondition struct{…}`

  - `Operands []any`

  - `Op AndEvaluationRunConditionOp`

    - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

### Auto Evaluation Agent Task Request With Item Locator

- `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

  - `Definition string`

  - `Name string`

  - `OutputRules []string`

  - `DataFields []string`

  - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

    - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

      - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

        - `Model string`

        - `Temperature float64`

      - `AgentName string`

        - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

    - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

      - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

        - `Model string`

      - `AgentName string`

        - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

    - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

      - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

        - `Model string`

      - `AgentName string`

        - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

    - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

      - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

        - `Model string`

      - `AgentName string`

        - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

  - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

    - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

    - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

    - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

    - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

  - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

    - `string`

    - `float64`

    - `bool`

  - `RubricID string`

  - `RubricVersion int64`

### Eq Evaluation Run Condition

- `type EqEvaluationRunCondition struct{…}`

  - `Left any`

  - `Right any`

  - `Op EqEvaluationRunConditionOp`

    - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

### Evaluation

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Evaluation Schema Response

- `type EvaluationSchemaResponse struct{…}`

  Schema information for an evaluation's item data structure

  - `EvaluationID string`

    The ID of the evaluation

  - `Fields []EvaluationSchemaResponseField`

    List of all discovered fields, ordered alphabetically by field_name

    - `DataType string`

      JSON type: 'string', 'number', 'boolean', 'object', 'array', or 'null'

    - `FieldName string`

      The flattened JSON key path (e.g., 'metadata.category')

    - `ItemCount int64`

      Number of evaluation items containing this field

    - `Source string`

      The source of the field: 'data' or 'task_result_cache'

      - `const EvaluationSchemaResponseFieldSourceData EvaluationSchemaResponseFieldSource = "data"`

      - `const EvaluationSchemaResponseFieldSourceTaskResultCache EvaluationSchemaResponseFieldSource = "task_result_cache"`

    - `Object string`

      - `const EvaluationSchemaResponseFieldObjectFieldSchema EvaluationSchemaResponseFieldObject = "field_schema"`

  - `TotalItems int64`

    Total number of evaluation items

  - `IsSampled bool`

    Whether schema was computed from a sample of items (for large evaluations)

  - `Object EvaluationSchemaResponseObject`

    - `const EvaluationSchemaResponseObjectEvaluationSchema EvaluationSchemaResponseObject = "evaluation_schema"`

  - `SampleSize int64`

    Number of items sampled for schema inference, if applicable

### Evaluation Task

- `type EvaluationTaskUnion interface{…}`

  - `type EvaluationTaskChatCompletion struct{…}`

    - `Configuration EvaluationTaskChatCompletionConfiguration`

      - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

        openai standard message format

        - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

        - `type ItemLocator string`

      - `Model string`

        model specified as `model_vendor/model`, for example `openai/gpt-4o`

      - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

        Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

        - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

        - `type ItemLocator string`

      - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

        Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

        - `float64`

        - `type ItemLocator string`

      - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

        Deprecated in favor of tool_choice. Controls which function is called by the model.

        - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

        - `type ItemLocator string`

      - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

        Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

        - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

        - `type ItemLocator string`

      - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

        Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

        - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

        - `type ItemLocator string`

      - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

        Whether to return log probabilities of the output tokens or not.

        - `bool`

        - `type ItemLocator string`

      - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

        An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

        - `int64`

        - `type ItemLocator string`

      - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

        Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

        - `int64`

        - `type ItemLocator string`

      - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

        Developer-defined tags and values used for filtering completions in the dashboard.

        - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

        - `type ItemLocator string`

      - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

        Output types that you would like the model to generate for this request.

        - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

        - `type ItemLocator string`

      - `N EvaluationTaskChatCompletionConfigurationNUnion`

        How many chat completion choices to generate for each input message.

        - `int64`

        - `type ItemLocator string`

      - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

        Whether to enable parallel function calling during tool use.

        - `bool`

        - `type ItemLocator string`

      - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

        Static predicted output content, such as the content of a text file being regenerated.

        - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

        - `type ItemLocator string`

      - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

        Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

        - `float64`

        - `type ItemLocator string`

      - `ReasoningEffort string`

        For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

      - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

        An object specifying the format that the model must output.

        - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

        - `type ItemLocator string`

      - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

        If specified, system will attempt to sample deterministically for repeated requests with same seed.

        - `int64`

        - `type ItemLocator string`

      - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

        Up to 4 sequences where the API will stop generating further tokens.

        - `string`

        - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

      - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

        Whether to store the output for use in model distillation or evals products.

        - `bool`

        - `type ItemLocator string`

      - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

        What sampling temperature to use. Higher values make output more random, lower more focused.

        - `float64`

        - `type ItemLocator string`

      - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

        Controls which tool is called by the model. Values: none, auto, required, or specific tool.

        - `string`

        - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

      - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

        A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

        - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

        - `type ItemLocator string`

      - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

        Only sample from the top K options for each subsequent token

        - `int64`

        - `type ItemLocator string`

      - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

        Number of most likely tokens to return at each position, with associated log probability.

        - `int64`

        - `type ItemLocator string`

      - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

        Alternative to temperature. Only tokens comprising top_p probability mass are considered.

        - `float64`

        - `type ItemLocator string`

    - `Alias string`

      Alias to title the results column. Defaults to the `chat_completion`

    - `TaskType string`

      - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

  - `type EvaluationTaskInference struct{…}`

    - `Configuration EvaluationTaskInferenceConfiguration`

      - `Model string`

        model specified as `vendor/name` (ex. openai/gpt-5)

      - `Args EvaluationTaskInferenceConfigurationArgsUnion`

        Arguments passed into model

        - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

        - `type ItemLocator string`

      - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

        Vendor specific configuration

        - `type LaunchInferenceConfiguration struct{…}`

          - `NumRetries int64`

          - `TimeoutSeconds int64`

        - `type ItemLocator string`

    - `Alias string`

      Alias to title the results column. Defaults to the `inference`

    - `TaskType string`

      - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

  - `type EvaluationTaskApplicationVariant struct{…}`

    - `Configuration EvaluationTaskApplicationVariantConfiguration`

      - `ApplicationVariantID string`

      - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

        Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

        - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

        - `type ItemLocator string`

      - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

        History of the application

        - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

          - `Request string`

            Request inputs

          - `Response string`

            Response outputs

          - `SessionData map[string, any]`

            Session data corresponding to the request response pair

        - `type ItemLocator string`

      - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

        Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

        - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

        - `type ItemLocator string`

      - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

        Optional overrides for the application

        - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

          Execution override options for agentic applications

          - `Concurrent bool`

          - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

            - `CurrentNode string`

            - `State map[string, any]`

          - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

            - `DurationMs int64`

            - `NodeID string`

            - `OperationInput string`

            - `OperationOutput string`

            - `OperationType string`

            - `StartTimestamp string`

            - `WorkflowID string`

            - `OperationMetadata map[string, any]`

          - `ReturnSpan bool`

          - `UseChannels bool`

        - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

          - `ArtifactIDsFilter []string`

          - `ArtifactNameRegex []string`

          - `Type string`

            - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

        - `type ItemLocator string`

    - `Alias string`

      Alias to title the results column. Defaults to the `application_variant`

    - `TaskType string`

      - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

  - `type EvaluationTaskAgentexOutput struct{…}`

    - `Configuration EvaluationTaskAgentexOutputConfiguration`

      - `AgentexAgentID string`

        The ID of the Agentex agent to use

      - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

        The dataset column to use as input for the agent

        - `string`

        - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

        - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

      - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

        Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

        - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

        - `type ItemLocator string`

      - `CompletionMode string`

        How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

        - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

        - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

      - `DeploymentID string`

        Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

      - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

        Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

        - `bool`

        - `type ItemLocator string`

      - `InputMode string`

        How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

        - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

        - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

      - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

        Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

        - `int64`

        - `type ItemLocator string`

      - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

        Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

        - `int64`

        - `type ItemLocator string`

    - `Alias string`

      Alias to title the results column. Defaults to the `agentex_output`

    - `TaskType string`

      - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

  - `type EvaluationTaskMetric struct{…}`

    - `Configuration EvaluationTaskMetricConfigurationUnion`

      - `type EvaluationTaskMetricConfigurationBleu struct{…}`

        - `Candidate string`

        - `Reference string`

        - `Type Bleu`

          - `const BleuBleu Bleu = "bleu"`

      - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

        - `Candidate string`

        - `Reference string`

        - `Type Meteor`

          - `const MeteorMeteor Meteor = "meteor"`

      - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

        - `Candidate string`

        - `Reference string`

        - `Type CosineSimilarity`

          - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

      - `type EvaluationTaskMetricConfigurationF1 struct{…}`

        - `Candidate string`

        - `Reference string`

        - `Type F1`

          - `const F1F1 F1 = "f1"`

      - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

        - `Candidate string`

        - `Reference string`

        - `Type Rouge1`

          - `const Rouge1Rouge1 Rouge1 = "rouge1"`

      - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

        - `Candidate string`

        - `Reference string`

        - `Type Rouge2`

          - `const Rouge2Rouge2 Rouge2 = "rouge2"`

      - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

        - `Candidate string`

        - `Reference string`

        - `Type RougeL`

          - `const RougeLRougeL RougeL = "rougeL"`

    - `Alias string`

      Alias to title the results column. Defaults to the metric type specified in the configuration

    - `TaskType string`

      - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

  - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

    - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

      - `Model string`

        model specified as `model_vendor/model_name`

      - `Prompt string`

      - `QuestionID string`

        question to be evaluated

    - `Alias string`

      Alias to title the results column. Defaults to the `auto_evaluation_question`

    - `TaskType string`

      - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

  - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

    - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

      - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `ResponseFormat map[string, any]`

          JSON schema used for structuring the model response

        - `InferenceArgs map[string, any]`

          Additional arguments to pass to the inference request

        - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

          - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

            - `Op string`

              - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

            - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

              - `string`

              - `float64`

              - `bool`

          - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

            - `Path string`

            - `Op string`

              - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

          - `type EqEvaluationRunCondition struct{…}`

            - `Left any`

            - `Right any`

            - `Op EqEvaluationRunConditionOp`

              - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

          - `type NeEvaluationRunCondition struct{…}`

            - `Left any`

            - `Right any`

            - `Op NeEvaluationRunConditionOp`

              - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

          - `type LtEvaluationRunCondition struct{…}`

            - `Left any`

            - `Right any`

            - `Op LtEvaluationRunConditionOp`

              - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

          - `type LteEvaluationRunCondition struct{…}`

            - `Left any`

            - `Right any`

            - `Op LteEvaluationRunConditionOp`

              - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

          - `type GtEvaluationRunCondition struct{…}`

            - `Left any`

            - `Right any`

            - `Op GtEvaluationRunConditionOp`

              - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

          - `type GteEvaluationRunCondition struct{…}`

            - `Left any`

            - `Right any`

            - `Op GteEvaluationRunConditionOp`

              - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

          - `type AndEvaluationRunCondition struct{…}`

            - `Operands []any`

            - `Op AndEvaluationRunConditionOp`

              - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

          - `type OrEvaluationRunCondition struct{…}`

            - `Operands []any`

            - `Op OrEvaluationRunConditionOp`

              - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

          - `type InEvaluationRunCondition struct{…}`

            - `Left any`

            - `Operands []any`

            - `Op InEvaluationRunConditionOp`

              - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

          - `type NotInEvaluationRunCondition struct{…}`

            - `Left any`

            - `Operands []any`

            - `Op NotInEvaluationRunConditionOp`

              - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

          - `type NotEvaluationRunCondition struct{…}`

            - `Operands []any`

            - `Op NotEvaluationRunConditionOp`

              - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

          - `type IsNullEvaluationRunCondition struct{…}`

            - `Operands []any`

            - `Op IsNullEvaluationRunConditionOp`

              - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

          - `type IsNotNullEvaluationRunCondition struct{…}`

            - `Operands []any`

            - `Op IsNotNullEvaluationRunConditionOp`

              - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

        - `SystemPrompt string`

      - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

        - `Choices []string`

          Choices array cannot be empty

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `InferenceArgs map[string, any]`

          Additional arguments to pass to the inference request

        - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

          - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

            - `Op string`

              - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

            - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

              - `string`

              - `float64`

              - `bool`

          - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

            - `Path string`

            - `Op string`

              - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

          - `type EqEvaluationRunCondition struct{…}`

          - `type NeEvaluationRunCondition struct{…}`

          - `type LtEvaluationRunCondition struct{…}`

          - `type LteEvaluationRunCondition struct{…}`

          - `type GtEvaluationRunCondition struct{…}`

          - `type GteEvaluationRunCondition struct{…}`

          - `type AndEvaluationRunCondition struct{…}`

          - `type OrEvaluationRunCondition struct{…}`

          - `type InEvaluationRunCondition struct{…}`

          - `type NotInEvaluationRunCondition struct{…}`

          - `type NotEvaluationRunCondition struct{…}`

          - `type IsNullEvaluationRunCondition struct{…}`

          - `type IsNotNullEvaluationRunCondition struct{…}`

        - `SystemPrompt string`

      - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

        - `Definition string`

        - `Name string`

        - `OutputRules []string`

        - `DataFields []string`

        - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

          - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

            - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

              - `Model string`

              - `Temperature float64`

            - `AgentName string`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

          - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

            - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

              - `Model string`

            - `AgentName string`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

          - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

            - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

              - `Model string`

            - `AgentName string`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

          - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

            - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

              - `Model string`

            - `AgentName string`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

        - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

          - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

          - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

          - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

          - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

        - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

          - `string`

          - `float64`

          - `bool`

        - `RubricID string`

        - `RubricVersion int64`

    - `Alias string`

      Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

    - `TaskType string`

      - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

  - `type EvaluationTaskAutoEvaluationAgent struct{…}`

    - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

    - `Alias string`

      Alias to title the results column. Defaults to the `auto_evaluation_agent`

    - `TaskType string`

      - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

  - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

    - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

      - `Layout Container`

        - `Children []ContainerChildUnion`

          The children to be displayed within the container

          - `type Container struct{…}`

          - `type Component struct{…}`

            - `Data ItemLocator`

              A pointer to the data in each evaluation item to be displayed within the component

            - `Label string`

        - `Direction ContainerDirection`

          The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

          - `const ContainerDirectionRow ContainerDirection = "row"`

          - `const ContainerDirectionColumn ContainerDirection = "column"`

      - `QuestionID string`

      - `PrefillFrom string`

        Dataset column to prefill contributor question task result

      - `QueueID string`

        The contributor annotation queue to include this task in. Defaults to `default`

      - `Required bool`

        Whether the question is required to be answered

      - `RubricID string`

        ID of the rubric to use for scoring this evaluation question

    - `Alias string`

      Alias to title the results column. Defaults to the `contributor_evaluation_question`

    - `TaskType string`

      - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

  - `type EvaluationTaskCustomFunction struct{…}`

    - `Configuration EvaluationTaskCustomFunctionConfiguration`

      Configuration for a custom Python function evaluation task.

      - `FunctionSource string`

        Python function source code

      - `ArgMapping map[string, string]`

        Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

      - `ConfigArgs map[string, any]`

        Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

      - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

        Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

        - `Path string`

          Dot path in the custom function return value to materialize.

        - `Alias string`

          Result column alias. Defaults to path with dots replaced by underscores.

    - `Alias string`

      Alias to title the results column. Defaults to the function name.

    - `TaskType string`

      - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Evaluation Tasks Progress Schema

- `type EvaluationTasksProgressSchema struct{…}`

  - `Items EvaluationTasksProgressSchemaItems`

    - `Failed int64`

    - `Pending int64`

    - `Successful int64`

    - `Total int64`

    - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

      - `ItemID string`

      - `Error string`

      - `ErrorType string`

  - `Workflows EvaluationTasksProgressSchemaWorkflows`

    - `Completed int64`

    - `Failed int64`

    - `Pending int64`

    - `Total int64`

### Evaluation Views

- `type EvaluationViews string`

  - `const EvaluationViewsTasks EvaluationViews = "tasks"`

### Gt Evaluation Run Condition

- `type GtEvaluationRunCondition struct{…}`

  - `Left any`

  - `Right any`

  - `Op GtEvaluationRunConditionOp`

    - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

### Gte Evaluation Run Condition

- `type GteEvaluationRunCondition struct{…}`

  - `Left any`

  - `Right any`

  - `Op GteEvaluationRunConditionOp`

    - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

### In Evaluation Run Condition

- `type InEvaluationRunCondition struct{…}`

  - `Left any`

  - `Operands []any`

  - `Op InEvaluationRunConditionOp`

    - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

### Is Not Null Evaluation Run Condition

- `type IsNotNullEvaluationRunCondition struct{…}`

  - `Operands []any`

  - `Op IsNotNullEvaluationRunConditionOp`

    - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

### Is Null Evaluation Run Condition

- `type IsNullEvaluationRunCondition struct{…}`

  - `Operands []any`

  - `Op IsNullEvaluationRunConditionOp`

    - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

### Item Locator

- `type ItemLocator string`

### Item Locator Template

- `type ItemLocatorTemplate string`

### Lt Evaluation Run Condition

- `type LtEvaluationRunCondition struct{…}`

  - `Left any`

  - `Right any`

  - `Op LtEvaluationRunConditionOp`

    - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

### Lte Evaluation Run Condition

- `type LteEvaluationRunCondition struct{…}`

  - `Left any`

  - `Right any`

  - `Op LteEvaluationRunConditionOp`

    - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

### Ne Evaluation Run Condition

- `type NeEvaluationRunCondition struct{…}`

  - `Left any`

  - `Right any`

  - `Op NeEvaluationRunConditionOp`

    - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

### Not Evaluation Run Condition

- `type NotEvaluationRunCondition struct{…}`

  - `Operands []any`

  - `Op NotEvaluationRunConditionOp`

    - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

### Not In Evaluation Run Condition

- `type NotInEvaluationRunCondition struct{…}`

  - `Left any`

  - `Operands []any`

  - `Op NotInEvaluationRunConditionOp`

    - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

### Or Evaluation Run Condition

- `type OrEvaluationRunCondition struct{…}`

  - `Operands []any`

  - `Op OrEvaluationRunConditionOp`

    - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

### Paginated List Evaluation

- `type PaginatedListEvaluation struct{…}`

  - `HasMore bool`

    Whether there are more items left to be fetched.

  - `Items []Evaluation`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

      - `ID string`

      - `Type IdentityType`

        - `const IdentityTypeUser IdentityType = "user"`

        - `const IdentityTypeServiceAccount IdentityType = "service_account"`

      - `Object IdentityObject`

        - `const IdentityObjectIdentity IdentityObject = "identity"`

    - `Datasets []Dataset`

      - `ID string`

        The unique identifier of the entity.

      - `CreatedAt Time`

        The date and time when the entity was created in ISO format.

      - `CreatedBy Identity`

        The identity that created the entity.

      - `CurrentVersionNum int64`

      - `Name string`

      - `Tags []string`

        The tags associated with the entity

      - `ArchivedAt Time`

        The date and time when the entity was archived in ISO format.

      - `Description string`

      - `Object DatasetObject`

        - `const DatasetObjectDataset DatasetObject = "dataset"`

    - `Name string`

    - `Status EvaluationStatus`

      - `const EvaluationStatusFailed EvaluationStatus = "failed"`

      - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

      - `const EvaluationStatusRunning EvaluationStatus = "running"`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `ErrorCount int64`

      Number of task errors across all items in this evaluation.

    - `Metadata map[string, any]`

      Metadata key-value pairs for the evaluation

    - `Object EvaluationObject`

      - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

    - `Progress EvaluationTasksProgressSchema`

      Progress of the evaluation's underlying async job

      - `Items EvaluationTasksProgressSchemaItems`

        - `Failed int64`

        - `Pending int64`

        - `Successful int64`

        - `Total int64`

        - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

          - `ItemID string`

          - `Error string`

          - `ErrorType string`

      - `Workflows EvaluationTasksProgressSchemaWorkflows`

        - `Completed int64`

        - `Failed int64`

        - `Pending int64`

        - `Total int64`

    - `StatusReason string`

      Reason for evaluation status

    - `Tasks []EvaluationTaskUnion`

      Tasks executed during evaluation. Populated with optional `task` view.

      - `type EvaluationTaskChatCompletion struct{…}`

        - `Configuration EvaluationTaskChatCompletionConfiguration`

          - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

            openai standard message format

            - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

            - `type ItemLocator string`

          - `Model string`

            model specified as `model_vendor/model`, for example `openai/gpt-4o`

          - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

            Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

            - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

            - `type ItemLocator string`

          - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

            Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

            - `float64`

            - `type ItemLocator string`

          - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

            Deprecated in favor of tool_choice. Controls which function is called by the model.

            - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

            - `type ItemLocator string`

          - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

            Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

            - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

            - `type ItemLocator string`

          - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

            Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

            - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

            - `type ItemLocator string`

          - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

            Whether to return log probabilities of the output tokens or not.

            - `bool`

            - `type ItemLocator string`

          - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

            An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

            - `int64`

            - `type ItemLocator string`

          - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

            Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

            - `int64`

            - `type ItemLocator string`

          - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

            Developer-defined tags and values used for filtering completions in the dashboard.

            - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

            - `type ItemLocator string`

          - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

            Output types that you would like the model to generate for this request.

            - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

            - `type ItemLocator string`

          - `N EvaluationTaskChatCompletionConfigurationNUnion`

            How many chat completion choices to generate for each input message.

            - `int64`

            - `type ItemLocator string`

          - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

            Whether to enable parallel function calling during tool use.

            - `bool`

            - `type ItemLocator string`

          - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

            Static predicted output content, such as the content of a text file being regenerated.

            - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

            - `type ItemLocator string`

          - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

            Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

            - `float64`

            - `type ItemLocator string`

          - `ReasoningEffort string`

            For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

          - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

            An object specifying the format that the model must output.

            - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

            - `type ItemLocator string`

          - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

            If specified, system will attempt to sample deterministically for repeated requests with same seed.

            - `int64`

            - `type ItemLocator string`

          - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

            Up to 4 sequences where the API will stop generating further tokens.

            - `string`

            - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

          - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

            Whether to store the output for use in model distillation or evals products.

            - `bool`

            - `type ItemLocator string`

          - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

            What sampling temperature to use. Higher values make output more random, lower more focused.

            - `float64`

            - `type ItemLocator string`

          - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

            Controls which tool is called by the model. Values: none, auto, required, or specific tool.

            - `string`

            - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

          - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

            A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

            - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

            - `type ItemLocator string`

          - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

            Only sample from the top K options for each subsequent token

            - `int64`

            - `type ItemLocator string`

          - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

            Number of most likely tokens to return at each position, with associated log probability.

            - `int64`

            - `type ItemLocator string`

          - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

            Alternative to temperature. Only tokens comprising top_p probability mass are considered.

            - `float64`

            - `type ItemLocator string`

        - `Alias string`

          Alias to title the results column. Defaults to the `chat_completion`

        - `TaskType string`

          - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

      - `type EvaluationTaskInference struct{…}`

        - `Configuration EvaluationTaskInferenceConfiguration`

          - `Model string`

            model specified as `vendor/name` (ex. openai/gpt-5)

          - `Args EvaluationTaskInferenceConfigurationArgsUnion`

            Arguments passed into model

            - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

            - `type ItemLocator string`

          - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

            Vendor specific configuration

            - `type LaunchInferenceConfiguration struct{…}`

              - `NumRetries int64`

              - `TimeoutSeconds int64`

            - `type ItemLocator string`

        - `Alias string`

          Alias to title the results column. Defaults to the `inference`

        - `TaskType string`

          - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

      - `type EvaluationTaskApplicationVariant struct{…}`

        - `Configuration EvaluationTaskApplicationVariantConfiguration`

          - `ApplicationVariantID string`

          - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

            Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

            - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

            - `type ItemLocator string`

          - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

            History of the application

            - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

              - `Request string`

                Request inputs

              - `Response string`

                Response outputs

              - `SessionData map[string, any]`

                Session data corresponding to the request response pair

            - `type ItemLocator string`

          - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

            Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

            - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

            - `type ItemLocator string`

          - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

            Optional overrides for the application

            - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

              Execution override options for agentic applications

              - `Concurrent bool`

              - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

                - `CurrentNode string`

                - `State map[string, any]`

              - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

                - `DurationMs int64`

                - `NodeID string`

                - `OperationInput string`

                - `OperationOutput string`

                - `OperationType string`

                - `StartTimestamp string`

                - `WorkflowID string`

                - `OperationMetadata map[string, any]`

              - `ReturnSpan bool`

              - `UseChannels bool`

            - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

              - `ArtifactIDsFilter []string`

              - `ArtifactNameRegex []string`

              - `Type string`

                - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

            - `type ItemLocator string`

        - `Alias string`

          Alias to title the results column. Defaults to the `application_variant`

        - `TaskType string`

          - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

      - `type EvaluationTaskAgentexOutput struct{…}`

        - `Configuration EvaluationTaskAgentexOutputConfiguration`

          - `AgentexAgentID string`

            The ID of the Agentex agent to use

          - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

            The dataset column to use as input for the agent

            - `string`

            - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

            - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

          - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

            Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

            - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

            - `type ItemLocator string`

          - `CompletionMode string`

            How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

            - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

            - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

          - `DeploymentID string`

            Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

          - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

            Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

            - `bool`

            - `type ItemLocator string`

          - `InputMode string`

            How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

            - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

            - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

          - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

            Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

            - `int64`

            - `type ItemLocator string`

          - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

            Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

            - `int64`

            - `type ItemLocator string`

        - `Alias string`

          Alias to title the results column. Defaults to the `agentex_output`

        - `TaskType string`

          - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

      - `type EvaluationTaskMetric struct{…}`

        - `Configuration EvaluationTaskMetricConfigurationUnion`

          - `type EvaluationTaskMetricConfigurationBleu struct{…}`

            - `Candidate string`

            - `Reference string`

            - `Type Bleu`

              - `const BleuBleu Bleu = "bleu"`

          - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

            - `Candidate string`

            - `Reference string`

            - `Type Meteor`

              - `const MeteorMeteor Meteor = "meteor"`

          - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

            - `Candidate string`

            - `Reference string`

            - `Type CosineSimilarity`

              - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

          - `type EvaluationTaskMetricConfigurationF1 struct{…}`

            - `Candidate string`

            - `Reference string`

            - `Type F1`

              - `const F1F1 F1 = "f1"`

          - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

            - `Candidate string`

            - `Reference string`

            - `Type Rouge1`

              - `const Rouge1Rouge1 Rouge1 = "rouge1"`

          - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

            - `Candidate string`

            - `Reference string`

            - `Type Rouge2`

              - `const Rouge2Rouge2 Rouge2 = "rouge2"`

          - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

            - `Candidate string`

            - `Reference string`

            - `Type RougeL`

              - `const RougeLRougeL RougeL = "rougeL"`

        - `Alias string`

          Alias to title the results column. Defaults to the metric type specified in the configuration

        - `TaskType string`

          - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

      - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

        - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `QuestionID string`

            question to be evaluated

        - `Alias string`

          Alias to title the results column. Defaults to the `auto_evaluation_question`

        - `TaskType string`

          - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

      - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

        - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

          - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

            - `Model string`

              model specified as `model_vendor/model_name`

            - `Prompt string`

            - `ResponseFormat map[string, any]`

              JSON schema used for structuring the model response

            - `InferenceArgs map[string, any]`

              Additional arguments to pass to the inference request

            - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

              - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

                - `Op string`

                  - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

                - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                  - `string`

                  - `float64`

                  - `bool`

              - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

                - `Path string`

                - `Op string`

                  - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

              - `type EqEvaluationRunCondition struct{…}`

                - `Left any`

                - `Right any`

                - `Op EqEvaluationRunConditionOp`

                  - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

              - `type NeEvaluationRunCondition struct{…}`

                - `Left any`

                - `Right any`

                - `Op NeEvaluationRunConditionOp`

                  - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

              - `type LtEvaluationRunCondition struct{…}`

                - `Left any`

                - `Right any`

                - `Op LtEvaluationRunConditionOp`

                  - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

              - `type LteEvaluationRunCondition struct{…}`

                - `Left any`

                - `Right any`

                - `Op LteEvaluationRunConditionOp`

                  - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

              - `type GtEvaluationRunCondition struct{…}`

                - `Left any`

                - `Right any`

                - `Op GtEvaluationRunConditionOp`

                  - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

              - `type GteEvaluationRunCondition struct{…}`

                - `Left any`

                - `Right any`

                - `Op GteEvaluationRunConditionOp`

                  - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

              - `type AndEvaluationRunCondition struct{…}`

                - `Operands []any`

                - `Op AndEvaluationRunConditionOp`

                  - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

              - `type OrEvaluationRunCondition struct{…}`

                - `Operands []any`

                - `Op OrEvaluationRunConditionOp`

                  - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

              - `type InEvaluationRunCondition struct{…}`

                - `Left any`

                - `Operands []any`

                - `Op InEvaluationRunConditionOp`

                  - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

              - `type NotInEvaluationRunCondition struct{…}`

                - `Left any`

                - `Operands []any`

                - `Op NotInEvaluationRunConditionOp`

                  - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

              - `type NotEvaluationRunCondition struct{…}`

                - `Operands []any`

                - `Op NotEvaluationRunConditionOp`

                  - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

              - `type IsNullEvaluationRunCondition struct{…}`

                - `Operands []any`

                - `Op IsNullEvaluationRunConditionOp`

                  - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

              - `type IsNotNullEvaluationRunCondition struct{…}`

                - `Operands []any`

                - `Op IsNotNullEvaluationRunConditionOp`

                  - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

            - `SystemPrompt string`

          - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

            - `Choices []string`

              Choices array cannot be empty

            - `Model string`

              model specified as `model_vendor/model_name`

            - `Prompt string`

            - `InferenceArgs map[string, any]`

              Additional arguments to pass to the inference request

            - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

              - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

                - `Op string`

                  - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

                - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                  - `string`

                  - `float64`

                  - `bool`

              - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

                - `Path string`

                - `Op string`

                  - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

              - `type EqEvaluationRunCondition struct{…}`

              - `type NeEvaluationRunCondition struct{…}`

              - `type LtEvaluationRunCondition struct{…}`

              - `type LteEvaluationRunCondition struct{…}`

              - `type GtEvaluationRunCondition struct{…}`

              - `type GteEvaluationRunCondition struct{…}`

              - `type AndEvaluationRunCondition struct{…}`

              - `type OrEvaluationRunCondition struct{…}`

              - `type InEvaluationRunCondition struct{…}`

              - `type NotInEvaluationRunCondition struct{…}`

              - `type NotEvaluationRunCondition struct{…}`

              - `type IsNullEvaluationRunCondition struct{…}`

              - `type IsNotNullEvaluationRunCondition struct{…}`

            - `SystemPrompt string`

          - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

            - `Definition string`

            - `Name string`

            - `OutputRules []string`

            - `DataFields []string`

            - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

              - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

                - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                  - `Model string`

                  - `Temperature float64`

                - `AgentName string`

                  - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

              - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

                - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                  - `Model string`

                - `AgentName string`

                  - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

              - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

                - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                  - `Model string`

                - `AgentName string`

                  - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

              - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

                - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                  - `Model string`

                - `AgentName string`

                  - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

            - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

              - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

            - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

              - `string`

              - `float64`

              - `bool`

            - `RubricID string`

            - `RubricVersion int64`

        - `Alias string`

          Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

        - `TaskType string`

          - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

      - `type EvaluationTaskAutoEvaluationAgent struct{…}`

        - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

        - `Alias string`

          Alias to title the results column. Defaults to the `auto_evaluation_agent`

        - `TaskType string`

          - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

      - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

        - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

          - `Layout Container`

            - `Children []ContainerChildUnion`

              The children to be displayed within the container

              - `type Container struct{…}`

              - `type Component struct{…}`

                - `Data ItemLocator`

                  A pointer to the data in each evaluation item to be displayed within the component

                - `Label string`

            - `Direction ContainerDirection`

              The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

              - `const ContainerDirectionRow ContainerDirection = "row"`

              - `const ContainerDirectionColumn ContainerDirection = "column"`

          - `QuestionID string`

          - `PrefillFrom string`

            Dataset column to prefill contributor question task result

          - `QueueID string`

            The contributor annotation queue to include this task in. Defaults to `default`

          - `Required bool`

            Whether the question is required to be answered

          - `RubricID string`

            ID of the rubric to use for scoring this evaluation question

        - `Alias string`

          Alias to title the results column. Defaults to the `contributor_evaluation_question`

        - `TaskType string`

          - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

      - `type EvaluationTaskCustomFunction struct{…}`

        - `Configuration EvaluationTaskCustomFunctionConfiguration`

          Configuration for a custom Python function evaluation task.

          - `FunctionSource string`

            Python function source code

          - `ArgMapping map[string, string]`

            Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

          - `ConfigArgs map[string, any]`

            Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

          - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

            Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

            - `Path string`

              Dot path in the custom function return value to materialize.

            - `Alias string`

              Result column alias. Defaults to path with dots replaced by underscores.

        - `Alias string`

          Alias to title the results column. Defaults to the function name.

        - `TaskType string`

          - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

  - `Total int64`

    The total of items that match the query. This is greater than or equal to the number of items returned.

  - `Limit int64`

    The maximum number of items to return.

  - `Object PaginatedListEvaluationObject`

    - `const PaginatedListEvaluationObjectList PaginatedListEvaluationObject = "list"`

# Tasks

## Add Test Criteria to Evaluation

`client.Evaluations.Tasks.Add(ctx, evaluationID, body) (*Evaluation, error)`

**post** `/v5/evaluations/{evaluation_id}/tasks`

Add a new test criteria to an existing evaluation.

Narrowed to contributor question tasks (`contributor_evaluation.question`); other task types
must be configured when the evaluation is first created and are rejected here. The request is
also rejected if the evaluation is archived, if a test criteria with the same alias already
exists, or if any contributor annotation task for the evaluation has already been claimed or
completed. Because only contributor question tasks are accepted, the added criteria is applied
synchronously and contributors answer it against the evaluation's existing items — no async job
or Temporal workflow is started.

### Parameters

- `evaluationID string`

- `body EvaluationTaskAddParams`

  - `Task param.Field[EvaluationTaskUnion]`

    New test criteria to add to the evaluation. Rejected when contributor annotation tasks for this evaluation have already been claimed or completed. Triggers a rerun so the new task executes against existing items.

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluation, err := client.Evaluations.Tasks.Add(
    context.TODO(),
    "evaluation_id",
    sgpdev.EvaluationTaskAddParams{
      Task: sgpdev.EvaluationTaskUnionParam{
        OfChatCompletion: &sgpdev.EvaluationTaskChatCompletionParam{
          Configuration: sgpdev.EvaluationTaskChatCompletionConfigurationParam{
            Messages: sgpdev.EvaluationTaskChatCompletionConfigurationMessagesUnionParam{
              OfMapOfAnyMap: []map[string]any{map[string]any{
              "foo": "bar",
              }},
            },
            Model: "model",
          },
          TaskType: "chat_completion",
        },
      },
    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluation.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Update Test Criteria Configuration

`client.Evaluations.Tasks.Update(ctx, alias, params) (*Evaluation, error)`

**patch** `/v5/evaluations/{evaluation_id}/tasks/{alias}`

Replace the full configuration of a single test criteria, identified by its alias.

The alias must match an existing test criteria on the evaluation, and the replacement
configuration is validated against the evaluation's current items before being applied. The
request is rejected if the evaluation is archived, if no test criteria matches the alias, or if
any contributor annotation task for the evaluation has already been claimed or completed — at
that point labelers are in-flight and mutating the task definition would corrupt their work.

### Parameters

- `alias string`

- `params EvaluationTaskUpdateParams`

  - `EvaluationID param.Field[string]`

    Path param

  - `Configuration param.Field[map[string, any]]`

    Body param: Full replacement for the test criteria's configuration JSON. Only allowed when no contributor annotation tasks for this evaluation have been claimed or completed.

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluation, err := client.Evaluations.Tasks.Update(
    context.TODO(),
    "alias",
    sgpdev.EvaluationTaskUpdateParams{
      EvaluationID: "evaluation_id",
      Configuration: map[string]any{
      "foo": "bar",
      },
    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluation.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```
