## Create Evaluation

`client.Evaluations.New(ctx, body) (*Evaluation, error)`

**post** `/v5/evaluations`

Create an evaluation together with its items, optionally running test criteria against them.

Accepts three request shapes: standalone (inline `data`), from an existing dataset
(`dataset_id` with optional per-item references), or with a new reusable dataset created inline
from `data`. When the evaluation includes tasks that require execution (for example an LLM judge
or custom function), an async job and a Temporal workflow are started and the evaluation is
returned immediately with status `running`; task results and `error_count` populate
asynchronously. When it includes only contributor tasks, taxonomy-only input, or no tasks, no
workflow runs and it is returned with status `completed`. Optional `tasks`, `metadata`, `tags`,
and `taxonomy_params` are persisted alongside the evaluation and its items.

### Parameters

- `body EvaluationNewParams`

  - `Evaluation param.Field[EvaluationNewParamsEvaluationUnion]`

    - `type EvaluationNewParamsEvaluationEvaluationStandaloneCreateRequest struct{…}`

      - `Data []map[string, any]`

        Items to be evaluated

      - `Name string`

      - `Description string`

      - `Files []map[string, string]`

        Files to be associated to the evaluation

      - `Metadata map[string, any]`

        Optional metadata key-value pairs for the evaluation

      - `SkipPrefilledRows bool`

        Do not queue a contributor task for prefilled questions

      - `Tags []string`

        The tags associated with the evaluation

      - `Tasks []EvaluationTaskUnion`

        Tasks allow you to augment and evaluate your data

        - `type EvaluationTaskChatCompletion struct{…}`

          - `Configuration EvaluationTaskChatCompletionConfiguration`

            - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

              openai standard message format

              - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

              - `type ItemLocator string`

            - `Model string`

              model specified as `model_vendor/model`, for example `openai/gpt-4o`

            - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

              Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

              - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

              - `type ItemLocator string`

            - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

              Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

              - `float64`

              - `type ItemLocator string`

            - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

              Deprecated in favor of tool_choice. Controls which function is called by the model.

              - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

              - `type ItemLocator string`

            - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

              Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

              - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

              - `type ItemLocator string`

            - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

              Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

              - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

              - `type ItemLocator string`

            - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

              Whether to return log probabilities of the output tokens or not.

              - `bool`

              - `type ItemLocator string`

            - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

              An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

              - `int64`

              - `type ItemLocator string`

            - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

              Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

              - `int64`

              - `type ItemLocator string`

            - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

              Developer-defined tags and values used for filtering completions in the dashboard.

              - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

              - `type ItemLocator string`

            - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

              Output types that you would like the model to generate for this request.

              - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

              - `type ItemLocator string`

            - `N EvaluationTaskChatCompletionConfigurationNUnion`

              How many chat completion choices to generate for each input message.

              - `int64`

              - `type ItemLocator string`

            - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

              Whether to enable parallel function calling during tool use.

              - `bool`

              - `type ItemLocator string`

            - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

              Static predicted output content, such as the content of a text file being regenerated.

              - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

              - `type ItemLocator string`

            - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

              Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

              - `float64`

              - `type ItemLocator string`

            - `ReasoningEffort string`

              For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

            - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

              An object specifying the format that the model must output.

              - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

              - `type ItemLocator string`

            - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

              If specified, system will attempt to sample deterministically for repeated requests with same seed.

              - `int64`

              - `type ItemLocator string`

            - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

              Up to 4 sequences where the API will stop generating further tokens.

              - `string`

              - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

            - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

              Whether to store the output for use in model distillation or evals products.

              - `bool`

              - `type ItemLocator string`

            - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

              What sampling temperature to use. Higher values make output more random, lower more focused.

              - `float64`

              - `type ItemLocator string`

            - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

              Controls which tool is called by the model. Values: none, auto, required, or specific tool.

              - `string`

              - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

            - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

              A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

              - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

              - `type ItemLocator string`

            - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

              Only sample from the top K options for each subsequent token

              - `int64`

              - `type ItemLocator string`

            - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

              Number of most likely tokens to return at each position, with associated log probability.

              - `int64`

              - `type ItemLocator string`

            - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

              Alternative to temperature. Only tokens comprising top_p probability mass are considered.

              - `float64`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `chat_completion`

          - `TaskType string`

            - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

        - `type EvaluationTaskInference struct{…}`

          - `Configuration EvaluationTaskInferenceConfiguration`

            - `Model string`

              model specified as `vendor/name` (ex. openai/gpt-5)

            - `Args EvaluationTaskInferenceConfigurationArgsUnion`

              Arguments passed into model

              - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

              - `type ItemLocator string`

            - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

              Vendor specific configuration

              - `type LaunchInferenceConfiguration struct{…}`

                - `NumRetries int64`

                - `TimeoutSeconds int64`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `inference`

          - `TaskType string`

            - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

        - `type EvaluationTaskApplicationVariant struct{…}`

          - `Configuration EvaluationTaskApplicationVariantConfiguration`

            - `ApplicationVariantID string`

            - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

              Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

              - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

              - `type ItemLocator string`

            - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

              History of the application

              - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

                - `Request string`

                  Request inputs

                - `Response string`

                  Response outputs

                - `SessionData map[string, any]`

                  Session data corresponding to the request response pair

              - `type ItemLocator string`

            - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

              Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

              - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

              - `type ItemLocator string`

            - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

              Optional overrides for the application

              - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

                Execution override options for agentic applications

                - `Concurrent bool`

                - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

                  - `CurrentNode string`

                  - `State map[string, any]`

                - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

                  - `DurationMs int64`

                  - `NodeID string`

                  - `OperationInput string`

                  - `OperationOutput string`

                  - `OperationType string`

                  - `StartTimestamp string`

                  - `WorkflowID string`

                  - `OperationMetadata map[string, any]`

                - `ReturnSpan bool`

                - `UseChannels bool`

              - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

                - `ArtifactIDsFilter []string`

                - `ArtifactNameRegex []string`

                - `Type string`

                  - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `application_variant`

          - `TaskType string`

            - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

        - `type EvaluationTaskAgentexOutput struct{…}`

          - `Configuration EvaluationTaskAgentexOutputConfiguration`

            - `AgentexAgentID string`

              The ID of the Agentex agent to use

            - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

              The dataset column to use as input for the agent

              - `string`

              - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

              - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

            - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

              Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

              - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

              - `type ItemLocator string`

            - `CompletionMode string`

              How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

              - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

              - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

            - `DeploymentID string`

              Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

            - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

              Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

              - `bool`

              - `type ItemLocator string`

            - `InputMode string`

              How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

              - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

              - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

            - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

              Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

              - `int64`

              - `type ItemLocator string`

            - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

              Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

              - `int64`

              - `type ItemLocator string`

          - `Alias string`

            Alias to title the results column. Defaults to the `agentex_output`

          - `TaskType string`

            - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

        - `type EvaluationTaskMetric struct{…}`

          - `Configuration EvaluationTaskMetricConfigurationUnion`

            - `type EvaluationTaskMetricConfigurationBleu struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Bleu`

                - `const BleuBleu Bleu = "bleu"`

            - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Meteor`

                - `const MeteorMeteor Meteor = "meteor"`

            - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type CosineSimilarity`

                - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

            - `type EvaluationTaskMetricConfigurationF1 struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type F1`

                - `const F1F1 F1 = "f1"`

            - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Rouge1`

                - `const Rouge1Rouge1 Rouge1 = "rouge1"`

            - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type Rouge2`

                - `const Rouge2Rouge2 Rouge2 = "rouge2"`

            - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

              - `Candidate string`

              - `Reference string`

              - `Type RougeL`

                - `const RougeLRougeL RougeL = "rougeL"`

          - `Alias string`

            Alias to title the results column. Defaults to the metric type specified in the configuration

          - `TaskType string`

            - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

        - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

          - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

            - `Model string`

              model specified as `model_vendor/model_name`

            - `Prompt string`

            - `QuestionID string`

              question to be evaluated

          - `Alias string`

            Alias to title the results column. Defaults to the `auto_evaluation_question`

          - `TaskType string`

            - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

        - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

          - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

              - `Model string`

                model specified as `model_vendor/model_name`

              - `Prompt string`

              - `ResponseFormat map[string, any]`

                JSON schema used for structuring the model response

              - `InferenceArgs map[string, any]`

                Additional arguments to pass to the inference request

              - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

                  - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                    - `string`

                    - `float64`

                    - `bool`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

                  - `Path string`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

                - `type EqEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op EqEvaluationRunConditionOp`

                    - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

                - `type NeEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op NeEvaluationRunConditionOp`

                    - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

                - `type LtEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op LtEvaluationRunConditionOp`

                    - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

                - `type LteEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op LteEvaluationRunConditionOp`

                    - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

                - `type GtEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op GtEvaluationRunConditionOp`

                    - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

                - `type GteEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Right any`

                  - `Op GteEvaluationRunConditionOp`

                    - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

                - `type AndEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op AndEvaluationRunConditionOp`

                    - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

                - `type OrEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op OrEvaluationRunConditionOp`

                    - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

                - `type InEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Operands []any`

                  - `Op InEvaluationRunConditionOp`

                    - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

                - `type NotInEvaluationRunCondition struct{…}`

                  - `Left any`

                  - `Operands []any`

                  - `Op NotInEvaluationRunConditionOp`

                    - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

                - `type NotEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op NotEvaluationRunConditionOp`

                    - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

                - `type IsNullEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op IsNullEvaluationRunConditionOp`

                    - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

                - `type IsNotNullEvaluationRunCondition struct{…}`

                  - `Operands []any`

                  - `Op IsNotNullEvaluationRunConditionOp`

                    - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

              - `SystemPrompt string`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

              - `Choices []string`

                Choices array cannot be empty

              - `Model string`

                model specified as `model_vendor/model_name`

              - `Prompt string`

              - `InferenceArgs map[string, any]`

                Additional arguments to pass to the inference request

              - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

                  - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                    - `string`

                    - `float64`

                    - `bool`

                - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

                  - `Path string`

                  - `Op string`

                    - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

                - `type EqEvaluationRunCondition struct{…}`

                - `type NeEvaluationRunCondition struct{…}`

                - `type LtEvaluationRunCondition struct{…}`

                - `type LteEvaluationRunCondition struct{…}`

                - `type GtEvaluationRunCondition struct{…}`

                - `type GteEvaluationRunCondition struct{…}`

                - `type AndEvaluationRunCondition struct{…}`

                - `type OrEvaluationRunCondition struct{…}`

                - `type InEvaluationRunCondition struct{…}`

                - `type NotInEvaluationRunCondition struct{…}`

                - `type NotEvaluationRunCondition struct{…}`

                - `type IsNullEvaluationRunCondition struct{…}`

                - `type IsNotNullEvaluationRunCondition struct{…}`

              - `SystemPrompt string`

            - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

              - `Definition string`

              - `Name string`

              - `OutputRules []string`

              - `DataFields []string`

              - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                    - `Model string`

                    - `Temperature float64`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                    - `Model string`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                    - `Model string`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

                - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

                  - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                    - `Model string`

                  - `AgentName string`

                    - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

              - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

              - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

                - `string`

                - `float64`

                - `bool`

              - `RubricID string`

              - `RubricVersion int64`

          - `Alias string`

            Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

          - `TaskType string`

            - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

        - `type EvaluationTaskAutoEvaluationAgent struct{…}`

          - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

          - `Alias string`

            Alias to title the results column. Defaults to the `auto_evaluation_agent`

          - `TaskType string`

            - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

        - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

          - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

            - `Layout Container`

              - `Children []ContainerChildUnion`

                The children to be displayed within the container

                - `type Container struct{…}`

                - `type Component struct{…}`

                  - `Data ItemLocator`

                    A pointer to the data in each evaluation item to be displayed within the component

                  - `Label string`

              - `Direction ContainerDirection`

                The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

                - `const ContainerDirectionRow ContainerDirection = "row"`

                - `const ContainerDirectionColumn ContainerDirection = "column"`

            - `QuestionID string`

            - `PrefillFrom string`

              Dataset column to prefill contributor question task result

            - `QueueID string`

              The contributor annotation queue to include this task in. Defaults to `default`

            - `Required bool`

              Whether the question is required to be answered

            - `RubricID string`

              ID of the rubric to use for scoring this evaluation question

          - `Alias string`

            Alias to title the results column. Defaults to the `contributor_evaluation_question`

          - `TaskType string`

            - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

        - `type EvaluationTaskCustomFunction struct{…}`

          - `Configuration EvaluationTaskCustomFunctionConfiguration`

            Configuration for a custom Python function evaluation task.

            - `FunctionSource string`

              Python function source code

            - `ArgMapping map[string, string]`

              Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

            - `ConfigArgs map[string, any]`

              Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

            - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

              Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

              - `Path string`

                Dot path in the custom function return value to materialize.

              - `Alias string`

                Result column alias. Defaults to path with dots replaced by underscores.

          - `Alias string`

            Alias to title the results column. Defaults to the function name.

          - `TaskType string`

            - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

      - `TaxonomyParams map[string, any]`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

    - `type EvaluationNewParamsEvaluationEvaluationFromDatasetCreateRequest struct{…}`

      - `DatasetID string`

        The ID of the dataset containing the items referenced by the `data` field

      - `Name string`

      - `Data []EvaluationNewParamsEvaluationEvaluationFromDatasetCreateRequestData`

        Items to be evaluated, including references to the input dataset

        - `DatasetItemID string`

      - `Description string`

      - `Metadata map[string, any]`

        Optional metadata key-value pairs for the evaluation

      - `SkipPrefilledRows bool`

        Do not queue a contributor task for prefilled questions

      - `Tags []string`

        The tags associated with the evaluation

      - `Tasks []EvaluationTaskUnion`

        Tasks allow you to augment and evaluate your data

        - `type EvaluationTaskChatCompletion struct{…}`

        - `type EvaluationTaskInference struct{…}`

        - `type EvaluationTaskApplicationVariant struct{…}`

        - `type EvaluationTaskAgentexOutput struct{…}`

        - `type EvaluationTaskMetric struct{…}`

        - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

        - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

        - `type EvaluationTaskAutoEvaluationAgent struct{…}`

        - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

        - `type EvaluationTaskCustomFunction struct{…}`

      - `TaxonomyParams map[string, any]`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

    - `type EvaluationNewParamsEvaluationEvaluationWithDatasetCreateRequest struct{…}`

      - `Data []map[string, any]`

        Items to be evaluated

      - `Dataset EvaluationNewParamsEvaluationEvaluationWithDatasetCreateRequestDataset`

        Create a reusable dataset from items in the `data` field

        - `Name string`

        - `Description string`

        - `Keys []string`

          Keys from items in the `data` field that should be included in the dataset. If not provided, all keys will be included.

        - `Tags []string`

          The tags associated with the entity

      - `Name string`

      - `Description string`

      - `Files []map[string, string]`

        Files to be associated to the evaluation

      - `Metadata map[string, any]`

        Optional metadata key-value pairs for the evaluation

      - `SkipPrefilledRows bool`

        Do not queue a contributor task for prefilled questions

      - `Tags []string`

        The tags associated with the evaluation

      - `Tasks []EvaluationTaskUnion`

        Tasks allow you to augment and evaluate your data

        - `type EvaluationTaskChatCompletion struct{…}`

        - `type EvaluationTaskInference struct{…}`

        - `type EvaluationTaskApplicationVariant struct{…}`

        - `type EvaluationTaskAgentexOutput struct{…}`

        - `type EvaluationTaskMetric struct{…}`

        - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

        - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

        - `type EvaluationTaskAutoEvaluationAgent struct{…}`

        - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

        - `type EvaluationTaskCustomFunction struct{…}`

      - `TaxonomyParams map[string, any]`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

### Returns

- `type Evaluation struct{…}`

  - `ID string`

    The unique identifier of the entity.

  - `CreatedAt Time`

    The date and time when the entity was created in ISO format.

  - `CreatedBy Identity`

    The identity that created the entity.

    - `ID string`

    - `Type IdentityType`

      - `const IdentityTypeUser IdentityType = "user"`

      - `const IdentityTypeServiceAccount IdentityType = "service_account"`

    - `Object IdentityObject`

      - `const IdentityObjectIdentity IdentityObject = "identity"`

  - `Datasets []Dataset`

    - `ID string`

      The unique identifier of the entity.

    - `CreatedAt Time`

      The date and time when the entity was created in ISO format.

    - `CreatedBy Identity`

      The identity that created the entity.

    - `CurrentVersionNum int64`

    - `Name string`

    - `Tags []string`

      The tags associated with the entity

    - `ArchivedAt Time`

      The date and time when the entity was archived in ISO format.

    - `Description string`

    - `Object DatasetObject`

      - `const DatasetObjectDataset DatasetObject = "dataset"`

  - `Name string`

  - `Status EvaluationStatus`

    - `const EvaluationStatusFailed EvaluationStatus = "failed"`

    - `const EvaluationStatusCompleted EvaluationStatus = "completed"`

    - `const EvaluationStatusRunning EvaluationStatus = "running"`

  - `Tags []string`

    The tags associated with the entity

  - `ArchivedAt Time`

    The date and time when the entity was archived in ISO format.

  - `Description string`

  - `ErrorCount int64`

    Number of task errors across all items in this evaluation.

  - `Metadata map[string, any]`

    Metadata key-value pairs for the evaluation

  - `Object EvaluationObject`

    - `const EvaluationObjectEvaluation EvaluationObject = "evaluation"`

  - `Progress EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `Items EvaluationTasksProgressSchemaItems`

      - `Failed int64`

      - `Pending int64`

      - `Successful int64`

      - `Total int64`

      - `FailedItems []EvaluationTasksProgressSchemaItemsFailedItem`

        - `ItemID string`

        - `Error string`

        - `ErrorType string`

    - `Workflows EvaluationTasksProgressSchemaWorkflows`

      - `Completed int64`

      - `Failed int64`

      - `Pending int64`

      - `Total int64`

  - `StatusReason string`

    Reason for evaluation status

  - `Tasks []EvaluationTaskUnion`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `type EvaluationTaskChatCompletion struct{…}`

      - `Configuration EvaluationTaskChatCompletionConfiguration`

        - `Messages EvaluationTaskChatCompletionConfigurationMessagesUnion`

          openai standard message format

          - `type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]`

          - `type ItemLocator string`

        - `Model string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `Audio EvaluationTaskChatCompletionConfigurationAudioUnion`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]`

          - `type ItemLocator string`

        - `FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnion`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]`

          - `type ItemLocator string`

        - `Functions EvaluationTaskChatCompletionConfigurationFunctionsUnion`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]`

          - `type ItemLocator string`

        - `LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnion`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]`

          - `type ItemLocator string`

        - `Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnion`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `type ItemLocator string`

        - `MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnion`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int64`

          - `type ItemLocator string`

        - `MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnion`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int64`

          - `type ItemLocator string`

        - `Metadata EvaluationTaskChatCompletionConfigurationMetadataUnion`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]`

          - `type ItemLocator string`

        - `Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnion`

          Output types that you would like the model to generate for this request.

          - `type EvaluationTaskChatCompletionConfigurationModalitiesArray []string`

          - `type ItemLocator string`

        - `N EvaluationTaskChatCompletionConfigurationNUnion`

          How many chat completion choices to generate for each input message.

          - `int64`

          - `type ItemLocator string`

        - `ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnion`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `type ItemLocator string`

        - `Prediction EvaluationTaskChatCompletionConfigurationPredictionUnion`

          Static predicted output content, such as the content of a text file being regenerated.

          - `type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]`

          - `type ItemLocator string`

        - `PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnion`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float64`

          - `type ItemLocator string`

        - `ReasoningEffort string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnion`

          An object specifying the format that the model must output.

          - `type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]`

          - `type ItemLocator string`

        - `Seed EvaluationTaskChatCompletionConfigurationSeedUnion`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int64`

          - `type ItemLocator string`

        - `Stop EvaluationTaskChatCompletionConfigurationStopUnion`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationStopArray []string`

        - `Store EvaluationTaskChatCompletionConfigurationStoreUnion`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `type ItemLocator string`

        - `Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnion`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float64`

          - `type ItemLocator string`

        - `ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnion`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]`

        - `Tools EvaluationTaskChatCompletionConfigurationToolsUnion`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]`

          - `type ItemLocator string`

        - `TopK EvaluationTaskChatCompletionConfigurationTopKUnion`

          Only sample from the top K options for each subsequent token

          - `int64`

          - `type ItemLocator string`

        - `TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnion`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int64`

          - `type ItemLocator string`

        - `TopP EvaluationTaskChatCompletionConfigurationTopPUnion`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `TaskType string`

        - `const EvaluationTaskChatCompletionTaskTypeChatCompletion EvaluationTaskChatCompletionTaskType = "chat_completion"`

    - `type EvaluationTaskInference struct{…}`

      - `Configuration EvaluationTaskInferenceConfiguration`

        - `Model string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `Args EvaluationTaskInferenceConfigurationArgsUnion`

          Arguments passed into model

          - `type EvaluationTaskInferenceConfigurationArgsMap map[string, any]`

          - `type ItemLocator string`

        - `InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnion`

          Vendor specific configuration

          - `type LaunchInferenceConfiguration struct{…}`

            - `NumRetries int64`

            - `TimeoutSeconds int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `inference`

      - `TaskType string`

        - `const EvaluationTaskInferenceTaskTypeInference EvaluationTaskInferenceTaskType = "inference"`

    - `type EvaluationTaskApplicationVariant struct{…}`

      - `Configuration EvaluationTaskApplicationVariantConfiguration`

        - `ApplicationVariantID string`

        - `Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]`

          - `type ItemLocator string`

        - `History EvaluationTaskApplicationVariantConfigurationHistoryUnion`

          History of the application

          - `type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem`

            - `Request string`

              Request inputs

            - `Response string`

              Response outputs

            - `SessionData map[string, any]`

              Session data corresponding to the request response pair

          - `type ItemLocator string`

        - `OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnion`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]`

          - `type ItemLocator string`

        - `OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnion`

          Optional overrides for the application

          - `type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}`

            Execution override options for agentic applications

            - `Concurrent bool`

            - `InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialState`

              - `CurrentNode string`

              - `State map[string, any]`

            - `PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTrace`

              - `DurationMs int64`

              - `NodeID string`

              - `OperationInput string`

              - `OperationOutput string`

              - `OperationType string`

              - `StartTimestamp string`

              - `WorkflowID string`

              - `OperationMetadata map[string, any]`

            - `ReturnSpan bool`

            - `UseChannels bool`

          - `type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]`

            - `ArtifactIDsFilter []string`

            - `ArtifactNameRegex []string`

            - `Type string`

              - `const EvaluationTaskApplicationVariantConfigurationOverridesMapItemTypeKnowledgeBaseSchema EvaluationTaskApplicationVariantConfigurationOverridesMapItemType = "knowledge_base_schema"`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `application_variant`

      - `TaskType string`

        - `const EvaluationTaskApplicationVariantTaskTypeApplicationVariant EvaluationTaskApplicationVariantTaskType = "application_variant"`

    - `type EvaluationTaskAgentexOutput struct{…}`

      - `Configuration EvaluationTaskAgentexOutputConfiguration`

        - `AgentexAgentID string`

          The ID of the Agentex agent to use

        - `InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion`

          The dataset column to use as input for the agent

          - `string`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]`

          - `type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any`

        - `AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnion`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]`

          - `type ItemLocator string`

        - `CompletionMode string`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"`

          - `const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"`

        - `DeploymentID string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnion`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `type ItemLocator string`

        - `InputMode string`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"`

          - `const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"`

        - `QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnion`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int64`

          - `type ItemLocator string`

        - `TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnion`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int64`

          - `type ItemLocator string`

      - `Alias string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `TaskType string`

        - `const EvaluationTaskAgentexOutputTaskTypeAgentexOutput EvaluationTaskAgentexOutputTaskType = "agentex_output"`

    - `type EvaluationTaskMetric struct{…}`

      - `Configuration EvaluationTaskMetricConfigurationUnion`

        - `type EvaluationTaskMetricConfigurationBleu struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Bleu`

            - `const BleuBleu Bleu = "bleu"`

        - `type EvaluationTaskMetricConfigurationMeteor struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Meteor`

            - `const MeteorMeteor Meteor = "meteor"`

        - `type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type CosineSimilarity`

            - `const CosineSimilarityCosineSimilarity CosineSimilarity = "cosine_similarity"`

        - `type EvaluationTaskMetricConfigurationF1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type F1`

            - `const F1F1 F1 = "f1"`

        - `type EvaluationTaskMetricConfigurationRouge1 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge1`

            - `const Rouge1Rouge1 Rouge1 = "rouge1"`

        - `type EvaluationTaskMetricConfigurationRouge2 struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type Rouge2`

            - `const Rouge2Rouge2 Rouge2 = "rouge2"`

        - `type EvaluationTaskMetricConfigurationRougeL struct{…}`

          - `Candidate string`

          - `Reference string`

          - `Type RougeL`

            - `const RougeLRougeL RougeL = "rougeL"`

      - `Alias string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `TaskType string`

        - `const EvaluationTaskMetricTaskTypeMetric EvaluationTaskMetricTaskType = "metric"`

    - `type EvaluationTaskAutoEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationQuestionConfiguration`

        - `Model string`

          model specified as `model_vendor/model_name`

        - `Prompt string`

        - `QuestionID string`

          question to be evaluated

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationQuestionTaskTypeAutoEvaluationQuestion EvaluationTaskAutoEvaluationQuestionTaskType = "auto_evaluation.question"`

    - `type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}`

      - `Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}`

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `ResponseFormat map[string, any]`

            JSON schema used for structuring the model response

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op EqEvaluationRunConditionOp`

                - `const EqEvaluationRunConditionOpEq EqEvaluationRunConditionOp = "eq"`

            - `type NeEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op NeEvaluationRunConditionOp`

                - `const NeEvaluationRunConditionOpNe NeEvaluationRunConditionOp = "ne"`

            - `type LtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LtEvaluationRunConditionOp`

                - `const LtEvaluationRunConditionOpLt LtEvaluationRunConditionOp = "lt"`

            - `type LteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op LteEvaluationRunConditionOp`

                - `const LteEvaluationRunConditionOpLte LteEvaluationRunConditionOp = "lte"`

            - `type GtEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GtEvaluationRunConditionOp`

                - `const GtEvaluationRunConditionOpGt GtEvaluationRunConditionOp = "gt"`

            - `type GteEvaluationRunCondition struct{…}`

              - `Left any`

              - `Right any`

              - `Op GteEvaluationRunConditionOp`

                - `const GteEvaluationRunConditionOpGte GteEvaluationRunConditionOp = "gte"`

            - `type AndEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op AndEvaluationRunConditionOp`

                - `const AndEvaluationRunConditionOpAnd AndEvaluationRunConditionOp = "and"`

            - `type OrEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op OrEvaluationRunConditionOp`

                - `const OrEvaluationRunConditionOpOr OrEvaluationRunConditionOp = "or"`

            - `type InEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op InEvaluationRunConditionOp`

                - `const InEvaluationRunConditionOpIn InEvaluationRunConditionOp = "in"`

            - `type NotInEvaluationRunCondition struct{…}`

              - `Left any`

              - `Operands []any`

              - `Op NotInEvaluationRunConditionOp`

                - `const NotInEvaluationRunConditionOpNotIn NotInEvaluationRunConditionOp = "not_in"`

            - `type NotEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op NotEvaluationRunConditionOp`

                - `const NotEvaluationRunConditionOpNot NotEvaluationRunConditionOp = "not"`

            - `type IsNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNullEvaluationRunConditionOp`

                - `const IsNullEvaluationRunConditionOpIsNull IsNullEvaluationRunConditionOp = "is_null"`

            - `type IsNotNullEvaluationRunCondition struct{…}`

              - `Operands []any`

              - `Op IsNotNullEvaluationRunConditionOp`

                - `const IsNotNullEvaluationRunConditionOpIsNotNull IsNotNullEvaluationRunConditionOp = "is_not_null"`

          - `SystemPrompt string`

        - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}`

          - `Choices []string`

            Choices array cannot be empty

          - `Model string`

            model specified as `model_vendor/model_name`

          - `Prompt string`

          - `InferenceArgs map[string, any]`

            Additional arguments to pass to the inference request

          - `RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnion`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOpConst EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstOp = "const"`

              - `Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnion`

                - `string`

                - `float64`

                - `bool`

            - `type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}`

              - `Path string`

              - `Op string`

                - `const EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOpVar EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarOp = "var"`

            - `type EqEvaluationRunCondition struct{…}`

            - `type NeEvaluationRunCondition struct{…}`

            - `type LtEvaluationRunCondition struct{…}`

            - `type LteEvaluationRunCondition struct{…}`

            - `type GtEvaluationRunCondition struct{…}`

            - `type GteEvaluationRunCondition struct{…}`

            - `type AndEvaluationRunCondition struct{…}`

            - `type OrEvaluationRunCondition struct{…}`

            - `type InEvaluationRunCondition struct{…}`

            - `type NotInEvaluationRunCondition struct{…}`

            - `type NotEvaluationRunCondition struct{…}`

            - `type IsNullEvaluationRunCondition struct{…}`

            - `type IsNotNullEvaluationRunCondition struct{…}`

          - `SystemPrompt string`

        - `type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}`

          - `Definition string`

          - `Name string`

          - `OutputRules []string`

          - `DataFields []string`

          - `DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnion`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig`

                - `Model string`

                - `Temperature float64`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentNameApeAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentAgentName = "APEAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentNameIfAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentAgentName = "IFAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentNameTruthfulnessAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentAgentName = "TruthfulnessAgent"`

            - `type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}`

              - `Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig`

                - `Model string`

              - `AgentName string`

                - `const AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentNameBaseAgent AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentAgentName = "BaseAgent"`

          - `OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputType`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"`

            - `const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"`

          - `OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnion`

            - `string`

            - `float64`

            - `bool`

          - `RubricID string`

          - `RubricVersion int64`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationGuidedDecodingTaskTypeAutoEvaluationGuidedDecoding EvaluationTaskAutoEvaluationGuidedDecodingTaskType = "auto_evaluation.guided_decoding"`

    - `type EvaluationTaskAutoEvaluationAgent struct{…}`

      - `Configuration AutoEvaluationAgentTaskRequestWithItemLocator`

      - `Alias string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `TaskType string`

        - `const EvaluationTaskAutoEvaluationAgentTaskTypeAutoEvaluationAgent EvaluationTaskAutoEvaluationAgentTaskType = "auto_evaluation.agent"`

    - `type EvaluationTaskContributorEvaluationQuestion struct{…}`

      - `Configuration EvaluationTaskContributorEvaluationQuestionConfiguration`

        - `Layout Container`

          - `Children []ContainerChildUnion`

            The children to be displayed within the container

            - `type Container struct{…}`

            - `type Component struct{…}`

              - `Data ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `Label string`

          - `Direction ContainerDirection`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `const ContainerDirectionRow ContainerDirection = "row"`

            - `const ContainerDirectionColumn ContainerDirection = "column"`

        - `QuestionID string`

        - `PrefillFrom string`

          Dataset column to prefill contributor question task result

        - `QueueID string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `Required bool`

          Whether the question is required to be answered

        - `RubricID string`

          ID of the rubric to use for scoring this evaluation question

      - `Alias string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `TaskType string`

        - `const EvaluationTaskContributorEvaluationQuestionTaskTypeContributorEvaluationQuestion EvaluationTaskContributorEvaluationQuestionTaskType = "contributor_evaluation.question"`

    - `type EvaluationTaskCustomFunction struct{…}`

      - `Configuration EvaluationTaskCustomFunctionConfiguration`

        Configuration for a custom Python function evaluation task.

        - `FunctionSource string`

          Python function source code

        - `ArgMapping map[string, string]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `ConfigArgs map[string, any]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `Outputs []EvaluationTaskCustomFunctionConfigurationOutput`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `Path string`

            Dot path in the custom function return value to materialize.

          - `Alias string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `Alias string`

        Alias to title the results column. Defaults to the function name.

      - `TaskType string`

        - `const EvaluationTaskCustomFunctionTaskTypeCustomFunction EvaluationTaskCustomFunctionTaskType = "custom_function"`

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  evaluation, err := client.Evaluations.New(context.TODO(), sgpdev.EvaluationNewParams{
    OfEvaluationStandaloneCreateRequest: &sgpdev.EvaluationNewParamsEvaluationEvaluationStandaloneCreateRequest{
      Data: []map[string]any{map[string]any{
      "foo": "bar",
      }},
      Name: "x",
    },
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", evaluation.ID)
}
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```
