# Evaluations

## Create Evaluation

`evaluations.create(EvaluationCreateParams**kwargs)  -> Evaluation`

**post** `/v5/evaluations`

Create an evaluation together with its items, optionally running test criteria against them.

Accepts three request shapes: standalone (inline `data`), from an existing dataset
(`dataset_id` with optional per-item references), or with a new reusable dataset created inline
from `data`. When the evaluation includes tasks that require execution (for example an LLM judge
or custom function), an async job and a Temporal workflow are started and the evaluation is
returned immediately with status `running`; task results and `error_count` populate
asynchronously. When it includes only contributor tasks, taxonomy-only input, or no tasks, no
workflow runs and it is returned with status `completed`. Optional `tasks`, `metadata`, `tags`,
and `taxonomy_params` are persisted alongside the evaluation and its items.

### Parameters

- `evaluation: Evaluation`

  - `class EvaluationEvaluationStandaloneCreateRequest: …`

    - `data: Iterable[Dict[str, object]]`

      Items to be evaluated

    - `name: str`

    - `description: Optional[str]`

    - `files: Optional[Iterable[Dict[str, str]]]`

      Files to be associated to the evaluation

    - `metadata: Optional[Dict[str, object]]`

      Optional metadata key-value pairs for the evaluation

    - `skip_prefilled_rows: Optional[bool]`

      Do not queue a contributor task for prefilled questions

    - `tags: Optional[Sequence[str]]`

      The tags associated with the evaluation

    - `tasks: Optional[Iterable[EvaluationTaskParam]]`

      Tasks allow you to augment and evaluate your data

      - `class ChatCompletionEvaluationTask: …`

        - `configuration: ChatCompletionEvaluationTaskConfiguration`

          - `messages: Union[List[Dict[str, object]], ItemLocator]`

            openai standard message format

            - `List[Dict[str, object]]`

            - `str`

          - `model: str`

            model specified as `model_vendor/model`, for example `openai/gpt-4o`

          - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

            - `Dict[str, object]`

            - `str`

          - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

            Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

            - `float`

            - `str`

          - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Deprecated in favor of tool_choice. Controls which function is called by the model.

            - `Dict[str, object]`

            - `str`

          - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

            Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

            - `List[Dict[str, object]]`

            - `str`

          - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

            Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

            - `Dict[str, int]`

            - `str`

          - `logprobs: Optional[Union[bool, ItemLocator, null]]`

            Whether to return log probabilities of the output tokens or not.

            - `bool`

            - `str`

          - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

            An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

            - `int`

            - `str`

          - `max_tokens: Optional[Union[int, ItemLocator, null]]`

            Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

            - `int`

            - `str`

          - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

            Developer-defined tags and values used for filtering completions in the dashboard.

            - `Dict[str, str]`

            - `str`

          - `modalities: Optional[Union[List[str], ItemLocator, null]]`

            Output types that you would like the model to generate for this request.

            - `List[str]`

            - `str`

          - `n: Optional[Union[int, ItemLocator, null]]`

            How many chat completion choices to generate for each input message.

            - `int`

            - `str`

          - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

            Whether to enable parallel function calling during tool use.

            - `bool`

            - `str`

          - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Static predicted output content, such as the content of a text file being regenerated.

            - `Dict[str, object]`

            - `str`

          - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

            Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

            - `float`

            - `str`

          - `reasoning_effort: Optional[str]`

            For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

          - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

            An object specifying the format that the model must output.

            - `Dict[str, object]`

            - `str`

          - `seed: Optional[Union[int, ItemLocator, null]]`

            If specified, system will attempt to sample deterministically for repeated requests with same seed.

            - `int`

            - `str`

          - `stop: Optional[Union[str, List[str], null]]`

            Up to 4 sequences where the API will stop generating further tokens.

            - `str`

            - `List[str]`

          - `store: Optional[Union[bool, ItemLocator, null]]`

            Whether to store the output for use in model distillation or evals products.

            - `bool`

            - `str`

          - `temperature: Optional[Union[float, ItemLocator, null]]`

            What sampling temperature to use. Higher values make output more random, lower more focused.

            - `float`

            - `str`

          - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

            Controls which tool is called by the model. Values: none, auto, required, or specific tool.

            - `str`

            - `Dict[str, object]`

          - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

            A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

            - `List[Dict[str, object]]`

            - `str`

          - `top_k: Optional[Union[int, ItemLocator, null]]`

            Only sample from the top K options for each subsequent token

            - `int`

            - `str`

          - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

            Number of most likely tokens to return at each position, with associated log probability.

            - `int`

            - `str`

          - `top_p: Optional[Union[float, ItemLocator, null]]`

            Alternative to temperature. Only tokens comprising top_p probability mass are considered.

            - `float`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `chat_completion`

        - `task_type: Optional[Literal["chat_completion"]]`

          - `"chat_completion"`

      - `class GenericInferenceEvaluationTask: …`

        - `configuration: GenericInferenceEvaluationTaskConfiguration`

          - `model: str`

            model specified as `vendor/name` (ex. openai/gpt-5)

          - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Arguments passed into model

            - `Dict[str, object]`

            - `str`

          - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

            Vendor specific configuration

            - `class LaunchInferenceConfiguration: …`

              - `num_retries: Optional[int]`

              - `timeout_seconds: Optional[int]`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `inference`

        - `task_type: Optional[Literal["inference"]]`

          - `"inference"`

      - `class ApplicationVariantV1EvaluationTask: …`

        - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

          - `application_variant_id: str`

          - `inputs: Union[Dict[str, object], ItemLocator]`

            Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

            - `Dict[str, object]`

            - `str`

          - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

            History of the application

            - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

              - `request: str`

                Request inputs

              - `response: str`

                Response outputs

              - `session_data: Optional[Dict[str, object]]`

                Session data corresponding to the request response pair

            - `str`

          - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

            - `Dict[str, object]`

            - `str`

          - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

            Optional overrides for the application

            - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

              Execution override options for agentic applications

              - `concurrent: Optional[bool]`

              - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

                - `current_node: str`

                - `state: Dict[str, object]`

              - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

                - `duration_ms: int`

                - `node_id: str`

                - `operation_input: str`

                - `operation_output: str`

                - `operation_type: str`

                - `start_timestamp: str`

                - `workflow_id: str`

                - `operation_metadata: Optional[Dict[str, object]]`

              - `return_span: Optional[bool]`

              - `use_channels: Optional[bool]`

            - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

              - `artifact_ids_filter: Optional[List[str]]`

              - `artifact_name_regex: Optional[List[str]]`

              - `type: Optional[Literal["knowledge_base_schema"]]`

                - `"knowledge_base_schema"`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `application_variant`

        - `task_type: Optional[Literal["application_variant"]]`

          - `"application_variant"`

      - `class AgentexOutputEvaluationTask: …`

        - `configuration: AgentexOutputEvaluationTaskConfiguration`

          - `agentex_agent_id: str`

            The ID of the Agentex agent to use

          - `input_column: Union[str, Dict[str, object], List[object]]`

            The dataset column to use as input for the agent

            - `str`

            - `Dict[str, object]`

            - `List[object]`

          - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

            - `Dict[str, object]`

            - `str`

          - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

            How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

            - `"first_message"`

            - `"turn_quiescence"`

          - `deployment_id: Optional[str]`

            Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

          - `include_traces: Optional[Union[bool, ItemLocator, null]]`

            Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

            - `bool`

            - `str`

          - `input_mode: Optional[Literal["text", "data"]]`

            How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

            - `"text"`

            - `"data"`

          - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

            Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

            - `int`

            - `str`

          - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

            Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

            - `int`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `agentex_output`

        - `task_type: Optional[Literal["agentex_output"]]`

          - `"agentex_output"`

      - `class MetricEvaluationTask: …`

        - `configuration: MetricEvaluationTaskConfiguration`

          - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["bleu"]`

              - `"bleu"`

          - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["meteor"]`

              - `"meteor"`

          - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["cosine_similarity"]`

              - `"cosine_similarity"`

          - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["f1"]`

              - `"f1"`

          - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["rouge1"]`

              - `"rouge1"`

          - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["rouge2"]`

              - `"rouge2"`

          - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["rougeL"]`

              - `"rougeL"`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the metric type specified in the configuration

        - `task_type: Optional[Literal["metric"]]`

          - `"metric"`

      - `class AutoEvaluationQuestionTask: …`

        - `configuration: AutoEvaluationQuestionTaskConfiguration`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `question_id: str`

            question to be evaluated

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `auto_evaluation_question`

        - `task_type: Optional[Literal["auto_evaluation.question"]]`

          - `"auto_evaluation.question"`

      - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

        - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

            - `model: str`

              model specified as `model_vendor/model_name`

            - `prompt: str`

            - `response_format: Dict[str, object]`

              JSON schema used for structuring the model response

            - `inference_args: Optional[Dict[str, object]]`

              Additional arguments to pass to the inference request

            - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

                - `op: Optional[Literal["const"]]`

                  - `"const"`

                - `value: Optional[Union[str, float, bool, null]]`

                  - `str`

                  - `float`

                  - `bool`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

                - `path: str`

                - `op: Optional[Literal["var"]]`

                  - `"var"`

              - `class EqEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["eq"]]`

                  - `"eq"`

              - `class NeEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["ne"]]`

                  - `"ne"`

              - `class LtEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["lt"]]`

                  - `"lt"`

              - `class LteEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["lte"]]`

                  - `"lte"`

              - `class GtEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["gt"]]`

                  - `"gt"`

              - `class GteEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["gte"]]`

                  - `"gte"`

              - `class AndEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["and"]]`

                  - `"and"`

              - `class OrEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["or"]]`

                  - `"or"`

              - `class InEvaluationRunCondition: …`

                - `left: object`

                - `operands: List[object]`

                - `op: Optional[Literal["in"]]`

                  - `"in"`

              - `class NotInEvaluationRunCondition: …`

                - `left: object`

                - `operands: List[object]`

                - `op: Optional[Literal["not_in"]]`

                  - `"not_in"`

              - `class NotEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["not"]]`

                  - `"not"`

              - `class IsNullEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["is_null"]]`

                  - `"is_null"`

              - `class IsNotNullEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["is_not_null"]]`

                  - `"is_not_null"`

            - `system_prompt: Optional[str]`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

            - `choices: List[str]`

              Choices array cannot be empty

            - `model: str`

              model specified as `model_vendor/model_name`

            - `prompt: str`

            - `inference_args: Optional[Dict[str, object]]`

              Additional arguments to pass to the inference request

            - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

                - `op: Optional[Literal["const"]]`

                  - `"const"`

                - `value: Optional[Union[str, float, bool, null]]`

                  - `str`

                  - `float`

                  - `bool`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

                - `path: str`

                - `op: Optional[Literal["var"]]`

                  - `"var"`

              - `class EqEvaluationRunCondition: …`

              - `class NeEvaluationRunCondition: …`

              - `class LtEvaluationRunCondition: …`

              - `class LteEvaluationRunCondition: …`

              - `class GtEvaluationRunCondition: …`

              - `class GteEvaluationRunCondition: …`

              - `class AndEvaluationRunCondition: …`

              - `class OrEvaluationRunCondition: …`

              - `class InEvaluationRunCondition: …`

              - `class NotInEvaluationRunCondition: …`

              - `class NotEvaluationRunCondition: …`

              - `class IsNullEvaluationRunCondition: …`

              - `class IsNotNullEvaluationRunCondition: …`

            - `system_prompt: Optional[str]`

          - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

            - `definition: str`

            - `name: str`

            - `output_rules: List[str]`

            - `data_fields: Optional[List[str]]`

            - `designated_to: Optional[DesignatedTo]`

              - `class DesignatedToApeAgent: …`

                - `config: DesignatedToApeAgentConfig`

                  - `model: Optional[str]`

                  - `temperature: Optional[float]`

                - `agent_name: Optional[Literal["APEAgent"]]`

                  - `"APEAgent"`

              - `class DesignatedToIfAgent: …`

                - `config: DesignatedToIfAgentConfig`

                  - `model: Optional[str]`

                - `agent_name: Optional[Literal["IFAgent"]]`

                  - `"IFAgent"`

              - `class DesignatedToTruthfulnessAgent: …`

                - `config: DesignatedToTruthfulnessAgentConfig`

                  - `model: Optional[str]`

                - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                  - `"TruthfulnessAgent"`

              - `class DesignatedToBaseAgent: …`

                - `config: DesignatedToBaseAgentConfig`

                  - `model: Optional[str]`

                - `agent_name: Optional[Literal["BaseAgent"]]`

                  - `"BaseAgent"`

            - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

              - `"text"`

              - `"integer"`

              - `"float"`

              - `"boolean"`

            - `output_values: Optional[List[Union[str, float, bool]]]`

              - `str`

              - `float`

              - `bool`

            - `rubric_id: Optional[str]`

            - `rubric_version: Optional[int]`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

        - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

          - `"auto_evaluation.guided_decoding"`

      - `class AutoEvaluationAgentEvaluationTask: …`

        - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `auto_evaluation_agent`

        - `task_type: Optional[Literal["auto_evaluation.agent"]]`

          - `"auto_evaluation.agent"`

      - `class ContributorEvaluationQuestionTask: …`

        - `configuration: ContributorEvaluationQuestionTaskConfiguration`

          - `layout: Container`

            - `children: List[Child]`

              The children to be displayed within the container

              - `class Container: …`

              - `class Component: …`

                - `data: ItemLocator`

                  A pointer to the data in each evaluation item to be displayed within the component

                - `label: Optional[str]`

            - `direction: Optional[Literal["row", "column"]]`

              The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

              - `"row"`

              - `"column"`

          - `question_id: str`

          - `prefill_from: Optional[str]`

            Dataset column to prefill contributor question task result

          - `queue_id: Optional[str]`

            The contributor annotation queue to include this task in. Defaults to `default`

          - `required: Optional[bool]`

            Whether the question is required to be answered

          - `rubric_id: Optional[str]`

            ID of the rubric to use for scoring this evaluation question

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `contributor_evaluation_question`

        - `task_type: Optional[Literal["contributor_evaluation.question"]]`

          - `"contributor_evaluation.question"`

      - `class CustomFunctionEvaluationTask: …`

        - `configuration: CustomFunctionEvaluationTaskConfiguration`

          Configuration for a custom Python function evaluation task.

          - `function_source: str`

            Python function source code

          - `arg_mapping: Optional[Dict[str, str]]`

            Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

          - `config_args: Optional[Dict[str, object]]`

            Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

          - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

            Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

            - `path: str`

              Dot path in the custom function return value to materialize.

            - `alias: Optional[str]`

              Result column alias. Defaults to path with dots replaced by underscores.

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the function name.

        - `task_type: Optional[Literal["custom_function"]]`

          - `"custom_function"`

    - `taxonomy_params: Optional[Dict[str, object]]`

      Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

  - `class EvaluationEvaluationFromDatasetCreateRequest: …`

    - `dataset_id: str`

      The ID of the dataset containing the items referenced by the `data` field

    - `name: str`

    - `data: Optional[Iterable[EvaluationEvaluationFromDatasetCreateRequestData]]`

      Items to be evaluated, including references to the input dataset

      - `dataset_item_id: str`

    - `description: Optional[str]`

    - `metadata: Optional[Dict[str, object]]`

      Optional metadata key-value pairs for the evaluation

    - `skip_prefilled_rows: Optional[bool]`

      Do not queue a contributor task for prefilled questions

    - `tags: Optional[Sequence[str]]`

      The tags associated with the evaluation

    - `tasks: Optional[Iterable[EvaluationTaskParam]]`

      Tasks allow you to augment and evaluate your data

      - `class ChatCompletionEvaluationTask: …`

      - `class GenericInferenceEvaluationTask: …`

      - `class ApplicationVariantV1EvaluationTask: …`

      - `class AgentexOutputEvaluationTask: …`

      - `class MetricEvaluationTask: …`

      - `class AutoEvaluationQuestionTask: …`

      - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `class AutoEvaluationAgentEvaluationTask: …`

      - `class ContributorEvaluationQuestionTask: …`

      - `class CustomFunctionEvaluationTask: …`

    - `taxonomy_params: Optional[Dict[str, object]]`

      Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

  - `class EvaluationEvaluationWithDatasetCreateRequest: …`

    - `data: Iterable[Dict[str, object]]`

      Items to be evaluated

    - `dataset: EvaluationEvaluationWithDatasetCreateRequestDataset`

      Create a reusable dataset from items in the `data` field

      - `name: str`

      - `description: Optional[str]`

      - `keys: Optional[Sequence[str]]`

        Keys from items in the `data` field that should be included in the dataset. If not provided, all keys will be included.

      - `tags: Optional[Sequence[str]]`

        The tags associated with the entity

    - `name: str`

    - `description: Optional[str]`

    - `files: Optional[Iterable[Dict[str, str]]]`

      Files to be associated to the evaluation

    - `metadata: Optional[Dict[str, object]]`

      Optional metadata key-value pairs for the evaluation

    - `skip_prefilled_rows: Optional[bool]`

      Do not queue a contributor task for prefilled questions

    - `tags: Optional[Sequence[str]]`

      The tags associated with the evaluation

    - `tasks: Optional[Iterable[EvaluationTaskParam]]`

      Tasks allow you to augment and evaluate your data

      - `class ChatCompletionEvaluationTask: …`

      - `class GenericInferenceEvaluationTask: …`

      - `class ApplicationVariantV1EvaluationTask: …`

      - `class AgentexOutputEvaluationTask: …`

      - `class MetricEvaluationTask: …`

      - `class AutoEvaluationQuestionTask: …`

      - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `class AutoEvaluationAgentEvaluationTask: …`

      - `class ContributorEvaluationQuestionTask: …`

      - `class CustomFunctionEvaluationTask: …`

    - `taxonomy_params: Optional[Dict[str, object]]`

      Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation = client.evaluations.create(
    evaluation={
        "data": [{
            "foo": "bar"
        }],
        "name": "x",
    },
)
print(evaluation.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## List Evaluations

`evaluations.list(EvaluationListParams**kwargs)  -> SyncCursorPage[Evaluation]`

**get** `/v5/evaluations`

List evaluations for the account, with pagination.

Supports filtering by case-insensitive name substring and by tags; archived
evaluations are excluded unless `include_archived` is set. Pass the `tasks` view to include each
evaluation's task configurations in the response. Use this for simple name or tag lookups;
to filter on metadata key-value pairs or status, use the filter endpoint instead.

### Parameters

- `ending_before: Optional[str]`

- `include_archived: Optional[bool]`

- `limit: Optional[int]`

- `name: Optional[str]`

- `sort_by: Optional[str]`

- `sort_order: Optional[SortOrder]`

  - `"asc"`

  - `"desc"`

- `starting_after: Optional[str]`

- `tags: Optional[Sequence[str]]`

- `views: Optional[List[EvaluationViews]]`

  - `"tasks"`

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
page = client.evaluations.list()
page = page.items[0]
print(page.id)
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Get Evaluation

`evaluations.retrieve(strevaluation_id, EvaluationRetrieveParams**kwargs)  -> Evaluation`

**get** `/v5/evaluations/{evaluation_id}`

Retrieve a single evaluation by ID.

Returns the evaluation with its datasets, async-job progress, metadata, and task-error count.
Archived evaluations are excluded unless `include_archived` is set. Pass the `tasks` view to
include the evaluation's task configurations in the response.

### Parameters

- `evaluation_id: str`

- `include_archived: Optional[bool]`

- `views: Optional[List[EvaluationViews]]`

  - `"tasks"`

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation = client.evaluations.retrieve(
    evaluation_id="evaluation_id",
)
print(evaluation.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Archive Evaluation

`evaluations.archive(strevaluation_id)  -> Evaluation`

**delete** `/v5/evaluations/{evaluation_id}`

Archive (soft-delete) an evaluation.

Sets the evaluation's archived timestamp rather than permanently deleting it, and cascades the
archive to the evaluation's items and dashboards while removing it from any evaluation groups.
The evaluation can later be brought back with a restore request to the update endpoint.

### Parameters

- `evaluation_id: str`

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation = client.evaluations.archive(
    "evaluation_id",
)
print(evaluation.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Update or Restore Evaluation

`evaluations.update(strevaluation_id, EvaluationUpdateParams**kwargs)  -> Evaluation`

**patch** `/v5/evaluations/{evaluation_id}`

Update an evaluation's mutable fields, or restore it from the archive.

The action is selected by the request body: a restore request un-archives the evaluation and
cascades the restore to its items and dashboards, while any other body applies a partial update
to fields such as name, description, tags, and metadata (metadata is applied as an RFC 7396
merge patch). Updating an already-archived evaluation is rejected — restore it first. The
evaluation row is locked for the duration of the write to avoid concurrent-update races.

### Parameters

- `evaluation_id: str`

- `evaluation: Evaluation`

  - `class EvaluationPartialEvaluationUpdateRequest: …`

    - `description: Optional[str]`

    - `metadata: Optional[Dict[str, object]]`

      Optional metadata key-value pairs for the evaluation

    - `name: Optional[str]`

    - `tags: Optional[Sequence[str]]`

      The tags associated with the evaluation

  - `class RestoreRequest: …`

    - `restore: Literal[true]`

      Set to true to restore the entity from the database.

      - `true`

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation = client.evaluations.update(
    evaluation_id="evaluation_id",
    evaluation={},
)
print(evaluation.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Get Evaluation Data Schema

`evaluations.retrieve_schema(strevaluation_id, EvaluationRetrieveSchemaParams**kwargs)  -> EvaluationSchemaResponse`

**get** `/v5/evaluations/{evaluation_id}/schema`

Describe the data schema of an evaluation's items.

Inspects the item `data` and task-result fields and returns each discovered field with its
flattened key path, JSON type, source, and the number of items containing it, ordered
alphabetically by field name. For large evaluations the schema may be inferred from a sample of
items, in which case `is_sampled` is set and `sample_size` reports how many were analyzed. Set
`include_archived` to include archived items in the analysis.

### Parameters

- `evaluation_id: str`

- `include_archived: Optional[bool]`

  Include archived items in schema analysis

### Returns

- `class EvaluationSchemaResponse: …`

  Schema information for an evaluation's item data structure

  - `evaluation_id: str`

    The ID of the evaluation

  - `fields: List[Field]`

    List of all discovered fields, ordered alphabetically by field_name

    - `data_type: str`

      JSON type: 'string', 'number', 'boolean', 'object', 'array', or 'null'

    - `field_name: str`

      The flattened JSON key path (e.g., 'metadata.category')

    - `item_count: int`

      Number of evaluation items containing this field

    - `source: Literal["data", "task_result_cache"]`

      The source of the field: 'data' or 'task_result_cache'

      - `"data"`

      - `"task_result_cache"`

    - `object: Optional[Literal["field_schema"]]`

      - `"field_schema"`

  - `total_items: int`

    Total number of evaluation items

  - `is_sampled: Optional[bool]`

    Whether schema was computed from a sample of items (for large evaluations)

  - `object: Optional[Literal["evaluation_schema"]]`

    - `"evaluation_schema"`

  - `sample_size: Optional[int]`

    Number of items sampled for schema inference, if applicable

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation_schema_response = client.evaluations.retrieve_schema(
    evaluation_id="evaluation_id",
)
print(evaluation_schema_response.evaluation_id)
```

#### Response

```json
{
  "evaluation_id": "evaluation_id",
  "fields": [
    {
      "data_type": "data_type",
      "field_name": "field_name",
      "item_count": 0,
      "source": "data",
      "object": "field_schema"
    }
  ],
  "total_items": 0,
  "is_sampled": true,
  "object": "evaluation_schema",
  "sample_size": 0
}
```

## Filter Evaluations

`evaluations.filter(EvaluationFilterParams**kwargs)  -> SyncCursorPage[Evaluation]`

**post** `/v5/evaluations/filter`

Filter evaluations by metadata, status, and tags.

Accepts up to 10 filters combined with AND logic, each comparing a key against a value with an
operator (`==`, `!=`, `>=`, `<=`, `IN`, `NOT_IN`). Filter on metadata keys returned by the
metadata-keys endpoint, plus the built-in `status` and `tag` keys. Archived evaluations are
excluded unless `include_archived` is set, and the `tasks` view includes task configurations in
each result. Use this for metadata or status filtering; for simple name or tag lookups the list
endpoint is sufficient.

### Parameters

- `filters: Iterable[Filter]`

  List of metadata filters to apply (maximum 10)

  - `key: str`

    The metadata key to filter on

  - `operator: Literal["==", "!=", ">=", 3 more]`

    The comparison operator to use

    - `"=="`

    - `"!="`

    - `">="`

    - `"<="`

    - `"IN"`

    - `"NOT_IN"`

  - `value: str`

    The value to compare against (string for all types)

  - `object: Optional[Literal["metadata_filter"]]`

    - `"metadata_filter"`

- `ending_before: Optional[str]`

- `include_archived: Optional[bool]`

- `limit: Optional[int]`

- `sort_by: Optional[str]`

- `sort_order: Optional[SortOrder]`

  - `"asc"`

  - `"desc"`

- `starting_after: Optional[str]`

- `views: Optional[List[EvaluationViews]]`

  - `"tasks"`

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
page = client.evaluations.filter(
    filters=[{
        "key": "key",
        "operator": "==",
        "value": "value",
    }],
)
page = page.items[0]
print(page.id)
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Get Evaluation Taxonomy

`evaluations.retrieve_taxonomy(strevaluation_id)  -> EvaluationRetrieveTaxonomyResponse`

**get** `/v5/evaluations/{evaluation_id}/taxonomy`

Get the taxonomy JSON for an evaluation's contributor question tasks.

Returns the raw taxonomy document stored for the evaluation. Responds with a not-found error if
the evaluation has no taxonomy.

### Parameters

- `evaluation_id: str`

### Returns

- `Dict[str, object]`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.evaluations.retrieve_taxonomy(
    "evaluation_id",
)
print(response)
```

#### Response

```json
{
  "foo": "bar"
}
```

## Domain Types

### And Evaluation Run Condition

- `class AndEvaluationRunCondition: …`

  - `operands: List[object]`

  - `op: Optional[Literal["and"]]`

    - `"and"`

### Auto Evaluation Agent Task Request With Item Locator

- `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

  - `definition: str`

  - `name: str`

  - `output_rules: List[str]`

  - `data_fields: Optional[List[str]]`

  - `designated_to: Optional[DesignatedTo]`

    - `class DesignatedToApeAgent: …`

      - `config: DesignatedToApeAgentConfig`

        - `model: Optional[str]`

        - `temperature: Optional[float]`

      - `agent_name: Optional[Literal["APEAgent"]]`

        - `"APEAgent"`

    - `class DesignatedToIfAgent: …`

      - `config: DesignatedToIfAgentConfig`

        - `model: Optional[str]`

      - `agent_name: Optional[Literal["IFAgent"]]`

        - `"IFAgent"`

    - `class DesignatedToTruthfulnessAgent: …`

      - `config: DesignatedToTruthfulnessAgentConfig`

        - `model: Optional[str]`

      - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

        - `"TruthfulnessAgent"`

    - `class DesignatedToBaseAgent: …`

      - `config: DesignatedToBaseAgentConfig`

        - `model: Optional[str]`

      - `agent_name: Optional[Literal["BaseAgent"]]`

        - `"BaseAgent"`

  - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

    - `"text"`

    - `"integer"`

    - `"float"`

    - `"boolean"`

  - `output_values: Optional[List[Union[str, float, bool]]]`

    - `str`

    - `float`

    - `bool`

  - `rubric_id: Optional[str]`

  - `rubric_version: Optional[int]`

### Eq Evaluation Run Condition

- `class EqEvaluationRunCondition: …`

  - `left: object`

  - `right: object`

  - `op: Optional[Literal["eq"]]`

    - `"eq"`

### Evaluation

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Evaluation Schema Response

- `class EvaluationSchemaResponse: …`

  Schema information for an evaluation's item data structure

  - `evaluation_id: str`

    The ID of the evaluation

  - `fields: List[Field]`

    List of all discovered fields, ordered alphabetically by field_name

    - `data_type: str`

      JSON type: 'string', 'number', 'boolean', 'object', 'array', or 'null'

    - `field_name: str`

      The flattened JSON key path (e.g., 'metadata.category')

    - `item_count: int`

      Number of evaluation items containing this field

    - `source: Literal["data", "task_result_cache"]`

      The source of the field: 'data' or 'task_result_cache'

      - `"data"`

      - `"task_result_cache"`

    - `object: Optional[Literal["field_schema"]]`

      - `"field_schema"`

  - `total_items: int`

    Total number of evaluation items

  - `is_sampled: Optional[bool]`

    Whether schema was computed from a sample of items (for large evaluations)

  - `object: Optional[Literal["evaluation_schema"]]`

    - `"evaluation_schema"`

  - `sample_size: Optional[int]`

    Number of items sampled for schema inference, if applicable

### Evaluation Task

- `EvaluationTask`

  - `class ChatCompletionEvaluationTask: …`

    - `configuration: ChatCompletionEvaluationTaskConfiguration`

      - `messages: Union[List[Dict[str, object]], ItemLocator]`

        openai standard message format

        - `List[Dict[str, object]]`

        - `str`

      - `model: str`

        model specified as `model_vendor/model`, for example `openai/gpt-4o`

      - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

        - `Dict[str, object]`

        - `str`

      - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

        Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

        - `float`

        - `str`

      - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Deprecated in favor of tool_choice. Controls which function is called by the model.

        - `Dict[str, object]`

        - `str`

      - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

        Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

        - `List[Dict[str, object]]`

        - `str`

      - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

        Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

        - `Dict[str, int]`

        - `str`

      - `logprobs: Optional[Union[bool, ItemLocator, null]]`

        Whether to return log probabilities of the output tokens or not.

        - `bool`

        - `str`

      - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

        An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

        - `int`

        - `str`

      - `max_tokens: Optional[Union[int, ItemLocator, null]]`

        Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

        - `int`

        - `str`

      - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

        Developer-defined tags and values used for filtering completions in the dashboard.

        - `Dict[str, str]`

        - `str`

      - `modalities: Optional[Union[List[str], ItemLocator, null]]`

        Output types that you would like the model to generate for this request.

        - `List[str]`

        - `str`

      - `n: Optional[Union[int, ItemLocator, null]]`

        How many chat completion choices to generate for each input message.

        - `int`

        - `str`

      - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

        Whether to enable parallel function calling during tool use.

        - `bool`

        - `str`

      - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Static predicted output content, such as the content of a text file being regenerated.

        - `Dict[str, object]`

        - `str`

      - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

        Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

        - `float`

        - `str`

      - `reasoning_effort: Optional[str]`

        For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

      - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

        An object specifying the format that the model must output.

        - `Dict[str, object]`

        - `str`

      - `seed: Optional[Union[int, ItemLocator, null]]`

        If specified, system will attempt to sample deterministically for repeated requests with same seed.

        - `int`

        - `str`

      - `stop: Optional[Union[str, List[str], null]]`

        Up to 4 sequences where the API will stop generating further tokens.

        - `str`

        - `List[str]`

      - `store: Optional[Union[bool, ItemLocator, null]]`

        Whether to store the output for use in model distillation or evals products.

        - `bool`

        - `str`

      - `temperature: Optional[Union[float, ItemLocator, null]]`

        What sampling temperature to use. Higher values make output more random, lower more focused.

        - `float`

        - `str`

      - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

        Controls which tool is called by the model. Values: none, auto, required, or specific tool.

        - `str`

        - `Dict[str, object]`

      - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

        A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

        - `List[Dict[str, object]]`

        - `str`

      - `top_k: Optional[Union[int, ItemLocator, null]]`

        Only sample from the top K options for each subsequent token

        - `int`

        - `str`

      - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

        Number of most likely tokens to return at each position, with associated log probability.

        - `int`

        - `str`

      - `top_p: Optional[Union[float, ItemLocator, null]]`

        Alternative to temperature. Only tokens comprising top_p probability mass are considered.

        - `float`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `chat_completion`

    - `task_type: Optional[Literal["chat_completion"]]`

      - `"chat_completion"`

  - `class GenericInferenceEvaluationTask: …`

    - `configuration: GenericInferenceEvaluationTaskConfiguration`

      - `model: str`

        model specified as `vendor/name` (ex. openai/gpt-5)

      - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Arguments passed into model

        - `Dict[str, object]`

        - `str`

      - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

        Vendor specific configuration

        - `class LaunchInferenceConfiguration: …`

          - `num_retries: Optional[int]`

          - `timeout_seconds: Optional[int]`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `inference`

    - `task_type: Optional[Literal["inference"]]`

      - `"inference"`

  - `class ApplicationVariantV1EvaluationTask: …`

    - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

      - `application_variant_id: str`

      - `inputs: Union[Dict[str, object], ItemLocator]`

        Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

        - `Dict[str, object]`

        - `str`

      - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

        History of the application

        - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

          - `request: str`

            Request inputs

          - `response: str`

            Response outputs

          - `session_data: Optional[Dict[str, object]]`

            Session data corresponding to the request response pair

        - `str`

      - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

        - `Dict[str, object]`

        - `str`

      - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

        Optional overrides for the application

        - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

          Execution override options for agentic applications

          - `concurrent: Optional[bool]`

          - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

            - `current_node: str`

            - `state: Dict[str, object]`

          - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

            - `duration_ms: int`

            - `node_id: str`

            - `operation_input: str`

            - `operation_output: str`

            - `operation_type: str`

            - `start_timestamp: str`

            - `workflow_id: str`

            - `operation_metadata: Optional[Dict[str, object]]`

          - `return_span: Optional[bool]`

          - `use_channels: Optional[bool]`

        - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

          - `artifact_ids_filter: Optional[List[str]]`

          - `artifact_name_regex: Optional[List[str]]`

          - `type: Optional[Literal["knowledge_base_schema"]]`

            - `"knowledge_base_schema"`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `application_variant`

    - `task_type: Optional[Literal["application_variant"]]`

      - `"application_variant"`

  - `class AgentexOutputEvaluationTask: …`

    - `configuration: AgentexOutputEvaluationTaskConfiguration`

      - `agentex_agent_id: str`

        The ID of the Agentex agent to use

      - `input_column: Union[str, Dict[str, object], List[object]]`

        The dataset column to use as input for the agent

        - `str`

        - `Dict[str, object]`

        - `List[object]`

      - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

        - `Dict[str, object]`

        - `str`

      - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

        How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

        - `"first_message"`

        - `"turn_quiescence"`

      - `deployment_id: Optional[str]`

        Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

      - `include_traces: Optional[Union[bool, ItemLocator, null]]`

        Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

        - `bool`

        - `str`

      - `input_mode: Optional[Literal["text", "data"]]`

        How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

        - `"text"`

        - `"data"`

      - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

        Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

        - `int`

        - `str`

      - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

        Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

        - `int`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `agentex_output`

    - `task_type: Optional[Literal["agentex_output"]]`

      - `"agentex_output"`

  - `class MetricEvaluationTask: …`

    - `configuration: MetricEvaluationTaskConfiguration`

      - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["bleu"]`

          - `"bleu"`

      - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["meteor"]`

          - `"meteor"`

      - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["cosine_similarity"]`

          - `"cosine_similarity"`

      - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["f1"]`

          - `"f1"`

      - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["rouge1"]`

          - `"rouge1"`

      - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["rouge2"]`

          - `"rouge2"`

      - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["rougeL"]`

          - `"rougeL"`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the metric type specified in the configuration

    - `task_type: Optional[Literal["metric"]]`

      - `"metric"`

  - `class AutoEvaluationQuestionTask: …`

    - `configuration: AutoEvaluationQuestionTaskConfiguration`

      - `model: str`

        model specified as `model_vendor/model_name`

      - `prompt: str`

      - `question_id: str`

        question to be evaluated

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `auto_evaluation_question`

    - `task_type: Optional[Literal["auto_evaluation.question"]]`

      - `"auto_evaluation.question"`

  - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

    - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

      - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `response_format: Dict[str, object]`

          JSON schema used for structuring the model response

        - `inference_args: Optional[Dict[str, object]]`

          Additional arguments to pass to the inference request

        - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

            - `op: Optional[Literal["const"]]`

              - `"const"`

            - `value: Optional[Union[str, float, bool, null]]`

              - `str`

              - `float`

              - `bool`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

            - `path: str`

            - `op: Optional[Literal["var"]]`

              - `"var"`

          - `class EqEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["eq"]]`

              - `"eq"`

          - `class NeEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["ne"]]`

              - `"ne"`

          - `class LtEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["lt"]]`

              - `"lt"`

          - `class LteEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["lte"]]`

              - `"lte"`

          - `class GtEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["gt"]]`

              - `"gt"`

          - `class GteEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["gte"]]`

              - `"gte"`

          - `class AndEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["and"]]`

              - `"and"`

          - `class OrEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["or"]]`

              - `"or"`

          - `class InEvaluationRunCondition: …`

            - `left: object`

            - `operands: List[object]`

            - `op: Optional[Literal["in"]]`

              - `"in"`

          - `class NotInEvaluationRunCondition: …`

            - `left: object`

            - `operands: List[object]`

            - `op: Optional[Literal["not_in"]]`

              - `"not_in"`

          - `class NotEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["not"]]`

              - `"not"`

          - `class IsNullEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["is_null"]]`

              - `"is_null"`

          - `class IsNotNullEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["is_not_null"]]`

              - `"is_not_null"`

        - `system_prompt: Optional[str]`

      - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

        - `choices: List[str]`

          Choices array cannot be empty

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `inference_args: Optional[Dict[str, object]]`

          Additional arguments to pass to the inference request

        - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

            - `op: Optional[Literal["const"]]`

              - `"const"`

            - `value: Optional[Union[str, float, bool, null]]`

              - `str`

              - `float`

              - `bool`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

            - `path: str`

            - `op: Optional[Literal["var"]]`

              - `"var"`

          - `class EqEvaluationRunCondition: …`

          - `class NeEvaluationRunCondition: …`

          - `class LtEvaluationRunCondition: …`

          - `class LteEvaluationRunCondition: …`

          - `class GtEvaluationRunCondition: …`

          - `class GteEvaluationRunCondition: …`

          - `class AndEvaluationRunCondition: …`

          - `class OrEvaluationRunCondition: …`

          - `class InEvaluationRunCondition: …`

          - `class NotInEvaluationRunCondition: …`

          - `class NotEvaluationRunCondition: …`

          - `class IsNullEvaluationRunCondition: …`

          - `class IsNotNullEvaluationRunCondition: …`

        - `system_prompt: Optional[str]`

      - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

        - `definition: str`

        - `name: str`

        - `output_rules: List[str]`

        - `data_fields: Optional[List[str]]`

        - `designated_to: Optional[DesignatedTo]`

          - `class DesignatedToApeAgent: …`

            - `config: DesignatedToApeAgentConfig`

              - `model: Optional[str]`

              - `temperature: Optional[float]`

            - `agent_name: Optional[Literal["APEAgent"]]`

              - `"APEAgent"`

          - `class DesignatedToIfAgent: …`

            - `config: DesignatedToIfAgentConfig`

              - `model: Optional[str]`

            - `agent_name: Optional[Literal["IFAgent"]]`

              - `"IFAgent"`

          - `class DesignatedToTruthfulnessAgent: …`

            - `config: DesignatedToTruthfulnessAgentConfig`

              - `model: Optional[str]`

            - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

              - `"TruthfulnessAgent"`

          - `class DesignatedToBaseAgent: …`

            - `config: DesignatedToBaseAgentConfig`

              - `model: Optional[str]`

            - `agent_name: Optional[Literal["BaseAgent"]]`

              - `"BaseAgent"`

        - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

          - `"text"`

          - `"integer"`

          - `"float"`

          - `"boolean"`

        - `output_values: Optional[List[Union[str, float, bool]]]`

          - `str`

          - `float`

          - `bool`

        - `rubric_id: Optional[str]`

        - `rubric_version: Optional[int]`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

    - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

      - `"auto_evaluation.guided_decoding"`

  - `class AutoEvaluationAgentEvaluationTask: …`

    - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `auto_evaluation_agent`

    - `task_type: Optional[Literal["auto_evaluation.agent"]]`

      - `"auto_evaluation.agent"`

  - `class ContributorEvaluationQuestionTask: …`

    - `configuration: ContributorEvaluationQuestionTaskConfiguration`

      - `layout: Container`

        - `children: List[Child]`

          The children to be displayed within the container

          - `class Container: …`

          - `class Component: …`

            - `data: ItemLocator`

              A pointer to the data in each evaluation item to be displayed within the component

            - `label: Optional[str]`

        - `direction: Optional[Literal["row", "column"]]`

          The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

          - `"row"`

          - `"column"`

      - `question_id: str`

      - `prefill_from: Optional[str]`

        Dataset column to prefill contributor question task result

      - `queue_id: Optional[str]`

        The contributor annotation queue to include this task in. Defaults to `default`

      - `required: Optional[bool]`

        Whether the question is required to be answered

      - `rubric_id: Optional[str]`

        ID of the rubric to use for scoring this evaluation question

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `contributor_evaluation_question`

    - `task_type: Optional[Literal["contributor_evaluation.question"]]`

      - `"contributor_evaluation.question"`

  - `class CustomFunctionEvaluationTask: …`

    - `configuration: CustomFunctionEvaluationTaskConfiguration`

      Configuration for a custom Python function evaluation task.

      - `function_source: str`

        Python function source code

      - `arg_mapping: Optional[Dict[str, str]]`

        Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

      - `config_args: Optional[Dict[str, object]]`

        Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

      - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

        Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

        - `path: str`

          Dot path in the custom function return value to materialize.

        - `alias: Optional[str]`

          Result column alias. Defaults to path with dots replaced by underscores.

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the function name.

    - `task_type: Optional[Literal["custom_function"]]`

      - `"custom_function"`

### Evaluation Tasks Progress Schema

- `class EvaluationTasksProgressSchema: …`

  - `items: Optional[Items]`

    - `failed: int`

    - `pending: int`

    - `successful: int`

    - `total: int`

    - `failed_items: Optional[List[ItemsFailedItem]]`

      - `item_id: str`

      - `error: Optional[str]`

      - `error_type: Optional[str]`

  - `workflows: Optional[Workflows]`

    - `completed: int`

    - `failed: int`

    - `pending: int`

    - `total: int`

### Evaluation Views

- `Literal["tasks"]`

  - `"tasks"`

### Gt Evaluation Run Condition

- `class GtEvaluationRunCondition: …`

  - `left: object`

  - `right: object`

  - `op: Optional[Literal["gt"]]`

    - `"gt"`

### Gte Evaluation Run Condition

- `class GteEvaluationRunCondition: …`

  - `left: object`

  - `right: object`

  - `op: Optional[Literal["gte"]]`

    - `"gte"`

### In Evaluation Run Condition

- `class InEvaluationRunCondition: …`

  - `left: object`

  - `operands: List[object]`

  - `op: Optional[Literal["in"]]`

    - `"in"`

### Is Not Null Evaluation Run Condition

- `class IsNotNullEvaluationRunCondition: …`

  - `operands: List[object]`

  - `op: Optional[Literal["is_not_null"]]`

    - `"is_not_null"`

### Is Null Evaluation Run Condition

- `class IsNullEvaluationRunCondition: …`

  - `operands: List[object]`

  - `op: Optional[Literal["is_null"]]`

    - `"is_null"`

### Item Locator

- `str`

### Item Locator Template

- `str`

### Lt Evaluation Run Condition

- `class LtEvaluationRunCondition: …`

  - `left: object`

  - `right: object`

  - `op: Optional[Literal["lt"]]`

    - `"lt"`

### Lte Evaluation Run Condition

- `class LteEvaluationRunCondition: …`

  - `left: object`

  - `right: object`

  - `op: Optional[Literal["lte"]]`

    - `"lte"`

### Ne Evaluation Run Condition

- `class NeEvaluationRunCondition: …`

  - `left: object`

  - `right: object`

  - `op: Optional[Literal["ne"]]`

    - `"ne"`

### Not Evaluation Run Condition

- `class NotEvaluationRunCondition: …`

  - `operands: List[object]`

  - `op: Optional[Literal["not"]]`

    - `"not"`

### Not In Evaluation Run Condition

- `class NotInEvaluationRunCondition: …`

  - `left: object`

  - `operands: List[object]`

  - `op: Optional[Literal["not_in"]]`

    - `"not_in"`

### Or Evaluation Run Condition

- `class OrEvaluationRunCondition: …`

  - `operands: List[object]`

  - `op: Optional[Literal["or"]]`

    - `"or"`

### Paginated List Evaluation

- `class PaginatedListEvaluation: …`

  - `has_more: bool`

    Whether there are more items left to be fetched.

  - `items: List[Evaluation]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

      - `id: str`

      - `type: Literal["user", "service_account"]`

        - `"user"`

        - `"service_account"`

      - `object: Optional[Literal["identity"]]`

        - `"identity"`

    - `datasets: Optional[List[Dataset]]`

      - `id: str`

        The unique identifier of the entity.

      - `created_at: datetime`

        The date and time when the entity was created in ISO format.

      - `created_by: Identity`

        The identity that created the entity.

      - `current_version_num: int`

      - `name: str`

      - `tags: Optional[List[str]]`

        The tags associated with the entity

      - `archived_at: Optional[datetime]`

        The date and time when the entity was archived in ISO format.

      - `description: Optional[str]`

      - `object: Optional[Literal["dataset"]]`

        - `"dataset"`

    - `name: str`

    - `status: Literal["failed", "completed", "running"]`

      - `"failed"`

      - `"completed"`

      - `"running"`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `error_count: Optional[int]`

      Number of task errors across all items in this evaluation.

    - `metadata: Optional[Dict[str, object]]`

      Metadata key-value pairs for the evaluation

    - `object: Optional[Literal["evaluation"]]`

      - `"evaluation"`

    - `progress: Optional[EvaluationTasksProgressSchema]`

      Progress of the evaluation's underlying async job

      - `items: Optional[Items]`

        - `failed: int`

        - `pending: int`

        - `successful: int`

        - `total: int`

        - `failed_items: Optional[List[ItemsFailedItem]]`

          - `item_id: str`

          - `error: Optional[str]`

          - `error_type: Optional[str]`

      - `workflows: Optional[Workflows]`

        - `completed: int`

        - `failed: int`

        - `pending: int`

        - `total: int`

    - `status_reason: Optional[str]`

      Reason for evaluation status

    - `tasks: Optional[List[EvaluationTask]]`

      Tasks executed during evaluation. Populated with optional `task` view.

      - `class ChatCompletionEvaluationTask: …`

        - `configuration: ChatCompletionEvaluationTaskConfiguration`

          - `messages: Union[List[Dict[str, object]], ItemLocator]`

            openai standard message format

            - `List[Dict[str, object]]`

            - `str`

          - `model: str`

            model specified as `model_vendor/model`, for example `openai/gpt-4o`

          - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

            - `Dict[str, object]`

            - `str`

          - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

            Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

            - `float`

            - `str`

          - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Deprecated in favor of tool_choice. Controls which function is called by the model.

            - `Dict[str, object]`

            - `str`

          - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

            Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

            - `List[Dict[str, object]]`

            - `str`

          - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

            Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

            - `Dict[str, int]`

            - `str`

          - `logprobs: Optional[Union[bool, ItemLocator, null]]`

            Whether to return log probabilities of the output tokens or not.

            - `bool`

            - `str`

          - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

            An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

            - `int`

            - `str`

          - `max_tokens: Optional[Union[int, ItemLocator, null]]`

            Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

            - `int`

            - `str`

          - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

            Developer-defined tags and values used for filtering completions in the dashboard.

            - `Dict[str, str]`

            - `str`

          - `modalities: Optional[Union[List[str], ItemLocator, null]]`

            Output types that you would like the model to generate for this request.

            - `List[str]`

            - `str`

          - `n: Optional[Union[int, ItemLocator, null]]`

            How many chat completion choices to generate for each input message.

            - `int`

            - `str`

          - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

            Whether to enable parallel function calling during tool use.

            - `bool`

            - `str`

          - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Static predicted output content, such as the content of a text file being regenerated.

            - `Dict[str, object]`

            - `str`

          - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

            Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

            - `float`

            - `str`

          - `reasoning_effort: Optional[str]`

            For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

          - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

            An object specifying the format that the model must output.

            - `Dict[str, object]`

            - `str`

          - `seed: Optional[Union[int, ItemLocator, null]]`

            If specified, system will attempt to sample deterministically for repeated requests with same seed.

            - `int`

            - `str`

          - `stop: Optional[Union[str, List[str], null]]`

            Up to 4 sequences where the API will stop generating further tokens.

            - `str`

            - `List[str]`

          - `store: Optional[Union[bool, ItemLocator, null]]`

            Whether to store the output for use in model distillation or evals products.

            - `bool`

            - `str`

          - `temperature: Optional[Union[float, ItemLocator, null]]`

            What sampling temperature to use. Higher values make output more random, lower more focused.

            - `float`

            - `str`

          - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

            Controls which tool is called by the model. Values: none, auto, required, or specific tool.

            - `str`

            - `Dict[str, object]`

          - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

            A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

            - `List[Dict[str, object]]`

            - `str`

          - `top_k: Optional[Union[int, ItemLocator, null]]`

            Only sample from the top K options for each subsequent token

            - `int`

            - `str`

          - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

            Number of most likely tokens to return at each position, with associated log probability.

            - `int`

            - `str`

          - `top_p: Optional[Union[float, ItemLocator, null]]`

            Alternative to temperature. Only tokens comprising top_p probability mass are considered.

            - `float`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `chat_completion`

        - `task_type: Optional[Literal["chat_completion"]]`

          - `"chat_completion"`

      - `class GenericInferenceEvaluationTask: …`

        - `configuration: GenericInferenceEvaluationTaskConfiguration`

          - `model: str`

            model specified as `vendor/name` (ex. openai/gpt-5)

          - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Arguments passed into model

            - `Dict[str, object]`

            - `str`

          - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

            Vendor specific configuration

            - `class LaunchInferenceConfiguration: …`

              - `num_retries: Optional[int]`

              - `timeout_seconds: Optional[int]`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `inference`

        - `task_type: Optional[Literal["inference"]]`

          - `"inference"`

      - `class ApplicationVariantV1EvaluationTask: …`

        - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

          - `application_variant_id: str`

          - `inputs: Union[Dict[str, object], ItemLocator]`

            Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

            - `Dict[str, object]`

            - `str`

          - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

            History of the application

            - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

              - `request: str`

                Request inputs

              - `response: str`

                Response outputs

              - `session_data: Optional[Dict[str, object]]`

                Session data corresponding to the request response pair

            - `str`

          - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

            - `Dict[str, object]`

            - `str`

          - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

            Optional overrides for the application

            - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

              Execution override options for agentic applications

              - `concurrent: Optional[bool]`

              - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

                - `current_node: str`

                - `state: Dict[str, object]`

              - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

                - `duration_ms: int`

                - `node_id: str`

                - `operation_input: str`

                - `operation_output: str`

                - `operation_type: str`

                - `start_timestamp: str`

                - `workflow_id: str`

                - `operation_metadata: Optional[Dict[str, object]]`

              - `return_span: Optional[bool]`

              - `use_channels: Optional[bool]`

            - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

              - `artifact_ids_filter: Optional[List[str]]`

              - `artifact_name_regex: Optional[List[str]]`

              - `type: Optional[Literal["knowledge_base_schema"]]`

                - `"knowledge_base_schema"`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `application_variant`

        - `task_type: Optional[Literal["application_variant"]]`

          - `"application_variant"`

      - `class AgentexOutputEvaluationTask: …`

        - `configuration: AgentexOutputEvaluationTaskConfiguration`

          - `agentex_agent_id: str`

            The ID of the Agentex agent to use

          - `input_column: Union[str, Dict[str, object], List[object]]`

            The dataset column to use as input for the agent

            - `str`

            - `Dict[str, object]`

            - `List[object]`

          - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

            Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

            - `Dict[str, object]`

            - `str`

          - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

            How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

            - `"first_message"`

            - `"turn_quiescence"`

          - `deployment_id: Optional[str]`

            Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

          - `include_traces: Optional[Union[bool, ItemLocator, null]]`

            Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

            - `bool`

            - `str`

          - `input_mode: Optional[Literal["text", "data"]]`

            How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

            - `"text"`

            - `"data"`

          - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

            Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

            - `int`

            - `str`

          - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

            Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

            - `int`

            - `str`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `agentex_output`

        - `task_type: Optional[Literal["agentex_output"]]`

          - `"agentex_output"`

      - `class MetricEvaluationTask: …`

        - `configuration: MetricEvaluationTaskConfiguration`

          - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["bleu"]`

              - `"bleu"`

          - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["meteor"]`

              - `"meteor"`

          - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["cosine_similarity"]`

              - `"cosine_similarity"`

          - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["f1"]`

              - `"f1"`

          - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["rouge1"]`

              - `"rouge1"`

          - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["rouge2"]`

              - `"rouge2"`

          - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

            - `candidate: str`

            - `reference: str`

            - `type: Literal["rougeL"]`

              - `"rougeL"`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the metric type specified in the configuration

        - `task_type: Optional[Literal["metric"]]`

          - `"metric"`

      - `class AutoEvaluationQuestionTask: …`

        - `configuration: AutoEvaluationQuestionTaskConfiguration`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `question_id: str`

            question to be evaluated

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `auto_evaluation_question`

        - `task_type: Optional[Literal["auto_evaluation.question"]]`

          - `"auto_evaluation.question"`

      - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

        - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

            - `model: str`

              model specified as `model_vendor/model_name`

            - `prompt: str`

            - `response_format: Dict[str, object]`

              JSON schema used for structuring the model response

            - `inference_args: Optional[Dict[str, object]]`

              Additional arguments to pass to the inference request

            - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

                - `op: Optional[Literal["const"]]`

                  - `"const"`

                - `value: Optional[Union[str, float, bool, null]]`

                  - `str`

                  - `float`

                  - `bool`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

                - `path: str`

                - `op: Optional[Literal["var"]]`

                  - `"var"`

              - `class EqEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["eq"]]`

                  - `"eq"`

              - `class NeEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["ne"]]`

                  - `"ne"`

              - `class LtEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["lt"]]`

                  - `"lt"`

              - `class LteEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["lte"]]`

                  - `"lte"`

              - `class GtEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["gt"]]`

                  - `"gt"`

              - `class GteEvaluationRunCondition: …`

                - `left: object`

                - `right: object`

                - `op: Optional[Literal["gte"]]`

                  - `"gte"`

              - `class AndEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["and"]]`

                  - `"and"`

              - `class OrEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["or"]]`

                  - `"or"`

              - `class InEvaluationRunCondition: …`

                - `left: object`

                - `operands: List[object]`

                - `op: Optional[Literal["in"]]`

                  - `"in"`

              - `class NotInEvaluationRunCondition: …`

                - `left: object`

                - `operands: List[object]`

                - `op: Optional[Literal["not_in"]]`

                  - `"not_in"`

              - `class NotEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["not"]]`

                  - `"not"`

              - `class IsNullEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["is_null"]]`

                  - `"is_null"`

              - `class IsNotNullEvaluationRunCondition: …`

                - `operands: List[object]`

                - `op: Optional[Literal["is_not_null"]]`

                  - `"is_not_null"`

            - `system_prompt: Optional[str]`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

            - `choices: List[str]`

              Choices array cannot be empty

            - `model: str`

              model specified as `model_vendor/model_name`

            - `prompt: str`

            - `inference_args: Optional[Dict[str, object]]`

              Additional arguments to pass to the inference request

            - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

                - `op: Optional[Literal["const"]]`

                  - `"const"`

                - `value: Optional[Union[str, float, bool, null]]`

                  - `str`

                  - `float`

                  - `bool`

              - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

                - `path: str`

                - `op: Optional[Literal["var"]]`

                  - `"var"`

              - `class EqEvaluationRunCondition: …`

              - `class NeEvaluationRunCondition: …`

              - `class LtEvaluationRunCondition: …`

              - `class LteEvaluationRunCondition: …`

              - `class GtEvaluationRunCondition: …`

              - `class GteEvaluationRunCondition: …`

              - `class AndEvaluationRunCondition: …`

              - `class OrEvaluationRunCondition: …`

              - `class InEvaluationRunCondition: …`

              - `class NotInEvaluationRunCondition: …`

              - `class NotEvaluationRunCondition: …`

              - `class IsNullEvaluationRunCondition: …`

              - `class IsNotNullEvaluationRunCondition: …`

            - `system_prompt: Optional[str]`

          - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

            - `definition: str`

            - `name: str`

            - `output_rules: List[str]`

            - `data_fields: Optional[List[str]]`

            - `designated_to: Optional[DesignatedTo]`

              - `class DesignatedToApeAgent: …`

                - `config: DesignatedToApeAgentConfig`

                  - `model: Optional[str]`

                  - `temperature: Optional[float]`

                - `agent_name: Optional[Literal["APEAgent"]]`

                  - `"APEAgent"`

              - `class DesignatedToIfAgent: …`

                - `config: DesignatedToIfAgentConfig`

                  - `model: Optional[str]`

                - `agent_name: Optional[Literal["IFAgent"]]`

                  - `"IFAgent"`

              - `class DesignatedToTruthfulnessAgent: …`

                - `config: DesignatedToTruthfulnessAgentConfig`

                  - `model: Optional[str]`

                - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                  - `"TruthfulnessAgent"`

              - `class DesignatedToBaseAgent: …`

                - `config: DesignatedToBaseAgentConfig`

                  - `model: Optional[str]`

                - `agent_name: Optional[Literal["BaseAgent"]]`

                  - `"BaseAgent"`

            - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

              - `"text"`

              - `"integer"`

              - `"float"`

              - `"boolean"`

            - `output_values: Optional[List[Union[str, float, bool]]]`

              - `str`

              - `float`

              - `bool`

            - `rubric_id: Optional[str]`

            - `rubric_version: Optional[int]`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

        - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

          - `"auto_evaluation.guided_decoding"`

      - `class AutoEvaluationAgentEvaluationTask: …`

        - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `auto_evaluation_agent`

        - `task_type: Optional[Literal["auto_evaluation.agent"]]`

          - `"auto_evaluation.agent"`

      - `class ContributorEvaluationQuestionTask: …`

        - `configuration: ContributorEvaluationQuestionTaskConfiguration`

          - `layout: Container`

            - `children: List[Child]`

              The children to be displayed within the container

              - `class Container: …`

              - `class Component: …`

                - `data: ItemLocator`

                  A pointer to the data in each evaluation item to be displayed within the component

                - `label: Optional[str]`

            - `direction: Optional[Literal["row", "column"]]`

              The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

              - `"row"`

              - `"column"`

          - `question_id: str`

          - `prefill_from: Optional[str]`

            Dataset column to prefill contributor question task result

          - `queue_id: Optional[str]`

            The contributor annotation queue to include this task in. Defaults to `default`

          - `required: Optional[bool]`

            Whether the question is required to be answered

          - `rubric_id: Optional[str]`

            ID of the rubric to use for scoring this evaluation question

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the `contributor_evaluation_question`

        - `task_type: Optional[Literal["contributor_evaluation.question"]]`

          - `"contributor_evaluation.question"`

      - `class CustomFunctionEvaluationTask: …`

        - `configuration: CustomFunctionEvaluationTaskConfiguration`

          Configuration for a custom Python function evaluation task.

          - `function_source: str`

            Python function source code

          - `arg_mapping: Optional[Dict[str, str]]`

            Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

          - `config_args: Optional[Dict[str, object]]`

            Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

          - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

            Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

            - `path: str`

              Dot path in the custom function return value to materialize.

            - `alias: Optional[str]`

              Result column alias. Defaults to path with dots replaced by underscores.

        - `alias: Optional[str]`

          Alias to title the results column. Defaults to the function name.

        - `task_type: Optional[Literal["custom_function"]]`

          - `"custom_function"`

  - `total: int`

    The total of items that match the query. This is greater than or equal to the number of items returned.

  - `limit: Optional[int]`

    The maximum number of items to return.

  - `object: Optional[Literal["list"]]`

    - `"list"`

### Evaluation Retrieve Taxonomy Response

- `Dict[str, object]`

# Tasks

## Add Test Criteria to Evaluation

`evaluations.tasks.add(strevaluation_id, TaskAddParams**kwargs)  -> Evaluation`

**post** `/v5/evaluations/{evaluation_id}/tasks`

Add a new test criteria to an existing evaluation.

Narrowed to contributor question tasks (`contributor_evaluation.question`); other task types
must be configured when the evaluation is first created and are rejected here. The request is
also rejected if the evaluation is archived, if a test criteria with the same alias already
exists, or if any contributor annotation task for the evaluation has already been claimed or
completed. Because only contributor question tasks are accepted, the added criteria is applied
synchronously and contributors answer it against the evaluation's existing items — no async job
or Temporal workflow is started.

### Parameters

- `evaluation_id: str`

- `task: EvaluationTaskParam`

  New test criteria to add to the evaluation. Rejected when contributor annotation tasks for this evaluation have already been claimed or completed. Triggers a rerun so the new task executes against existing items.

  - `class ChatCompletionEvaluationTask: …`

    - `configuration: ChatCompletionEvaluationTaskConfiguration`

      - `messages: Union[List[Dict[str, object]], ItemLocator]`

        openai standard message format

        - `List[Dict[str, object]]`

        - `str`

      - `model: str`

        model specified as `model_vendor/model`, for example `openai/gpt-4o`

      - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

        - `Dict[str, object]`

        - `str`

      - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

        Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

        - `float`

        - `str`

      - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Deprecated in favor of tool_choice. Controls which function is called by the model.

        - `Dict[str, object]`

        - `str`

      - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

        Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

        - `List[Dict[str, object]]`

        - `str`

      - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

        Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

        - `Dict[str, int]`

        - `str`

      - `logprobs: Optional[Union[bool, ItemLocator, null]]`

        Whether to return log probabilities of the output tokens or not.

        - `bool`

        - `str`

      - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

        An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

        - `int`

        - `str`

      - `max_tokens: Optional[Union[int, ItemLocator, null]]`

        Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

        - `int`

        - `str`

      - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

        Developer-defined tags and values used for filtering completions in the dashboard.

        - `Dict[str, str]`

        - `str`

      - `modalities: Optional[Union[List[str], ItemLocator, null]]`

        Output types that you would like the model to generate for this request.

        - `List[str]`

        - `str`

      - `n: Optional[Union[int, ItemLocator, null]]`

        How many chat completion choices to generate for each input message.

        - `int`

        - `str`

      - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

        Whether to enable parallel function calling during tool use.

        - `bool`

        - `str`

      - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Static predicted output content, such as the content of a text file being regenerated.

        - `Dict[str, object]`

        - `str`

      - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

        Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

        - `float`

        - `str`

      - `reasoning_effort: Optional[str]`

        For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

      - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

        An object specifying the format that the model must output.

        - `Dict[str, object]`

        - `str`

      - `seed: Optional[Union[int, ItemLocator, null]]`

        If specified, system will attempt to sample deterministically for repeated requests with same seed.

        - `int`

        - `str`

      - `stop: Optional[Union[str, List[str], null]]`

        Up to 4 sequences where the API will stop generating further tokens.

        - `str`

        - `List[str]`

      - `store: Optional[Union[bool, ItemLocator, null]]`

        Whether to store the output for use in model distillation or evals products.

        - `bool`

        - `str`

      - `temperature: Optional[Union[float, ItemLocator, null]]`

        What sampling temperature to use. Higher values make output more random, lower more focused.

        - `float`

        - `str`

      - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

        Controls which tool is called by the model. Values: none, auto, required, or specific tool.

        - `str`

        - `Dict[str, object]`

      - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

        A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

        - `List[Dict[str, object]]`

        - `str`

      - `top_k: Optional[Union[int, ItemLocator, null]]`

        Only sample from the top K options for each subsequent token

        - `int`

        - `str`

      - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

        Number of most likely tokens to return at each position, with associated log probability.

        - `int`

        - `str`

      - `top_p: Optional[Union[float, ItemLocator, null]]`

        Alternative to temperature. Only tokens comprising top_p probability mass are considered.

        - `float`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `chat_completion`

    - `task_type: Optional[Literal["chat_completion"]]`

      - `"chat_completion"`

  - `class GenericInferenceEvaluationTask: …`

    - `configuration: GenericInferenceEvaluationTaskConfiguration`

      - `model: str`

        model specified as `vendor/name` (ex. openai/gpt-5)

      - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Arguments passed into model

        - `Dict[str, object]`

        - `str`

      - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

        Vendor specific configuration

        - `class LaunchInferenceConfiguration: …`

          - `num_retries: Optional[int]`

          - `timeout_seconds: Optional[int]`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `inference`

    - `task_type: Optional[Literal["inference"]]`

      - `"inference"`

  - `class ApplicationVariantV1EvaluationTask: …`

    - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

      - `application_variant_id: str`

      - `inputs: Union[Dict[str, object], ItemLocator]`

        Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

        - `Dict[str, object]`

        - `str`

      - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

        History of the application

        - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

          - `request: str`

            Request inputs

          - `response: str`

            Response outputs

          - `session_data: Optional[Dict[str, object]]`

            Session data corresponding to the request response pair

        - `str`

      - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

        - `Dict[str, object]`

        - `str`

      - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

        Optional overrides for the application

        - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

          Execution override options for agentic applications

          - `concurrent: Optional[bool]`

          - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

            - `current_node: str`

            - `state: Dict[str, object]`

          - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

            - `duration_ms: int`

            - `node_id: str`

            - `operation_input: str`

            - `operation_output: str`

            - `operation_type: str`

            - `start_timestamp: str`

            - `workflow_id: str`

            - `operation_metadata: Optional[Dict[str, object]]`

          - `return_span: Optional[bool]`

          - `use_channels: Optional[bool]`

        - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

          - `artifact_ids_filter: Optional[List[str]]`

          - `artifact_name_regex: Optional[List[str]]`

          - `type: Optional[Literal["knowledge_base_schema"]]`

            - `"knowledge_base_schema"`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `application_variant`

    - `task_type: Optional[Literal["application_variant"]]`

      - `"application_variant"`

  - `class AgentexOutputEvaluationTask: …`

    - `configuration: AgentexOutputEvaluationTaskConfiguration`

      - `agentex_agent_id: str`

        The ID of the Agentex agent to use

      - `input_column: Union[str, Dict[str, object], List[object]]`

        The dataset column to use as input for the agent

        - `str`

        - `Dict[str, object]`

        - `List[object]`

      - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

        Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

        - `Dict[str, object]`

        - `str`

      - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

        How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

        - `"first_message"`

        - `"turn_quiescence"`

      - `deployment_id: Optional[str]`

        Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

      - `include_traces: Optional[Union[bool, ItemLocator, null]]`

        Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

        - `bool`

        - `str`

      - `input_mode: Optional[Literal["text", "data"]]`

        How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

        - `"text"`

        - `"data"`

      - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

        Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

        - `int`

        - `str`

      - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

        Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

        - `int`

        - `str`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `agentex_output`

    - `task_type: Optional[Literal["agentex_output"]]`

      - `"agentex_output"`

  - `class MetricEvaluationTask: …`

    - `configuration: MetricEvaluationTaskConfiguration`

      - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["bleu"]`

          - `"bleu"`

      - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["meteor"]`

          - `"meteor"`

      - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["cosine_similarity"]`

          - `"cosine_similarity"`

      - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["f1"]`

          - `"f1"`

      - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["rouge1"]`

          - `"rouge1"`

      - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["rouge2"]`

          - `"rouge2"`

      - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

        - `candidate: str`

        - `reference: str`

        - `type: Literal["rougeL"]`

          - `"rougeL"`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the metric type specified in the configuration

    - `task_type: Optional[Literal["metric"]]`

      - `"metric"`

  - `class AutoEvaluationQuestionTask: …`

    - `configuration: AutoEvaluationQuestionTaskConfiguration`

      - `model: str`

        model specified as `model_vendor/model_name`

      - `prompt: str`

      - `question_id: str`

        question to be evaluated

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `auto_evaluation_question`

    - `task_type: Optional[Literal["auto_evaluation.question"]]`

      - `"auto_evaluation.question"`

  - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

    - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

      - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `response_format: Dict[str, object]`

          JSON schema used for structuring the model response

        - `inference_args: Optional[Dict[str, object]]`

          Additional arguments to pass to the inference request

        - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

            - `op: Optional[Literal["const"]]`

              - `"const"`

            - `value: Optional[Union[str, float, bool, null]]`

              - `str`

              - `float`

              - `bool`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

            - `path: str`

            - `op: Optional[Literal["var"]]`

              - `"var"`

          - `class EqEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["eq"]]`

              - `"eq"`

          - `class NeEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["ne"]]`

              - `"ne"`

          - `class LtEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["lt"]]`

              - `"lt"`

          - `class LteEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["lte"]]`

              - `"lte"`

          - `class GtEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["gt"]]`

              - `"gt"`

          - `class GteEvaluationRunCondition: …`

            - `left: object`

            - `right: object`

            - `op: Optional[Literal["gte"]]`

              - `"gte"`

          - `class AndEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["and"]]`

              - `"and"`

          - `class OrEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["or"]]`

              - `"or"`

          - `class InEvaluationRunCondition: …`

            - `left: object`

            - `operands: List[object]`

            - `op: Optional[Literal["in"]]`

              - `"in"`

          - `class NotInEvaluationRunCondition: …`

            - `left: object`

            - `operands: List[object]`

            - `op: Optional[Literal["not_in"]]`

              - `"not_in"`

          - `class NotEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["not"]]`

              - `"not"`

          - `class IsNullEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["is_null"]]`

              - `"is_null"`

          - `class IsNotNullEvaluationRunCondition: …`

            - `operands: List[object]`

            - `op: Optional[Literal["is_not_null"]]`

              - `"is_not_null"`

        - `system_prompt: Optional[str]`

      - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

        - `choices: List[str]`

          Choices array cannot be empty

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `inference_args: Optional[Dict[str, object]]`

          Additional arguments to pass to the inference request

        - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

            - `op: Optional[Literal["const"]]`

              - `"const"`

            - `value: Optional[Union[str, float, bool, null]]`

              - `str`

              - `float`

              - `bool`

          - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

            - `path: str`

            - `op: Optional[Literal["var"]]`

              - `"var"`

          - `class EqEvaluationRunCondition: …`

          - `class NeEvaluationRunCondition: …`

          - `class LtEvaluationRunCondition: …`

          - `class LteEvaluationRunCondition: …`

          - `class GtEvaluationRunCondition: …`

          - `class GteEvaluationRunCondition: …`

          - `class AndEvaluationRunCondition: …`

          - `class OrEvaluationRunCondition: …`

          - `class InEvaluationRunCondition: …`

          - `class NotInEvaluationRunCondition: …`

          - `class NotEvaluationRunCondition: …`

          - `class IsNullEvaluationRunCondition: …`

          - `class IsNotNullEvaluationRunCondition: …`

        - `system_prompt: Optional[str]`

      - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

        - `definition: str`

        - `name: str`

        - `output_rules: List[str]`

        - `data_fields: Optional[List[str]]`

        - `designated_to: Optional[DesignatedTo]`

          - `class DesignatedToApeAgent: …`

            - `config: DesignatedToApeAgentConfig`

              - `model: Optional[str]`

              - `temperature: Optional[float]`

            - `agent_name: Optional[Literal["APEAgent"]]`

              - `"APEAgent"`

          - `class DesignatedToIfAgent: …`

            - `config: DesignatedToIfAgentConfig`

              - `model: Optional[str]`

            - `agent_name: Optional[Literal["IFAgent"]]`

              - `"IFAgent"`

          - `class DesignatedToTruthfulnessAgent: …`

            - `config: DesignatedToTruthfulnessAgentConfig`

              - `model: Optional[str]`

            - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

              - `"TruthfulnessAgent"`

          - `class DesignatedToBaseAgent: …`

            - `config: DesignatedToBaseAgentConfig`

              - `model: Optional[str]`

            - `agent_name: Optional[Literal["BaseAgent"]]`

              - `"BaseAgent"`

        - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

          - `"text"`

          - `"integer"`

          - `"float"`

          - `"boolean"`

        - `output_values: Optional[List[Union[str, float, bool]]]`

          - `str`

          - `float`

          - `bool`

        - `rubric_id: Optional[str]`

        - `rubric_version: Optional[int]`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

    - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

      - `"auto_evaluation.guided_decoding"`

  - `class AutoEvaluationAgentEvaluationTask: …`

    - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `auto_evaluation_agent`

    - `task_type: Optional[Literal["auto_evaluation.agent"]]`

      - `"auto_evaluation.agent"`

  - `class ContributorEvaluationQuestionTask: …`

    - `configuration: ContributorEvaluationQuestionTaskConfiguration`

      - `layout: Container`

        - `children: List[Child]`

          The children to be displayed within the container

          - `class Container: …`

          - `class Component: …`

            - `data: ItemLocator`

              A pointer to the data in each evaluation item to be displayed within the component

            - `label: Optional[str]`

        - `direction: Optional[Literal["row", "column"]]`

          The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

          - `"row"`

          - `"column"`

      - `question_id: str`

      - `prefill_from: Optional[str]`

        Dataset column to prefill contributor question task result

      - `queue_id: Optional[str]`

        The contributor annotation queue to include this task in. Defaults to `default`

      - `required: Optional[bool]`

        Whether the question is required to be answered

      - `rubric_id: Optional[str]`

        ID of the rubric to use for scoring this evaluation question

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the `contributor_evaluation_question`

    - `task_type: Optional[Literal["contributor_evaluation.question"]]`

      - `"contributor_evaluation.question"`

  - `class CustomFunctionEvaluationTask: …`

    - `configuration: CustomFunctionEvaluationTaskConfiguration`

      Configuration for a custom Python function evaluation task.

      - `function_source: str`

        Python function source code

      - `arg_mapping: Optional[Dict[str, str]]`

        Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

      - `config_args: Optional[Dict[str, object]]`

        Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

      - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

        Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

        - `path: str`

          Dot path in the custom function return value to materialize.

        - `alias: Optional[str]`

          Result column alias. Defaults to path with dots replaced by underscores.

    - `alias: Optional[str]`

      Alias to title the results column. Defaults to the function name.

    - `task_type: Optional[Literal["custom_function"]]`

      - `"custom_function"`

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation = client.evaluations.tasks.add(
    evaluation_id="evaluation_id",
    task={
        "configuration": {
            "messages": [{
                "foo": "bar"
            }],
            "model": "model",
        },
        "task_type": "chat_completion",
    },
)
print(evaluation.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```

## Update Test Criteria Configuration

`evaluations.tasks.update(stralias, TaskUpdateParams**kwargs)  -> Evaluation`

**patch** `/v5/evaluations/{evaluation_id}/tasks/{alias}`

Replace the full configuration of a single test criteria, identified by its alias.

The alias must match an existing test criteria on the evaluation, and the replacement
configuration is validated against the evaluation's current items before being applied. The
request is rejected if the evaluation is archived, if no test criteria matches the alias, or if
any contributor annotation task for the evaluation has already been claimed or completed — at
that point labelers are in-flight and mutating the task definition would corrupt their work.

### Parameters

- `evaluation_id: str`

- `alias: str`

- `configuration: Dict[str, object]`

  Full replacement for the test criteria's configuration JSON. Only allowed when no contributor annotation tasks for this evaluation have been claimed or completed.

### Returns

- `class Evaluation: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `datasets: Optional[List[Dataset]]`

    - `id: str`

      The unique identifier of the entity.

    - `created_at: datetime`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: int`

    - `name: str`

    - `tags: Optional[List[str]]`

      The tags associated with the entity

    - `archived_at: Optional[datetime]`

      The date and time when the entity was archived in ISO format.

    - `description: Optional[str]`

    - `object: Optional[Literal["dataset"]]`

      - `"dataset"`

  - `name: str`

  - `status: Literal["failed", "completed", "running"]`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `error_count: Optional[int]`

    Number of task errors across all items in this evaluation.

  - `metadata: Optional[Dict[str, object]]`

    Metadata key-value pairs for the evaluation

  - `object: Optional[Literal["evaluation"]]`

    - `"evaluation"`

  - `progress: Optional[EvaluationTasksProgressSchema]`

    Progress of the evaluation's underlying async job

    - `items: Optional[Items]`

      - `failed: int`

      - `pending: int`

      - `successful: int`

      - `total: int`

      - `failed_items: Optional[List[ItemsFailedItem]]`

        - `item_id: str`

        - `error: Optional[str]`

        - `error_type: Optional[str]`

    - `workflows: Optional[Workflows]`

      - `completed: int`

      - `failed: int`

      - `pending: int`

      - `total: int`

  - `status_reason: Optional[str]`

    Reason for evaluation status

  - `tasks: Optional[List[EvaluationTask]]`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `class ChatCompletionEvaluationTask: …`

      - `configuration: ChatCompletionEvaluationTaskConfiguration`

        - `messages: Union[List[Dict[str, object]], ItemLocator]`

          openai standard message format

          - `List[Dict[str, object]]`

          - `str`

        - `model: str`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Dict[str, object]`

          - `str`

        - `frequency_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `float`

          - `str`

        - `function_call: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Dict[str, object]`

          - `str`

        - `functions: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `List[Dict[str, object]]`

          - `str`

        - `logit_bias: Optional[Union[Dict[str, int], ItemLocator, null]]`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Dict[str, int]`

          - `str`

        - `logprobs: Optional[Union[bool, ItemLocator, null]]`

          Whether to return log probabilities of the output tokens or not.

          - `bool`

          - `str`

        - `max_completion_tokens: Optional[Union[int, ItemLocator, null]]`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `int`

          - `str`

        - `max_tokens: Optional[Union[int, ItemLocator, null]]`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `int`

          - `str`

        - `metadata: Optional[Union[Dict[str, str], ItemLocator, null]]`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Dict[str, str]`

          - `str`

        - `modalities: Optional[Union[List[str], ItemLocator, null]]`

          Output types that you would like the model to generate for this request.

          - `List[str]`

          - `str`

        - `n: Optional[Union[int, ItemLocator, null]]`

          How many chat completion choices to generate for each input message.

          - `int`

          - `str`

        - `parallel_tool_calls: Optional[Union[bool, ItemLocator, null]]`

          Whether to enable parallel function calling during tool use.

          - `bool`

          - `str`

        - `prediction: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Dict[str, object]`

          - `str`

        - `presence_penalty: Optional[Union[float, ItemLocator, null]]`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `float`

          - `str`

        - `reasoning_effort: Optional[str]`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format: Optional[Union[Dict[str, object], ItemLocator, null]]`

          An object specifying the format that the model must output.

          - `Dict[str, object]`

          - `str`

        - `seed: Optional[Union[int, ItemLocator, null]]`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `int`

          - `str`

        - `stop: Optional[Union[str, List[str], null]]`

          Up to 4 sequences where the API will stop generating further tokens.

          - `str`

          - `List[str]`

        - `store: Optional[Union[bool, ItemLocator, null]]`

          Whether to store the output for use in model distillation or evals products.

          - `bool`

          - `str`

        - `temperature: Optional[Union[float, ItemLocator, null]]`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `float`

          - `str`

        - `tool_choice: Optional[Union[str, Dict[str, object], null]]`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `str`

          - `Dict[str, object]`

        - `tools: Optional[Union[List[Dict[str, object]], ItemLocator, null]]`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `List[Dict[str, object]]`

          - `str`

        - `top_k: Optional[Union[int, ItemLocator, null]]`

          Only sample from the top K options for each subsequent token

          - `int`

          - `str`

        - `top_logprobs: Optional[Union[int, ItemLocator, null]]`

          Number of most likely tokens to return at each position, with associated log probability.

          - `int`

          - `str`

        - `top_p: Optional[Union[float, ItemLocator, null]]`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `float`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type: Optional[Literal["chat_completion"]]`

        - `"chat_completion"`

    - `class GenericInferenceEvaluationTask: …`

      - `configuration: GenericInferenceEvaluationTaskConfiguration`

        - `model: str`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arguments passed into model

          - `Dict[str, object]`

          - `str`

        - `inference_configuration: Optional[GenericInferenceEvaluationTaskConfigurationInferenceConfiguration]`

          Vendor specific configuration

          - `class LaunchInferenceConfiguration: …`

            - `num_retries: Optional[int]`

            - `timeout_seconds: Optional[int]`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `inference`

      - `task_type: Optional[Literal["inference"]]`

        - `"inference"`

    - `class ApplicationVariantV1EvaluationTask: …`

      - `configuration: ApplicationVariantV1EvaluationTaskConfiguration`

        - `application_variant_id: str`

        - `inputs: Union[Dict[str, object], ItemLocator]`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Dict[str, object]`

          - `str`

        - `history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]`

          History of the application

          - `List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray]`

            - `request: str`

              Request inputs

            - `response: str`

              Response outputs

            - `session_data: Optional[Dict[str, object]]`

              Session data corresponding to the request response pair

          - `str`

        - `operation_metadata: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Dict[str, object]`

          - `str`

        - `overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]`

          Optional overrides for the application

          - `class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …`

            Execution override options for agentic applications

            - `concurrent: Optional[bool]`

            - `initial_state: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesInitialState]`

              - `current_node: str`

              - `state: Dict[str, object]`

            - `partial_trace: Optional[List[ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverridesPartialTrace]]`

              - `duration_ms: int`

              - `node_id: str`

              - `operation_input: str`

              - `operation_output: str`

              - `operation_type: str`

              - `start_timestamp: str`

              - `workflow_id: str`

              - `operation_metadata: Optional[Dict[str, object]]`

            - `return_span: Optional[bool]`

            - `use_channels: Optional[bool]`

          - `Dict[str, ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1ApplicationVariantV1EvaluationTaskConfigurationOverridesUnionMember1Item]`

            - `artifact_ids_filter: Optional[List[str]]`

            - `artifact_name_regex: Optional[List[str]]`

            - `type: Optional[Literal["knowledge_base_schema"]]`

              - `"knowledge_base_schema"`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type: Optional[Literal["application_variant"]]`

        - `"application_variant"`

    - `class AgentexOutputEvaluationTask: …`

      - `configuration: AgentexOutputEvaluationTaskConfiguration`

        - `agentex_agent_id: str`

          The ID of the Agentex agent to use

        - `input_column: Union[str, Dict[str, object], List[object]]`

          The dataset column to use as input for the agent

          - `str`

          - `Dict[str, object]`

          - `List[object]`

        - `agent_task_params: Optional[Union[Dict[str, object], ItemLocator, null]]`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Dict[str, object]`

          - `str`

        - `completion_mode: Optional[Literal["first_message", "turn_quiescence"]]`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id: Optional[str]`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces: Optional[Union[bool, ItemLocator, null]]`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `bool`

          - `str`

        - `input_mode: Optional[Literal["text", "data"]]`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds: Optional[Union[int, ItemLocator, null]]`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `int`

          - `str`

        - `timeout_seconds: Optional[Union[int, ItemLocator, null]]`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `int`

          - `str`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type: Optional[Literal["agentex_output"]]`

        - `"agentex_output"`

    - `class MetricEvaluationTask: …`

      - `configuration: MetricEvaluationTaskConfiguration`

        - `class MetricEvaluationTaskConfigurationBleuScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["bleu"]`

            - `"bleu"`

        - `class MetricEvaluationTaskConfigurationMeteorScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["meteor"]`

            - `"meteor"`

        - `class MetricEvaluationTaskConfigurationCosineSimilarityScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["cosine_similarity"]`

            - `"cosine_similarity"`

        - `class MetricEvaluationTaskConfigurationF1ScorerConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["f1"]`

            - `"f1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer1ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge1"]`

            - `"rouge1"`

        - `class MetricEvaluationTaskConfigurationRougeScorer2ConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rouge2"]`

            - `"rouge2"`

        - `class MetricEvaluationTaskConfigurationRougeScorerLConfigWithItemLocator: …`

          - `candidate: str`

          - `reference: str`

          - `type: Literal["rougeL"]`

            - `"rougeL"`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type: Optional[Literal["metric"]]`

        - `"metric"`

    - `class AutoEvaluationQuestionTask: …`

      - `configuration: AutoEvaluationQuestionTaskConfiguration`

        - `model: str`

          model specified as `model_vendor/model_name`

        - `prompt: str`

        - `question_id: str`

          question to be evaluated

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type: Optional[Literal["auto_evaluation.question"]]`

        - `"auto_evaluation.question"`

    - `class AutoEvaluationGuidedDecodingEvaluationTask: …`

      - `configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …`

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `response_format: Dict[str, object]`

            JSON schema used for structuring the model response

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["eq"]]`

                - `"eq"`

            - `class NeEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["ne"]]`

                - `"ne"`

            - `class LtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lt"]]`

                - `"lt"`

            - `class LteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["lte"]]`

                - `"lte"`

            - `class GtEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gt"]]`

                - `"gt"`

            - `class GteEvaluationRunCondition: …`

              - `left: object`

              - `right: object`

              - `op: Optional[Literal["gte"]]`

                - `"gte"`

            - `class AndEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["and"]]`

                - `"and"`

            - `class OrEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["or"]]`

                - `"or"`

            - `class InEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["in"]]`

                - `"in"`

            - `class NotInEvaluationRunCondition: …`

              - `left: object`

              - `operands: List[object]`

              - `op: Optional[Literal["not_in"]]`

                - `"not_in"`

            - `class NotEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["not"]]`

                - `"not"`

            - `class IsNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_null"]]`

                - `"is_null"`

            - `class IsNotNullEvaluationRunCondition: …`

              - `operands: List[object]`

              - `op: Optional[Literal["is_not_null"]]`

                - `"is_not_null"`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …`

          - `choices: List[str]`

            Choices array cannot be empty

          - `model: str`

            model specified as `model_vendor/model_name`

          - `prompt: str`

          - `inference_args: Optional[Dict[str, object]]`

            Additional arguments to pass to the inference request

          - `run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …`

              - `op: Optional[Literal["const"]]`

                - `"const"`

              - `value: Optional[Union[str, float, bool, null]]`

                - `str`

                - `float`

                - `bool`

            - `class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …`

              - `path: str`

              - `op: Optional[Literal["var"]]`

                - `"var"`

            - `class EqEvaluationRunCondition: …`

            - `class NeEvaluationRunCondition: …`

            - `class LtEvaluationRunCondition: …`

            - `class LteEvaluationRunCondition: …`

            - `class GtEvaluationRunCondition: …`

            - `class GteEvaluationRunCondition: …`

            - `class AndEvaluationRunCondition: …`

            - `class OrEvaluationRunCondition: …`

            - `class InEvaluationRunCondition: …`

            - `class NotInEvaluationRunCondition: …`

            - `class NotEvaluationRunCondition: …`

            - `class IsNullEvaluationRunCondition: …`

            - `class IsNotNullEvaluationRunCondition: …`

          - `system_prompt: Optional[str]`

        - `class AutoEvaluationAgentTaskRequestWithItemLocator: …`

          - `definition: str`

          - `name: str`

          - `output_rules: List[str]`

          - `data_fields: Optional[List[str]]`

          - `designated_to: Optional[DesignatedTo]`

            - `class DesignatedToApeAgent: …`

              - `config: DesignatedToApeAgentConfig`

                - `model: Optional[str]`

                - `temperature: Optional[float]`

              - `agent_name: Optional[Literal["APEAgent"]]`

                - `"APEAgent"`

            - `class DesignatedToIfAgent: …`

              - `config: DesignatedToIfAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["IFAgent"]]`

                - `"IFAgent"`

            - `class DesignatedToTruthfulnessAgent: …`

              - `config: DesignatedToTruthfulnessAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["TruthfulnessAgent"]]`

                - `"TruthfulnessAgent"`

            - `class DesignatedToBaseAgent: …`

              - `config: DesignatedToBaseAgentConfig`

                - `model: Optional[str]`

              - `agent_name: Optional[Literal["BaseAgent"]]`

                - `"BaseAgent"`

          - `output_type: Optional[Literal["text", "integer", "float", "boolean"]]`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values: Optional[List[Union[str, float, bool]]]`

            - `str`

            - `float`

            - `bool`

          - `rubric_id: Optional[str]`

          - `rubric_version: Optional[int]`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type: Optional[Literal["auto_evaluation.guided_decoding"]]`

        - `"auto_evaluation.guided_decoding"`

    - `class AutoEvaluationAgentEvaluationTask: …`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type: Optional[Literal["auto_evaluation.agent"]]`

        - `"auto_evaluation.agent"`

    - `class ContributorEvaluationQuestionTask: …`

      - `configuration: ContributorEvaluationQuestionTaskConfiguration`

        - `layout: Container`

          - `children: List[Child]`

            The children to be displayed within the container

            - `class Container: …`

            - `class Component: …`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label: Optional[str]`

          - `direction: Optional[Literal["row", "column"]]`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: str`

        - `prefill_from: Optional[str]`

          Dataset column to prefill contributor question task result

        - `queue_id: Optional[str]`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required: Optional[bool]`

          Whether the question is required to be answered

        - `rubric_id: Optional[str]`

          ID of the rubric to use for scoring this evaluation question

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type: Optional[Literal["contributor_evaluation.question"]]`

        - `"contributor_evaluation.question"`

    - `class CustomFunctionEvaluationTask: …`

      - `configuration: CustomFunctionEvaluationTaskConfiguration`

        Configuration for a custom Python function evaluation task.

        - `function_source: str`

          Python function source code

        - `arg_mapping: Optional[Dict[str, str]]`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args: Optional[Dict[str, object]]`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs: Optional[List[CustomFunctionEvaluationTaskConfigurationOutput]]`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: str`

            Dot path in the custom function return value to materialize.

          - `alias: Optional[str]`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias: Optional[str]`

        Alias to title the results column. Defaults to the function name.

      - `task_type: Optional[Literal["custom_function"]]`

        - `"custom_function"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
evaluation = client.evaluations.tasks.update(
    alias="alias",
    evaluation_id="evaluation_id",
    configuration={
        "foo": "bar"
    },
)
print(evaluation.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```
