## Create Evaluation

`client.evaluations.create(EvaluationCreateParamsparams, RequestOptionsoptions?): Evaluation`

**post** `/v5/evaluations`

Create an evaluation together with its items, optionally running test criteria against them.

Accepts three request shapes: standalone (inline `data`), from an existing dataset
(`dataset_id` with optional per-item references), or with a new reusable dataset created inline
from `data`. When the evaluation includes tasks that require execution (for example an LLM judge
or custom function), an async job and a Temporal workflow are started and the evaluation is
returned immediately with status `running`; task results and `error_count` populate
asynchronously. When it includes only contributor tasks, taxonomy-only input, or no tasks, no
workflow runs and it is returned with status `completed`. Optional `tasks`, `metadata`, `tags`,
and `taxonomy_params` are persisted alongside the evaluation and its items.

### Parameters

- `params: EvaluationCreateParams`

  - `evaluation: EvaluationStandaloneCreateRequest | EvaluationFromDatasetCreateRequest | EvaluationWithDatasetCreateRequest`

    - `EvaluationStandaloneCreateRequest`

      - `data: Array<Record<string, unknown>>`

        Items to be evaluated

      - `name: string`

      - `description?: string`

      - `files?: Array<Record<string, string>>`

        Files to be associated to the evaluation

      - `metadata?: Record<string, unknown>`

        Optional metadata key-value pairs for the evaluation

      - `skip_prefilled_rows?: boolean`

        Do not queue a contributor task for prefilled questions

      - `tags?: Array<string>`

        The tags associated with the evaluation

      - `tasks?: Array<EvaluationTask>`

        Tasks allow you to augment and evaluate your data

        - `ChatCompletionEvaluationTask`

          - `configuration: Configuration`

            - `messages: Array<Record<string, unknown>> | ItemLocator`

              openai standard message format

              - `Array<Record<string, unknown>>`

              - `ItemLocator = string`

            - `model: string`

              model specified as `model_vendor/model`, for example `openai/gpt-4o`

            - `audio?: Record<string, unknown> | ItemLocator`

              Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `frequency_penalty?: number | ItemLocator`

              Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

              - `number`

              - `ItemLocator = string`

            - `function_call?: Record<string, unknown> | ItemLocator`

              Deprecated in favor of tool_choice. Controls which function is called by the model.

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `functions?: Array<Record<string, unknown>> | ItemLocator`

              Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

              - `Array<Record<string, unknown>>`

              - `ItemLocator = string`

            - `logit_bias?: Record<string, number> | ItemLocator`

              Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

              - `Record<string, number>`

              - `ItemLocator = string`

            - `logprobs?: boolean | ItemLocator`

              Whether to return log probabilities of the output tokens or not.

              - `boolean`

              - `ItemLocator = string`

            - `max_completion_tokens?: number | ItemLocator`

              An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

              - `number`

              - `ItemLocator = string`

            - `max_tokens?: number | ItemLocator`

              Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

              - `number`

              - `ItemLocator = string`

            - `metadata?: Record<string, string> | ItemLocator`

              Developer-defined tags and values used for filtering completions in the dashboard.

              - `Record<string, string>`

              - `ItemLocator = string`

            - `modalities?: Array<string> | ItemLocator`

              Output types that you would like the model to generate for this request.

              - `Array<string>`

              - `ItemLocator = string`

            - `n?: number | ItemLocator`

              How many chat completion choices to generate for each input message.

              - `number`

              - `ItemLocator = string`

            - `parallel_tool_calls?: boolean | ItemLocator`

              Whether to enable parallel function calling during tool use.

              - `boolean`

              - `ItemLocator = string`

            - `prediction?: Record<string, unknown> | ItemLocator`

              Static predicted output content, such as the content of a text file being regenerated.

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `presence_penalty?: number | ItemLocator`

              Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

              - `number`

              - `ItemLocator = string`

            - `reasoning_effort?: string`

              For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

            - `response_format?: Record<string, unknown> | ItemLocator`

              An object specifying the format that the model must output.

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `seed?: number | ItemLocator`

              If specified, system will attempt to sample deterministically for repeated requests with same seed.

              - `number`

              - `ItemLocator = string`

            - `stop?: string | Array<string>`

              Up to 4 sequences where the API will stop generating further tokens.

              - `string`

              - `Array<string>`

            - `store?: boolean | ItemLocator`

              Whether to store the output for use in model distillation or evals products.

              - `boolean`

              - `ItemLocator = string`

            - `temperature?: number | ItemLocator`

              What sampling temperature to use. Higher values make output more random, lower more focused.

              - `number`

              - `ItemLocator = string`

            - `tool_choice?: string | Record<string, unknown>`

              Controls which tool is called by the model. Values: none, auto, required, or specific tool.

              - `string`

              - `Record<string, unknown>`

            - `tools?: Array<Record<string, unknown>> | ItemLocator`

              A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

              - `Array<Record<string, unknown>>`

              - `ItemLocator = string`

            - `top_k?: number | ItemLocator`

              Only sample from the top K options for each subsequent token

              - `number`

              - `ItemLocator = string`

            - `top_logprobs?: number | ItemLocator`

              Number of most likely tokens to return at each position, with associated log probability.

              - `number`

              - `ItemLocator = string`

            - `top_p?: number | ItemLocator`

              Alternative to temperature. Only tokens comprising top_p probability mass are considered.

              - `number`

              - `ItemLocator = string`

          - `alias?: string`

            Alias to title the results column. Defaults to the `chat_completion`

          - `task_type?: "chat_completion"`

            - `"chat_completion"`

        - `GenericInferenceEvaluationTask`

          - `configuration: Configuration`

            - `model: string`

              model specified as `vendor/name` (ex. openai/gpt-5)

            - `args?: Record<string, unknown> | ItemLocator`

              Arguments passed into model

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `inference_configuration?: LaunchInferenceConfiguration | ItemLocator`

              Vendor specific configuration

              - `LaunchInferenceConfiguration`

                - `num_retries?: number`

                - `timeout_seconds?: number`

              - `ItemLocator = string`

          - `alias?: string`

            Alias to title the results column. Defaults to the `inference`

          - `task_type?: "inference"`

            - `"inference"`

        - `ApplicationVariantV1EvaluationTask`

          - `configuration: Configuration`

            - `application_variant_id: string`

            - `inputs: Record<string, unknown> | ItemLocator`

              Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `history?: Array<ApplicationRequestResponsePairArray> | ItemLocator`

              History of the application

              - `Array<ApplicationRequestResponsePairArray>`

                - `request: string`

                  Request inputs

                - `response: string`

                  Response outputs

                - `session_data?: Record<string, unknown>`

                  Session data corresponding to the request response pair

              - `ItemLocator = string`

            - `operation_metadata?: Record<string, unknown> | ItemLocator`

              Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `overrides?: AgenticApplicationOverrides | Record<string, KnowledgeBaseNodeOverride> | ItemLocator`

              Optional overrides for the application

              - `AgenticApplicationOverrides`

                Execution override options for agentic applications

                - `concurrent?: boolean`

                - `initial_state?: InitialState`

                  - `current_node: string`

                  - `state: Record<string, unknown>`

                - `partial_trace?: Array<PartialTrace>`

                  - `duration_ms: number`

                  - `node_id: string`

                  - `operation_input: string`

                  - `operation_output: string`

                  - `operation_type: string`

                  - `start_timestamp: string`

                  - `workflow_id: string`

                  - `operation_metadata?: Record<string, unknown>`

                - `return_span?: boolean`

                - `use_channels?: boolean`

              - `Record<string, KnowledgeBaseNodeOverride>`

                - `artifact_ids_filter?: Array<string>`

                - `artifact_name_regex?: Array<string>`

                - `type?: "knowledge_base_schema"`

                  - `"knowledge_base_schema"`

              - `ItemLocator = string`

          - `alias?: string`

            Alias to title the results column. Defaults to the `application_variant`

          - `task_type?: "application_variant"`

            - `"application_variant"`

        - `AgentexOutputEvaluationTask`

          - `configuration: Configuration`

            - `agentex_agent_id: string`

              The ID of the Agentex agent to use

            - `input_column: string | Record<string, unknown> | Array<unknown>`

              The dataset column to use as input for the agent

              - `string`

              - `Record<string, unknown>`

              - `Array<unknown>`

            - `agent_task_params?: Record<string, unknown> | ItemLocator`

              Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

              - `Record<string, unknown>`

              - `ItemLocator = string`

            - `completion_mode?: "first_message" | "turn_quiescence"`

              How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

              - `"first_message"`

              - `"turn_quiescence"`

            - `deployment_id?: string`

              Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

            - `include_traces?: boolean | ItemLocator`

              Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

              - `boolean`

              - `ItemLocator = string`

            - `input_mode?: "text" | "data"`

              How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

              - `"text"`

              - `"data"`

            - `quiescence_seconds?: number | ItemLocator`

              Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

              - `number`

              - `ItemLocator = string`

            - `timeout_seconds?: number | ItemLocator`

              Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

              - `number`

              - `ItemLocator = string`

          - `alias?: string`

            Alias to title the results column. Defaults to the `agentex_output`

          - `task_type?: "agentex_output"`

            - `"agentex_output"`

        - `MetricEvaluationTask`

          - `configuration: BleuScorerConfigWithItemLocator | MeteorScorerConfigWithItemLocator | CosineSimilarityScorerConfigWithItemLocator | 4 more`

            - `BleuScorerConfigWithItemLocator`

              - `candidate: string`

              - `reference: string`

              - `type: "bleu"`

                - `"bleu"`

            - `MeteorScorerConfigWithItemLocator`

              - `candidate: string`

              - `reference: string`

              - `type: "meteor"`

                - `"meteor"`

            - `CosineSimilarityScorerConfigWithItemLocator`

              - `candidate: string`

              - `reference: string`

              - `type: "cosine_similarity"`

                - `"cosine_similarity"`

            - `F1ScorerConfigWithItemLocator`

              - `candidate: string`

              - `reference: string`

              - `type: "f1"`

                - `"f1"`

            - `RougeScorer1ConfigWithItemLocator`

              - `candidate: string`

              - `reference: string`

              - `type: "rouge1"`

                - `"rouge1"`

            - `RougeScorer2ConfigWithItemLocator`

              - `candidate: string`

              - `reference: string`

              - `type: "rouge2"`

                - `"rouge2"`

            - `RougeScorerLConfigWithItemLocator`

              - `candidate: string`

              - `reference: string`

              - `type: "rougeL"`

                - `"rougeL"`

          - `alias?: string`

            Alias to title the results column. Defaults to the metric type specified in the configuration

          - `task_type?: "metric"`

            - `"metric"`

        - `AutoEvaluationQuestionTask`

          - `configuration: Configuration`

            - `model: string`

              model specified as `model_vendor/model_name`

            - `prompt: string`

            - `question_id: string`

              question to be evaluated

          - `alias?: string`

            Alias to title the results column. Defaults to the `auto_evaluation_question`

          - `task_type?: "auto_evaluation.question"`

            - `"auto_evaluation.question"`

        - `AutoEvaluationGuidedDecodingEvaluationTask`

          - `configuration: AutoEvaluationStructuredOutputTaskRequestWithItemLocator | AutoEvaluationGuidedDecodingTaskRequestWithItemLocator | AutoEvaluationAgentTaskRequestWithItemLocator`

            - `AutoEvaluationStructuredOutputTaskRequestWithItemLocator`

              - `model: string`

                model specified as `model_vendor/model_name`

              - `prompt: string`

              - `response_format: Record<string, unknown>`

                JSON schema used for structuring the model response

              - `inference_args?: Record<string, unknown>`

                Additional arguments to pass to the inference request

              - `run_condition?: ConstEvaluationRunCondition | VarEvaluationRunCondition | EqEvaluationRunCondition | 12 more`

                - `ConstEvaluationRunCondition`

                  - `op?: "const"`

                    - `"const"`

                  - `value?: string | number | boolean`

                    - `string`

                    - `number`

                    - `boolean`

                - `VarEvaluationRunCondition`

                  - `path: string`

                  - `op?: "var"`

                    - `"var"`

                - `EqEvaluationRunCondition`

                  - `left: unknown`

                  - `right: unknown`

                  - `op?: "eq"`

                    - `"eq"`

                - `NeEvaluationRunCondition`

                  - `left: unknown`

                  - `right: unknown`

                  - `op?: "ne"`

                    - `"ne"`

                - `LtEvaluationRunCondition`

                  - `left: unknown`

                  - `right: unknown`

                  - `op?: "lt"`

                    - `"lt"`

                - `LteEvaluationRunCondition`

                  - `left: unknown`

                  - `right: unknown`

                  - `op?: "lte"`

                    - `"lte"`

                - `GtEvaluationRunCondition`

                  - `left: unknown`

                  - `right: unknown`

                  - `op?: "gt"`

                    - `"gt"`

                - `GteEvaluationRunCondition`

                  - `left: unknown`

                  - `right: unknown`

                  - `op?: "gte"`

                    - `"gte"`

                - `AndEvaluationRunCondition`

                  - `operands: Array<unknown>`

                  - `op?: "and"`

                    - `"and"`

                - `OrEvaluationRunCondition`

                  - `operands: Array<unknown>`

                  - `op?: "or"`

                    - `"or"`

                - `InEvaluationRunCondition`

                  - `left: unknown`

                  - `operands: Array<unknown>`

                  - `op?: "in"`

                    - `"in"`

                - `NotInEvaluationRunCondition`

                  - `left: unknown`

                  - `operands: Array<unknown>`

                  - `op?: "not_in"`

                    - `"not_in"`

                - `NotEvaluationRunCondition`

                  - `operands: Array<unknown>`

                  - `op?: "not"`

                    - `"not"`

                - `IsNullEvaluationRunCondition`

                  - `operands: Array<unknown>`

                  - `op?: "is_null"`

                    - `"is_null"`

                - `IsNotNullEvaluationRunCondition`

                  - `operands: Array<unknown>`

                  - `op?: "is_not_null"`

                    - `"is_not_null"`

              - `system_prompt?: string`

            - `AutoEvaluationGuidedDecodingTaskRequestWithItemLocator`

              - `choices: Array<string>`

                Choices array cannot be empty

              - `model: string`

                model specified as `model_vendor/model_name`

              - `prompt: string`

              - `inference_args?: Record<string, unknown>`

                Additional arguments to pass to the inference request

              - `run_condition?: ConstEvaluationRunCondition | VarEvaluationRunCondition | EqEvaluationRunCondition | 12 more`

                - `ConstEvaluationRunCondition`

                  - `op?: "const"`

                    - `"const"`

                  - `value?: string | number | boolean`

                    - `string`

                    - `number`

                    - `boolean`

                - `VarEvaluationRunCondition`

                  - `path: string`

                  - `op?: "var"`

                    - `"var"`

                - `EqEvaluationRunCondition`

                - `NeEvaluationRunCondition`

                - `LtEvaluationRunCondition`

                - `LteEvaluationRunCondition`

                - `GtEvaluationRunCondition`

                - `GteEvaluationRunCondition`

                - `AndEvaluationRunCondition`

                - `OrEvaluationRunCondition`

                - `InEvaluationRunCondition`

                - `NotInEvaluationRunCondition`

                - `NotEvaluationRunCondition`

                - `IsNullEvaluationRunCondition`

                - `IsNotNullEvaluationRunCondition`

              - `system_prompt?: string`

            - `AutoEvaluationAgentTaskRequestWithItemLocator`

              - `definition: string`

              - `name: string`

              - `output_rules: Array<string>`

              - `data_fields?: Array<string>`

              - `designated_to?: ApeAgent | IfAgent | TruthfulnessAgent | BaseAgent`

                - `ApeAgent`

                  - `config: Config`

                    - `model?: string`

                    - `temperature?: number`

                  - `agent_name?: "APEAgent"`

                    - `"APEAgent"`

                - `IfAgent`

                  - `config: Config`

                    - `model?: string`

                  - `agent_name?: "IFAgent"`

                    - `"IFAgent"`

                - `TruthfulnessAgent`

                  - `config: Config`

                    - `model?: string`

                  - `agent_name?: "TruthfulnessAgent"`

                    - `"TruthfulnessAgent"`

                - `BaseAgent`

                  - `config: Config`

                    - `model?: string`

                  - `agent_name?: "BaseAgent"`

                    - `"BaseAgent"`

              - `output_type?: "text" | "integer" | "float" | "boolean"`

                - `"text"`

                - `"integer"`

                - `"float"`

                - `"boolean"`

              - `output_values?: Array<string | number | boolean>`

                - `string`

                - `number`

                - `boolean`

              - `rubric_id?: string`

              - `rubric_version?: number`

          - `alias?: string`

            Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

          - `task_type?: "auto_evaluation.guided_decoding"`

            - `"auto_evaluation.guided_decoding"`

        - `AutoEvaluationAgentEvaluationTask`

          - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

          - `alias?: string`

            Alias to title the results column. Defaults to the `auto_evaluation_agent`

          - `task_type?: "auto_evaluation.agent"`

            - `"auto_evaluation.agent"`

        - `ContributorEvaluationQuestionTask`

          - `configuration: Configuration`

            - `layout: Container`

              - `children: Array<Container | Component>`

                The children to be displayed within the container

                - `Container`

                - `Component`

                  - `data: ItemLocator`

                    A pointer to the data in each evaluation item to be displayed within the component

                  - `label?: string`

              - `direction?: "row" | "column"`

                The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

                - `"row"`

                - `"column"`

            - `question_id: string`

            - `prefill_from?: string`

              Dataset column to prefill contributor question task result

            - `queue_id?: string`

              The contributor annotation queue to include this task in. Defaults to `default`

            - `required?: boolean`

              Whether the question is required to be answered

            - `rubric_id?: string`

              ID of the rubric to use for scoring this evaluation question

          - `alias?: string`

            Alias to title the results column. Defaults to the `contributor_evaluation_question`

          - `task_type?: "contributor_evaluation.question"`

            - `"contributor_evaluation.question"`

        - `CustomFunctionEvaluationTask`

          - `configuration: Configuration`

            Configuration for a custom Python function evaluation task.

            - `function_source: string`

              Python function source code

            - `arg_mapping?: Record<string, string>`

              Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

            - `config_args?: Record<string, unknown>`

              Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

            - `outputs?: Array<Output>`

              Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

              - `path: string`

                Dot path in the custom function return value to materialize.

              - `alias?: string`

                Result column alias. Defaults to path with dots replaced by underscores.

          - `alias?: string`

            Alias to title the results column. Defaults to the function name.

          - `task_type?: "custom_function"`

            - `"custom_function"`

      - `taxonomy_params?: Record<string, unknown>`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

    - `EvaluationFromDatasetCreateRequest`

      - `dataset_id: string`

        The ID of the dataset containing the items referenced by the `data` field

      - `name: string`

      - `data?: Array<Data>`

        Items to be evaluated, including references to the input dataset

        - `dataset_item_id: string`

      - `description?: string`

      - `metadata?: Record<string, unknown>`

        Optional metadata key-value pairs for the evaluation

      - `skip_prefilled_rows?: boolean`

        Do not queue a contributor task for prefilled questions

      - `tags?: Array<string>`

        The tags associated with the evaluation

      - `tasks?: Array<EvaluationTask>`

        Tasks allow you to augment and evaluate your data

        - `ChatCompletionEvaluationTask`

        - `GenericInferenceEvaluationTask`

        - `ApplicationVariantV1EvaluationTask`

        - `AgentexOutputEvaluationTask`

        - `MetricEvaluationTask`

        - `AutoEvaluationQuestionTask`

        - `AutoEvaluationGuidedDecodingEvaluationTask`

        - `AutoEvaluationAgentEvaluationTask`

        - `ContributorEvaluationQuestionTask`

        - `CustomFunctionEvaluationTask`

      - `taxonomy_params?: Record<string, unknown>`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

    - `EvaluationWithDatasetCreateRequest`

      - `data: Array<Record<string, unknown>>`

        Items to be evaluated

      - `dataset: Dataset`

        Create a reusable dataset from items in the `data` field

        - `name: string`

        - `description?: string`

        - `keys?: Array<string>`

          Keys from items in the `data` field that should be included in the dataset. If not provided, all keys will be included.

        - `tags?: Array<string>`

          The tags associated with the entity

      - `name: string`

      - `description?: string`

      - `files?: Array<Record<string, string>>`

        Files to be associated to the evaluation

      - `metadata?: Record<string, unknown>`

        Optional metadata key-value pairs for the evaluation

      - `skip_prefilled_rows?: boolean`

        Do not queue a contributor task for prefilled questions

      - `tags?: Array<string>`

        The tags associated with the evaluation

      - `tasks?: Array<EvaluationTask>`

        Tasks allow you to augment and evaluate your data

        - `ChatCompletionEvaluationTask`

        - `GenericInferenceEvaluationTask`

        - `ApplicationVariantV1EvaluationTask`

        - `AgentexOutputEvaluationTask`

        - `MetricEvaluationTask`

        - `AutoEvaluationQuestionTask`

        - `AutoEvaluationGuidedDecodingEvaluationTask`

        - `AutoEvaluationAgentEvaluationTask`

        - `ContributorEvaluationQuestionTask`

        - `CustomFunctionEvaluationTask`

      - `taxonomy_params?: Record<string, unknown>`

        Taxonomy params from the task builder. When provided, stores directly as evaluation taxonomy.

### Returns

- `Evaluation`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: string`

    - `type: "user" | "service_account"`

      - `"user"`

      - `"service_account"`

    - `object?: "identity"`

      - `"identity"`

  - `datasets: Array<Dataset> | null`

    - `id: string`

      The unique identifier of the entity.

    - `created_at: string`

      The date and time when the entity was created in ISO format.

    - `created_by: Identity`

      The identity that created the entity.

    - `current_version_num: number`

    - `name: string`

    - `tags: Array<string> | null`

      The tags associated with the entity

    - `archived_at?: string`

      The date and time when the entity was archived in ISO format.

    - `description?: string`

    - `object?: "dataset"`

      - `"dataset"`

  - `name: string`

  - `status: "failed" | "completed" | "running"`

    - `"failed"`

    - `"completed"`

    - `"running"`

  - `tags: Array<string> | null`

    The tags associated with the entity

  - `archived_at?: string`

    The date and time when the entity was archived in ISO format.

  - `description?: string`

  - `error_count?: number`

    Number of task errors across all items in this evaluation.

  - `metadata?: Record<string, unknown>`

    Metadata key-value pairs for the evaluation

  - `object?: "evaluation"`

    - `"evaluation"`

  - `progress?: EvaluationTasksProgressSchema`

    Progress of the evaluation's underlying async job

    - `items?: Items`

      - `failed: number`

      - `pending: number`

      - `successful: number`

      - `total: number`

      - `failed_items?: Array<FailedItem>`

        - `item_id: string`

        - `error?: string`

        - `error_type?: string`

    - `workflows?: Workflows`

      - `completed: number`

      - `failed: number`

      - `pending: number`

      - `total: number`

  - `status_reason?: string`

    Reason for evaluation status

  - `tasks?: Array<EvaluationTask>`

    Tasks executed during evaluation. Populated with optional `task` view.

    - `ChatCompletionEvaluationTask`

      - `configuration: Configuration`

        - `messages: Array<Record<string, unknown>> | ItemLocator`

          openai standard message format

          - `Array<Record<string, unknown>>`

          - `ItemLocator = string`

        - `model: string`

          model specified as `model_vendor/model`, for example `openai/gpt-4o`

        - `audio?: Record<string, unknown> | ItemLocator`

          Parameters for audio output. Required when audio output is requested with modalities: ['audio'].

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `frequency_penalty?: number | ItemLocator`

          Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

          - `number`

          - `ItemLocator = string`

        - `function_call?: Record<string, unknown> | ItemLocator`

          Deprecated in favor of tool_choice. Controls which function is called by the model.

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `functions?: Array<Record<string, unknown>> | ItemLocator`

          Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

          - `Array<Record<string, unknown>>`

          - `ItemLocator = string`

        - `logit_bias?: Record<string, number> | ItemLocator`

          Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

          - `Record<string, number>`

          - `ItemLocator = string`

        - `logprobs?: boolean | ItemLocator`

          Whether to return log probabilities of the output tokens or not.

          - `boolean`

          - `ItemLocator = string`

        - `max_completion_tokens?: number | ItemLocator`

          An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

          - `number`

          - `ItemLocator = string`

        - `max_tokens?: number | ItemLocator`

          Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

          - `number`

          - `ItemLocator = string`

        - `metadata?: Record<string, string> | ItemLocator`

          Developer-defined tags and values used for filtering completions in the dashboard.

          - `Record<string, string>`

          - `ItemLocator = string`

        - `modalities?: Array<string> | ItemLocator`

          Output types that you would like the model to generate for this request.

          - `Array<string>`

          - `ItemLocator = string`

        - `n?: number | ItemLocator`

          How many chat completion choices to generate for each input message.

          - `number`

          - `ItemLocator = string`

        - `parallel_tool_calls?: boolean | ItemLocator`

          Whether to enable parallel function calling during tool use.

          - `boolean`

          - `ItemLocator = string`

        - `prediction?: Record<string, unknown> | ItemLocator`

          Static predicted output content, such as the content of a text file being regenerated.

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `presence_penalty?: number | ItemLocator`

          Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

          - `number`

          - `ItemLocator = string`

        - `reasoning_effort?: string`

          For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

        - `response_format?: Record<string, unknown> | ItemLocator`

          An object specifying the format that the model must output.

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `seed?: number | ItemLocator`

          If specified, system will attempt to sample deterministically for repeated requests with same seed.

          - `number`

          - `ItemLocator = string`

        - `stop?: string | Array<string>`

          Up to 4 sequences where the API will stop generating further tokens.

          - `string`

          - `Array<string>`

        - `store?: boolean | ItemLocator`

          Whether to store the output for use in model distillation or evals products.

          - `boolean`

          - `ItemLocator = string`

        - `temperature?: number | ItemLocator`

          What sampling temperature to use. Higher values make output more random, lower more focused.

          - `number`

          - `ItemLocator = string`

        - `tool_choice?: string | Record<string, unknown>`

          Controls which tool is called by the model. Values: none, auto, required, or specific tool.

          - `string`

          - `Record<string, unknown>`

        - `tools?: Array<Record<string, unknown>> | ItemLocator`

          A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

          - `Array<Record<string, unknown>>`

          - `ItemLocator = string`

        - `top_k?: number | ItemLocator`

          Only sample from the top K options for each subsequent token

          - `number`

          - `ItemLocator = string`

        - `top_logprobs?: number | ItemLocator`

          Number of most likely tokens to return at each position, with associated log probability.

          - `number`

          - `ItemLocator = string`

        - `top_p?: number | ItemLocator`

          Alternative to temperature. Only tokens comprising top_p probability mass are considered.

          - `number`

          - `ItemLocator = string`

      - `alias?: string`

        Alias to title the results column. Defaults to the `chat_completion`

      - `task_type?: "chat_completion"`

        - `"chat_completion"`

    - `GenericInferenceEvaluationTask`

      - `configuration: Configuration`

        - `model: string`

          model specified as `vendor/name` (ex. openai/gpt-5)

        - `args?: Record<string, unknown> | ItemLocator`

          Arguments passed into model

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `inference_configuration?: LaunchInferenceConfiguration | ItemLocator`

          Vendor specific configuration

          - `LaunchInferenceConfiguration`

            - `num_retries?: number`

            - `timeout_seconds?: number`

          - `ItemLocator = string`

      - `alias?: string`

        Alias to title the results column. Defaults to the `inference`

      - `task_type?: "inference"`

        - `"inference"`

    - `ApplicationVariantV1EvaluationTask`

      - `configuration: Configuration`

        - `application_variant_id: string`

        - `inputs: Record<string, unknown> | ItemLocator`

          Input data for the application. For agents service variants, you must provide inputs as a mapping from `{input_name: input_value}`. For V0 variants, you must specify the node your input should be passed to, structuring your input as `{node_id: {input_name: input_value}}`.

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `history?: Array<ApplicationRequestResponsePairArray> | ItemLocator`

          History of the application

          - `Array<ApplicationRequestResponsePairArray>`

            - `request: string`

              Request inputs

            - `response: string`

              Response outputs

            - `session_data?: Record<string, unknown>`

              Session data corresponding to the request response pair

          - `ItemLocator = string`

        - `operation_metadata?: Record<string, unknown> | ItemLocator`

          Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `overrides?: AgenticApplicationOverrides | Record<string, KnowledgeBaseNodeOverride> | ItemLocator`

          Optional overrides for the application

          - `AgenticApplicationOverrides`

            Execution override options for agentic applications

            - `concurrent?: boolean`

            - `initial_state?: InitialState`

              - `current_node: string`

              - `state: Record<string, unknown>`

            - `partial_trace?: Array<PartialTrace>`

              - `duration_ms: number`

              - `node_id: string`

              - `operation_input: string`

              - `operation_output: string`

              - `operation_type: string`

              - `start_timestamp: string`

              - `workflow_id: string`

              - `operation_metadata?: Record<string, unknown>`

            - `return_span?: boolean`

            - `use_channels?: boolean`

          - `Record<string, KnowledgeBaseNodeOverride>`

            - `artifact_ids_filter?: Array<string>`

            - `artifact_name_regex?: Array<string>`

            - `type?: "knowledge_base_schema"`

              - `"knowledge_base_schema"`

          - `ItemLocator = string`

      - `alias?: string`

        Alias to title the results column. Defaults to the `application_variant`

      - `task_type?: "application_variant"`

        - `"application_variant"`

    - `AgentexOutputEvaluationTask`

      - `configuration: Configuration`

        - `agentex_agent_id: string`

          The ID of the Agentex agent to use

        - `input_column: string | Record<string, unknown> | Array<unknown>`

          The dataset column to use as input for the agent

          - `string`

          - `Record<string, unknown>`

          - `Array<unknown>`

        - `agent_task_params?: Record<string, unknown> | ItemLocator`

          Extra params merged into the Agentex `task/create` call's `params` object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation -- the golden agent, for example, rejects any task whose params omit `config_id`. SGP always pins `is_eval: true`; a caller-supplied `description` overrides the SGP default. Nested `item.`-prefixed strings and `{{item.x}}` templates are resolved per evaluation item, so a per-row `config_id` can come from a dataset column.

          - `Record<string, unknown>`

          - `ItemLocator = string`

        - `completion_mode?: "first_message" | "turn_quiescence"`

          How the agent's first turn is judged finished. `first_message` (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. `turn_quiescence` keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for `quiescence_seconds` -- the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and `timeout_seconds` always bounds it.

          - `"first_message"`

          - `"turn_quiescence"`

        - `deployment_id?: string`

          Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent's default RPC endpoint, which resolves through the agent's current routing rules on the Agentex side.

        - `include_traces?: boolean | ItemLocator`

          Whether to include trace data in the evaluation results. Traces are read from SGP's own span store for the agent's trace, not from Agentex.

          - `boolean`

          - `ItemLocator = string`

        - `input_mode?: "text" | "data"`

          How the resolved `input_column` is delivered to the agent. `text` (the default) sends a TextContent message with the value stringified. `data` sends a DataContent message whose `data` is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject `data`.

          - `"text"`

          - `"data"`

        - `quiescence_seconds?: number | ItemLocator`

          Seconds of no new messages before `completion_mode: turn_quiescence` considers the turn finished. Ignored in `first_message` mode. Should exceed the agent's longest expected gap between messages (a slow tool call), or the turn is graded early.

          - `number`

          - `ItemLocator = string`

        - `timeout_seconds?: number | ItemLocator`

          Maximum seconds to wait for the agent's first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity's 1800s start-to-close budget.

          - `number`

          - `ItemLocator = string`

      - `alias?: string`

        Alias to title the results column. Defaults to the `agentex_output`

      - `task_type?: "agentex_output"`

        - `"agentex_output"`

    - `MetricEvaluationTask`

      - `configuration: BleuScorerConfigWithItemLocator | MeteorScorerConfigWithItemLocator | CosineSimilarityScorerConfigWithItemLocator | 4 more`

        - `BleuScorerConfigWithItemLocator`

          - `candidate: string`

          - `reference: string`

          - `type: "bleu"`

            - `"bleu"`

        - `MeteorScorerConfigWithItemLocator`

          - `candidate: string`

          - `reference: string`

          - `type: "meteor"`

            - `"meteor"`

        - `CosineSimilarityScorerConfigWithItemLocator`

          - `candidate: string`

          - `reference: string`

          - `type: "cosine_similarity"`

            - `"cosine_similarity"`

        - `F1ScorerConfigWithItemLocator`

          - `candidate: string`

          - `reference: string`

          - `type: "f1"`

            - `"f1"`

        - `RougeScorer1ConfigWithItemLocator`

          - `candidate: string`

          - `reference: string`

          - `type: "rouge1"`

            - `"rouge1"`

        - `RougeScorer2ConfigWithItemLocator`

          - `candidate: string`

          - `reference: string`

          - `type: "rouge2"`

            - `"rouge2"`

        - `RougeScorerLConfigWithItemLocator`

          - `candidate: string`

          - `reference: string`

          - `type: "rougeL"`

            - `"rougeL"`

      - `alias?: string`

        Alias to title the results column. Defaults to the metric type specified in the configuration

      - `task_type?: "metric"`

        - `"metric"`

    - `AutoEvaluationQuestionTask`

      - `configuration: Configuration`

        - `model: string`

          model specified as `model_vendor/model_name`

        - `prompt: string`

        - `question_id: string`

          question to be evaluated

      - `alias?: string`

        Alias to title the results column. Defaults to the `auto_evaluation_question`

      - `task_type?: "auto_evaluation.question"`

        - `"auto_evaluation.question"`

    - `AutoEvaluationGuidedDecodingEvaluationTask`

      - `configuration: AutoEvaluationStructuredOutputTaskRequestWithItemLocator | AutoEvaluationGuidedDecodingTaskRequestWithItemLocator | AutoEvaluationAgentTaskRequestWithItemLocator`

        - `AutoEvaluationStructuredOutputTaskRequestWithItemLocator`

          - `model: string`

            model specified as `model_vendor/model_name`

          - `prompt: string`

          - `response_format: Record<string, unknown>`

            JSON schema used for structuring the model response

          - `inference_args?: Record<string, unknown>`

            Additional arguments to pass to the inference request

          - `run_condition?: ConstEvaluationRunCondition | VarEvaluationRunCondition | EqEvaluationRunCondition | 12 more`

            - `ConstEvaluationRunCondition`

              - `op?: "const"`

                - `"const"`

              - `value?: string | number | boolean`

                - `string`

                - `number`

                - `boolean`

            - `VarEvaluationRunCondition`

              - `path: string`

              - `op?: "var"`

                - `"var"`

            - `EqEvaluationRunCondition`

              - `left: unknown`

              - `right: unknown`

              - `op?: "eq"`

                - `"eq"`

            - `NeEvaluationRunCondition`

              - `left: unknown`

              - `right: unknown`

              - `op?: "ne"`

                - `"ne"`

            - `LtEvaluationRunCondition`

              - `left: unknown`

              - `right: unknown`

              - `op?: "lt"`

                - `"lt"`

            - `LteEvaluationRunCondition`

              - `left: unknown`

              - `right: unknown`

              - `op?: "lte"`

                - `"lte"`

            - `GtEvaluationRunCondition`

              - `left: unknown`

              - `right: unknown`

              - `op?: "gt"`

                - `"gt"`

            - `GteEvaluationRunCondition`

              - `left: unknown`

              - `right: unknown`

              - `op?: "gte"`

                - `"gte"`

            - `AndEvaluationRunCondition`

              - `operands: Array<unknown>`

              - `op?: "and"`

                - `"and"`

            - `OrEvaluationRunCondition`

              - `operands: Array<unknown>`

              - `op?: "or"`

                - `"or"`

            - `InEvaluationRunCondition`

              - `left: unknown`

              - `operands: Array<unknown>`

              - `op?: "in"`

                - `"in"`

            - `NotInEvaluationRunCondition`

              - `left: unknown`

              - `operands: Array<unknown>`

              - `op?: "not_in"`

                - `"not_in"`

            - `NotEvaluationRunCondition`

              - `operands: Array<unknown>`

              - `op?: "not"`

                - `"not"`

            - `IsNullEvaluationRunCondition`

              - `operands: Array<unknown>`

              - `op?: "is_null"`

                - `"is_null"`

            - `IsNotNullEvaluationRunCondition`

              - `operands: Array<unknown>`

              - `op?: "is_not_null"`

                - `"is_not_null"`

          - `system_prompt?: string`

        - `AutoEvaluationGuidedDecodingTaskRequestWithItemLocator`

          - `choices: Array<string>`

            Choices array cannot be empty

          - `model: string`

            model specified as `model_vendor/model_name`

          - `prompt: string`

          - `inference_args?: Record<string, unknown>`

            Additional arguments to pass to the inference request

          - `run_condition?: ConstEvaluationRunCondition | VarEvaluationRunCondition | EqEvaluationRunCondition | 12 more`

            - `ConstEvaluationRunCondition`

              - `op?: "const"`

                - `"const"`

              - `value?: string | number | boolean`

                - `string`

                - `number`

                - `boolean`

            - `VarEvaluationRunCondition`

              - `path: string`

              - `op?: "var"`

                - `"var"`

            - `EqEvaluationRunCondition`

            - `NeEvaluationRunCondition`

            - `LtEvaluationRunCondition`

            - `LteEvaluationRunCondition`

            - `GtEvaluationRunCondition`

            - `GteEvaluationRunCondition`

            - `AndEvaluationRunCondition`

            - `OrEvaluationRunCondition`

            - `InEvaluationRunCondition`

            - `NotInEvaluationRunCondition`

            - `NotEvaluationRunCondition`

            - `IsNullEvaluationRunCondition`

            - `IsNotNullEvaluationRunCondition`

          - `system_prompt?: string`

        - `AutoEvaluationAgentTaskRequestWithItemLocator`

          - `definition: string`

          - `name: string`

          - `output_rules: Array<string>`

          - `data_fields?: Array<string>`

          - `designated_to?: ApeAgent | IfAgent | TruthfulnessAgent | BaseAgent`

            - `ApeAgent`

              - `config: Config`

                - `model?: string`

                - `temperature?: number`

              - `agent_name?: "APEAgent"`

                - `"APEAgent"`

            - `IfAgent`

              - `config: Config`

                - `model?: string`

              - `agent_name?: "IFAgent"`

                - `"IFAgent"`

            - `TruthfulnessAgent`

              - `config: Config`

                - `model?: string`

              - `agent_name?: "TruthfulnessAgent"`

                - `"TruthfulnessAgent"`

            - `BaseAgent`

              - `config: Config`

                - `model?: string`

              - `agent_name?: "BaseAgent"`

                - `"BaseAgent"`

          - `output_type?: "text" | "integer" | "float" | "boolean"`

            - `"text"`

            - `"integer"`

            - `"float"`

            - `"boolean"`

          - `output_values?: Array<string | number | boolean>`

            - `string`

            - `number`

            - `boolean`

          - `rubric_id?: string`

          - `rubric_version?: number`

      - `alias?: string`

        Alias to title the results column. Defaults to the `auto_evaluation_guided_decoding`

      - `task_type?: "auto_evaluation.guided_decoding"`

        - `"auto_evaluation.guided_decoding"`

    - `AutoEvaluationAgentEvaluationTask`

      - `configuration: AutoEvaluationAgentTaskRequestWithItemLocator`

      - `alias?: string`

        Alias to title the results column. Defaults to the `auto_evaluation_agent`

      - `task_type?: "auto_evaluation.agent"`

        - `"auto_evaluation.agent"`

    - `ContributorEvaluationQuestionTask`

      - `configuration: Configuration`

        - `layout: Container`

          - `children: Array<Container | Component>`

            The children to be displayed within the container

            - `Container`

            - `Component`

              - `data: ItemLocator`

                A pointer to the data in each evaluation item to be displayed within the component

              - `label?: string`

          - `direction?: "row" | "column"`

            The axis that children are placed in the container. Based on CSS `flex-direction` (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

            - `"row"`

            - `"column"`

        - `question_id: string`

        - `prefill_from?: string`

          Dataset column to prefill contributor question task result

        - `queue_id?: string`

          The contributor annotation queue to include this task in. Defaults to `default`

        - `required?: boolean`

          Whether the question is required to be answered

        - `rubric_id?: string`

          ID of the rubric to use for scoring this evaluation question

      - `alias?: string`

        Alias to title the results column. Defaults to the `contributor_evaluation_question`

      - `task_type?: "contributor_evaluation.question"`

        - `"contributor_evaluation.question"`

    - `CustomFunctionEvaluationTask`

      - `configuration: Configuration`

        Configuration for a custom Python function evaluation task.

        - `function_source: string`

          Python function source code

        - `arg_mapping?: Record<string, string>`

          Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

        - `config_args?: Record<string, unknown>`

          Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

        - `outputs?: Array<Output>`

          Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

          - `path: string`

            Dot path in the custom function return value to materialize.

          - `alias?: string`

            Result column alias. Defaults to path with dots replaced by underscores.

      - `alias?: string`

        Alias to title the results column. Defaults to the function name.

      - `task_type?: "custom_function"`

        - `"custom_function"`

### Example

```typescript
import SGPClient from 'scale-gp';

const client = new SGPClient({
  accountID: 'My Account ID',
  apiKey: process.env['SGP_API_KEY'], // This is the default and can be omitted
});

const evaluation = await client.evaluations.create({
  evaluation: { data: [{ foo: 'bar' }], name: 'x' },
});

console.log(evaluation.id);
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "datasets": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "name": "name",
  "status": "failed",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "error_count": 0,
  "metadata": {
    "foo": "bar"
  },
  "object": "evaluation",
  "progress": {
    "items": {
      "failed": 0,
      "pending": 0,
      "successful": 0,
      "total": 0,
      "failed_items": [
        {
          "item_id": "item_id",
          "error": "error",
          "error_type": "error_type"
        }
      ]
    },
    "workflows": {
      "completed": 0,
      "failed": 0,
      "pending": 0,
      "total": 0
    }
  },
  "status_reason": "status_reason",
  "tasks": [
    {
      "configuration": {
        "messages": [
          {
            "foo": "bar"
          }
        ],
        "model": "model",
        "audio": {
          "foo": "bar"
        },
        "frequency_penalty": -2,
        "function_call": {
          "foo": "bar"
        },
        "functions": [
          {
            "foo": "bar"
          }
        ],
        "logit_bias": {
          "foo": 0
        },
        "logprobs": true,
        "max_completion_tokens": 0,
        "max_tokens": 0,
        "metadata": {
          "foo": "string"
        },
        "modalities": [
          "string"
        ],
        "n": 0,
        "parallel_tool_calls": true,
        "prediction": {
          "foo": "bar"
        },
        "presence_penalty": -2,
        "reasoning_effort": "reasoning_effort",
        "response_format": {
          "foo": "bar"
        },
        "seed": 0,
        "stop": "string",
        "store": true,
        "temperature": 0,
        "tool_choice": "string",
        "tools": [
          {
            "foo": "bar"
          }
        ],
        "top_k": 0,
        "top_logprobs": 0,
        "top_p": 0
      },
      "alias": "alias",
      "task_type": "chat_completion"
    }
  ]
}
```
