Add Test Criteria to Evaluation
Add a new test criteria to an existing evaluation.
Narrowed to contributor question tasks (contributor_evaluation.question); other task types
must be configured when the evaluation is first created and are rejected here. The request is
also rejected if the evaluation is archived, if a test criteria with the same alias already
exists, or if any contributor annotation task for the evaluation has already been claimed or
completed. Because only contributor question tasks are accepted, the added criteria is applied
synchronously and contributors answer it against the evaluation’s existing items — no async job
or Temporal workflow is started.
ParametersExpand Collapse
New test criteria to add to the evaluation. Rejected when contributor annotation tasks for this evaluation have already been claimed or completed. Triggers a rerun so the new task executes against existing items.
New test criteria to add to the evaluation. Rejected when contributor annotation tasks for this evaluation have already been claimed or completed. Triggers a rerun so the new task executes against existing items.
class ChatCompletionEvaluationTask: …
configuration: ChatCompletionEvaluationTaskConfiguration
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
For o1 models only. Constrains effort on reasoning. Values: low, medium, high.
stop: Optional[Union[str, List[str], null]]Up to 4 sequences where the API will stop generating further tokens.
Up to 4 sequences where the API will stop generating further tokens.
tool_choice: Optional[Union[str, Dict[str, object], null]]Controls which tool is called by the model. Values: none, auto, required, or specific tool.
Controls which tool is called by the model. Values: none, auto, required, or specific tool.
class GenericInferenceEvaluationTask: …
class ApplicationVariantV1EvaluationTask: …
configuration: ApplicationVariantV1EvaluationTaskConfiguration
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]History of the application
History of the application
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]Optional overrides for the application
Optional overrides for the application
class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …Execution override options for agentic applications
Execution override options for agentic applications
class AgentexOutputEvaluationTask: …
configuration: AgentexOutputEvaluationTaskConfiguration
input_column: Union[str, Dict[str, object], List[object]]The dataset column to use as input for the agent
The dataset column to use as input for the agent
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
completion_mode: Optional[Literal["first_message", "turn_quiescence"]]How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
input_mode: Optional[Literal["text", "data"]]How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
class MetricEvaluationTask: …
configuration: MetricEvaluationTaskConfiguration
class AutoEvaluationGuidedDecodingEvaluationTask: …
configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …
run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …
run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
class AutoEvaluationAgentEvaluationTask: …
class ContributorEvaluationQuestionTask: …
configuration: ContributorEvaluationQuestionTaskConfiguration
children: List[Child]The children to be displayed within the container
The children to be displayed within the container
direction: Optional[Literal["row", "column"]]The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
class CustomFunctionEvaluationTask: …
configuration: CustomFunctionEvaluationTaskConfigurationConfiguration for a custom Python function evaluation task.
Configuration for a custom Python function evaluation task.
Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.
Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.
ReturnsExpand Collapse
class Evaluation: …
The date and time when the entity was archived in ISO format.
Tasks executed during evaluation. Populated with optional task view.
Tasks executed during evaluation. Populated with optional task view.
class ChatCompletionEvaluationTask: …
configuration: ChatCompletionEvaluationTaskConfiguration
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.
For o1 models only. Constrains effort on reasoning. Values: low, medium, high.
stop: Optional[Union[str, List[str], null]]Up to 4 sequences where the API will stop generating further tokens.
Up to 4 sequences where the API will stop generating further tokens.
tool_choice: Optional[Union[str, Dict[str, object], null]]Controls which tool is called by the model. Values: none, auto, required, or specific tool.
Controls which tool is called by the model. Values: none, auto, required, or specific tool.
class GenericInferenceEvaluationTask: …
class ApplicationVariantV1EvaluationTask: …
configuration: ApplicationVariantV1EvaluationTaskConfiguration
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.
history: Optional[Union[List[ApplicationVariantV1EvaluationTaskConfigurationHistoryApplicationRequestResponsePairArray], ItemLocator, null]]History of the application
History of the application
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.
overrides: Optional[ApplicationVariantV1EvaluationTaskConfigurationOverrides]Optional overrides for the application
Optional overrides for the application
class ApplicationVariantV1EvaluationTaskConfigurationOverridesAgenticApplicationOverrides: …Execution override options for agentic applications
Execution override options for agentic applications
class AgentexOutputEvaluationTask: …
configuration: AgentexOutputEvaluationTaskConfiguration
input_column: Union[str, Dict[str, object], List[object]]The dataset column to use as input for the agent
The dataset column to use as input for the agent
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.
completion_mode: Optional[Literal["first_message", "turn_quiescence"]]How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.
Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.
input_mode: Optional[Literal["text", "data"]]How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.
class MetricEvaluationTask: …
configuration: MetricEvaluationTaskConfiguration
class AutoEvaluationGuidedDecodingEvaluationTask: …
configuration: AutoEvaluationGuidedDecodingEvaluationTaskConfiguration
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator: …
run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunCondition]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator: …
run_condition: Optional[AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunCondition]
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstEvaluationRunCondition: …
class AutoEvaluationGuidedDecodingEvaluationTaskConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVarEvaluationRunCondition: …
class AutoEvaluationAgentEvaluationTask: …
class ContributorEvaluationQuestionTask: …
configuration: ContributorEvaluationQuestionTaskConfiguration
children: List[Child]The children to be displayed within the container
The children to be displayed within the container
direction: Optional[Literal["row", "column"]]The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)
class CustomFunctionEvaluationTask: …
configuration: CustomFunctionEvaluationTaskConfigurationConfiguration for a custom Python function evaluation task.
Configuration for a custom Python function evaluation task.
Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.
Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.
Add Test Criteria to Evaluation
import os
from scale_gp_beta import SGPClient
client = SGPClient(
api_key=os.environ.get("SGP_API_KEY"), # This is the default and can be omitted
)
evaluation = client.evaluations.tasks.add(
evaluation_id="evaluation_id",
task={
"configuration": {
"messages": [{
"foo": "bar"
}],
"model": "model",
},
"task_type": "chat_completion",
},
)
print(evaluation.id){
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"datasets": [
{
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"current_version_num": 0,
"name": "name",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"object": "dataset"
}
],
"name": "name",
"status": "failed",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"error_count": 0,
"metadata": {
"foo": "bar"
},
"object": "evaluation",
"progress": {
"items": {
"failed": 0,
"pending": 0,
"successful": 0,
"total": 0,
"failed_items": [
{
"item_id": "item_id",
"error": "error",
"error_type": "error_type"
}
]
},
"workflows": {
"completed": 0,
"failed": 0,
"pending": 0,
"total": 0
}
},
"status_reason": "status_reason",
"tasks": [
{
"configuration": {
"messages": [
{
"foo": "bar"
}
],
"model": "model",
"audio": {
"foo": "bar"
},
"frequency_penalty": -2,
"function_call": {
"foo": "bar"
},
"functions": [
{
"foo": "bar"
}
],
"logit_bias": {
"foo": 0
},
"logprobs": true,
"max_completion_tokens": 0,
"max_tokens": 0,
"metadata": {
"foo": "string"
},
"modalities": [
"string"
],
"n": 0,
"parallel_tool_calls": true,
"prediction": {
"foo": "bar"
},
"presence_penalty": -2,
"reasoning_effort": "reasoning_effort",
"response_format": {
"foo": "bar"
},
"seed": 0,
"stop": "string",
"store": true,
"temperature": 0,
"tool_choice": "string",
"tools": [
{
"foo": "bar"
}
],
"top_k": 0,
"top_logprobs": 0,
"top_p": 0
},
"alias": "alias",
"task_type": "chat_completion"
}
]
}Returns Examples
{
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"datasets": [
{
"id": "id",
"created_at": "2019-12-27T18:11:19.117Z",
"created_by": {
"id": "id",
"type": "user",
"object": "identity"
},
"current_version_num": 0,
"name": "name",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"object": "dataset"
}
],
"name": "name",
"status": "failed",
"tags": [
"string"
],
"archived_at": "2019-12-27T18:11:19.117Z",
"description": "description",
"error_count": 0,
"metadata": {
"foo": "bar"
},
"object": "evaluation",
"progress": {
"items": {
"failed": 0,
"pending": 0,
"successful": 0,
"total": 0,
"failed_items": [
{
"item_id": "item_id",
"error": "error",
"error_type": "error_type"
}
]
},
"workflows": {
"completed": 0,
"failed": 0,
"pending": 0,
"total": 0
}
},
"status_reason": "status_reason",
"tasks": [
{
"configuration": {
"messages": [
{
"foo": "bar"
}
],
"model": "model",
"audio": {
"foo": "bar"
},
"frequency_penalty": -2,
"function_call": {
"foo": "bar"
},
"functions": [
{
"foo": "bar"
}
],
"logit_bias": {
"foo": 0
},
"logprobs": true,
"max_completion_tokens": 0,
"max_tokens": 0,
"metadata": {
"foo": "string"
},
"modalities": [
"string"
],
"n": 0,
"parallel_tool_calls": true,
"prediction": {
"foo": "bar"
},
"presence_penalty": -2,
"reasoning_effort": "reasoning_effort",
"response_format": {
"foo": "bar"
},
"seed": 0,
"stop": "string",
"store": true,
"temperature": 0,
"tool_choice": "string",
"tools": [
{
"foo": "bar"
}
],
"top_k": 0,
"top_logprobs": 0,
"top_p": 0
},
"alias": "alias",
"task_type": "chat_completion"
}
]
}