Skip to content

Filter Evaluations

client.Evaluations.Filter(ctx, params) (*CursorPage[Evaluation], error)
POST/v5/evaluations/filter

Filter evaluations by metadata, status, and tags.

Accepts up to 10 filters combined with AND logic, each comparing a key against a value with an operator (==, !=, >=, <=, IN, NOT_IN). Filter on metadata keys returned by the metadata-keys endpoint, plus the built-in status and tag keys. Archived evaluations are excluded unless include_archived is set, and the tasks view includes task configurations in each result. Use this for metadata or status filtering; for simple name or tag lookups the list endpoint is sufficient.

ParametersExpand Collapse
params EvaluationFilterParams
Filters param.Field[[]EvaluationFilterParamsFilter]

Body param: List of metadata filters to apply (maximum 10)

Key string

The metadata key to filter on

Operator string

The comparison operator to use

One of the following:
const EvaluationFilterParamsFilterOperatorEquals EvaluationFilterParamsFilterOperator = "=="
const EvaluationFilterParamsFilterOperatorNotEquals EvaluationFilterParamsFilterOperator = "!="
const EvaluationFilterParamsFilterOperatorGreaterOrEquals EvaluationFilterParamsFilterOperator = ">="
const EvaluationFilterParamsFilterOperatorLessOrEquals EvaluationFilterParamsFilterOperator = "<="
const EvaluationFilterParamsFilterOperatorIn EvaluationFilterParamsFilterOperator = "IN"
const EvaluationFilterParamsFilterOperatorNotIn EvaluationFilterParamsFilterOperator = "NOT_IN"
Value string

The value to compare against (string for all types)

Object stringOptional
EndingBefore param.Field[string]Optional

Query param

IncludeArchived param.Field[bool]Optional

Query param

Limit param.Field[int64]Optional

Query param

maximum10000
minimum1
SortBy param.Field[string]Optional

Query param

SortOrder param.Field[SortOrder]Optional

Query param

StartingAfter param.Field[string]Optional

Query param

Views param.Field[[]EvaluationViews]Optional

Query param

const EvaluationViewsTasks EvaluationViews = "tasks"
ReturnsExpand Collapse
type Evaluation struct{…}
ID string

The unique identifier of the entity.

CreatedAt Time

The date and time when the entity was created in ISO format.

formatdate-time
CreatedBy Identity

The identity that created the entity.

ID string
Type IdentityType
One of the following:
const IdentityTypeUser IdentityType = "user"
const IdentityTypeServiceAccount IdentityType = "service_account"
Object IdentityObjectOptional
Datasets []Dataset
ID string

The unique identifier of the entity.

CreatedAt Time

The date and time when the entity was created in ISO format.

formatdate-time
CreatedBy Identity

The identity that created the entity.

ID string
Type IdentityType
One of the following:
const IdentityTypeUser IdentityType = "user"
const IdentityTypeServiceAccount IdentityType = "service_account"
Object IdentityObjectOptional
CurrentVersionNum int64
Name string
Tags []string

The tags associated with the entity

ArchivedAt TimeOptional

The date and time when the entity was archived in ISO format.

formatdate-time
Description stringOptional
Object DatasetObjectOptional
Name string
Status EvaluationStatus
One of the following:
const EvaluationStatusFailed EvaluationStatus = "failed"
const EvaluationStatusCompleted EvaluationStatus = "completed"
const EvaluationStatusRunning EvaluationStatus = "running"
Tags []string

The tags associated with the entity

ArchivedAt TimeOptional

The date and time when the entity was archived in ISO format.

formatdate-time
Description stringOptional
ErrorCount int64Optional

Number of task errors across all items in this evaluation.

Metadata map[string, any]Optional

Metadata key-value pairs for the evaluation

Object EvaluationObjectOptional

Progress of the evaluation’s underlying async job

Items EvaluationTasksProgressSchemaItemsOptional
Failed int64
Pending int64
Successful int64
Total int64
FailedItems []EvaluationTasksProgressSchemaItemsFailedItemOptional
ItemID string
Error stringOptional
ErrorType stringOptional
Workflows EvaluationTasksProgressSchemaWorkflowsOptional
Completed int64
Failed int64
Pending int64
Total int64
StatusReason stringOptional

Reason for evaluation status

Tasks []EvaluationTaskUnionOptional

Tasks executed during evaluation. Populated with optional task view.

One of the following:
type EvaluationTaskChatCompletion struct{…}
Configuration EvaluationTaskChatCompletionConfiguration
Messages EvaluationTaskChatCompletionConfigurationMessagesUnion

openai standard message format

One of the following:
type EvaluationTaskChatCompletionConfigurationMessagesArray []map[string, any]
type ItemLocator string
Model string

model specified as model_vendor/model, for example openai/gpt-4o

Audio EvaluationTaskChatCompletionConfigurationAudioUnionOptional

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
type EvaluationTaskChatCompletionConfigurationAudioMap map[string, any]
type ItemLocator string
FrequencyPenalty EvaluationTaskChatCompletionConfigurationFrequencyPenaltyUnionOptional

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
float64
type ItemLocator string
FunctionCall EvaluationTaskChatCompletionConfigurationFunctionCallUnionOptional

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
type EvaluationTaskChatCompletionConfigurationFunctionCallMap map[string, any]
type ItemLocator string
Functions EvaluationTaskChatCompletionConfigurationFunctionsUnionOptional

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
type EvaluationTaskChatCompletionConfigurationFunctionsArray []map[string, any]
type ItemLocator string
LogitBias EvaluationTaskChatCompletionConfigurationLogitBiasUnionOptional

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
type EvaluationTaskChatCompletionConfigurationLogitBiasMap map[string, int64]
type ItemLocator string
Logprobs EvaluationTaskChatCompletionConfigurationLogprobsUnionOptional

Whether to return log probabilities of the output tokens or not.

One of the following:
bool
type ItemLocator string
MaxCompletionTokens EvaluationTaskChatCompletionConfigurationMaxCompletionTokensUnionOptional

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
int64
type ItemLocator string
MaxTokens EvaluationTaskChatCompletionConfigurationMaxTokensUnionOptional

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
int64
type ItemLocator string
Metadata EvaluationTaskChatCompletionConfigurationMetadataUnionOptional

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
type EvaluationTaskChatCompletionConfigurationMetadataMap map[string, string]
type ItemLocator string
Modalities EvaluationTaskChatCompletionConfigurationModalitiesUnionOptional

Output types that you would like the model to generate for this request.

One of the following:
type EvaluationTaskChatCompletionConfigurationModalitiesArray []string
type ItemLocator string
N EvaluationTaskChatCompletionConfigurationNUnionOptional

How many chat completion choices to generate for each input message.

One of the following:
int64
type ItemLocator string
ParallelToolCalls EvaluationTaskChatCompletionConfigurationParallelToolCallsUnionOptional

Whether to enable parallel function calling during tool use.

One of the following:
bool
type ItemLocator string
Prediction EvaluationTaskChatCompletionConfigurationPredictionUnionOptional

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
type EvaluationTaskChatCompletionConfigurationPredictionMap map[string, any]
type ItemLocator string
PresencePenalty EvaluationTaskChatCompletionConfigurationPresencePenaltyUnionOptional

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
float64
type ItemLocator string
ReasoningEffort stringOptional

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

ResponseFormat EvaluationTaskChatCompletionConfigurationResponseFormatUnionOptional

An object specifying the format that the model must output.

One of the following:
type EvaluationTaskChatCompletionConfigurationResponseFormatMap map[string, any]
type ItemLocator string
Seed EvaluationTaskChatCompletionConfigurationSeedUnionOptional

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
int64
type ItemLocator string
Stop EvaluationTaskChatCompletionConfigurationStopUnionOptional

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
string
type EvaluationTaskChatCompletionConfigurationStopArray []string
Store EvaluationTaskChatCompletionConfigurationStoreUnionOptional

Whether to store the output for use in model distillation or evals products.

One of the following:
bool
type ItemLocator string
Temperature EvaluationTaskChatCompletionConfigurationTemperatureUnionOptional

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
float64
type ItemLocator string
ToolChoice EvaluationTaskChatCompletionConfigurationToolChoiceUnionOptional

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
string
type EvaluationTaskChatCompletionConfigurationToolChoiceMap map[string, any]
Tools EvaluationTaskChatCompletionConfigurationToolsUnionOptional

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
type EvaluationTaskChatCompletionConfigurationToolsArray []map[string, any]
type ItemLocator string
TopK EvaluationTaskChatCompletionConfigurationTopKUnionOptional

Only sample from the top K options for each subsequent token

One of the following:
int64
type ItemLocator string
TopLogprobs EvaluationTaskChatCompletionConfigurationTopLogprobsUnionOptional

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
int64
type ItemLocator string
TopP EvaluationTaskChatCompletionConfigurationTopPUnionOptional

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
float64
type ItemLocator string
Alias stringOptional

Alias to title the results column. Defaults to the chat_completion

TaskType stringOptional
type EvaluationTaskInference struct{…}
Configuration EvaluationTaskInferenceConfiguration
Model string

model specified as vendor/name (ex. openai/gpt-5)

Args EvaluationTaskInferenceConfigurationArgsUnionOptional

Arguments passed into model

One of the following:
type EvaluationTaskInferenceConfigurationArgsMap map[string, any]
type ItemLocator string
InferenceConfiguration EvaluationTaskInferenceConfigurationInferenceConfigurationUnionOptional

Vendor specific configuration

One of the following:
type LaunchInferenceConfiguration struct{…}
NumRetries int64Optional
TimeoutSeconds int64Optional
type ItemLocator string
Alias stringOptional

Alias to title the results column. Defaults to the inference

TaskType stringOptional
type EvaluationTaskApplicationVariant struct{…}
Configuration EvaluationTaskApplicationVariantConfiguration
ApplicationVariantID string
Inputs EvaluationTaskApplicationVariantConfigurationInputsUnion

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
type EvaluationTaskApplicationVariantConfigurationInputsMap map[string, any]
type ItemLocator string
History EvaluationTaskApplicationVariantConfigurationHistoryUnionOptional

History of the application

One of the following:
type EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArray []EvaluationTaskApplicationVariantConfigurationHistoryApplicationRequestResponsePairArrayItem
Request string

Request inputs

Response string

Response outputs

SessionData map[string, any]Optional

Session data corresponding to the request response pair

type ItemLocator string
OperationMetadata EvaluationTaskApplicationVariantConfigurationOperationMetadataUnionOptional

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
type EvaluationTaskApplicationVariantConfigurationOperationMetadataMap map[string, any]
type ItemLocator string
OverridesProperty EvaluationTaskApplicationVariantConfigurationOverridesUnionOptional

Optional overrides for the application

One of the following:
type EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverrides struct{…}

Execution override options for agentic applications

Concurrent boolOptional
InitialState EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesInitialStateOptional
CurrentNode string
State map[string, any]
PartialTrace []EvaluationTaskApplicationVariantConfigurationOverridesAgenticApplicationOverridesPartialTraceOptional
DurationMs int64
NodeID string
OperationInput string
OperationOutput string
OperationType string
StartTimestamp string
WorkflowID string
OperationMetadata map[string, any]Optional
ReturnSpan boolOptional
UseChannels boolOptional
type EvaluationTaskApplicationVariantConfigurationOverridesMap map[string, EvaluationTaskApplicationVariantConfigurationOverridesMapItem]
ArtifactIDsFilter []stringOptional
ArtifactNameRegex []stringOptional
Type stringOptional
type ItemLocator string
Alias stringOptional

Alias to title the results column. Defaults to the application_variant

TaskType stringOptional
type EvaluationTaskAgentexOutput struct{…}
Configuration EvaluationTaskAgentexOutputConfiguration
AgentexAgentID string

The ID of the Agentex agent to use

InputColumn EvaluationTaskAgentexOutputConfigurationInputColumnUnion

The dataset column to use as input for the agent

One of the following:
string
type EvaluationTaskAgentexOutputConfigurationInputColumnMap map[string, any]
type EvaluationTaskAgentexOutputConfigurationInputColumnArray []any
AgentTaskParams EvaluationTaskAgentexOutputConfigurationAgentTaskParamsUnionOptional

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
type EvaluationTaskAgentexOutputConfigurationAgentTaskParamsMap map[string, any]
type ItemLocator string
CompletionMode stringOptional

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
const EvaluationTaskAgentexOutputConfigurationCompletionModeFirstMessage EvaluationTaskAgentexOutputConfigurationCompletionMode = "first_message"
const EvaluationTaskAgentexOutputConfigurationCompletionModeTurnQuiescence EvaluationTaskAgentexOutputConfigurationCompletionMode = "turn_quiescence"
DeploymentID stringOptional

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

IncludeTraces EvaluationTaskAgentexOutputConfigurationIncludeTracesUnionOptional

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
bool
type ItemLocator string
InputMode stringOptional

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
const EvaluationTaskAgentexOutputConfigurationInputModeText EvaluationTaskAgentexOutputConfigurationInputMode = "text"
const EvaluationTaskAgentexOutputConfigurationInputModeData EvaluationTaskAgentexOutputConfigurationInputMode = "data"
QuiescenceSeconds EvaluationTaskAgentexOutputConfigurationQuiescenceSecondsUnionOptional

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
int64
type ItemLocator string
TimeoutSeconds EvaluationTaskAgentexOutputConfigurationTimeoutSecondsUnionOptional

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
int64
type ItemLocator string
Alias stringOptional

Alias to title the results column. Defaults to the agentex_output

TaskType stringOptional
type EvaluationTaskMetric struct{…}
Configuration EvaluationTaskMetricConfigurationUnion
One of the following:
type EvaluationTaskMetricConfigurationBleu struct{…}
Candidate string
Reference string
Type Bleu
type EvaluationTaskMetricConfigurationMeteor struct{…}
Candidate string
Reference string
Type Meteor
type EvaluationTaskMetricConfigurationCosineSimilarity struct{…}
Candidate string
Reference string
Type CosineSimilarity
type EvaluationTaskMetricConfigurationF1 struct{…}
Candidate string
Reference string
Type F1
type EvaluationTaskMetricConfigurationRouge1 struct{…}
Candidate string
Reference string
Type Rouge1
type EvaluationTaskMetricConfigurationRouge2 struct{…}
Candidate string
Reference string
Type Rouge2
type EvaluationTaskMetricConfigurationRougeL struct{…}
Candidate string
Reference string
Type RougeL
Alias stringOptional

Alias to title the results column. Defaults to the metric type specified in the configuration

TaskType stringOptional
type EvaluationTaskAutoEvaluationQuestion struct{…}
Configuration EvaluationTaskAutoEvaluationQuestionConfiguration
Model string

model specified as model_vendor/model_name

Prompt string
QuestionID string

question to be evaluated

Alias stringOptional

Alias to title the results column. Defaults to the auto_evaluation_question

TaskType stringOptional
type EvaluationTaskAutoEvaluationGuidedDecoding struct{…}
Configuration EvaluationTaskAutoEvaluationGuidedDecodingConfigurationUnion
One of the following:
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocator struct{…}
Model string

model specified as model_vendor/model_name

Prompt string
ResponseFormat map[string, any]

JSON schema used for structuring the model response

InferenceArgs map[string, any]Optional

Additional arguments to pass to the inference request

RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionUnionOptional
One of the following:
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConst struct{…}
Op stringOptional
Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionConstValueUnionOptional
One of the following:
string
float64
bool
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationStructuredOutputTaskRequestWithItemLocatorRunConditionVar struct{…}
Path string
Op stringOptional
type EqEvaluationRunCondition struct{…}
Left any
Right any
Op EqEvaluationRunConditionOpOptional
type NeEvaluationRunCondition struct{…}
Left any
Right any
Op NeEvaluationRunConditionOpOptional
type LtEvaluationRunCondition struct{…}
Left any
Right any
Op LtEvaluationRunConditionOpOptional
type LteEvaluationRunCondition struct{…}
Left any
Right any
Op LteEvaluationRunConditionOpOptional
type GtEvaluationRunCondition struct{…}
Left any
Right any
Op GtEvaluationRunConditionOpOptional
type GteEvaluationRunCondition struct{…}
Left any
Right any
Op GteEvaluationRunConditionOpOptional
type AndEvaluationRunCondition struct{…}
Operands []any
Op AndEvaluationRunConditionOpOptional
type OrEvaluationRunCondition struct{…}
Operands []any
Op OrEvaluationRunConditionOpOptional
type InEvaluationRunCondition struct{…}
Left any
Operands []any
Op InEvaluationRunConditionOpOptional
type NotInEvaluationRunCondition struct{…}
Left any
Operands []any
Op NotInEvaluationRunConditionOpOptional
type NotEvaluationRunCondition struct{…}
Operands []any
Op NotEvaluationRunConditionOpOptional
type IsNullEvaluationRunCondition struct{…}
Operands []any
Op IsNullEvaluationRunConditionOpOptional
type IsNotNullEvaluationRunCondition struct{…}
Operands []any
Op IsNotNullEvaluationRunConditionOpOptional
SystemPrompt stringOptional
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocator struct{…}
Choices []string

Choices array cannot be empty

Model string

model specified as model_vendor/model_name

Prompt string
InferenceArgs map[string, any]Optional

Additional arguments to pass to the inference request

RunCondition EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionUnionOptional
One of the following:
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConst struct{…}
Op stringOptional
Value EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionConstValueUnionOptional
One of the following:
string
float64
bool
type EvaluationTaskAutoEvaluationGuidedDecodingConfigurationAutoEvaluationGuidedDecodingTaskRequestWithItemLocatorRunConditionVar struct{…}
Path string
Op stringOptional
type EqEvaluationRunCondition struct{…}
Left any
Right any
Op EqEvaluationRunConditionOpOptional
type NeEvaluationRunCondition struct{…}
Left any
Right any
Op NeEvaluationRunConditionOpOptional
type LtEvaluationRunCondition struct{…}
Left any
Right any
Op LtEvaluationRunConditionOpOptional
type LteEvaluationRunCondition struct{…}
Left any
Right any
Op LteEvaluationRunConditionOpOptional
type GtEvaluationRunCondition struct{…}
Left any
Right any
Op GtEvaluationRunConditionOpOptional
type GteEvaluationRunCondition struct{…}
Left any
Right any
Op GteEvaluationRunConditionOpOptional
type AndEvaluationRunCondition struct{…}
Operands []any
Op AndEvaluationRunConditionOpOptional
type OrEvaluationRunCondition struct{…}
Operands []any
Op OrEvaluationRunConditionOpOptional
type InEvaluationRunCondition struct{…}
Left any
Operands []any
Op InEvaluationRunConditionOpOptional
type NotInEvaluationRunCondition struct{…}
Left any
Operands []any
Op NotInEvaluationRunConditionOpOptional
type NotEvaluationRunCondition struct{…}
Operands []any
Op NotEvaluationRunConditionOpOptional
type IsNullEvaluationRunCondition struct{…}
Operands []any
Op IsNullEvaluationRunConditionOpOptional
type IsNotNullEvaluationRunCondition struct{…}
Operands []any
Op IsNotNullEvaluationRunConditionOpOptional
SystemPrompt stringOptional
type AutoEvaluationAgentTaskRequestWithItemLocator struct{…}
Definition string
Name string
OutputRules []string
DataFields []stringOptional
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
One of the following:
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig
Model stringOptional
Temperature float64Optional
AgentName stringOptional
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig
Model stringOptional
AgentName stringOptional
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig
Model stringOptional
AgentName stringOptional
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig
Model stringOptional
AgentName stringOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
One of the following:
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"
OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnionOptional
One of the following:
string
float64
bool
RubricID stringOptional
RubricVersion int64Optional
Alias stringOptional

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

TaskType stringOptional
type EvaluationTaskAutoEvaluationAgent struct{…}
Definition string
Name string
OutputRules []string
DataFields []stringOptional
DesignatedTo AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToUnionOptional
One of the following:
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToApeAgentConfig
Model stringOptional
Temperature float64Optional
AgentName stringOptional
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToIfAgentConfig
Model stringOptional
AgentName stringOptional
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToTruthfulnessAgentConfig
Model stringOptional
AgentName stringOptional
type AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgent struct{…}
Config AutoEvaluationAgentTaskRequestWithItemLocatorDesignatedToBaseAgentConfig
Model stringOptional
AgentName stringOptional
OutputType AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeOptional
One of the following:
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeText AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "text"
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeInteger AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "integer"
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeFloat AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "float"
const AutoEvaluationAgentTaskRequestWithItemLocatorOutputTypeBoolean AutoEvaluationAgentTaskRequestWithItemLocatorOutputType = "boolean"
OutputValues []AutoEvaluationAgentTaskRequestWithItemLocatorOutputValueUnionOptional
One of the following:
string
float64
bool
RubricID stringOptional
RubricVersion int64Optional
Alias stringOptional

Alias to title the results column. Defaults to the auto_evaluation_agent

TaskType stringOptional
type EvaluationTaskContributorEvaluationQuestion struct{…}
Configuration EvaluationTaskContributorEvaluationQuestionConfiguration
Layout Container
Children []ContainerChildUnion

The children to be displayed within the container

One of the following:
type Container Container
type Component struct{…}

A pointer to the data in each evaluation item to be displayed within the component

Label stringOptional
Direction ContainerDirectionOptional

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
const ContainerDirectionRow ContainerDirection = "row"
const ContainerDirectionColumn ContainerDirection = "column"
QuestionID string
PrefillFrom stringOptional

Dataset column to prefill contributor question task result

minLength1
QueueID stringOptional

The contributor annotation queue to include this task in. Defaults to default

maxLength100
Required boolOptional

Whether the question is required to be answered

RubricID stringOptional

ID of the rubric to use for scoring this evaluation question

Alias stringOptional

Alias to title the results column. Defaults to the contributor_evaluation_question

TaskType stringOptional
type EvaluationTaskCustomFunction struct{…}
Configuration EvaluationTaskCustomFunctionConfiguration

Configuration for a custom Python function evaluation task.

FunctionSource string

Python function source code

maxLength10000
ArgMapping map[string, string]Optional

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

ConfigArgs map[string, any]Optional

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

Outputs []EvaluationTaskCustomFunctionConfigurationOutputOptional

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

Path string

Dot path in the custom function return value to materialize.

minLength1
Alias stringOptional

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
Alias stringOptional

Alias to title the results column. Defaults to the function name.

TaskType stringOptional

Filter Evaluations

package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  page, err := client.Evaluations.Filter(context.TODO(), sgpdev.EvaluationFilterParams{
    Filters: []sgpdev.EvaluationFilterParamsFilter{sgpdev.EvaluationFilterParamsFilter{
      Key: "key",
      Operator: "==",
      Value: "value",
    }},
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", page)
}
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
Returns Examples
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}