Skip to content

Filter Evaluations

POST/v5/evaluations/filter

Filter evaluations by metadata, status, and tags.

Accepts up to 10 filters combined with AND logic, each comparing a key against a value with an operator (==, !=, >=, <=, IN, NOT_IN). Filter on metadata keys returned by the metadata-keys endpoint, plus the built-in status and tag keys. Archived evaluations are excluded unless include_archived is set, and the tasks view includes task configurations in each result. Use this for metadata or status filtering; for simple name or tag lookups the list endpoint is sufficient.

Query ParametersExpand Collapse
ending_before: optional string
include_archived: optional boolean
limit: optional number
maximum10000
minimum1
sort_by: optional string
sort_order: optional SortOrder
One of the following:
"asc"
"desc"
starting_after: optional string
views: optional array of EvaluationViews
Body ParametersJSONExpand Collapse
filters: array of object { key, operator, value, object }

List of metadata filters to apply (maximum 10)

key: string

The metadata key to filter on

operator: "==" or "!=" or ">=" or 3 more

The comparison operator to use

One of the following:
"=="
"!="
">="
"<="
"IN"
"NOT_IN"
value: string

The value to compare against (string for all types)

object: optional "metadata_filter"
ReturnsExpand Collapse
PaginatedListEvaluation object { has_more, items, total, 2 more }
has_more: boolean

Whether there are more items left to be fetched.

items: array of Evaluation { id, created_at, created_by, 12 more }
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity { id, type, object }

The identity that created the entity.

id: string
type: "user" or "service_account"
One of the following:
"user"
"service_account"
object: optional "identity"
datasets: array of Dataset { id, created_at, created_by, 6 more }
id: string

The unique identifier of the entity.

created_at: string

The date and time when the entity was created in ISO format.

formatdate-time
created_by: Identity { id, type, object }

The identity that created the entity.

id: string
type: "user" or "service_account"
One of the following:
"user"
"service_account"
object: optional "identity"
current_version_num: number
name: string
tags: array of string

The tags associated with the entity

archived_at: optional string

The date and time when the entity was archived in ISO format.

formatdate-time
description: optional string
object: optional "dataset"
name: string
status: "failed" or "completed" or "running"
One of the following:
"failed"
"completed"
"running"
tags: array of string

The tags associated with the entity

archived_at: optional string

The date and time when the entity was archived in ISO format.

formatdate-time
description: optional string
error_count: optional number

Number of task errors across all items in this evaluation.

metadata: optional map[unknown]

Metadata key-value pairs for the evaluation

object: optional "evaluation"
progress: optional EvaluationTasksProgressSchema { items, workflows }

Progress of the evaluation’s underlying async job

items: optional object { failed, pending, successful, 2 more }
failed: number
pending: number
successful: number
total: number
failed_items: optional array of object { item_id, error, error_type }
item_id: string
error: optional string
error_type: optional string
workflows: optional object { completed, failed, pending, total }
completed: number
failed: number
pending: number
total: number
status_reason: optional string

Reason for evaluation status

tasks: optional array of EvaluationTask

Tasks executed during evaluation. Populated with optional task view.

One of the following:
ChatCompletion object { configuration, alias, task_type }
configuration: object { messages, model, audio, 24 more }
messages: array of map[unknown] or ItemLocator

openai standard message format

One of the following:
array of map[unknown]
ItemLocator = string
model: string

model specified as model_vendor/model, for example openai/gpt-4o

audio: optional map[unknown] or ItemLocator

Parameters for audio output. Required when audio output is requested with modalities: [‘audio’].

One of the following:
map[unknown]
ItemLocator = string
frequency_penalty: optional number or ItemLocator

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.

One of the following:
number
ItemLocator = string
function_call: optional map[unknown] or ItemLocator

Deprecated in favor of tool_choice. Controls which function is called by the model.

One of the following:
map[unknown]
ItemLocator = string
functions: optional array of map[unknown] or ItemLocator

Deprecated in favor of tools. A list of functions the model may generate JSON inputs for.

One of the following:
array of map[unknown]
ItemLocator = string
logit_bias: optional map[number] or ItemLocator

Modify the likelihood of specified tokens appearing in the completion. Maps tokens to bias values from -100 to 100.

One of the following:
map[number]
ItemLocator = string
logprobs: optional boolean or ItemLocator

Whether to return log probabilities of the output tokens or not.

One of the following:
boolean
ItemLocator = string
max_completion_tokens: optional number or ItemLocator

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens.

One of the following:
number
ItemLocator = string
max_tokens: optional number or ItemLocator

Deprecated in favor of max_completion_tokens. The maximum number of tokens to generate.

One of the following:
number
ItemLocator = string
metadata: optional map[string] or ItemLocator

Developer-defined tags and values used for filtering completions in the dashboard.

One of the following:
map[string]
ItemLocator = string
modalities: optional array of string or ItemLocator

Output types that you would like the model to generate for this request.

One of the following:
array of string
ItemLocator = string
n: optional number or ItemLocator

How many chat completion choices to generate for each input message.

One of the following:
number
ItemLocator = string
parallel_tool_calls: optional boolean or ItemLocator

Whether to enable parallel function calling during tool use.

One of the following:
boolean
ItemLocator = string
prediction: optional map[unknown] or ItemLocator

Static predicted output content, such as the content of a text file being regenerated.

One of the following:
map[unknown]
ItemLocator = string
presence_penalty: optional number or ItemLocator

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they appear in the text so far.

One of the following:
number
ItemLocator = string
reasoning_effort: optional string

For o1 models only. Constrains effort on reasoning. Values: low, medium, high.

response_format: optional map[unknown] or ItemLocator

An object specifying the format that the model must output.

One of the following:
map[unknown]
ItemLocator = string
seed: optional number or ItemLocator

If specified, system will attempt to sample deterministically for repeated requests with same seed.

One of the following:
number
ItemLocator = string
stop: optional string or array of string

Up to 4 sequences where the API will stop generating further tokens.

One of the following:
string
array of string
store: optional boolean or ItemLocator

Whether to store the output for use in model distillation or evals products.

One of the following:
boolean
ItemLocator = string
temperature: optional number or ItemLocator

What sampling temperature to use. Higher values make output more random, lower more focused.

One of the following:
number
ItemLocator = string
tool_choice: optional string or map[unknown]

Controls which tool is called by the model. Values: none, auto, required, or specific tool.

One of the following:
string
map[unknown]
tools: optional array of map[unknown] or ItemLocator

A list of tools the model may call. Currently, only functions are supported. Max 128 functions.

One of the following:
array of map[unknown]
ItemLocator = string
top_k: optional number or ItemLocator

Only sample from the top K options for each subsequent token

One of the following:
number
ItemLocator = string
top_logprobs: optional number or ItemLocator

Number of most likely tokens to return at each position, with associated log probability.

One of the following:
number
ItemLocator = string
top_p: optional number or ItemLocator

Alternative to temperature. Only tokens comprising top_p probability mass are considered.

One of the following:
number
ItemLocator = string
alias: optional string

Alias to title the results column. Defaults to the chat_completion

task_type: optional "chat_completion"
Inference object { configuration, alias, task_type }
configuration: object { model, args, inference_configuration }
model: string

model specified as vendor/name (ex. openai/gpt-5)

args: optional map[unknown] or ItemLocator

Arguments passed into model

One of the following:
map[unknown]
ItemLocator = string
inference_configuration: optional LaunchInferenceConfiguration { num_retries, timeout_seconds } or ItemLocator

Vendor specific configuration

One of the following:
LaunchInferenceConfiguration object { num_retries, timeout_seconds }
num_retries: optional number
timeout_seconds: optional number
ItemLocator = string
alias: optional string

Alias to title the results column. Defaults to the inference

task_type: optional "inference"
ApplicationVariant object { configuration, alias, task_type }
configuration: object { application_variant_id, inputs, history, 2 more }
application_variant_id: string
inputs: map[unknown] or ItemLocator

Input data for the application. For agents service variants, you must provide inputs as a mapping from {input_name: input_value}. For V0 variants, you must specify the node your input should be passed to, structuring your input as {node_id: {input_name: input_value}}.

One of the following:
map[unknown]
ItemLocator = string
history: optional array of object { request, response, session_data } or ItemLocator

History of the application

One of the following:
ApplicationRequestResponsePairArray = array of object { request, response, session_data }
request: string

Request inputs

response: string

Response outputs

session_data: optional map[unknown]

Session data corresponding to the request response pair

ItemLocator = string
operation_metadata: optional map[unknown] or ItemLocator

Arbitrary user-defined metadata that can be attached to the process operations and will be registered in the interaction.

One of the following:
map[unknown]
ItemLocator = string
overrides: optional object { concurrent, initial_state, partial_trace, 2 more } or map[object { artifact_ids_filter, artifact_name_regex, type } ] or ItemLocator

Optional overrides for the application

One of the following:
AgenticApplicationOverrides object { concurrent, initial_state, partial_trace, 2 more }

Execution override options for agentic applications

concurrent: optional boolean
initial_state: optional object { current_node, state }
current_node: string
state: map[unknown]
partial_trace: optional array of object { duration_ms, node_id, operation_input, 5 more }
duration_ms: number
node_id: string
operation_input: string
operation_output: string
operation_type: string
start_timestamp: string
workflow_id: string
operation_metadata: optional map[unknown]
return_span: optional boolean
use_channels: optional boolean
map[object { artifact_ids_filter, artifact_name_regex, type } ]
artifact_ids_filter: optional array of string
artifact_name_regex: optional array of string
type: optional "knowledge_base_schema"
ItemLocator = string
alias: optional string

Alias to title the results column. Defaults to the application_variant

task_type: optional "application_variant"
AgentexOutput object { configuration, alias, task_type }
configuration: object { agentex_agent_id, input_column, agent_task_params, 6 more }
agentex_agent_id: string

The ID of the Agentex agent to use

input_column: string or map[unknown] or array of unknown

The dataset column to use as input for the agent

One of the following:
string
map[unknown]
array of unknown
agent_task_params: optional map[unknown] or ItemLocator

Extra params merged into the Agentex task/create call’s params object and forwarded verbatim to the agent. Required by agents that demand configuration at task creation — the golden agent, for example, rejects any task whose params omit config_id. SGP always pins is_eval: true; a caller-supplied description overrides the SGP default. Nested item.-prefixed strings and {{item.x}} templates are resolved per evaluation item, so a per-row config_id can come from a dataset column.

One of the following:
map[unknown]
ItemLocator = string
completion_mode: optional "first_message" or "turn_quiescence"

How the agent’s first turn is judged finished. first_message (the default) grades the first non-empty agent text message after the input, which is cheap but grades a streaming harness on whatever text block streamed first. turn_quiescence keeps listening while the agent is still producing messages and grades once at least one agent text message exists and nothing new has arrived for quiescence_seconds — the right choice for tool-using agents. Neither mode requires the agent to mark the task complete; a terminal task status always ends the wait, and timeout_seconds always bounds it.

One of the following:
"first_message"
"turn_quiescence"
deployment_id: optional string

Optional Agentex deployment ID to pin the eval to a specific deployment. When set, RPC traffic routes through /agents/{agent_id}/deployments/{deployment_id}/rpc. When unset, traffic uses the agent’s default RPC endpoint, which resolves through the agent’s current routing rules on the Agentex side.

include_traces: optional boolean or ItemLocator

Whether to include trace data in the evaluation results. Traces are read from SGP’s own span store for the agent’s trace, not from Agentex.

One of the following:
boolean
ItemLocator = string
input_mode: optional "text" or "data"

How the resolved input_column is delivered to the agent. text (the default) sends a TextContent message with the value stringified. data sends a DataContent message whose data is the value as a JSON object; the resolved value must be an object, or a string that parses to one. Most agents accept text only and reject data.

One of the following:
"text"
"data"
quiescence_seconds: optional number or ItemLocator

Seconds of no new messages before completion_mode: turn_quiescence considers the turn finished. Ignored in first_message mode. Should exceed the agent’s longest expected gap between messages (a slow tool call), or the turn is graded early.

One of the following:
number
ItemLocator = string
timeout_seconds: optional number or ItemLocator

Maximum seconds to wait for the agent’s first-turn response per item. If not set, the server-side default of 600s applies. Capped at 1500s to stay within the evaluation item activity’s 1800s start-to-close budget.

One of the following:
number
ItemLocator = string
alias: optional string

Alias to title the results column. Defaults to the agentex_output

task_type: optional "agentex_output"
Metric object { configuration, alias, task_type }
configuration: object { candidate, reference, type } or object { candidate, reference, type } or object { candidate, reference, type } or 4 more
One of the following:
Bleu object { candidate, reference, type }
candidate: string
reference: string
type: "bleu"
Meteor object { candidate, reference, type }
candidate: string
reference: string
type: "meteor"
CosineSimilarity object { candidate, reference, type }
candidate: string
reference: string
type: "cosine_similarity"
F1 object { candidate, reference, type }
candidate: string
reference: string
type: "f1"
Rouge1 object { candidate, reference, type }
candidate: string
reference: string
type: "rouge1"
Rouge2 object { candidate, reference, type }
candidate: string
reference: string
type: "rouge2"
RougeL object { candidate, reference, type }
candidate: string
reference: string
type: "rougeL"
alias: optional string

Alias to title the results column. Defaults to the metric type specified in the configuration

task_type: optional "metric"
AutoEvaluationQuestion object { configuration, alias, task_type }
configuration: object { model, prompt, question_id }
model: string

model specified as model_vendor/model_name

prompt: string
question_id: string

question to be evaluated

alias: optional string

Alias to title the results column. Defaults to the auto_evaluation_question

task_type: optional "auto_evaluation.question"
AutoEvaluationGuidedDecoding object { configuration, alias, task_type }
configuration: object { model, prompt, response_format, 3 more } or object { choices, model, prompt, 3 more } or AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
One of the following:
AutoEvaluationStructuredOutputTaskRequestWithItemLocator object { model, prompt, response_format, 3 more }
model: string

model specified as model_vendor/model_name

prompt: string
response_format: map[unknown]

JSON schema used for structuring the model response

inference_args: optional map[unknown]

Additional arguments to pass to the inference request

run_condition: optional object { op, value } or object { path, op } or EqEvaluationRunCondition { left, right, op } or 12 more
One of the following:
Const object { op, value }
op: optional "const"
value: optional string or number or boolean
One of the following:
string
number
boolean
Var object { path, op }
path: string
op: optional "var"
EqEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "eq"
NeEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "ne"
LtEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "lt"
LteEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "lte"
GtEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "gt"
GteEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "gte"
AndEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "and"
OrEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "or"
InEvaluationRunCondition object { left, operands, op }
left: unknown
operands: array of unknown
op: optional "in"
NotInEvaluationRunCondition object { left, operands, op }
left: unknown
operands: array of unknown
op: optional "not_in"
NotEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "not"
IsNullEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "is_null"
IsNotNullEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "is_not_null"
system_prompt: optional string
AutoEvaluationGuidedDecodingTaskRequestWithItemLocator object { choices, model, prompt, 3 more }
choices: array of string

Choices array cannot be empty

model: string

model specified as model_vendor/model_name

prompt: string
inference_args: optional map[unknown]

Additional arguments to pass to the inference request

run_condition: optional object { op, value } or object { path, op } or EqEvaluationRunCondition { left, right, op } or 12 more
One of the following:
Const object { op, value }
op: optional "const"
value: optional string or number or boolean
One of the following:
string
number
boolean
Var object { path, op }
path: string
op: optional "var"
EqEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "eq"
NeEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "ne"
LtEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "lt"
LteEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "lte"
GtEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "gt"
GteEvaluationRunCondition object { left, right, op }
left: unknown
right: unknown
op: optional "gte"
AndEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "and"
OrEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "or"
InEvaluationRunCondition object { left, operands, op }
left: unknown
operands: array of unknown
op: optional "in"
NotInEvaluationRunCondition object { left, operands, op }
left: unknown
operands: array of unknown
op: optional "not_in"
NotEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "not"
IsNullEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "is_null"
IsNotNullEvaluationRunCondition object { operands, op }
operands: array of unknown
op: optional "is_not_null"
system_prompt: optional string
AutoEvaluationAgentTaskRequestWithItemLocator object { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: array of string
data_fields: optional array of string
designated_to: optional object { config, agent_name } or object { config, agent_name } or object { config, agent_name } or object { config, agent_name }
One of the following:
ApeAgent object { config, agent_name }
config: object { model, temperature }
model: optional string
temperature: optional number
agent_name: optional "APEAgent"
IfAgent object { config, agent_name }
config: object { model }
model: optional string
agent_name: optional "IFAgent"
TruthfulnessAgent object { config, agent_name }
config: object { model }
model: optional string
agent_name: optional "TruthfulnessAgent"
BaseAgent object { config, agent_name }
config: object { model }
model: optional string
agent_name: optional "BaseAgent"
output_type: optional "text" or "integer" or "float" or "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: optional array of string or number or boolean
One of the following:
string
number
boolean
rubric_id: optional string
rubric_version: optional number
alias: optional string

Alias to title the results column. Defaults to the auto_evaluation_guided_decoding

task_type: optional "auto_evaluation.guided_decoding"
AutoEvaluationAgent object { configuration, alias, task_type }
configuration: AutoEvaluationAgentTaskRequestWithItemLocator { definition, name, output_rules, 6 more }
definition: string
name: string
output_rules: array of string
data_fields: optional array of string
designated_to: optional object { config, agent_name } or object { config, agent_name } or object { config, agent_name } or object { config, agent_name }
One of the following:
ApeAgent object { config, agent_name }
config: object { model, temperature }
model: optional string
temperature: optional number
agent_name: optional "APEAgent"
IfAgent object { config, agent_name }
config: object { model }
model: optional string
agent_name: optional "IFAgent"
TruthfulnessAgent object { config, agent_name }
config: object { model }
model: optional string
agent_name: optional "TruthfulnessAgent"
BaseAgent object { config, agent_name }
config: object { model }
model: optional string
agent_name: optional "BaseAgent"
output_type: optional "text" or "integer" or "float" or "boolean"
One of the following:
"text"
"integer"
"float"
"boolean"
output_values: optional array of string or number or boolean
One of the following:
string
number
boolean
rubric_id: optional string
rubric_version: optional number
alias: optional string

Alias to title the results column. Defaults to the auto_evaluation_agent

task_type: optional "auto_evaluation.agent"
ContributorEvaluationQuestion object { configuration, alias, task_type }
configuration: object { layout, question_id, prefill_from, 3 more }
layout: Container { children, direction }
children: array of Container { children, direction } or Component { data, label }

The children to be displayed within the container

One of the following:
Container object { children, direction }
Component object { data, label }

A pointer to the data in each evaluation item to be displayed within the component

label: optional string
direction: optional "row" or "column"

The axis that children are placed in the container. Based on CSS flex-direction (see: https://developer.mozilla.org/en-US/docs/Web/CSS/flex-direction)

One of the following:
"row"
"column"
question_id: string
prefill_from: optional string

Dataset column to prefill contributor question task result

minLength1
queue_id: optional string

The contributor annotation queue to include this task in. Defaults to default

maxLength100
required: optional boolean

Whether the question is required to be answered

rubric_id: optional string

ID of the rubric to use for scoring this evaluation question

alias: optional string

Alias to title the results column. Defaults to the contributor_evaluation_question

task_type: optional "contributor_evaluation.question"
CustomFunction object { configuration, alias, task_type }
configuration: object { function_source, arg_mapping, config_args, outputs }

Configuration for a custom Python function evaluation task.

function_source: string

Python function source code

maxLength10000
arg_mapping: optional map[string]

Mapping of function parameter names to item locators (e.g. item.field). Auto-derived from function signature if not provided.

config_args: optional map[unknown]

Literal argument values for function parameters, such as thresholds or RNG seeds. Serialized JSON must be at most 10000 characters.

outputs: optional array of object { path, alias }

Optional output paths to materialize as separate result columns. If omitted, the function return value is stored only under the task alias/data key.

path: string

Dot path in the custom function return value to materialize.

minLength1
alias: optional string

Result column alias. Defaults to path with dots replaced by underscores.

minLength1
alias: optional string

Alias to title the results column. Defaults to the function name.

task_type: optional "custom_function"
total: number

The total of items that match the query. This is greater than or equal to the number of items returned.

limit: optional number

The maximum number of items to return.

object: optional "list"

Filter Evaluations

curl https://api.egp.scale.com/v5/evaluations/filter \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "filters": [
            {
              "key": "key",
              "operator": "==",
              "value": "value"
            }
          ]
        }'
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
Returns Examples
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "datasets": [
        {
          "id": "id",
          "created_at": "2019-12-27T18:11:19.117Z",
          "created_by": {
            "id": "id",
            "type": "user",
            "object": "identity"
          },
          "current_version_num": 0,
          "name": "name",
          "tags": [
            "string"
          ],
          "archived_at": "2019-12-27T18:11:19.117Z",
          "description": "description",
          "object": "dataset"
        }
      ],
      "name": "name",
      "status": "failed",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "error_count": 0,
      "metadata": {
        "foo": "bar"
      },
      "object": "evaluation",
      "progress": {
        "items": {
          "failed": 0,
          "pending": 0,
          "successful": 0,
          "total": 0,
          "failed_items": [
            {
              "item_id": "item_id",
              "error": "error",
              "error_type": "error_type"
            }
          ]
        },
        "workflows": {
          "completed": 0,
          "failed": 0,
          "pending": 0,
          "total": 0
        }
      },
      "status_reason": "status_reason",
      "tasks": [
        {
          "configuration": {
            "messages": [
              {
                "foo": "bar"
              }
            ],
            "model": "model",
            "audio": {
              "foo": "bar"
            },
            "frequency_penalty": -2,
            "function_call": {
              "foo": "bar"
            },
            "functions": [
              {
                "foo": "bar"
              }
            ],
            "logit_bias": {
              "foo": 0
            },
            "logprobs": true,
            "max_completion_tokens": 0,
            "max_tokens": 0,
            "metadata": {
              "foo": "string"
            },
            "modalities": [
              "string"
            ],
            "n": 0,
            "parallel_tool_calls": true,
            "prediction": {
              "foo": "bar"
            },
            "presence_penalty": -2,
            "reasoning_effort": "reasoning_effort",
            "response_format": {
              "foo": "bar"
            },
            "seed": 0,
            "stop": "string",
            "store": true,
            "temperature": 0,
            "tool_choice": "string",
            "tools": [
              {
                "foo": "bar"
              }
            ],
            "top_k": 0,
            "top_logprobs": 0,
            "top_p": 0
          },
          "alias": "alias",
          "task_type": "chat_completion"
        }
      ]
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}