# Datasets

## Create a dataset

**post** `/v5/datasets`

Create a dataset and populate it with its initial items in a single call.

Use this when the dataset does not exist yet: it creates the dataset and seeds it with the
items you provide. To add more items to a dataset that already exists, use
POST /v5/dataset-items/batch instead. A name is required, and the request must include at
least one item in `data`; the optional `files` field associates files with the created items.
The provided items are inserted as dataset items as a side effect of creation, and any `tags`
are attached to the dataset. The returned dataset is at version 1.

### Body Parameters

- `data: array of map[unknown]`

  Items to be included in the dataset

- `name: string`

- `description: optional string`

- `files: optional array of map[string]`

  Files to be associated to the dataset

- `tags: optional array of string`

  The tags associated with the entity

### Returns

- `Dataset object { id, created_at, created_by, 6 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: string`

    - `type: "user" or "service_account"`

      - `"user"`

      - `"service_account"`

    - `object: optional "identity"`

      - `"identity"`

  - `current_version_num: number`

  - `name: string`

  - `tags: array of string`

    The tags associated with the entity

  - `archived_at: optional string`

    The date and time when the entity was archived in ISO format.

  - `description: optional string`

  - `object: optional "dataset"`

    - `"dataset"`

### Example

```http
curl https://api.egp.scale.com/v5/datasets \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "data": [
            {
              "foo": "bar"
            }
          ],
          "name": "name"
        }'
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "current_version_num": 0,
  "name": "name",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "object": "dataset"
}
```

## List datasets

**get** `/v5/datasets`

Return a paginated list of the account's datasets.

Archived datasets are excluded unless `include_archived` is set to true. Results can be
narrowed by `tags` and by the standard filterable parameters; a `name` filter matches as a
case-insensitive substring rather than an exact match.

### Query Parameters

- `ending_before: optional string`

- `include_archived: optional boolean`

- `limit: optional number`

- `name: optional string`

- `sort_by: optional string`

- `sort_order: optional SortOrder`

  - `"asc"`

  - `"desc"`

- `starting_after: optional string`

- `tags: optional array of string`

### Returns

- `has_more: boolean`

  Whether there are more items left to be fetched.

- `items: array of Dataset`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: string`

    - `type: "user" or "service_account"`

      - `"user"`

      - `"service_account"`

    - `object: optional "identity"`

      - `"identity"`

  - `current_version_num: number`

  - `name: string`

  - `tags: array of string`

    The tags associated with the entity

  - `archived_at: optional string`

    The date and time when the entity was archived in ISO format.

  - `description: optional string`

  - `object: optional "dataset"`

    - `"dataset"`

- `total: number`

  The total of items that match the query. This is greater than or equal to the number of items returned.

- `limit: optional number`

  The maximum number of items to return.

- `object: optional "list"`

  - `"list"`

### Example

```http
curl https://api.egp.scale.com/v5/datasets \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Archive a dataset

**delete** `/v5/datasets/{dataset_id}`

Soft-delete a dataset by archiving it; the row is retained and can be restored later.

This sets the dataset's `archived_at` timestamp rather than permanently removing it, and
cascades the same archival to the dataset's items and versions. Archiving is not idempotent:
if the dataset is already archived the call fails. The dataset can be brought back with the
PATCH restore request. The response reports the dataset id with `deleted` set to true.

### Path Parameters

- `dataset_id: string`

### Returns

- `id: string`

- `deleted: boolean`

- `object: optional "dataset"`

  - `"dataset"`

### Example

```http
curl https://api.egp.scale.com/v5/datasets/$DATASET_ID \
    -X DELETE \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "id": "id",
  "deleted": true,
  "object": "dataset"
}
```

## Update or restore a dataset

**patch** `/v5/datasets/{dataset_id}`

Update a dataset's fields, or restore a previously archived dataset.

The request body is a union discriminated by its contents: a body of `{"restore": true}`
triggers a restore, and any other body is treated as a partial update of the dataset's `name`,
`description`, and `tags`. Restore clears `archived_at` on the dataset and cascades the
un-archival to its items and versions; restoring a dataset that is not archived returns it
unchanged. A partial update on an archived dataset fails, since archived datasets cannot be
modified.

### Path Parameters

- `dataset_id: string`

### Body Parameters

- `dataset: object { description, name, tags }  or RestoreRequest`

  - `PartialDatasetRequestBase object { description, name, tags }`

    - `description: optional string`

    - `name: optional string`

    - `tags: optional array of string`

      The tags associated with the entity

  - `RestoreRequest object { restore }`

    - `restore: true`

      Set to true to restore the entity from the database.

      - `true`

### Returns

- `Dataset object { id, created_at, created_by, 6 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: string`

    - `type: "user" or "service_account"`

      - `"user"`

      - `"service_account"`

    - `object: optional "identity"`

      - `"identity"`

  - `current_version_num: number`

  - `name: string`

  - `tags: array of string`

    The tags associated with the entity

  - `archived_at: optional string`

    The date and time when the entity was archived in ISO format.

  - `description: optional string`

  - `object: optional "dataset"`

    - `"dataset"`

### Example

```http
curl https://api.egp.scale.com/v5/datasets/$DATASET_ID \
    -X PATCH \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "description": "description",
          "name": "name",
          "tags": [
            "string"
          ]
        }'
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "current_version_num": 0,
  "name": "name",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "object": "dataset"
}
```

## Get a dataset

**get** `/v5/datasets/{dataset_id}`

Retrieve a single dataset by its id.

By default an archived dataset is not returned; set `include_archived` to true to fetch a
dataset regardless of whether it has been archived.

### Path Parameters

- `dataset_id: string`

### Query Parameters

- `include_archived: optional boolean`

### Returns

- `Dataset object { id, created_at, created_by, 6 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: string`

    - `type: "user" or "service_account"`

      - `"user"`

      - `"service_account"`

    - `object: optional "identity"`

      - `"identity"`

  - `current_version_num: number`

  - `name: string`

  - `tags: array of string`

    The tags associated with the entity

  - `archived_at: optional string`

    The date and time when the entity was archived in ISO format.

  - `description: optional string`

  - `object: optional "dataset"`

    - `"dataset"`

### Example

```http
curl https://api.egp.scale.com/v5/datasets/$DATASET_ID \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "current_version_num": 0,
  "name": "name",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "object": "dataset"
}
```

## Domain Types

### Dataset

- `Dataset object { id, created_at, created_by, 6 more }`

  - `id: string`

    The unique identifier of the entity.

  - `created_at: string`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: string`

    - `type: "user" or "service_account"`

      - `"user"`

      - `"service_account"`

    - `object: optional "identity"`

      - `"identity"`

  - `current_version_num: number`

  - `name: string`

  - `tags: array of string`

    The tags associated with the entity

  - `archived_at: optional string`

    The date and time when the entity was archived in ISO format.

  - `description: optional string`

  - `object: optional "dataset"`

    - `"dataset"`

### Dataset Archive Response

- `DatasetArchiveResponse object { id, deleted, object }`

  - `id: string`

  - `deleted: boolean`

  - `object: optional "dataset"`

    - `"dataset"`
