# Datasets

## Create a dataset

`datasets.create(DatasetCreateParams**kwargs)  -> Dataset`

**post** `/v5/datasets`

Create a dataset and populate it with its initial items in a single call.

Use this when the dataset does not exist yet: it creates the dataset and seeds it with the
items you provide. To add more items to a dataset that already exists, use
POST /v5/dataset-items/batch instead. A name is required, and the request must include at
least one item in `data`; the optional `files` field associates files with the created items.
The provided items are inserted as dataset items as a side effect of creation, and any `tags`
are attached to the dataset. The returned dataset is at version 1.

### Parameters

- `data: Iterable[Dict[str, object]]`

  Items to be included in the dataset

- `name: str`

- `description: Optional[str]`

- `files: Optional[Iterable[Dict[str, str]]]`

  Files to be associated to the dataset

- `tags: Optional[Sequence[str]]`

  The tags associated with the entity

### Returns

- `class Dataset: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `current_version_num: int`

  - `name: str`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `object: Optional[Literal["dataset"]]`

    - `"dataset"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
dataset = client.datasets.create(
    data=[{
        "foo": "bar"
    }],
    name="name",
)
print(dataset.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "current_version_num": 0,
  "name": "name",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "object": "dataset"
}
```

## List datasets

`datasets.list(DatasetListParams**kwargs)  -> SyncCursorPage[Dataset]`

**get** `/v5/datasets`

Return a paginated list of the account's datasets.

Archived datasets are excluded unless `include_archived` is set to true. Results can be
narrowed by `tags` and by the standard filterable parameters; a `name` filter matches as a
case-insensitive substring rather than an exact match.

### Parameters

- `ending_before: Optional[str]`

- `include_archived: Optional[bool]`

- `limit: Optional[int]`

- `name: Optional[str]`

- `sort_by: Optional[str]`

- `sort_order: Optional[SortOrder]`

  - `"asc"`

  - `"desc"`

- `starting_after: Optional[str]`

- `tags: Optional[Sequence[str]]`

### Returns

- `class Dataset: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `current_version_num: int`

  - `name: str`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `object: Optional[Literal["dataset"]]`

    - `"dataset"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
page = client.datasets.list()
page = page.items[0]
print(page.id)
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "created_by": {
        "id": "id",
        "type": "user",
        "object": "identity"
      },
      "current_version_num": 0,
      "name": "name",
      "tags": [
        "string"
      ],
      "archived_at": "2019-12-27T18:11:19.117Z",
      "description": "description",
      "object": "dataset"
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Archive a dataset

`datasets.archive(strdataset_id)  -> DatasetArchiveResponse`

**delete** `/v5/datasets/{dataset_id}`

Soft-delete a dataset by archiving it; the row is retained and can be restored later.

This sets the dataset's `archived_at` timestamp rather than permanently removing it, and
cascades the same archival to the dataset's items and versions. Archiving is not idempotent:
if the dataset is already archived the call fails. The dataset can be brought back with the
PATCH restore request. The response reports the dataset id with `deleted` set to true.

### Parameters

- `dataset_id: str`

### Returns

- `class DatasetArchiveResponse: …`

  - `id: str`

  - `deleted: bool`

  - `object: Optional[Literal["dataset"]]`

    - `"dataset"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.datasets.archive(
    "dataset_id",
)
print(response.id)
```

#### Response

```json
{
  "id": "id",
  "deleted": true,
  "object": "dataset"
}
```

## Update or restore a dataset

`datasets.update(strdataset_id, DatasetUpdateParams**kwargs)  -> Dataset`

**patch** `/v5/datasets/{dataset_id}`

Update a dataset's fields, or restore a previously archived dataset.

The request body is a union discriminated by its contents: a body of `{"restore": true}`
triggers a restore, and any other body is treated as a partial update of the dataset's `name`,
`description`, and `tags`. Restore clears `archived_at` on the dataset and cascades the
un-archival to its items and versions; restoring a dataset that is not archived returns it
unchanged. A partial update on an archived dataset fails, since archived datasets cannot be
modified.

### Parameters

- `dataset_id: str`

- `dataset: Dataset`

  - `class DatasetPartialDatasetRequestBase: …`

    - `description: Optional[str]`

    - `name: Optional[str]`

    - `tags: Optional[Sequence[str]]`

      The tags associated with the entity

  - `class RestoreRequest: …`

    - `restore: Literal[true]`

      Set to true to restore the entity from the database.

      - `true`

### Returns

- `class Dataset: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `current_version_num: int`

  - `name: str`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `object: Optional[Literal["dataset"]]`

    - `"dataset"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
dataset = client.datasets.update(
    dataset_id="dataset_id",
    dataset={},
)
print(dataset.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "current_version_num": 0,
  "name": "name",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "object": "dataset"
}
```

## Get a dataset

`datasets.retrieve(strdataset_id, DatasetRetrieveParams**kwargs)  -> Dataset`

**get** `/v5/datasets/{dataset_id}`

Retrieve a single dataset by its id.

By default an archived dataset is not returned; set `include_archived` to true to fetch a
dataset regardless of whether it has been archived.

### Parameters

- `dataset_id: str`

- `include_archived: Optional[bool]`

### Returns

- `class Dataset: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `current_version_num: int`

  - `name: str`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `object: Optional[Literal["dataset"]]`

    - `"dataset"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
dataset = client.datasets.retrieve(
    dataset_id="dataset_id",
)
print(dataset.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "created_by": {
    "id": "id",
    "type": "user",
    "object": "identity"
  },
  "current_version_num": 0,
  "name": "name",
  "tags": [
    "string"
  ],
  "archived_at": "2019-12-27T18:11:19.117Z",
  "description": "description",
  "object": "dataset"
}
```

## Domain Types

### Dataset

- `class Dataset: …`

  - `id: str`

    The unique identifier of the entity.

  - `created_at: datetime`

    The date and time when the entity was created in ISO format.

  - `created_by: Identity`

    The identity that created the entity.

    - `id: str`

    - `type: Literal["user", "service_account"]`

      - `"user"`

      - `"service_account"`

    - `object: Optional[Literal["identity"]]`

      - `"identity"`

  - `current_version_num: int`

  - `name: str`

  - `tags: Optional[List[str]]`

    The tags associated with the entity

  - `archived_at: Optional[datetime]`

    The date and time when the entity was archived in ISO format.

  - `description: Optional[str]`

  - `object: Optional[Literal["dataset"]]`

    - `"dataset"`

### Dataset Archive Response

- `class DatasetArchiveResponse: …`

  - `id: str`

  - `deleted: bool`

  - `object: Optional[Literal["dataset"]]`

    - `"dataset"`
