# Vector Stores

## List Vector Stores

`vector_stores.list(VectorStoreListParams**kwargs)  -> SyncCursorPageByName[VectorStore]`

**get** `/v5/vector-stores`

List all vector stores in your account with pagination.

Returns vector stores sorted by creation date (newest first). Each store includes its configuration,
embedding model, dimensions, indexed fields, and timestamps.

### Parameters

- `ending_before: Optional[str]`

- `limit: Optional[int]`

- `sort_by: Optional[str]`

- `sort_order: Optional[SortOrder]`

  - `"asc"`

  - `"desc"`

- `starting_after: Optional[str]`

### Returns

- `class VectorStore: …`

  Response model for vector store operations.

  - `id: str`

    The unique identifier of the vector store

  - `created_at: datetime`

    Timestamp of creation

  - `embedding_dimensions: int`

    Dimensionality of the embedding vectors

  - `name: str`

    The name of the vector store

  - `updated_at: datetime`

    Timestamp of last update

  - `embedding_config: Optional[EmbeddingConfig]`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `class EmbeddingConfigModelsAPI: …`

      - `model_deployment_id: str`

        The ID of the deployment of the created model in the Models API V3.

      - `type: Literal["models_api"]`

        The type of the embedding configuration.

        - `"models_api"`

    - `class EmbeddingConfigBase: …`

      - `embedding_model: Union[EmbeddingModelName, str]`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `str`

      - `type: Optional[Literal["base"]]`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: Optional[Dict[str, Literal["string", "number", "boolean"]]]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
page = client.vector_stores.list()
page = page.items[0]
print(page.id)
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "embedding_dimensions": 0,
      "name": "name",
      "updated_at": "2019-12-27T18:11:19.117Z",
      "embedding_config": {
        "model_deployment_id": "model_deployment_id",
        "type": "models_api"
      },
      "indexed_metadata_fields": {
        "foo": "string"
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Create Vector Store

`vector_stores.create(VectorStoreCreateParams**kwargs)  -> VectorStore`

**post** `/v5/vector-stores/create`

Create a new vector store for storing and querying document embeddings.

The vector store name must be unique within your account and follow naming conventions (3-63 characters,
alphanumeric with hyphens/underscores). Once created, the embedding configuration and dimensions
are immutable and cannot be changed. To use a different model, you must create a new vector store.

**Embedding Configuration:** Provide `embedding_config` (for base or custom model deployments),
`embedding_model` (shorthand for a base model), or `dimensions` only (raw embeddings).

- With `embedding_config` or `embedding_model`: dimensions are auto-derived, and documents can be
  upserted with text content (auto-embedded) or with pre-computed embeddings.
- With `dimensions` only: the store accepts only pre-computed embeddings. Semantic/hybrid queries
  are not supported (lexical search only).

**Indexed Fields:** Optionally specify metadata fields to index at creation time. Only indexed fields
can be used for filtering -- indexing is required, not just a performance optimization. Additional indexed
fields can be added later using the configure endpoint, but cannot be removed once added. Keep in mind
that each indexed field increases write latency and storage overhead, so only index fields you actively filter on.

### Parameters

- `name: str`

  A unique name for the vector store within the account

- `dimensions: Optional[int]`

  Dimension size of embedding vectors. Required when neither 'embedding_config' nor 'embedding_model' is set. Automatically derived when an embedding model is provided.

- `embedding_config: Optional[EmbeddingConfigParam]`

  The embedding configuration. Either 'base' type with an embedding_model, or 'models_api' type with a model_deployment_id for custom models.

  - `class EmbeddingConfigModelsAPI: …`

    - `model_deployment_id: str`

      The ID of the deployment of the created model in the Models API V3.

    - `type: Literal["models_api"]`

      The type of the embedding configuration.

      - `"models_api"`

  - `class EmbeddingConfigBase: …`

    - `embedding_model: Union[EmbeddingModelName, str]`

      The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

      - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

        - `"sentence-transformers/all-MiniLM-L12-v2"`

        - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

        - `"openai/text-embedding-ada-002"`

        - `"openai/text-embedding-3-small"`

        - `"openai/text-embedding-3-large"`

        - `"embed-english-v3.0"`

        - `"embed-english-light-v3.0"`

        - `"embed-multilingual-v3.0"`

        - `"gemini/text-embedding-005"`

        - `"gemini/text-multilingual-embedding-002"`

        - `"gemini/gemini-embedding-001"`

      - `str`

    - `type: Optional[Literal["base"]]`

      The type of the embedding configuration.

      - `"base"`

- `embedding_model: Optional[EmbeddingModelName]`

  The base embedding model to use. Shorthand for embedding_config with type 'base'. Provide either embedding_config or embedding_model, not both.

  - `"sentence-transformers/all-MiniLM-L12-v2"`

  - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

  - `"openai/text-embedding-ada-002"`

  - `"openai/text-embedding-3-small"`

  - `"openai/text-embedding-3-large"`

  - `"embed-english-v3.0"`

  - `"embed-english-light-v3.0"`

  - `"embed-multilingual-v3.0"`

  - `"gemini/text-embedding-005"`

  - `"gemini/text-multilingual-embedding-002"`

  - `"gemini/gemini-embedding-001"`

- `indexed_metadata_fields: Optional[Dict[str, Literal["string", "number", "boolean"]]]`

  Dictionary mapping metadata field names to their types for efficient filtering. Only STRING, NUMBER, and BOOLEAN types can be indexed.

  - `"string"`

  - `"number"`

  - `"boolean"`

### Returns

- `class VectorStore: …`

  Response model for vector store operations.

  - `id: str`

    The unique identifier of the vector store

  - `created_at: datetime`

    Timestamp of creation

  - `embedding_dimensions: int`

    Dimensionality of the embedding vectors

  - `name: str`

    The name of the vector store

  - `updated_at: datetime`

    Timestamp of last update

  - `embedding_config: Optional[EmbeddingConfig]`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `class EmbeddingConfigModelsAPI: …`

      - `model_deployment_id: str`

        The ID of the deployment of the created model in the Models API V3.

      - `type: Literal["models_api"]`

        The type of the embedding configuration.

        - `"models_api"`

    - `class EmbeddingConfigBase: …`

      - `embedding_model: Union[EmbeddingModelName, str]`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `str`

      - `type: Optional[Literal["base"]]`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: Optional[Dict[str, Literal["string", "number", "boolean"]]]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
vector_store = client.vector_stores.create(
    name="name",
)
print(vector_store.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "embedding_dimensions": 0,
  "name": "name",
  "updated_at": "2019-12-27T18:11:19.117Z",
  "embedding_config": {
    "model_deployment_id": "model_deployment_id",
    "type": "models_api"
  },
  "indexed_metadata_fields": {
    "foo": "string"
  }
}
```

## Get Vector Store

`vector_stores.retrieve(strvector_store_name)  -> VectorStore`

**get** `/v5/vector-stores/{vector_store_name}`

Retrieve detailed configuration and metadata for a specific vector store.

Returns the store's embedding model, dimensions, indexed metadata field definitions,
creation timestamp, and last update timestamp. Use this to verify store settings before
performing operations or to display store information in your application.

### Parameters

- `vector_store_name: str`

  The name of the vector store

### Returns

- `class VectorStore: …`

  Response model for vector store operations.

  - `id: str`

    The unique identifier of the vector store

  - `created_at: datetime`

    Timestamp of creation

  - `embedding_dimensions: int`

    Dimensionality of the embedding vectors

  - `name: str`

    The name of the vector store

  - `updated_at: datetime`

    Timestamp of last update

  - `embedding_config: Optional[EmbeddingConfig]`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `class EmbeddingConfigModelsAPI: …`

      - `model_deployment_id: str`

        The ID of the deployment of the created model in the Models API V3.

      - `type: Literal["models_api"]`

        The type of the embedding configuration.

        - `"models_api"`

    - `class EmbeddingConfigBase: …`

      - `embedding_model: Union[EmbeddingModelName, str]`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `str`

      - `type: Optional[Literal["base"]]`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: Optional[Dict[str, Literal["string", "number", "boolean"]]]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
vector_store = client.vector_stores.retrieve(
    "vector_store_name",
)
print(vector_store.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "embedding_dimensions": 0,
  "name": "name",
  "updated_at": "2019-12-27T18:11:19.117Z",
  "embedding_config": {
    "model_deployment_id": "model_deployment_id",
    "type": "models_api"
  },
  "indexed_metadata_fields": {
    "foo": "string"
  }
}
```

## Configure Vector Store

`vector_stores.configure(strvector_store_name, VectorStoreConfigureParams**kwargs)  -> VectorStore`

**post** `/v5/vector-stores/{vector_store_name}/configure`

Update the indexed metadata fields configuration for a vector store.

This replaces the current set of indexed metadata fields. Only indexed fields can be used for
filtering during query, list, and count operations; non-indexed fields are still stored and
returned, but cannot be filtered on.

**Field Types:** Only STRING, NUMBER, and BOOLEAN fields can be indexed (maximum 20 fields).
OBJECT and LIST types are stored but cannot be indexed for filtering.

**Adding Fields:** New indexed fields can be added at any time. They are indexed for documents
upserted after the change; to make existing documents filterable on a new field, re-upsert them.

**Removing Fields:** Omitting a field removes it from this configuration, so it can no longer be
filtered on. The underlying index is append-only, so removal does not reclaim storage or reduce
write overhead; the field stays in the physical index until the store is recreated. Prefer
indexing only the fields you filter on.

**Note:** The `name` and `embedding_config` are immutable after creation.

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `indexed_metadata_fields: Dict[str, Literal["string", "number", "boolean"]]`

  Dictionary mapping metadata field names to their types. Only STRING, NUMBER, and BOOLEAN types can be indexed.

  - `"string"`

  - `"number"`

  - `"boolean"`

### Returns

- `class VectorStore: …`

  Response model for vector store operations.

  - `id: str`

    The unique identifier of the vector store

  - `created_at: datetime`

    Timestamp of creation

  - `embedding_dimensions: int`

    Dimensionality of the embedding vectors

  - `name: str`

    The name of the vector store

  - `updated_at: datetime`

    Timestamp of last update

  - `embedding_config: Optional[EmbeddingConfig]`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `class EmbeddingConfigModelsAPI: …`

      - `model_deployment_id: str`

        The ID of the deployment of the created model in the Models API V3.

      - `type: Literal["models_api"]`

        The type of the embedding configuration.

        - `"models_api"`

    - `class EmbeddingConfigBase: …`

      - `embedding_model: Union[EmbeddingModelName, str]`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `str`

      - `type: Optional[Literal["base"]]`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: Optional[Dict[str, Literal["string", "number", "boolean"]]]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
vector_store = client.vector_stores.configure(
    vector_store_name="vector_store_name",
    indexed_metadata_fields={
        "foo": "string"
    },
)
print(vector_store.id)
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "embedding_dimensions": 0,
  "name": "name",
  "updated_at": "2019-12-27T18:11:19.117Z",
  "embedding_config": {
    "model_deployment_id": "model_deployment_id",
    "type": "models_api"
  },
  "indexed_metadata_fields": {
    "foo": "string"
  }
}
```

## Drop Vector Store

`vector_stores.drop(strvector_store_name)  -> VectorStoreDropResponse`

**post** `/v5/vector-stores/{vector_store_name}/drop`

Permanently delete a vector store and all its contents.

**⚠️ WARNING:** This is a destructive operation that cannot be undone. All documents, embeddings, metadata,
and index configurations will be permanently deleted. Data recovery is not possible after deletion.

### Parameters

- `vector_store_name: str`

  The name of the vector store

### Returns

- `class VectorStoreDropResponse: …`

  Response for vector store deletion.

  - `name: str`

    The name of the deleted vector store

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.vector_stores.drop(
    "vector_store_name",
)
print(response.name)
```

#### Response

```json
{
  "name": "name"
}
```

## Upsert Vectors

`vector_stores.upsert(strvector_store_name, VectorStoreUpsertParams**kwargs)  -> VectorStoreUpsertResponse`

**post** `/v5/vector-stores/{vector_store_name}/upsert`

Insert new documents or update existing documents in a vector store.

**Upsert Behavior:** If a document ID already exists, it will be completely replaced with the new content
and metadata. The previous document's text, embedding, and all metadata fields are discarded. If the ID
does not exist, a new document is created.

**Document Content:** Each document supports several modes:

- `content` only: text is automatically embedded using the store's configured model.
- `embedding` only: pre-computed embedding vector is used directly. Dimension must match the store's configuration.
- Both `content` and `embedding`: the pre-computed embedding is stored and text is kept for retrieval/search.
- Neither (metadata-only): only metadata is updated on an existing document without re-embedding.
  If the document does not exist, it will appear as a failure in the batch response.

A store created without an embedding model (dimensions-only) only accepts documents with pre-computed `embedding`.

**Batch Operations:** This endpoint supports batch operations with partial success handling and mixed
document types (some with raw embeddings, some with content) in the same call.

**Metadata:** Supports nested metadata with string, number, boolean, object, and array types. Null values
are not permitted—omit the field or use an empty string instead.

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `vectors: Iterable[Vector]`

  Array of documents to upsert

  - `id: str`

    Unique document ID

  - `content: Optional[TextContentParam]`

    Text content for documents.

    - `text: str`

      Text content to be embedded

    - `type: Optional[Literal["text"]]`

      Content type identifier

      - `"text"`

  - `embedding: Optional[Iterable[float]]`

    Pre-computed embedding vector

  - `metadata: Optional[Dict[str, object]]`

    Key-value metadata

### Returns

- `class VectorStoreUpsertResponse: …`

  Response for batch insert/upsert operations.

  - `failure_count: int`

    Number of failed documents

  - `success_count: int`

    Number of successfully processed documents

  - `failed: Optional[List[Failed]]`

    Failed documents with their error messages

    - `id: str`

      Document ID

    - `error: str`

      Error message describing why the document failed

  - `succeeded: Optional[List[str]]`

    IDs of successfully processed documents

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.vector_stores.upsert(
    vector_store_name="vector_store_name",
    vectors=[{
        "id": "id"
    }],
)
print(response.failure_count)
```

#### Response

```json
{
  "failure_count": 0,
  "success_count": 0,
  "failed": [
    {
      "id": "id",
      "error": "error"
    }
  ],
  "succeeded": [
    "string"
  ]
}
```

## Delete Vectors

`vector_stores.delete(strvector_store_name, VectorStoreDeleteParams**kwargs)  -> VectorStoreDeleteResponse`

**post** `/v5/vector-stores/{vector_store_name}/delete`

Delete documents from a vector store by document IDs or metadata filter criteria.

**Delete by IDs:** Provide an array of document IDs to delete specific documents. Non-existent documents
are silently skipped.

**Delete by Filter:** Use metadata filters to delete all documents matching the specified criteria (e.g.,
delete all documents where `status: "archived"`). The filter must specify at least one condition and cannot
be empty. To delete all documents, use the drop endpoint instead.

**Filter Operators:** Supports MongoDB-style operators including equality (`{"field": "value"}`),
comparison (`$gt`, `$gte`, `$lt`, `$lte`, `$eq`, `$ne`), logical (`$and`, `$or`, `$not`),
and membership (`$in`, `$nin`). Only indexed metadata fields can be used for filtering.

**Best Practice:** Use the count endpoint with the same filter to preview the number of documents that
will be deleted before executing the deletion operation.

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `filter: Optional[Dict[str, object]]`

  Metadata filter expression for deletion

- `ids: Optional[Sequence[str]]`

  Array of document IDs to delete

### Returns

- `class VectorStoreDeleteResponse: …`

  Response for delete operation.

  - `deleted_count: int`

    Number of documents deleted

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
vector_store = client.vector_stores.delete(
    vector_store_name="vector_store_name",
)
print(vector_store.deleted_count)
```

#### Response

```json
{
  "deleted_count": 0
}
```

## Count Vectors

`vector_stores.count(strvector_store_name, VectorStoreCountParams**kwargs)  -> VectorStoreCountResponse`

**post** `/v5/vector-stores/{vector_store_name}/count`

Count documents in a vector store, optionally filtered by metadata.

**Use Cases:**

- Monitor vector store size and growth over time
- Preview the number of documents matching a filter before deletion
- Validate data ingestion by comparing expected versus actual document counts
- Analyze document distribution across metadata categories

**Filtering:** Apply the same metadata filter syntax as delete and list operations. Only indexed fields
can be used for filtering. An empty filter counts all documents in the store.

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `filter: Optional[Dict[str, object]]`

  Metadata filter expression

### Returns

- `class VectorStoreCountResponse: …`

  Response for count operation.

  - `count: int`

    Number of documents matching the criteria

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.vector_stores.count(
    vector_store_name="vector_store_name",
)
print(response.count)
```

#### Response

```json
{
  "count": 0
}
```

## Query Vectors

`vector_stores.query(strvector_store_name, VectorStoreQueryParams**kwargs)  -> VectorStoreQueryResponse`

**post** `/v5/vector-stores/{vector_store_name}/query`

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation)
applications. Returns documents ranked by relevance to the query text with similarity scores.

**Query Types:**

- `semantic` (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search,
  and finding semantically related content without requiring exact keyword matches.
- `lexical`: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns,
  and scenarios where keyword presence is more important than semantic similarity.
- `hybrid`: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by
  identifying documents matching either semantically or lexically.

**Metadata Filtering:** Narrow the search scope by applying metadata filters (e.g., search only documents
where `category: "technical"`). Only indexed fields can be used for filtering.
Filters are applied before similarity search for optimal efficiency.

**Reranking (Advanced):** Optionally enhance result quality using a cross-encoder reranking model.
The reranker rescores the initial results using a more sophisticated model that evaluates the complete
query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves
precision for high-stakes applications.

**Reranking Strategy:** Set `top_k` higher than the desired final count (e.g., 50) to retrieve more
candidates from the initial search. Then configure `rerank_top_n` to the desired final count (e.g., 10)
to return only the most relevant documents after reranking. This two-stage approach maximizes both recall
and precision.

**Performance Metrics:** The response includes detailed timing breakdowns (embedding generation time,
index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

**Similarity Scores:** Each result includes a `score` field indicating relevance. Higher scores indicate
greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity,
lexical scores use BM25, hybrid scores combine both approaches).

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `content: TextContentParam`

  Text content for documents.

  - `text: str`

    Text content to be embedded

  - `type: Optional[Literal["text"]]`

    Content type identifier

    - `"text"`

- `filter: Optional[Dict[str, object]]`

  Metadata filter expression

- `include_vectors: Optional[bool]`

  Include embedding vectors in response

- `query_type: Optional[Literal["semantic", "lexical", "hybrid"]]`

  Query type: semantic, lexical, or hybrid

  - `"semantic"`

  - `"lexical"`

  - `"hybrid"`

- `rerank: Optional[bool]`

  [Deprecated: use rerank_config] Enable reranking of search results

- `rerank_config: Optional[RerankConfig]`

  Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

  - `instruction: Optional[str]`

    Custom instruction for the reranking model (e.g., 'Given a medical question, retrieve relevant clinical passages'). Only applies to instruction-following rerankers like Qwen3.

  - `model: Optional[str]`

    Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. 'cross-encoder/ms-marco-MiniLM-L-12-v2'), Vertex semantic-ranker names (e.g. 'semantic-ranker-default-004') when provider='vertex', or any model id the inference proxy serves when provider='proxy'.

  - `provider: Optional[Literal["launch", "vertex", "proxy"]]`

    Reranking provider to use. When omitted, the deployment default is used ('launch', or 'proxy' on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. 'vertex') to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

    - `"launch"`

    - `"vertex"`

    - `"proxy"`

  - `top_n: Optional[int]`

    Number of results to keep after reranking (defaults to top_k)

  - `type: Optional[Literal["base"]]`

    Reranking configuration type. Currently only 'base' is supported.

    - `"base"`

- `rerank_instruction: Optional[str]`

  [Deprecated: use rerank_config.instruction] Custom instruction for reranker

- `rerank_model: Optional[str]`

  [Deprecated: use rerank_config.model] Reranking model to use

- `rerank_top_n: Optional[int]`

  [Deprecated: use rerank_config.top_n] Number of results after reranking

- `top_k: Optional[int]`

  Number of search results to return

### Returns

- `class VectorStoreQueryResponse: …`

  Response for query operation.

  - `metadata: Metadata`

    Query execution metadata

    - `search_type: str`

      Type of search performed (semantic, lexical, hybrid)

    - `total_query_time_ms: int`

      Total end-to-end query execution time in milliseconds

    - `embedding_config: Optional[EmbeddingConfig]`

      Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

      - `class EmbeddingConfigModelsAPI: …`

        - `model_deployment_id: str`

          The ID of the deployment of the created model in the Models API V3.

        - `type: Literal["models_api"]`

          The type of the embedding configuration.

          - `"models_api"`

      - `class EmbeddingConfigBase: …`

        - `embedding_model: Union[EmbeddingModelName, str]`

          The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

          - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

            - `"sentence-transformers/all-MiniLM-L12-v2"`

            - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

            - `"openai/text-embedding-ada-002"`

            - `"openai/text-embedding-3-small"`

            - `"openai/text-embedding-3-large"`

            - `"embed-english-v3.0"`

            - `"embed-english-light-v3.0"`

            - `"embed-multilingual-v3.0"`

            - `"gemini/text-embedding-005"`

            - `"gemini/text-multilingual-embedding-002"`

            - `"gemini/gemini-embedding-001"`

          - `str`

        - `type: Optional[Literal["base"]]`

          The type of the embedding configuration.

          - `"base"`

    - `embedding_time_ms: Optional[int]`

      Time spent generating embeddings in milliseconds (None for lexical queries)

    - `index_query_time_ms: Optional[int]`

      Time spent querying the vector index (OpenSearch) in milliseconds

    - `reranking_model: Optional[str]`

      Reranking model used (None if reranking not enabled)

    - `reranking_time_ms: Optional[int]`

      Time spent reranking results in milliseconds (None if reranking not enabled)

  - `vectors: List[Vector]`

    Array of matching documents

    - `id: str`

      Document ID

    - `score: float`

      Similarity score indicating relevance

    - `content: Optional[TextContent]`

      Text content for documents.

      - `text: str`

        Text content to be embedded

      - `type: Optional[Literal["text"]]`

        Content type identifier

        - `"text"`

    - `metadata: Optional[Dict[str, object]]`

      Key-value metadata

    - `vector: Optional[List[float]]`

      Embedding vector (if requested)

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.vector_stores.query(
    vector_store_name="vector_store_name",
    content={
        "text": "text"
    },
)
print(response.metadata)
```

#### Response

```json
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
```

## Domain Types

### Embedding Config

- `EmbeddingConfig`

  - `class EmbeddingConfigModelsAPI: …`

    - `model_deployment_id: str`

      The ID of the deployment of the created model in the Models API V3.

    - `type: Literal["models_api"]`

      The type of the embedding configuration.

      - `"models_api"`

  - `class EmbeddingConfigBase: …`

    - `embedding_model: Union[EmbeddingModelName, str]`

      The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

      - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

        - `"sentence-transformers/all-MiniLM-L12-v2"`

        - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

        - `"openai/text-embedding-ada-002"`

        - `"openai/text-embedding-3-small"`

        - `"openai/text-embedding-3-large"`

        - `"embed-english-v3.0"`

        - `"embed-english-light-v3.0"`

        - `"embed-multilingual-v3.0"`

        - `"gemini/text-embedding-005"`

        - `"gemini/text-multilingual-embedding-002"`

        - `"gemini/gemini-embedding-001"`

      - `str`

    - `type: Optional[Literal["base"]]`

      The type of the embedding configuration.

      - `"base"`

### Embedding Config Base

- `class EmbeddingConfigBase: …`

  - `embedding_model: Union[EmbeddingModelName, str]`

    The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

    - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

      - `"sentence-transformers/all-MiniLM-L12-v2"`

      - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

      - `"openai/text-embedding-ada-002"`

      - `"openai/text-embedding-3-small"`

      - `"openai/text-embedding-3-large"`

      - `"embed-english-v3.0"`

      - `"embed-english-light-v3.0"`

      - `"embed-multilingual-v3.0"`

      - `"gemini/text-embedding-005"`

      - `"gemini/text-multilingual-embedding-002"`

      - `"gemini/gemini-embedding-001"`

    - `str`

  - `type: Optional[Literal["base"]]`

    The type of the embedding configuration.

    - `"base"`

### Embedding Config Models API

- `class EmbeddingConfigModelsAPI: …`

  - `model_deployment_id: str`

    The ID of the deployment of the created model in the Models API V3.

  - `type: Literal["models_api"]`

    The type of the embedding configuration.

    - `"models_api"`

### Embedding Model Name

- `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

  - `"sentence-transformers/all-MiniLM-L12-v2"`

  - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

  - `"openai/text-embedding-ada-002"`

  - `"openai/text-embedding-3-small"`

  - `"openai/text-embedding-3-large"`

  - `"embed-english-v3.0"`

  - `"embed-english-light-v3.0"`

  - `"embed-multilingual-v3.0"`

  - `"gemini/text-embedding-005"`

  - `"gemini/text-multilingual-embedding-002"`

  - `"gemini/gemini-embedding-001"`

### Text Content

- `class TextContent: …`

  Text content for documents.

  - `text: str`

    Text content to be embedded

  - `type: Optional[Literal["text"]]`

    Content type identifier

    - `"text"`

### Vector Store

- `class VectorStore: …`

  Response model for vector store operations.

  - `id: str`

    The unique identifier of the vector store

  - `created_at: datetime`

    Timestamp of creation

  - `embedding_dimensions: int`

    Dimensionality of the embedding vectors

  - `name: str`

    The name of the vector store

  - `updated_at: datetime`

    Timestamp of last update

  - `embedding_config: Optional[EmbeddingConfig]`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `class EmbeddingConfigModelsAPI: …`

      - `model_deployment_id: str`

        The ID of the deployment of the created model in the Models API V3.

      - `type: Literal["models_api"]`

        The type of the embedding configuration.

        - `"models_api"`

    - `class EmbeddingConfigBase: …`

      - `embedding_model: Union[EmbeddingModelName, str]`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `str`

      - `type: Optional[Literal["base"]]`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: Optional[Dict[str, Literal["string", "number", "boolean"]]]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Vector Store Drop Response

- `class VectorStoreDropResponse: …`

  Response for vector store deletion.

  - `name: str`

    The name of the deleted vector store

### Vector Store Upsert Response

- `class VectorStoreUpsertResponse: …`

  Response for batch insert/upsert operations.

  - `failure_count: int`

    Number of failed documents

  - `success_count: int`

    Number of successfully processed documents

  - `failed: Optional[List[Failed]]`

    Failed documents with their error messages

    - `id: str`

      Document ID

    - `error: str`

      Error message describing why the document failed

  - `succeeded: Optional[List[str]]`

    IDs of successfully processed documents

### Vector Store Delete Response

- `class VectorStoreDeleteResponse: …`

  Response for delete operation.

  - `deleted_count: int`

    Number of documents deleted

### Vector Store Count Response

- `class VectorStoreCountResponse: …`

  Response for count operation.

  - `count: int`

    Number of documents matching the criteria

### Vector Store Query Response

- `class VectorStoreQueryResponse: …`

  Response for query operation.

  - `metadata: Metadata`

    Query execution metadata

    - `search_type: str`

      Type of search performed (semantic, lexical, hybrid)

    - `total_query_time_ms: int`

      Total end-to-end query execution time in milliseconds

    - `embedding_config: Optional[EmbeddingConfig]`

      Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

      - `class EmbeddingConfigModelsAPI: …`

        - `model_deployment_id: str`

          The ID of the deployment of the created model in the Models API V3.

        - `type: Literal["models_api"]`

          The type of the embedding configuration.

          - `"models_api"`

      - `class EmbeddingConfigBase: …`

        - `embedding_model: Union[EmbeddingModelName, str]`

          The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

          - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

            - `"sentence-transformers/all-MiniLM-L12-v2"`

            - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

            - `"openai/text-embedding-ada-002"`

            - `"openai/text-embedding-3-small"`

            - `"openai/text-embedding-3-large"`

            - `"embed-english-v3.0"`

            - `"embed-english-light-v3.0"`

            - `"embed-multilingual-v3.0"`

            - `"gemini/text-embedding-005"`

            - `"gemini/text-multilingual-embedding-002"`

            - `"gemini/gemini-embedding-001"`

          - `str`

        - `type: Optional[Literal["base"]]`

          The type of the embedding configuration.

          - `"base"`

    - `embedding_time_ms: Optional[int]`

      Time spent generating embeddings in milliseconds (None for lexical queries)

    - `index_query_time_ms: Optional[int]`

      Time spent querying the vector index (OpenSearch) in milliseconds

    - `reranking_model: Optional[str]`

      Reranking model used (None if reranking not enabled)

    - `reranking_time_ms: Optional[int]`

      Time spent reranking results in milliseconds (None if reranking not enabled)

  - `vectors: List[Vector]`

    Array of matching documents

    - `id: str`

      Document ID

    - `score: float`

      Similarity score indicating relevance

    - `content: Optional[TextContent]`

      Text content for documents.

      - `text: str`

        Text content to be embedded

      - `type: Optional[Literal["text"]]`

        Content type identifier

        - `"text"`

    - `metadata: Optional[Dict[str, object]]`

      Key-value metadata

    - `vector: Optional[List[float]]`

      Embedding vector (if requested)

# Vectors

## List Vectors

`vector_stores.vectors.list(strvector_store_name, VectorListParams**kwargs)  -> SyncCursorPageVectors[VectorDocument]`

**get** `/v5/vector-stores/{vector_store_name}/vectors`

List documents in a vector store with cursor-based pagination.

**Use Cases:** Browse documents, export content, audit stored data, or retrieve documents by metadata
without semantic search.

**Ordering:** Documents are returned in storage order (insertion order), not ranked by similarity.
For similarity-based retrieval, use the query endpoint.

**Filtering:** Apply metadata filters to narrow results to specific subsets (e.g., all documents
where `category: "research"`). Only indexed fields can be used for filtering.

**Pagination:** Uses cursor-based pagination for efficient traversal of large datasets. Pass the
`next_cursor` from each response as `starting_after` to retrieve the next page, or `prev_cursor`
as `ending_before` to retrieve the previous page. A null cursor indicates no further pages exist.

**Embedding Vectors:** Setting `include_vectors=true` includes the full embedding vector arrays in
the response. This significantly increases payload size and reduces the maximum page size from 1000 to
100 documents. Enable only when raw vectors are required for external processing.

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `cursor: Optional[str]`

  Alias for starting_after. Use starting_after instead.

- `ending_before: Optional[str]`

- `filter: Optional[str]`

  Metadata filter expression (JSON)

- `include_vectors: Optional[bool]`

  Include embedding vectors

- `limit: Optional[int]`

- `sort_by: Optional[str]`

- `sort_order: Optional[SortOrder]`

  - `"asc"`

  - `"desc"`

- `starting_after: Optional[str]`

### Returns

- `class VectorDocument: …`

  A document returned from direct lookups (get/list operations).

  - `id: str`

    Document ID

  - `content: Optional[TextContent]`

    Text content for documents.

    - `text: str`

      Text content to be embedded

    - `type: Optional[Literal["text"]]`

      Content type identifier

      - `"text"`

  - `metadata: Optional[Dict[str, object]]`

    Key-value metadata

  - `vector: Optional[List[float]]`

    Embedding vector (if requested)

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
page = client.vector_stores.vectors.list(
    vector_store_name="vector_store_name",
)
page = page.vectors[0]
print(page.id)
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ],
  "total": 0,
  "vectors": [
    {
      "id": "id",
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ],
  "limit": 0,
  "next_cursor": "next_cursor",
  "object": "list",
  "prev_cursor": "prev_cursor"
}
```

## Get Vector

`vector_stores.vectors.retrieve(strvector_id, VectorRetrieveParams**kwargs)  -> VectorDocument`

**get** `/v5/vector-stores/{vector_store_name}/vectors/{vector_id}`

Retrieve a single document by its unique ID.

Returns the document's full content, metadata, and optionally its embedding vector. Use this endpoint
for direct lookups when the exact document ID is known. For content similarity search,
use the query endpoint.

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `vector_id: str`

  The ID of the vector to retrieve

- `include_vectors: Optional[bool]`

  Include embedding vectors

### Returns

- `class VectorDocument: …`

  A document returned from direct lookups (get/list operations).

  - `id: str`

    Document ID

  - `content: Optional[TextContent]`

    Text content for documents.

    - `text: str`

      Text content to be embedded

    - `type: Optional[Literal["text"]]`

      Content type identifier

      - `"text"`

  - `metadata: Optional[Dict[str, object]]`

    Key-value metadata

  - `vector: Optional[List[float]]`

    Embedding vector (if requested)

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
vector_document = client.vector_stores.vectors.retrieve(
    vector_id="vector_id",
    vector_store_name="vector_store_name",
)
print(vector_document.id)
```

#### Response

```json
{
  "id": "id",
  "content": {
    "text": "text",
    "type": "text"
  },
  "metadata": {
    "foo": "bar"
  },
  "vector": [
    0
  ]
}
```

## Domain Types

### Vector Document

- `class VectorDocument: …`

  A document returned from direct lookups (get/list operations).

  - `id: str`

    Document ID

  - `content: Optional[TextContent]`

    Text content for documents.

    - `text: str`

      Text content to be embedded

    - `type: Optional[Literal["text"]]`

      Content type identifier

      - `"text"`

  - `metadata: Optional[Dict[str, object]]`

    Key-value metadata

  - `vector: Optional[List[float]]`

    Embedding vector (if requested)
