# Vector Stores

## List Vector Stores

**get** `/v5/vector-stores`

List all vector stores in your account with pagination.

Returns vector stores sorted by creation date (newest first). Each store includes its configuration,
embedding model, dimensions, indexed fields, and timestamps.

### Query Parameters

- `ending_before: optional string`

- `limit: optional number`

- `sort_by: optional string`

- `sort_order: optional SortOrder`

  - `"asc"`

  - `"desc"`

- `starting_after: optional string`

### Returns

- `has_more: boolean`

  Whether there are more items left to be fetched.

- `items: array of VectorStore`

  - `id: string`

    The unique identifier of the vector store

  - `created_at: string`

    Timestamp of creation

  - `embedding_dimensions: number`

    Dimensionality of the embedding vectors

  - `name: string`

    The name of the vector store

  - `updated_at: string`

    Timestamp of last update

  - `embedding_config: optional EmbeddingConfig`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

      - `model_deployment_id: string`

        The ID of the deployment of the created model in the Models API V3.

      - `type: "models_api"`

        The type of the embedding configuration.

        - `"models_api"`

    - `EmbeddingConfigBase object { embedding_model, type }`

      - `embedding_model: EmbeddingModelName or string`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `string`

      - `type: optional "base"`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: optional map["string" or "number" or "boolean"]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

- `total: number`

  The total of items that match the query. This is greater than or equal to the number of items returned.

- `limit: optional number`

  The maximum number of items to return.

- `object: optional "list"`

  - `"list"`

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "created_at": "2019-12-27T18:11:19.117Z",
      "embedding_dimensions": 0,
      "name": "name",
      "updated_at": "2019-12-27T18:11:19.117Z",
      "embedding_config": {
        "model_deployment_id": "model_deployment_id",
        "type": "models_api"
      },
      "indexed_metadata_fields": {
        "foo": "string"
      }
    }
  ],
  "total": 0,
  "limit": 0,
  "object": "list"
}
```

## Create Vector Store

**post** `/v5/vector-stores/create`

Create a new vector store for storing and querying document embeddings.

The vector store name must be unique within your account and follow naming conventions (3-63 characters,
alphanumeric with hyphens/underscores). Once created, the embedding configuration and dimensions
are immutable and cannot be changed. To use a different model, you must create a new vector store.

**Embedding Configuration:** Provide `embedding_config` (for base or custom model deployments),
`embedding_model` (shorthand for a base model), or `dimensions` only (raw embeddings).

- With `embedding_config` or `embedding_model`: dimensions are auto-derived, and documents can be
  upserted with text content (auto-embedded) or with pre-computed embeddings.
- With `dimensions` only: the store accepts only pre-computed embeddings. Semantic/hybrid queries
  are not supported (lexical search only).

**Indexed Fields:** Optionally specify metadata fields to index at creation time. Only indexed fields
can be used for filtering -- indexing is required, not just a performance optimization. Additional indexed
fields can be added later using the configure endpoint, but cannot be removed once added. Keep in mind
that each indexed field increases write latency and storage overhead, so only index fields you actively filter on.

### Body Parameters

- `name: string`

  A unique name for the vector store within the account

- `dimensions: optional number`

  Dimension size of embedding vectors. Required when neither 'embedding_config' nor 'embedding_model' is set. Automatically derived when an embedding model is provided.

- `embedding_config: optional EmbeddingConfig`

  The embedding configuration. Either 'base' type with an embedding_model, or 'models_api' type with a model_deployment_id for custom models.

  - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

    - `model_deployment_id: string`

      The ID of the deployment of the created model in the Models API V3.

    - `type: "models_api"`

      The type of the embedding configuration.

      - `"models_api"`

  - `EmbeddingConfigBase object { embedding_model, type }`

    - `embedding_model: EmbeddingModelName or string`

      The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

      - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

        - `"sentence-transformers/all-MiniLM-L12-v2"`

        - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

        - `"openai/text-embedding-ada-002"`

        - `"openai/text-embedding-3-small"`

        - `"openai/text-embedding-3-large"`

        - `"embed-english-v3.0"`

        - `"embed-english-light-v3.0"`

        - `"embed-multilingual-v3.0"`

        - `"gemini/text-embedding-005"`

        - `"gemini/text-multilingual-embedding-002"`

        - `"gemini/gemini-embedding-001"`

      - `string`

    - `type: optional "base"`

      The type of the embedding configuration.

      - `"base"`

- `embedding_model: optional EmbeddingModelName`

  The base embedding model to use. Shorthand for embedding_config with type 'base'. Provide either embedding_config or embedding_model, not both.

- `indexed_metadata_fields: optional map["string" or "number" or "boolean"]`

  Dictionary mapping metadata field names to their types for efficient filtering. Only STRING, NUMBER, and BOOLEAN types can be indexed.

  - `"string"`

  - `"number"`

  - `"boolean"`

### Returns

- `VectorStore object { id, created_at, embedding_dimensions, 4 more }`

  Response model for vector store operations.

  - `id: string`

    The unique identifier of the vector store

  - `created_at: string`

    Timestamp of creation

  - `embedding_dimensions: number`

    Dimensionality of the embedding vectors

  - `name: string`

    The name of the vector store

  - `updated_at: string`

    Timestamp of last update

  - `embedding_config: optional EmbeddingConfig`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

      - `model_deployment_id: string`

        The ID of the deployment of the created model in the Models API V3.

      - `type: "models_api"`

        The type of the embedding configuration.

        - `"models_api"`

    - `EmbeddingConfigBase object { embedding_model, type }`

      - `embedding_model: EmbeddingModelName or string`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `string`

      - `type: optional "base"`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: optional map["string" or "number" or "boolean"]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/create \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "name": "name"
        }'
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "embedding_dimensions": 0,
  "name": "name",
  "updated_at": "2019-12-27T18:11:19.117Z",
  "embedding_config": {
    "model_deployment_id": "model_deployment_id",
    "type": "models_api"
  },
  "indexed_metadata_fields": {
    "foo": "string"
  }
}
```

## Get Vector Store

**get** `/v5/vector-stores/{vector_store_name}`

Retrieve detailed configuration and metadata for a specific vector store.

Returns the store's embedding model, dimensions, indexed metadata field definitions,
creation timestamp, and last update timestamp. Use this to verify store settings before
performing operations or to display store information in your application.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Returns

- `VectorStore object { id, created_at, embedding_dimensions, 4 more }`

  Response model for vector store operations.

  - `id: string`

    The unique identifier of the vector store

  - `created_at: string`

    Timestamp of creation

  - `embedding_dimensions: number`

    Dimensionality of the embedding vectors

  - `name: string`

    The name of the vector store

  - `updated_at: string`

    Timestamp of last update

  - `embedding_config: optional EmbeddingConfig`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

      - `model_deployment_id: string`

        The ID of the deployment of the created model in the Models API V3.

      - `type: "models_api"`

        The type of the embedding configuration.

        - `"models_api"`

    - `EmbeddingConfigBase object { embedding_model, type }`

      - `embedding_model: EmbeddingModelName or string`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `string`

      - `type: optional "base"`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: optional map["string" or "number" or "boolean"]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "embedding_dimensions": 0,
  "name": "name",
  "updated_at": "2019-12-27T18:11:19.117Z",
  "embedding_config": {
    "model_deployment_id": "model_deployment_id",
    "type": "models_api"
  },
  "indexed_metadata_fields": {
    "foo": "string"
  }
}
```

## Configure Vector Store

**post** `/v5/vector-stores/{vector_store_name}/configure`

Update the indexed metadata fields configuration for a vector store.

This replaces the current set of indexed metadata fields. Only indexed fields can be used for
filtering during query, list, and count operations; non-indexed fields are still stored and
returned, but cannot be filtered on.

**Field Types:** Only STRING, NUMBER, and BOOLEAN fields can be indexed (maximum 20 fields).
OBJECT and LIST types are stored but cannot be indexed for filtering.

**Adding Fields:** New indexed fields can be added at any time. They are indexed for documents
upserted after the change; to make existing documents filterable on a new field, re-upsert them.

**Removing Fields:** Omitting a field removes it from this configuration, so it can no longer be
filtered on. The underlying index is append-only, so removal does not reclaim storage or reduce
write overhead; the field stays in the physical index until the store is recreated. Prefer
indexing only the fields you filter on.

**Note:** The `name` and `embedding_config` are immutable after creation.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Body Parameters

- `indexed_metadata_fields: map["string" or "number" or "boolean"]`

  Dictionary mapping metadata field names to their types. Only STRING, NUMBER, and BOOLEAN types can be indexed.

  - `"string"`

  - `"number"`

  - `"boolean"`

### Returns

- `VectorStore object { id, created_at, embedding_dimensions, 4 more }`

  Response model for vector store operations.

  - `id: string`

    The unique identifier of the vector store

  - `created_at: string`

    Timestamp of creation

  - `embedding_dimensions: number`

    Dimensionality of the embedding vectors

  - `name: string`

    The name of the vector store

  - `updated_at: string`

    Timestamp of last update

  - `embedding_config: optional EmbeddingConfig`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

      - `model_deployment_id: string`

        The ID of the deployment of the created model in the Models API V3.

      - `type: "models_api"`

        The type of the embedding configuration.

        - `"models_api"`

    - `EmbeddingConfigBase object { embedding_model, type }`

      - `embedding_model: EmbeddingModelName or string`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `string`

      - `type: optional "base"`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: optional map["string" or "number" or "boolean"]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/configure \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "indexed_metadata_fields": {
            "foo": "string"
          }
        }'
```

#### Response

```json
{
  "id": "id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "embedding_dimensions": 0,
  "name": "name",
  "updated_at": "2019-12-27T18:11:19.117Z",
  "embedding_config": {
    "model_deployment_id": "model_deployment_id",
    "type": "models_api"
  },
  "indexed_metadata_fields": {
    "foo": "string"
  }
}
```

## Drop Vector Store

**post** `/v5/vector-stores/{vector_store_name}/drop`

Permanently delete a vector store and all its contents.

**⚠️ WARNING:** This is a destructive operation that cannot be undone. All documents, embeddings, metadata,
and index configurations will be permanently deleted. Data recovery is not possible after deletion.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Returns

- `name: string`

  The name of the deleted vector store

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/drop \
    -X POST \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "name": "name"
}
```

## Upsert Vectors

**post** `/v5/vector-stores/{vector_store_name}/upsert`

Insert new documents or update existing documents in a vector store.

**Upsert Behavior:** If a document ID already exists, it will be completely replaced with the new content
and metadata. The previous document's text, embedding, and all metadata fields are discarded. If the ID
does not exist, a new document is created.

**Document Content:** Each document supports several modes:

- `content` only: text is automatically embedded using the store's configured model.
- `embedding` only: pre-computed embedding vector is used directly. Dimension must match the store's configuration.
- Both `content` and `embedding`: the pre-computed embedding is stored and text is kept for retrieval/search.
- Neither (metadata-only): only metadata is updated on an existing document without re-embedding.
  If the document does not exist, it will appear as a failure in the batch response.

A store created without an embedding model (dimensions-only) only accepts documents with pre-computed `embedding`.

**Batch Operations:** This endpoint supports batch operations with partial success handling and mixed
document types (some with raw embeddings, some with content) in the same call.

**Metadata:** Supports nested metadata with string, number, boolean, object, and array types. Null values
are not permitted—omit the field or use an empty string instead.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Body Parameters

- `vectors: array of object { id, content, embedding, metadata }`

  Array of documents to upsert

  - `id: string`

    Unique document ID

  - `content: optional TextContent`

    Text content for documents.

    - `text: string`

      Text content to be embedded

    - `type: optional "text"`

      Content type identifier

      - `"text"`

  - `embedding: optional array of number`

    Pre-computed embedding vector

  - `metadata: optional map[unknown]`

    Key-value metadata

### Returns

- `failure_count: number`

  Number of failed documents

- `success_count: number`

  Number of successfully processed documents

- `failed: optional array of object { id, error }`

  Failed documents with their error messages

  - `id: string`

    Document ID

  - `error: string`

    Error message describing why the document failed

- `succeeded: optional array of string`

  IDs of successfully processed documents

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/upsert \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "vectors": [
            {
              "id": "id"
            }
          ]
        }'
```

#### Response

```json
{
  "failure_count": 0,
  "success_count": 0,
  "failed": [
    {
      "id": "id",
      "error": "error"
    }
  ],
  "succeeded": [
    "string"
  ]
}
```

## Delete Vectors

**post** `/v5/vector-stores/{vector_store_name}/delete`

Delete documents from a vector store by document IDs or metadata filter criteria.

**Delete by IDs:** Provide an array of document IDs to delete specific documents. Non-existent documents
are silently skipped.

**Delete by Filter:** Use metadata filters to delete all documents matching the specified criteria (e.g.,
delete all documents where `status: "archived"`). The filter must specify at least one condition and cannot
be empty. To delete all documents, use the drop endpoint instead.

**Filter Operators:** Supports MongoDB-style operators including equality (`{"field": "value"}`),
comparison (`$gt`, `$gte`, `$lt`, `$lte`, `$eq`, `$ne`), logical (`$and`, `$or`, `$not`),
and membership (`$in`, `$nin`). Only indexed metadata fields can be used for filtering.

**Best Practice:** Use the count endpoint with the same filter to preview the number of documents that
will be deleted before executing the deletion operation.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Body Parameters

- `filter: optional map[unknown]`

  Metadata filter expression for deletion

- `ids: optional array of string`

  Array of document IDs to delete

### Returns

- `deleted_count: number`

  Number of documents deleted

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/delete \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{}'
```

#### Response

```json
{
  "deleted_count": 0
}
```

## Count Vectors

**post** `/v5/vector-stores/{vector_store_name}/count`

Count documents in a vector store, optionally filtered by metadata.

**Use Cases:**

- Monitor vector store size and growth over time
- Preview the number of documents matching a filter before deletion
- Validate data ingestion by comparing expected versus actual document counts
- Analyze document distribution across metadata categories

**Filtering:** Apply the same metadata filter syntax as delete and list operations. Only indexed fields
can be used for filtering. An empty filter counts all documents in the store.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Body Parameters

- `filter: optional map[unknown]`

  Metadata filter expression

### Returns

- `count: number`

  Number of documents matching the criteria

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/count \
    -X POST \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "count": 0
}
```

## Query Vectors

**post** `/v5/vector-stores/{vector_store_name}/query`

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation)
applications. Returns documents ranked by relevance to the query text with similarity scores.

**Query Types:**

- `semantic` (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search,
  and finding semantically related content without requiring exact keyword matches.
- `lexical`: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns,
  and scenarios where keyword presence is more important than semantic similarity.
- `hybrid`: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by
  identifying documents matching either semantically or lexically.

**Metadata Filtering:** Narrow the search scope by applying metadata filters (e.g., search only documents
where `category: "technical"`). Only indexed fields can be used for filtering.
Filters are applied before similarity search for optimal efficiency.

**Reranking (Advanced):** Optionally enhance result quality using a cross-encoder reranking model.
The reranker rescores the initial results using a more sophisticated model that evaluates the complete
query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves
precision for high-stakes applications.

**Reranking Strategy:** Set `top_k` higher than the desired final count (e.g., 50) to retrieve more
candidates from the initial search. Then configure `rerank_top_n` to the desired final count (e.g., 10)
to return only the most relevant documents after reranking. This two-stage approach maximizes both recall
and precision.

**Performance Metrics:** The response includes detailed timing breakdowns (embedding generation time,
index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

**Similarity Scores:** Each result includes a `score` field indicating relevance. Higher scores indicate
greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity,
lexical scores use BM25, hybrid scores combine both approaches).

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Body Parameters

- `content: TextContent`

  Text content for documents.

  - `text: string`

    Text content to be embedded

  - `type: optional "text"`

    Content type identifier

    - `"text"`

- `filter: optional map[unknown]`

  Metadata filter expression

- `include_vectors: optional boolean`

  Include embedding vectors in response

- `query_type: optional "semantic" or "lexical" or "hybrid"`

  Query type: semantic, lexical, or hybrid

  - `"semantic"`

  - `"lexical"`

  - `"hybrid"`

- `rerank: optional boolean`

  [Deprecated: use rerank_config] Enable reranking of search results

- `rerank_config: optional object { instruction, model, provider, 2 more }`

  Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

  - `instruction: optional string`

    Custom instruction for the reranking model (e.g., 'Given a medical question, retrieve relevant clinical passages'). Only applies to instruction-following rerankers like Qwen3.

  - `model: optional string`

    Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. 'cross-encoder/ms-marco-MiniLM-L-12-v2'), Vertex semantic-ranker names (e.g. 'semantic-ranker-default-004') when provider='vertex', or any model id the inference proxy serves when provider='proxy'.

  - `provider: optional "launch" or "vertex" or "proxy"`

    Reranking provider to use. When omitted, the deployment default is used ('launch', or 'proxy' on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. 'vertex') to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

    - `"launch"`

    - `"vertex"`

    - `"proxy"`

  - `top_n: optional number`

    Number of results to keep after reranking (defaults to top_k)

  - `type: optional "base"`

    Reranking configuration type. Currently only 'base' is supported.

    - `"base"`

- `rerank_instruction: optional string`

  [Deprecated: use rerank_config.instruction] Custom instruction for reranker

- `rerank_model: optional string`

  [Deprecated: use rerank_config.model] Reranking model to use

- `rerank_top_n: optional number`

  [Deprecated: use rerank_config.top_n] Number of results after reranking

- `top_k: optional number`

  Number of search results to return

### Returns

- `metadata: object { search_type, total_query_time_ms, embedding_config, 4 more }`

  Query execution metadata

  - `search_type: string`

    Type of search performed (semantic, lexical, hybrid)

  - `total_query_time_ms: number`

    Total end-to-end query execution time in milliseconds

  - `embedding_config: optional EmbeddingConfig`

    Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

    - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

      - `model_deployment_id: string`

        The ID of the deployment of the created model in the Models API V3.

      - `type: "models_api"`

        The type of the embedding configuration.

        - `"models_api"`

    - `EmbeddingConfigBase object { embedding_model, type }`

      - `embedding_model: EmbeddingModelName or string`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `string`

      - `type: optional "base"`

        The type of the embedding configuration.

        - `"base"`

  - `embedding_time_ms: optional number`

    Time spent generating embeddings in milliseconds (None for lexical queries)

  - `index_query_time_ms: optional number`

    Time spent querying the vector index (OpenSearch) in milliseconds

  - `reranking_model: optional string`

    Reranking model used (None if reranking not enabled)

  - `reranking_time_ms: optional number`

    Time spent reranking results in milliseconds (None if reranking not enabled)

- `vectors: array of object { id, score, content, 2 more }`

  Array of matching documents

  - `id: string`

    Document ID

  - `score: number`

    Similarity score indicating relevance

  - `content: optional TextContent`

    Text content for documents.

    - `text: string`

      Text content to be embedded

    - `type: optional "text"`

      Content type identifier

      - `"text"`

  - `metadata: optional map[unknown]`

    Key-value metadata

  - `vector: optional array of number`

    Embedding vector (if requested)

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/query \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "content": {
            "text": "text"
          }
        }'
```

#### Response

```json
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
```

## Domain Types

### Embedding Config

- `EmbeddingConfig = EmbeddingConfigModelsAPI or EmbeddingConfigBase`

  - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

    - `model_deployment_id: string`

      The ID of the deployment of the created model in the Models API V3.

    - `type: "models_api"`

      The type of the embedding configuration.

      - `"models_api"`

  - `EmbeddingConfigBase object { embedding_model, type }`

    - `embedding_model: EmbeddingModelName or string`

      The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

      - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

        - `"sentence-transformers/all-MiniLM-L12-v2"`

        - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

        - `"openai/text-embedding-ada-002"`

        - `"openai/text-embedding-3-small"`

        - `"openai/text-embedding-3-large"`

        - `"embed-english-v3.0"`

        - `"embed-english-light-v3.0"`

        - `"embed-multilingual-v3.0"`

        - `"gemini/text-embedding-005"`

        - `"gemini/text-multilingual-embedding-002"`

        - `"gemini/gemini-embedding-001"`

      - `string`

    - `type: optional "base"`

      The type of the embedding configuration.

      - `"base"`

### Embedding Config Base

- `EmbeddingConfigBase object { embedding_model, type }`

  - `embedding_model: EmbeddingModelName or string`

    The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

    - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

      - `"sentence-transformers/all-MiniLM-L12-v2"`

      - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

      - `"openai/text-embedding-ada-002"`

      - `"openai/text-embedding-3-small"`

      - `"openai/text-embedding-3-large"`

      - `"embed-english-v3.0"`

      - `"embed-english-light-v3.0"`

      - `"embed-multilingual-v3.0"`

      - `"gemini/text-embedding-005"`

      - `"gemini/text-multilingual-embedding-002"`

      - `"gemini/gemini-embedding-001"`

    - `string`

  - `type: optional "base"`

    The type of the embedding configuration.

    - `"base"`

### Embedding Config Models API

- `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

  - `model_deployment_id: string`

    The ID of the deployment of the created model in the Models API V3.

  - `type: "models_api"`

    The type of the embedding configuration.

    - `"models_api"`

### Embedding Model Name

- `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

  - `"sentence-transformers/all-MiniLM-L12-v2"`

  - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

  - `"openai/text-embedding-ada-002"`

  - `"openai/text-embedding-3-small"`

  - `"openai/text-embedding-3-large"`

  - `"embed-english-v3.0"`

  - `"embed-english-light-v3.0"`

  - `"embed-multilingual-v3.0"`

  - `"gemini/text-embedding-005"`

  - `"gemini/text-multilingual-embedding-002"`

  - `"gemini/gemini-embedding-001"`

### Text Content

- `TextContent object { text, type }`

  Text content for documents.

  - `text: string`

    Text content to be embedded

  - `type: optional "text"`

    Content type identifier

    - `"text"`

### Vector Store

- `VectorStore object { id, created_at, embedding_dimensions, 4 more }`

  Response model for vector store operations.

  - `id: string`

    The unique identifier of the vector store

  - `created_at: string`

    Timestamp of creation

  - `embedding_dimensions: number`

    Dimensionality of the embedding vectors

  - `name: string`

    The name of the vector store

  - `updated_at: string`

    Timestamp of last update

  - `embedding_config: optional EmbeddingConfig`

    Embedding configuration identifying the model and its type. None for raw-embedding-only stores.

    - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

      - `model_deployment_id: string`

        The ID of the deployment of the created model in the Models API V3.

      - `type: "models_api"`

        The type of the embedding configuration.

        - `"models_api"`

    - `EmbeddingConfigBase object { embedding_model, type }`

      - `embedding_model: EmbeddingModelName or string`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `string`

      - `type: optional "base"`

        The type of the embedding configuration.

        - `"base"`

  - `indexed_metadata_fields: optional map["string" or "number" or "boolean"]`

    Dictionary mapping metadata field names to their types

    - `"string"`

    - `"number"`

    - `"boolean"`

### Vector Store Drop Response

- `VectorStoreDropResponse object { name }`

  Response for vector store deletion.

  - `name: string`

    The name of the deleted vector store

### Vector Store Upsert Response

- `VectorStoreUpsertResponse object { failure_count, success_count, failed, succeeded }`

  Response for batch insert/upsert operations.

  - `failure_count: number`

    Number of failed documents

  - `success_count: number`

    Number of successfully processed documents

  - `failed: optional array of object { id, error }`

    Failed documents with their error messages

    - `id: string`

      Document ID

    - `error: string`

      Error message describing why the document failed

  - `succeeded: optional array of string`

    IDs of successfully processed documents

### Vector Store Delete Response

- `VectorStoreDeleteResponse object { deleted_count }`

  Response for delete operation.

  - `deleted_count: number`

    Number of documents deleted

### Vector Store Count Response

- `VectorStoreCountResponse object { count }`

  Response for count operation.

  - `count: number`

    Number of documents matching the criteria

### Vector Store Query Response

- `VectorStoreQueryResponse object { metadata, vectors }`

  Response for query operation.

  - `metadata: object { search_type, total_query_time_ms, embedding_config, 4 more }`

    Query execution metadata

    - `search_type: string`

      Type of search performed (semantic, lexical, hybrid)

    - `total_query_time_ms: number`

      Total end-to-end query execution time in milliseconds

    - `embedding_config: optional EmbeddingConfig`

      Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

      - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

        - `model_deployment_id: string`

          The ID of the deployment of the created model in the Models API V3.

        - `type: "models_api"`

          The type of the embedding configuration.

          - `"models_api"`

      - `EmbeddingConfigBase object { embedding_model, type }`

        - `embedding_model: EmbeddingModelName or string`

          The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

          - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

            - `"sentence-transformers/all-MiniLM-L12-v2"`

            - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

            - `"openai/text-embedding-ada-002"`

            - `"openai/text-embedding-3-small"`

            - `"openai/text-embedding-3-large"`

            - `"embed-english-v3.0"`

            - `"embed-english-light-v3.0"`

            - `"embed-multilingual-v3.0"`

            - `"gemini/text-embedding-005"`

            - `"gemini/text-multilingual-embedding-002"`

            - `"gemini/gemini-embedding-001"`

          - `string`

        - `type: optional "base"`

          The type of the embedding configuration.

          - `"base"`

    - `embedding_time_ms: optional number`

      Time spent generating embeddings in milliseconds (None for lexical queries)

    - `index_query_time_ms: optional number`

      Time spent querying the vector index (OpenSearch) in milliseconds

    - `reranking_model: optional string`

      Reranking model used (None if reranking not enabled)

    - `reranking_time_ms: optional number`

      Time spent reranking results in milliseconds (None if reranking not enabled)

  - `vectors: array of object { id, score, content, 2 more }`

    Array of matching documents

    - `id: string`

      Document ID

    - `score: number`

      Similarity score indicating relevance

    - `content: optional TextContent`

      Text content for documents.

      - `text: string`

        Text content to be embedded

      - `type: optional "text"`

        Content type identifier

        - `"text"`

    - `metadata: optional map[unknown]`

      Key-value metadata

    - `vector: optional array of number`

      Embedding vector (if requested)

# Vectors

## List Vectors

**get** `/v5/vector-stores/{vector_store_name}/vectors`

List documents in a vector store with cursor-based pagination.

**Use Cases:** Browse documents, export content, audit stored data, or retrieve documents by metadata
without semantic search.

**Ordering:** Documents are returned in storage order (insertion order), not ranked by similarity.
For similarity-based retrieval, use the query endpoint.

**Filtering:** Apply metadata filters to narrow results to specific subsets (e.g., all documents
where `category: "research"`). Only indexed fields can be used for filtering.

**Pagination:** Uses cursor-based pagination for efficient traversal of large datasets. Pass the
`next_cursor` from each response as `starting_after` to retrieve the next page, or `prev_cursor`
as `ending_before` to retrieve the previous page. A null cursor indicates no further pages exist.

**Embedding Vectors:** Setting `include_vectors=true` includes the full embedding vector arrays in
the response. This significantly increases payload size and reduces the maximum page size from 1000 to
100 documents. Enable only when raw vectors are required for external processing.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Query Parameters

- `cursor: optional string`

  Alias for starting_after. Use starting_after instead.

- `ending_before: optional string`

- `filter: optional string`

  Metadata filter expression (JSON)

- `include_vectors: optional boolean`

  Include embedding vectors

- `limit: optional number`

- `sort_by: optional string`

- `sort_order: optional SortOrder`

  - `"asc"`

  - `"desc"`

- `starting_after: optional string`

### Returns

- `has_more: boolean`

  Whether there are more items left to be fetched.

- `items: array of VectorDocument`

  Array of documents

  - `id: string`

    Document ID

  - `content: optional TextContent`

    Text content for documents.

    - `text: string`

      Text content to be embedded

    - `type: optional "text"`

      Content type identifier

      - `"text"`

  - `metadata: optional map[unknown]`

    Key-value metadata

  - `vector: optional array of number`

    Embedding vector (if requested)

- `total: number`

  The total of items that match the query. This is greater than or equal to the number of items returned.

- `vectors: array of VectorDocument`

  Array of documents. Deprecated: use `items` instead.

  - `id: string`

    Document ID

  - `content: optional TextContent`

    Text content for documents.

  - `metadata: optional map[unknown]`

    Key-value metadata

  - `vector: optional array of number`

    Embedding vector (if requested)

- `limit: optional number`

  The maximum number of items to return.

- `next_cursor: optional string`

  Pass as starting_after to fetch the next page. None when there is no next page.

- `object: optional "list"`

  - `"list"`

- `prev_cursor: optional string`

  Pass as ending_before to fetch the previous page. None when on the first page.

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/vectors \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "has_more": true,
  "items": [
    {
      "id": "id",
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ],
  "total": 0,
  "vectors": [
    {
      "id": "id",
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ],
  "limit": 0,
  "next_cursor": "next_cursor",
  "object": "list",
  "prev_cursor": "prev_cursor"
}
```

## Get Vector

**get** `/v5/vector-stores/{vector_store_name}/vectors/{vector_id}`

Retrieve a single document by its unique ID.

Returns the document's full content, metadata, and optionally its embedding vector. Use this endpoint
for direct lookups when the exact document ID is known. For content similarity search,
use the query endpoint.

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

- `vector_id: string`

  The ID of the vector to retrieve

### Query Parameters

- `include_vectors: optional boolean`

  Include embedding vectors

### Returns

- `VectorDocument object { id, content, metadata, vector }`

  A document returned from direct lookups (get/list operations).

  - `id: string`

    Document ID

  - `content: optional TextContent`

    Text content for documents.

    - `text: string`

      Text content to be embedded

    - `type: optional "text"`

      Content type identifier

      - `"text"`

  - `metadata: optional map[unknown]`

    Key-value metadata

  - `vector: optional array of number`

    Embedding vector (if requested)

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/vectors/$VECTOR_ID \
    -H "x-api-key: $SGP_API_KEY"
```

#### Response

```json
{
  "id": "id",
  "content": {
    "text": "text",
    "type": "text"
  },
  "metadata": {
    "foo": "bar"
  },
  "vector": [
    0
  ]
}
```

## Domain Types

### Vector Document

- `VectorDocument object { id, content, metadata, vector }`

  A document returned from direct lookups (get/list operations).

  - `id: string`

    Document ID

  - `content: optional TextContent`

    Text content for documents.

    - `text: string`

      Text content to be embedded

    - `type: optional "text"`

      Content type identifier

      - `"text"`

  - `metadata: optional map[unknown]`

    Key-value metadata

  - `vector: optional array of number`

    Embedding vector (if requested)
