## Query Vectors

**post** `/v5/vector-stores/{vector_store_name}/query`

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation)
applications. Returns documents ranked by relevance to the query text with similarity scores.

**Query Types:**

- `semantic` (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search,
  and finding semantically related content without requiring exact keyword matches.
- `lexical`: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns,
  and scenarios where keyword presence is more important than semantic similarity.
- `hybrid`: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by
  identifying documents matching either semantically or lexically.

**Metadata Filtering:** Narrow the search scope by applying metadata filters (e.g., search only documents
where `category: "technical"`). Only indexed fields can be used for filtering.
Filters are applied before similarity search for optimal efficiency.

**Reranking (Advanced):** Optionally enhance result quality using a cross-encoder reranking model.
The reranker rescores the initial results using a more sophisticated model that evaluates the complete
query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves
precision for high-stakes applications.

**Reranking Strategy:** Set `top_k` higher than the desired final count (e.g., 50) to retrieve more
candidates from the initial search. Then configure `rerank_top_n` to the desired final count (e.g., 10)
to return only the most relevant documents after reranking. This two-stage approach maximizes both recall
and precision.

**Performance Metrics:** The response includes detailed timing breakdowns (embedding generation time,
index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

**Similarity Scores:** Each result includes a `score` field indicating relevance. Higher scores indicate
greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity,
lexical scores use BM25, hybrid scores combine both approaches).

### Path Parameters

- `vector_store_name: string`

  The name of the vector store

### Body Parameters

- `content: TextContent`

  Text content for documents.

  - `text: string`

    Text content to be embedded

  - `type: optional "text"`

    Content type identifier

    - `"text"`

- `filter: optional map[unknown]`

  Metadata filter expression

- `include_vectors: optional boolean`

  Include embedding vectors in response

- `query_type: optional "semantic" or "lexical" or "hybrid"`

  Query type: semantic, lexical, or hybrid

  - `"semantic"`

  - `"lexical"`

  - `"hybrid"`

- `rerank: optional boolean`

  [Deprecated: use rerank_config] Enable reranking of search results

- `rerank_config: optional object { instruction, model, provider, 2 more }`

  Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

  - `instruction: optional string`

    Custom instruction for the reranking model (e.g., 'Given a medical question, retrieve relevant clinical passages'). Only applies to instruction-following rerankers like Qwen3.

  - `model: optional string`

    Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. 'cross-encoder/ms-marco-MiniLM-L-12-v2'), Vertex semantic-ranker names (e.g. 'semantic-ranker-default-004') when provider='vertex', or any model id the inference proxy serves when provider='proxy'.

  - `provider: optional "launch" or "vertex" or "proxy"`

    Reranking provider to use. When omitted, the deployment default is used ('launch', or 'proxy' on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. 'vertex') to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

    - `"launch"`

    - `"vertex"`

    - `"proxy"`

  - `top_n: optional number`

    Number of results to keep after reranking (defaults to top_k)

  - `type: optional "base"`

    Reranking configuration type. Currently only 'base' is supported.

    - `"base"`

- `rerank_instruction: optional string`

  [Deprecated: use rerank_config.instruction] Custom instruction for reranker

- `rerank_model: optional string`

  [Deprecated: use rerank_config.model] Reranking model to use

- `rerank_top_n: optional number`

  [Deprecated: use rerank_config.top_n] Number of results after reranking

- `top_k: optional number`

  Number of search results to return

### Returns

- `metadata: object { search_type, total_query_time_ms, embedding_config, 4 more }`

  Query execution metadata

  - `search_type: string`

    Type of search performed (semantic, lexical, hybrid)

  - `total_query_time_ms: number`

    Total end-to-end query execution time in milliseconds

  - `embedding_config: optional EmbeddingConfig`

    Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

    - `EmbeddingConfigModelsAPI object { model_deployment_id, type }`

      - `model_deployment_id: string`

        The ID of the deployment of the created model in the Models API V3.

      - `type: "models_api"`

        The type of the embedding configuration.

        - `"models_api"`

    - `EmbeddingConfigBase object { embedding_model, type }`

      - `embedding_model: EmbeddingModelName or string`

        The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

        - `EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more`

          - `"sentence-transformers/all-MiniLM-L12-v2"`

          - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

          - `"openai/text-embedding-ada-002"`

          - `"openai/text-embedding-3-small"`

          - `"openai/text-embedding-3-large"`

          - `"embed-english-v3.0"`

          - `"embed-english-light-v3.0"`

          - `"embed-multilingual-v3.0"`

          - `"gemini/text-embedding-005"`

          - `"gemini/text-multilingual-embedding-002"`

          - `"gemini/gemini-embedding-001"`

        - `string`

      - `type: optional "base"`

        The type of the embedding configuration.

        - `"base"`

  - `embedding_time_ms: optional number`

    Time spent generating embeddings in milliseconds (None for lexical queries)

  - `index_query_time_ms: optional number`

    Time spent querying the vector index (OpenSearch) in milliseconds

  - `reranking_model: optional string`

    Reranking model used (None if reranking not enabled)

  - `reranking_time_ms: optional number`

    Time spent reranking results in milliseconds (None if reranking not enabled)

- `vectors: array of object { id, score, content, 2 more }`

  Array of matching documents

  - `id: string`

    Document ID

  - `score: number`

    Similarity score indicating relevance

  - `content: optional TextContent`

    Text content for documents.

    - `text: string`

      Text content to be embedded

    - `type: optional "text"`

      Content type identifier

      - `"text"`

  - `metadata: optional map[unknown]`

    Key-value metadata

  - `vector: optional array of number`

    Embedding vector (if requested)

### Example

```http
curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/query \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "content": {
            "text": "text"
          }
        }'
```

#### Response

```json
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
```
