## Query Vectors

`vector_stores.query(strvector_store_name, VectorStoreQueryParams**kwargs)  -> VectorStoreQueryResponse`

**post** `/v5/vector-stores/{vector_store_name}/query`

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation)
applications. Returns documents ranked by relevance to the query text with similarity scores.

**Query Types:**

- `semantic` (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search,
  and finding semantically related content without requiring exact keyword matches.
- `lexical`: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns,
  and scenarios where keyword presence is more important than semantic similarity.
- `hybrid`: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by
  identifying documents matching either semantically or lexically.

**Metadata Filtering:** Narrow the search scope by applying metadata filters (e.g., search only documents
where `category: "technical"`). Only indexed fields can be used for filtering.
Filters are applied before similarity search for optimal efficiency.

**Reranking (Advanced):** Optionally enhance result quality using a cross-encoder reranking model.
The reranker rescores the initial results using a more sophisticated model that evaluates the complete
query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves
precision for high-stakes applications.

**Reranking Strategy:** Set `top_k` higher than the desired final count (e.g., 50) to retrieve more
candidates from the initial search. Then configure `rerank_top_n` to the desired final count (e.g., 10)
to return only the most relevant documents after reranking. This two-stage approach maximizes both recall
and precision.

**Performance Metrics:** The response includes detailed timing breakdowns (embedding generation time,
index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

**Similarity Scores:** Each result includes a `score` field indicating relevance. Higher scores indicate
greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity,
lexical scores use BM25, hybrid scores combine both approaches).

### Parameters

- `vector_store_name: str`

  The name of the vector store

- `content: TextContentParam`

  Text content for documents.

  - `text: str`

    Text content to be embedded

  - `type: Optional[Literal["text"]]`

    Content type identifier

    - `"text"`

- `filter: Optional[Dict[str, object]]`

  Metadata filter expression

- `include_vectors: Optional[bool]`

  Include embedding vectors in response

- `query_type: Optional[Literal["semantic", "lexical", "hybrid"]]`

  Query type: semantic, lexical, or hybrid

  - `"semantic"`

  - `"lexical"`

  - `"hybrid"`

- `rerank: Optional[bool]`

  [Deprecated: use rerank_config] Enable reranking of search results

- `rerank_config: Optional[RerankConfig]`

  Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

  - `instruction: Optional[str]`

    Custom instruction for the reranking model (e.g., 'Given a medical question, retrieve relevant clinical passages'). Only applies to instruction-following rerankers like Qwen3.

  - `model: Optional[str]`

    Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. 'cross-encoder/ms-marco-MiniLM-L-12-v2'), Vertex semantic-ranker names (e.g. 'semantic-ranker-default-004') when provider='vertex', or any model id the inference proxy serves when provider='proxy'.

  - `provider: Optional[Literal["launch", "vertex", "proxy"]]`

    Reranking provider to use. When omitted, the deployment default is used ('launch', or 'proxy' on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. 'vertex') to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

    - `"launch"`

    - `"vertex"`

    - `"proxy"`

  - `top_n: Optional[int]`

    Number of results to keep after reranking (defaults to top_k)

  - `type: Optional[Literal["base"]]`

    Reranking configuration type. Currently only 'base' is supported.

    - `"base"`

- `rerank_instruction: Optional[str]`

  [Deprecated: use rerank_config.instruction] Custom instruction for reranker

- `rerank_model: Optional[str]`

  [Deprecated: use rerank_config.model] Reranking model to use

- `rerank_top_n: Optional[int]`

  [Deprecated: use rerank_config.top_n] Number of results after reranking

- `top_k: Optional[int]`

  Number of search results to return

### Returns

- `class VectorStoreQueryResponse: …`

  Response for query operation.

  - `metadata: Metadata`

    Query execution metadata

    - `search_type: str`

      Type of search performed (semantic, lexical, hybrid)

    - `total_query_time_ms: int`

      Total end-to-end query execution time in milliseconds

    - `embedding_config: Optional[EmbeddingConfig]`

      Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

      - `class EmbeddingConfigModelsAPI: …`

        - `model_deployment_id: str`

          The ID of the deployment of the created model in the Models API V3.

        - `type: Literal["models_api"]`

          The type of the embedding configuration.

          - `"models_api"`

      - `class EmbeddingConfigBase: …`

        - `embedding_model: Union[EmbeddingModelName, str]`

          The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

          - `Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]`

            - `"sentence-transformers/all-MiniLM-L12-v2"`

            - `"sentence-transformers/multi-qa-distilbert-cos-v1"`

            - `"openai/text-embedding-ada-002"`

            - `"openai/text-embedding-3-small"`

            - `"openai/text-embedding-3-large"`

            - `"embed-english-v3.0"`

            - `"embed-english-light-v3.0"`

            - `"embed-multilingual-v3.0"`

            - `"gemini/text-embedding-005"`

            - `"gemini/text-multilingual-embedding-002"`

            - `"gemini/gemini-embedding-001"`

          - `str`

        - `type: Optional[Literal["base"]]`

          The type of the embedding configuration.

          - `"base"`

    - `embedding_time_ms: Optional[int]`

      Time spent generating embeddings in milliseconds (None for lexical queries)

    - `index_query_time_ms: Optional[int]`

      Time spent querying the vector index (OpenSearch) in milliseconds

    - `reranking_model: Optional[str]`

      Reranking model used (None if reranking not enabled)

    - `reranking_time_ms: Optional[int]`

      Time spent reranking results in milliseconds (None if reranking not enabled)

  - `vectors: List[Vector]`

    Array of matching documents

    - `id: str`

      Document ID

    - `score: float`

      Similarity score indicating relevance

    - `content: Optional[TextContent]`

      Text content for documents.

      - `text: str`

        Text content to be embedded

      - `type: Optional[Literal["text"]]`

        Content type identifier

        - `"text"`

    - `metadata: Optional[Dict[str, object]]`

      Key-value metadata

    - `vector: Optional[List[float]]`

      Embedding vector (if requested)

### Example

```python
import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.vector_stores.query(
    vector_store_name="vector_store_name",
    content={
        "text": "text"
    },
)
print(response.metadata)
```

#### Response

```json
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
```
