Skip to content

Query Vectors

vector_stores.query(strvector_store_name, VectorStoreQueryParams**kwargs) -> VectorStoreQueryResponse
POST/v5/vector-stores/{vector_store_name}/query

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation) applications. Returns documents ranked by relevance to the query text with similarity scores.

Query Types:

  • semantic (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search, and finding semantically related content without requiring exact keyword matches.
  • lexical: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns, and scenarios where keyword presence is more important than semantic similarity.
  • hybrid: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by identifying documents matching either semantically or lexically.

Metadata Filtering: Narrow the search scope by applying metadata filters (e.g., search only documents where category: "technical"). Only indexed fields can be used for filtering. Filters are applied before similarity search for optimal efficiency.

Reranking (Advanced): Optionally enhance result quality using a cross-encoder reranking model. The reranker rescores the initial results using a more sophisticated model that evaluates the complete query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves precision for high-stakes applications.

Reranking Strategy: Set top_k higher than the desired final count (e.g., 50) to retrieve more candidates from the initial search. Then configure rerank_top_n to the desired final count (e.g., 10) to return only the most relevant documents after reranking. This two-stage approach maximizes both recall and precision.

Performance Metrics: The response includes detailed timing breakdowns (embedding generation time, index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

Similarity Scores: Each result includes a score field indicating relevance. Higher scores indicate greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity, lexical scores use BM25, hybrid scores combine both approaches).

ParametersExpand Collapse
vector_store_name: str

The name of the vector store

Text content for documents.

text: str

Text content to be embedded

type: Optional[Literal["text"]]

Content type identifier

filter: Optional[Dict[str, object]]

Metadata filter expression

include_vectors: Optional[bool]

Include embedding vectors in response

query_type: Optional[Literal["semantic", "lexical", "hybrid"]]

Query type: semantic, lexical, or hybrid

One of the following:
"semantic"
"lexical"
"hybrid"
rerank: Optional[bool]

[Deprecated: use rerank_config] Enable reranking of search results

rerank_config: Optional[RerankConfig]

Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

instruction: Optional[str]

Custom instruction for the reranking model (e.g., ‘Given a medical question, retrieve relevant clinical passages’). Only applies to instruction-following rerankers like Qwen3.

model: Optional[str]

Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. ‘cross-encoder/ms-marco-MiniLM-L-12-v2’), Vertex semantic-ranker names (e.g. ‘semantic-ranker-default-004’) when provider=‘vertex’, or any model id the inference proxy serves when provider=‘proxy’.

provider: Optional[Literal["launch", "vertex", "proxy"]]

Reranking provider to use. When omitted, the deployment default is used (‘launch’, or ‘proxy’ on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. ‘vertex’) to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

One of the following:
"launch"
"vertex"
"proxy"
top_n: Optional[int]

Number of results to keep after reranking (defaults to top_k)

minimum1
type: Optional[Literal["base"]]

Reranking configuration type. Currently only ‘base’ is supported.

rerank_instruction: Optional[str]

[Deprecated: use rerank_config.instruction] Custom instruction for reranker

rerank_model: Optional[str]

[Deprecated: use rerank_config.model] Reranking model to use

rerank_top_n: Optional[int]

[Deprecated: use rerank_config.top_n] Number of results after reranking

minimum1
top_k: Optional[int]

Number of search results to return

minimum1
ReturnsExpand Collapse
class VectorStoreQueryResponse: …

Response for query operation.

metadata: Metadata

Query execution metadata

search_type: str

Type of search performed (semantic, lexical, hybrid)

total_query_time_ms: int

Total end-to-end query execution time in milliseconds

embedding_config: Optional[EmbeddingConfig]

Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

One of the following:
class EmbeddingConfigModelsAPI: …
model_deployment_id: str

The ID of the deployment of the created model in the Models API V3.

type: Literal["models_api"]

The type of the embedding configuration.

class EmbeddingConfigBase: …
embedding_model: Union[EmbeddingModelName, str]

The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. ‘nomic-embed-text-v1.5’). For fully custom deployments, use type ‘models_api’ with a model_deployment_id.

One of the following:
Literal["sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/multi-qa-distilbert-cos-v1", "openai/text-embedding-ada-002", 8 more]
One of the following:
"sentence-transformers/all-MiniLM-L12-v2"
"sentence-transformers/multi-qa-distilbert-cos-v1"
"openai/text-embedding-ada-002"
"openai/text-embedding-3-small"
"openai/text-embedding-3-large"
"embed-english-v3.0"
"embed-english-light-v3.0"
"embed-multilingual-v3.0"
"gemini/text-embedding-005"
"gemini/text-multilingual-embedding-002"
"gemini/gemini-embedding-001"
str
type: Optional[Literal["base"]]

The type of the embedding configuration.

embedding_time_ms: Optional[int]

Time spent generating embeddings in milliseconds (None for lexical queries)

index_query_time_ms: Optional[int]

Time spent querying the vector index (OpenSearch) in milliseconds

reranking_model: Optional[str]

Reranking model used (None if reranking not enabled)

reranking_time_ms: Optional[int]

Time spent reranking results in milliseconds (None if reranking not enabled)

vectors: List[Vector]

Array of matching documents

id: str

Document ID

score: float

Similarity score indicating relevance

content: Optional[TextContent]

Text content for documents.

text: str

Text content to be embedded

type: Optional[Literal["text"]]

Content type identifier

metadata: Optional[Dict[str, object]]

Key-value metadata

vector: Optional[List[float]]

Embedding vector (if requested)

Query Vectors

import os
from scale_gp_beta import SGPClient

client = SGPClient(
    api_key=os.environ.get("SGP_API_KEY"),  # This is the default and can be omitted
)
response = client.vector_stores.query(
    vector_store_name="vector_store_name",
    content={
        "text": "text"
    },
)
print(response.metadata)
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
Returns Examples
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}