Skip to content

Query Vectors

POST/v5/vector-stores/{vector_store_name}/query

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation) applications. Returns documents ranked by relevance to the query text with similarity scores.

Query Types:

  • semantic (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search, and finding semantically related content without requiring exact keyword matches.
  • lexical: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns, and scenarios where keyword presence is more important than semantic similarity.
  • hybrid: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by identifying documents matching either semantically or lexically.

Metadata Filtering: Narrow the search scope by applying metadata filters (e.g., search only documents where category: "technical"). Only indexed fields can be used for filtering. Filters are applied before similarity search for optimal efficiency.

Reranking (Advanced): Optionally enhance result quality using a cross-encoder reranking model. The reranker rescores the initial results using a more sophisticated model that evaluates the complete query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves precision for high-stakes applications.

Reranking Strategy: Set top_k higher than the desired final count (e.g., 50) to retrieve more candidates from the initial search. Then configure rerank_top_n to the desired final count (e.g., 10) to return only the most relevant documents after reranking. This two-stage approach maximizes both recall and precision.

Performance Metrics: The response includes detailed timing breakdowns (embedding generation time, index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

Similarity Scores: Each result includes a score field indicating relevance. Higher scores indicate greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity, lexical scores use BM25, hybrid scores combine both approaches).

Path ParametersExpand Collapse
vector_store_name: string

The name of the vector store

Body ParametersJSONExpand Collapse
content: TextContent { text, type }

Text content for documents.

text: string

Text content to be embedded

type: optional "text"

Content type identifier

filter: optional map[unknown]

Metadata filter expression

include_vectors: optional boolean

Include embedding vectors in response

query_type: optional "semantic" or "lexical" or "hybrid"

Query type: semantic, lexical, or hybrid

One of the following:
"semantic"
"lexical"
"hybrid"
rerank: optional boolean

[Deprecated: use rerank_config] Enable reranking of search results

rerank_config: optional object { instruction, model, provider, 2 more }

Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

instruction: optional string

Custom instruction for the reranking model (e.g., ‘Given a medical question, retrieve relevant clinical passages’). Only applies to instruction-following rerankers like Qwen3.

model: optional string

Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. ‘cross-encoder/ms-marco-MiniLM-L-12-v2’), Vertex semantic-ranker names (e.g. ‘semantic-ranker-default-004’) when provider=‘vertex’, or any model id the inference proxy serves when provider=‘proxy’.

provider: optional "launch" or "vertex" or "proxy"

Reranking provider to use. When omitted, the deployment default is used (‘launch’, or ‘proxy’ on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. ‘vertex’) to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

One of the following:
"launch"
"vertex"
"proxy"
top_n: optional number

Number of results to keep after reranking (defaults to top_k)

minimum1
type: optional "base"

Reranking configuration type. Currently only ‘base’ is supported.

rerank_instruction: optional string

[Deprecated: use rerank_config.instruction] Custom instruction for reranker

rerank_model: optional string

[Deprecated: use rerank_config.model] Reranking model to use

rerank_top_n: optional number

[Deprecated: use rerank_config.top_n] Number of results after reranking

minimum1
top_k: optional number

Number of search results to return

minimum1
ReturnsExpand Collapse
metadata: object { search_type, total_query_time_ms, embedding_config, 4 more }

Query execution metadata

search_type: string

Type of search performed (semantic, lexical, hybrid)

total_query_time_ms: number

Total end-to-end query execution time in milliseconds

embedding_config: optional EmbeddingConfig

Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

One of the following:
EmbeddingConfigModelsAPI object { model_deployment_id, type }
model_deployment_id: string

The ID of the deployment of the created model in the Models API V3.

type: "models_api"

The type of the embedding configuration.

EmbeddingConfigBase object { embedding_model, type }
embedding_model: EmbeddingModelName or string

The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. ‘nomic-embed-text-v1.5’). For fully custom deployments, use type ‘models_api’ with a model_deployment_id.

One of the following:
EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2" or "sentence-transformers/multi-qa-distilbert-cos-v1" or "openai/text-embedding-ada-002" or 8 more
One of the following:
"sentence-transformers/all-MiniLM-L12-v2"
"sentence-transformers/multi-qa-distilbert-cos-v1"
"openai/text-embedding-ada-002"
"openai/text-embedding-3-small"
"openai/text-embedding-3-large"
"embed-english-v3.0"
"embed-english-light-v3.0"
"embed-multilingual-v3.0"
"gemini/text-embedding-005"
"gemini/text-multilingual-embedding-002"
"gemini/gemini-embedding-001"
string
type: optional "base"

The type of the embedding configuration.

embedding_time_ms: optional number

Time spent generating embeddings in milliseconds (None for lexical queries)

index_query_time_ms: optional number

Time spent querying the vector index (OpenSearch) in milliseconds

reranking_model: optional string

Reranking model used (None if reranking not enabled)

reranking_time_ms: optional number

Time spent reranking results in milliseconds (None if reranking not enabled)

vectors: array of object { id, score, content, 2 more }

Array of matching documents

id: string

Document ID

score: number

Similarity score indicating relevance

content: optional TextContent { text, type }

Text content for documents.

text: string

Text content to be embedded

type: optional "text"

Content type identifier

metadata: optional map[unknown]

Key-value metadata

vector: optional array of number

Embedding vector (if requested)

Query Vectors

curl https://api.egp.scale.com/v5/vector-stores/$VECTOR_STORE_NAME/query \
    -H 'Content-Type: application/json' \
    -H "x-api-key: $SGP_API_KEY" \
    -d '{
          "content": {
            "text": "text"
          }
        }'
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
Returns Examples
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}