Skip to content

Query Vectors

client.VectorStores.Query(ctx, vectorStoreName, body) (*VectorStoreQueryResponse, error)
POST/v5/vector-stores/{vector_store_name}/query

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation) applications. Returns documents ranked by relevance to the query text with similarity scores.

Query Types:

  • semantic (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search, and finding semantically related content without requiring exact keyword matches.
  • lexical: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns, and scenarios where keyword presence is more important than semantic similarity.
  • hybrid: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by identifying documents matching either semantically or lexically.

Metadata Filtering: Narrow the search scope by applying metadata filters (e.g., search only documents where category: "technical"). Only indexed fields can be used for filtering. Filters are applied before similarity search for optimal efficiency.

Reranking (Advanced): Optionally enhance result quality using a cross-encoder reranking model. The reranker rescores the initial results using a more sophisticated model that evaluates the complete query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves precision for high-stakes applications.

Reranking Strategy: Set top_k higher than the desired final count (e.g., 50) to retrieve more candidates from the initial search. Then configure rerank_top_n to the desired final count (e.g., 10) to return only the most relevant documents after reranking. This two-stage approach maximizes both recall and precision.

Performance Metrics: The response includes detailed timing breakdowns (embedding generation time, index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

Similarity Scores: Each result includes a score field indicating relevance. Higher scores indicate greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity, lexical scores use BM25, hybrid scores combine both approaches).

ParametersExpand Collapse
vectorStoreName string

The name of the vector store

body VectorStoreQueryParams
Content param.Field[TextContent]

Text content for documents.

Filter param.Field[map[string, any]]Optional

Metadata filter expression

IncludeVectors param.Field[bool]Optional

Include embedding vectors in response

QueryType param.Field[VectorStoreQueryParamsQueryType]Optional

Query type: semantic, lexical, or hybrid

const VectorStoreQueryParamsQueryTypeSemantic VectorStoreQueryParamsQueryType = "semantic"
const VectorStoreQueryParamsQueryTypeLexical VectorStoreQueryParamsQueryType = "lexical"
const VectorStoreQueryParamsQueryTypeHybrid VectorStoreQueryParamsQueryType = "hybrid"
Rerank param.Field[bool]Optional

[Deprecated: use rerank_config] Enable reranking of search results

RerankConfig param.Field[VectorStoreQueryParamsRerankConfig]Optional

Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

Instruction stringOptional

Custom instruction for the reranking model (e.g., ‘Given a medical question, retrieve relevant clinical passages’). Only applies to instruction-following rerankers like Qwen3.

Model stringOptional

Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. ‘cross-encoder/ms-marco-MiniLM-L-12-v2’), Vertex semantic-ranker names (e.g. ‘semantic-ranker-default-004’) when provider=‘vertex’, or any model id the inference proxy serves when provider=‘proxy’.

Provider stringOptional

Reranking provider to use. When omitted, the deployment default is used (‘launch’, or ‘proxy’ on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. ‘vertex’) to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

One of the following:
const VectorStoreQueryParamsRerankConfigProviderLaunch VectorStoreQueryParamsRerankConfigProvider = "launch"
const VectorStoreQueryParamsRerankConfigProviderVertex VectorStoreQueryParamsRerankConfigProvider = "vertex"
const VectorStoreQueryParamsRerankConfigProviderProxy VectorStoreQueryParamsRerankConfigProvider = "proxy"
TopN int64Optional

Number of results to keep after reranking (defaults to top_k)

minimum1
Type stringOptional

Reranking configuration type. Currently only ‘base’ is supported.

RerankInstruction param.Field[string]Optional

[Deprecated: use rerank_config.instruction] Custom instruction for reranker

RerankModel param.Field[string]Optional

[Deprecated: use rerank_config.model] Reranking model to use

RerankTopN param.Field[int64]Optional

[Deprecated: use rerank_config.top_n] Number of results after reranking

minimum1
TopK param.Field[int64]Optional

Number of search results to return

minimum1
ReturnsExpand Collapse
type VectorStoreQueryResponse struct{…}

Response for query operation.

Metadata VectorStoreQueryResponseMetadata

Query execution metadata

SearchType string

Type of search performed (semantic, lexical, hybrid)

TotalQueryTimeMs int64

Total end-to-end query execution time in milliseconds

EmbeddingConfig EmbeddingConfigUnionOptional

Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

One of the following:
type EmbeddingConfigModelsAPI struct{…}
ModelDeploymentID string

The ID of the deployment of the created model in the Models API V3.

Type ModelsAPI

The type of the embedding configuration.

type EmbeddingConfigBase struct{…}
EmbeddingModel EmbeddingModelName

The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. ‘nomic-embed-text-v1.5’). For fully custom deployments, use type ‘models_api’ with a model_deployment_id.

One of the following:
type EmbeddingModelName string
One of the following:
const EmbeddingModelNameSentenceTransformersAllMiniLmL12V2 EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2"
const EmbeddingModelNameSentenceTransformersMultiQaDistilbertCosV1 EmbeddingModelName = "sentence-transformers/multi-qa-distilbert-cos-v1"
const EmbeddingModelNameOpenAITextEmbeddingAda002 EmbeddingModelName = "openai/text-embedding-ada-002"
const EmbeddingModelNameOpenAITextEmbedding3Small EmbeddingModelName = "openai/text-embedding-3-small"
const EmbeddingModelNameOpenAITextEmbedding3Large EmbeddingModelName = "openai/text-embedding-3-large"
const EmbeddingModelNameEmbedEnglishV3_0 EmbeddingModelName = "embed-english-v3.0"
const EmbeddingModelNameEmbedEnglishLightV3_0 EmbeddingModelName = "embed-english-light-v3.0"
const EmbeddingModelNameEmbedMultilingualV3_0 EmbeddingModelName = "embed-multilingual-v3.0"
const EmbeddingModelNameGeminiTextEmbedding005 EmbeddingModelName = "gemini/text-embedding-005"
const EmbeddingModelNameGeminiTextMultilingualEmbedding002 EmbeddingModelName = "gemini/text-multilingual-embedding-002"
const EmbeddingModelNameGeminiGeminiEmbedding001 EmbeddingModelName = "gemini/gemini-embedding-001"
string
Type EmbeddingConfigBaseTypeOptional

The type of the embedding configuration.

EmbeddingTimeMs int64Optional

Time spent generating embeddings in milliseconds (None for lexical queries)

IndexQueryTimeMs int64Optional

Time spent querying the vector index (OpenSearch) in milliseconds

RerankingModel stringOptional

Reranking model used (None if reranking not enabled)

RerankingTimeMs int64Optional

Time spent reranking results in milliseconds (None if reranking not enabled)

Vectors []VectorStoreQueryResponseVector

Array of matching documents

ID string

Document ID

Score float64

Similarity score indicating relevance

Content TextContentOptional

Text content for documents.

Text string

Text content to be embedded

Type TextContentTypeOptional

Content type identifier

Metadata map[string, any]Optional

Key-value metadata

Vector []float64Optional

Embedding vector (if requested)

Query Vectors

package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  response, err := client.VectorStores.Query(
    context.TODO(),
    "vector_store_name",
    sgpdev.VectorStoreQueryParams{
      Content: sgpdev.TextContentParam{
        Text: "text",
      },
    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", response.Metadata)
}
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
Returns Examples
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}