## Query Vectors

`client.VectorStores.Query(ctx, vectorStoreName, body) (*VectorStoreQueryResponse, error)`

**post** `/v5/vector-stores/{vector_store_name}/query`

Query documents using similarity search with optional reranking.

Primary endpoint for semantic search, question-answering, and RAG (Retrieval-Augmented Generation)
applications. Returns documents ranked by relevance to the query text with similarity scores.

**Query Types:**

- `semantic` (default): Approximate nearest-neighbor search using HNSW over cosine similarity of document embeddings. Optimal for question-answering, conceptual search,
  and finding semantically related content without requiring exact keyword matches.
- `lexical`: Keyword-based text search (BM25 algorithm). Optimal for exact phrase matching, proper nouns,
  and scenarios where keyword presence is more important than semantic similarity.
- `hybrid`: Combines semantic and lexical approaches with weighted scoring. Provides maximum recall by
  identifying documents matching either semantically or lexically.

**Metadata Filtering:** Narrow the search scope by applying metadata filters (e.g., search only documents
where `category: "technical"`). Only indexed fields can be used for filtering.
Filters are applied before similarity search for optimal efficiency.

**Reranking (Advanced):** Optionally enhance result quality using a cross-encoder reranking model.
The reranker rescores the initial results using a more sophisticated model that evaluates the complete
query-document pair (not solely embeddings). This adds 100-500ms latency but significantly improves
precision for high-stakes applications.

**Reranking Strategy:** Set `top_k` higher than the desired final count (e.g., 50) to retrieve more
candidates from the initial search. Then configure `rerank_top_n` to the desired final count (e.g., 10)
to return only the most relevant documents after reranking. This two-stage approach maximizes both recall
and precision.

**Performance Metrics:** The response includes detailed timing breakdowns (embedding generation time,
index query time, reranking time) to facilitate search pipeline optimization and latency analysis.

**Similarity Scores:** Each result includes a `score` field indicating relevance. Higher scores indicate
greater relevance. Score ranges and semantics vary by query type (semantic scores use cosine similarity,
lexical scores use BM25, hybrid scores combine both approaches).

### Parameters

- `vectorStoreName string`

  The name of the vector store

- `body VectorStoreQueryParams`

  - `Content param.Field[TextContent]`

    Text content for documents.

  - `Filter param.Field[map[string, any]]`

    Metadata filter expression

  - `IncludeVectors param.Field[bool]`

    Include embedding vectors in response

  - `QueryType param.Field[VectorStoreQueryParamsQueryType]`

    Query type: semantic, lexical, or hybrid

    - `const VectorStoreQueryParamsQueryTypeSemantic VectorStoreQueryParamsQueryType = "semantic"`

    - `const VectorStoreQueryParamsQueryTypeLexical VectorStoreQueryParamsQueryType = "lexical"`

    - `const VectorStoreQueryParamsQueryTypeHybrid VectorStoreQueryParamsQueryType = "hybrid"`

  - `Rerank param.Field[bool]`

    [Deprecated: use rerank_config] Enable reranking of search results

  - `RerankConfig param.Field[VectorStoreQueryParamsRerankConfig]`

    Reranking configuration. Presence enables reranking; omit to disable. Pass an empty object ({}) to enable reranking with system defaults.

    - `Instruction string`

      Custom instruction for the reranking model (e.g., 'Given a medical question, retrieve relevant clinical passages'). Only applies to instruction-following rerankers like Qwen3.

    - `Model string`

      Reranking model to use (uses system default if not specified). Supported values depend on the selected provider: Launch cross-encoder names (e.g. 'cross-encoder/ms-marco-MiniLM-L-12-v2'), Vertex semantic-ranker names (e.g. 'semantic-ranker-default-004') when provider='vertex', or any model id the inference proxy serves when provider='proxy'.

    - `Provider string`

      Reranking provider to use. When omitted, the deployment default is used ('launch', or 'proxy' on ray-serve deployments configured for the OpenAI-compatible inference proxy). Set explicitly (e.g. 'vertex') to route to a specific provider on a deployment that has more than one configured. Requesting a provider that is not configured on the deployment returns a 400.

      - `const VectorStoreQueryParamsRerankConfigProviderLaunch VectorStoreQueryParamsRerankConfigProvider = "launch"`

      - `const VectorStoreQueryParamsRerankConfigProviderVertex VectorStoreQueryParamsRerankConfigProvider = "vertex"`

      - `const VectorStoreQueryParamsRerankConfigProviderProxy VectorStoreQueryParamsRerankConfigProvider = "proxy"`

    - `TopN int64`

      Number of results to keep after reranking (defaults to top_k)

    - `Type string`

      Reranking configuration type. Currently only 'base' is supported.

      - `const VectorStoreQueryParamsRerankConfigTypeBase VectorStoreQueryParamsRerankConfigType = "base"`

  - `RerankInstruction param.Field[string]`

    [Deprecated: use rerank_config.instruction] Custom instruction for reranker

  - `RerankModel param.Field[string]`

    [Deprecated: use rerank_config.model] Reranking model to use

  - `RerankTopN param.Field[int64]`

    [Deprecated: use rerank_config.top_n] Number of results after reranking

  - `TopK param.Field[int64]`

    Number of search results to return

### Returns

- `type VectorStoreQueryResponse struct{…}`

  Response for query operation.

  - `Metadata VectorStoreQueryResponseMetadata`

    Query execution metadata

    - `SearchType string`

      Type of search performed (semantic, lexical, hybrid)

    - `TotalQueryTimeMs int64`

      Total end-to-end query execution time in milliseconds

    - `EmbeddingConfig EmbeddingConfigUnion`

      Embedding configuration used for query vectorization. None for lexical queries on model-less stores.

      - `type EmbeddingConfigModelsAPI struct{…}`

        - `ModelDeploymentID string`

          The ID of the deployment of the created model in the Models API V3.

        - `Type ModelsAPI`

          The type of the embedding configuration.

          - `const ModelsAPIModelsAPI ModelsAPI = "models_api"`

      - `type EmbeddingConfigBase struct{…}`

        - `EmbeddingModel EmbeddingModelName`

          The name of the base embedding model to use. Either a known base model (EmbeddingModelName) or, in ray-serve deployments with NATIVE_OPENAI_EMBEDDING_GATEWAY enabled, any model id served by the OpenAI-compatible inference proxy (e.g. 'nomic-embed-text-v1.5'). For fully custom deployments, use type 'models_api' with a model_deployment_id.

          - `type EmbeddingModelName string`

            - `const EmbeddingModelNameSentenceTransformersAllMiniLmL12V2 EmbeddingModelName = "sentence-transformers/all-MiniLM-L12-v2"`

            - `const EmbeddingModelNameSentenceTransformersMultiQaDistilbertCosV1 EmbeddingModelName = "sentence-transformers/multi-qa-distilbert-cos-v1"`

            - `const EmbeddingModelNameOpenAITextEmbeddingAda002 EmbeddingModelName = "openai/text-embedding-ada-002"`

            - `const EmbeddingModelNameOpenAITextEmbedding3Small EmbeddingModelName = "openai/text-embedding-3-small"`

            - `const EmbeddingModelNameOpenAITextEmbedding3Large EmbeddingModelName = "openai/text-embedding-3-large"`

            - `const EmbeddingModelNameEmbedEnglishV3_0 EmbeddingModelName = "embed-english-v3.0"`

            - `const EmbeddingModelNameEmbedEnglishLightV3_0 EmbeddingModelName = "embed-english-light-v3.0"`

            - `const EmbeddingModelNameEmbedMultilingualV3_0 EmbeddingModelName = "embed-multilingual-v3.0"`

            - `const EmbeddingModelNameGeminiTextEmbedding005 EmbeddingModelName = "gemini/text-embedding-005"`

            - `const EmbeddingModelNameGeminiTextMultilingualEmbedding002 EmbeddingModelName = "gemini/text-multilingual-embedding-002"`

            - `const EmbeddingModelNameGeminiGeminiEmbedding001 EmbeddingModelName = "gemini/gemini-embedding-001"`

          - `string`

        - `Type EmbeddingConfigBaseType`

          The type of the embedding configuration.

          - `const EmbeddingConfigBaseTypeBase EmbeddingConfigBaseType = "base"`

    - `EmbeddingTimeMs int64`

      Time spent generating embeddings in milliseconds (None for lexical queries)

    - `IndexQueryTimeMs int64`

      Time spent querying the vector index (OpenSearch) in milliseconds

    - `RerankingModel string`

      Reranking model used (None if reranking not enabled)

    - `RerankingTimeMs int64`

      Time spent reranking results in milliseconds (None if reranking not enabled)

  - `Vectors []VectorStoreQueryResponseVector`

    Array of matching documents

    - `ID string`

      Document ID

    - `Score float64`

      Similarity score indicating relevance

    - `Content TextContent`

      Text content for documents.

      - `Text string`

        Text content to be embedded

      - `Type TextContentType`

        Content type identifier

        - `const TextContentTypeText TextContentType = "text"`

    - `Metadata map[string, any]`

      Key-value metadata

    - `Vector []float64`

      Embedding vector (if requested)

### Example

```go
package main

import (
  "context"
  "fmt"

  "github.com/scaleapi/sgp-dev-go"
  "github.com/scaleapi/sgp-dev-go/option"
)

func main() {
  client := sgpdev.NewClient(
    option.WithAPIKey("My API Key"),
    option.WithAccountID("My Account ID"),
  )
  response, err := client.VectorStores.Query(
    context.TODO(),
    "vector_store_name",
    sgpdev.VectorStoreQueryParams{
      Content: sgpdev.TextContentParam{
        Text: "text",
      },
    },
  )
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", response.Metadata)
}
```

#### Response

```json
{
  "metadata": {
    "search_type": "search_type",
    "total_query_time_ms": 0,
    "embedding_config": {
      "model_deployment_id": "model_deployment_id",
      "type": "models_api"
    },
    "embedding_time_ms": 0,
    "index_query_time_ms": 0,
    "reranking_model": "reranking_model",
    "reranking_time_ms": 0
  },
  "vectors": [
    {
      "id": "id",
      "score": 0,
      "content": {
        "text": "text",
        "type": "text"
      },
      "metadata": {
        "foo": "bar"
      },
      "vector": [
        0
      ]
    }
  ]
}
```
