scan_factor and max_candidates, that let you trade off recall (search quality) for lower latency and higher throughput. By default, queries use internal heuristics that favor recall. If your application is latency-sensitive or needs higher QPS, you can tune these parameters to reduce the work done per query, or increase them for higher recall.
These parameters only take effect on dedicated read nodes indexes with dense vectors. On-demand indexes accept the parameters but ignore them. On indexes that store only sparse vectors, specifying either parameter returns an error. Using these parameters requires API version 2025-10 or later.
How scan_factor and max_candidates work
Dense vector search on dedicated read nodes uses a two-stage pipeline:- Scanning: Controlled by
scan_factor. For IVF-based indexes, the system scans a fraction of partitions determined byscan_factor / sqrt(num_partitions). A lowerscan_factorscans fewer partitions, producing fewer candidates faster. This parameter only affects IVF-based slabs; for other index architectures (e.g., smaller indexes using flat search),scan_factorhas no effect. - Reranking: Controlled by
max_candidates. The top candidates from the scanning stage are reranked by computing exact distances. More reranking improves recall but increases latency. This parameter applies to all index architectures.
You can set one or both per query. Omitting both preserves the current default behavior, so existing applications are unaffected.
Default max_candidates behavior
When max_candidates isn’t set, the system calculates an effective value using the following formula:
- If
top_k<= 1000:min(top_k * 10, 1000) - If
top_k> 1000:top_k - Then, a floor of 2500 is applied (the effective value is at least 2500)
top_k <= 2500), the effective default is 2500. This isn’t the maximum possible value. You can raise max_candidates (up to 100,000) to increase recall, or lower it (down to your query’s top_k) to reduce latency.
When you explicitly set max_candidates, the value you provide is used directly, bypassing the formula and the floor.
Impact on recall and performance
Lowerscan_factor or max_candidates values reduce the work done per query, which improves latency and throughput but may reduce recall. The tables below summarize benchmarked behavior on a 2.68M-vector index (1536 dimensions, cosine similarity). Actual results are dataset-dependent.
scan_factor benchmarks
Starting from the default (4.0), lowering scan_factor reduces the fraction of IVF partitions scanned:
Testing shows that lower
scan_factor values can reduce p50 and p99 latency by 30–50% or more.
Tuning max_candidates
Higher max_candidates improves recall by reranking more candidates but increases latency and reduces throughput; lower values favor speed. For guidance on choosing values, see Tuning guidance. We recommend benchmarking on your own dataset and workload to find the right balance. Use the Test your workload process to validate latency and recall.
Tuning guidance
Start with the defaults and adjust based on your workload requirements:- To optimize for throughput/latency: Lower
scan_factorfirst (from the default of 4.0). This has the most impact on IVF-based indexes. If you need further improvement, lowermax_candidatesbelow the default of 2500 (down to your query’stop_kvalue). - To optimize for recall: Raise
max_candidatesabove the default of 2500 (up to 100,000). This reranks more candidate vectors at the cost of higher latency.
Trade-offs to consider
- One parameter at a time:
scan_factorcontrols the scanning stage (IVF only) andmax_candidatescontrols the reranking stage (all index types). Tuning them independently makes it easier to isolate the effect. - Safe defaults: Omitting both parameters preserves existing behavior, so existing queries aren’t affected.
- Cost reduction: By achieving higher throughput per node, you may be able to serve the same query rate with fewer replicas.
Behavior by vector type
API and SDK examples
Both parameters are optional fields on thePOST /query request.
Validation errors
scan_factor and max_candidates don’t affect billing. Read costs for dedicated read nodes are based on provisioned capacity (node type, shards, and replicas), not per-query effort. By tuning these parameters to achieve higher throughput, you may be able to serve the same query rate with fewer provisioned replicas, reducing your overall cost.