Route, plan, filter-aware traverse, merge
The request first lands on whichever node the client is connected to. That node acts as the coordinator for this request. It parses the query, validates the filter, and determines which shards are involved. If the collection uses custom sharding and the request specifies a shard key, the coordinator routes to that one shard. Otherwise it fans out to all shards of the collection. Each shard then runs the query locally: it consults its payload indexes to plan the filter, prunes segments that cannot contain matching points, and runs a filter-aware HNSW traversal on the remaining segments. The filter is applied during the traversal, not after it, so the search expands its candidate list until it has enough points that pass the filter. Each segment returns its local top-k, each shard merges its segments' results, and the coordinator collects the per-shard results and merges them into a global top-k, which it returns to the client.
The most important internal detail is the interaction between the filter and the HNSW traversal. If Qdrant applied the filter after the ANN search, it could return fewer than k results when the filter is selective, and it would waste work on points that do not match. Instead, the filter is evaluated during graph traversal: when the search visits a candidate, it checks the filter, and only matching points are added to the result set. But a filter that rejects most candidates means the traversal has to explore more of the graph to find k matches, so the effective search breadth must increase. This is why filter-aware search can be slower and can drop recall if ef is not raised. Qdrant uses payload indexes to help: for a filter on an indexed field, the planner can identify a candidate set or a bitmap of matching points and use it to guide the traversal. For a filter on an unindexed field, the traversal has to check the payload for every visited point, which is more expensive. The second important detail is segment pruning: if a segment's payload index shows that no point in it can match the filter, the coordinator skips it entirely, which is a significant saving when the filter is selective.
Coordinator: parses the query, plans the filter, decides shard routing (single shard if shard key is given, fan-out otherwise).
Shard: plans the filter against payload indexes, prunes segments that cannot match, runs filter-aware HNSW on the rest.
Filter-aware traversal: the filter is applied during graph search, and ef is effectively expanded to find k matches.
Segment merge: each shard merges its segments' local top-k.
Global merge: the coordinator merges per-shard results into a global top-k, respecting offset and limit.
The trade-off is between filter selectivity and search cost. A highly selective filter on an unindexed field is the worst case: the traversal has to explore a large portion of the graph to find enough matches, and the cost can approach a full scan. A selective filter on an indexed field is much cheaper because the planner can use the index to guide or prune. The common mistake is assuming the filter is free. It is not - it changes the traversal cost and can reduce recall unless you raise ef. The second mistake is not creating a payload index on a field you filter on frequently, which forces the traversal to check the payload for every candidate. The third mistake is not accounting for fan-out: a query without a shard key on a cluster with many shards pays the cost of searching every shard, and the coordinator's merge cost grows with shard count. Version note: the query planner, the filter-aware traversal, and the way payload indexes are used have all evolved; the exact behavior under a selective filter differs between versions, and the introduction of the prefetch/fusion API changed how multi-stage requests are planned and merged.
Version-dependent: the query planner, filter-aware traversal behavior, and the exact way ef is interpreted under a filter have changed across releases. Some versions expand ef automatically when a filter is present; others require you to set it explicitly. The prefetch/fusion API, which lets you express multi-stage queries in a single request, is a recent addition and changes how the coordinator plans and merges. If you are tuning filtered search performance, benchmark on your version with your actual filter selectivity rather than relying on general guidance.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience