跳到主要内容
版本:最新版

Rerank documents

Vector search finds documents that are similar to a request; a reranker reads the request and each document together and scores how well the document answers it. Vela 1.0 Reranker reorders the candidates of the RAG plugin so the most relevant ones reach the prompt.

Turn it on​

Describe the reranker deployment, bind it as rag.reranker (it reads relevance scores, relevance_scores.v1) and opt a route's RAG plugin into rerank:

global:
model_catalog:
deployments:
document-ranker:
provider: model_runtime
artifact: vllm-sr/Vela-1.0-Encoder-307M-Reranker
device: cpu
input:
max_tokens: 4096
overflow: reject
routing:
model_bindings:
rag.reranker:
deployment: document-ranker
contract: relevance_scores.v1
pair_scorer:
layer: 22
dimension: 768
decisions:
- name: answer-from-docs
priority: 100
rules:
operator: AND
conditions: []
modelRefs:
- model: answer-model
plugins:
- type: rag
configuration:
enabled: true
backend: vectorstore
backend_config:
vector_store_id: vs-your-documents
top_k: 10
rerank:
top_k: 3
on_failure: block

top_k retrieves ten candidates and rerank.top_k keeps the best three. pair_scorer selects the size of the model: the full 22 layers and 768 dimensions are the most accurate; layers 3, 6 and 11 and smaller dimensions are faster.

The input limit covers the request and one document together. A pair over the limit is rejected, never shortened, and the plugin's on_failure decides what happens to the request.

Check it​

vllm-sr serve vllm-sr/Vela-1.0-Encoder-307M-Reranker --device cpu --port 8100
curl -s localhost:8100/v1/rerank -H 'content-type: application/json' -d '{
"query": "How do I reset my password?",
"documents": ["Our offices are closed on Sunday.", "Open Settings, then Security, then Reset password."],
"top_n": 2
}'

The results come back most relevant first, each with the document's index, its relevance_score and the raw logit. Scores rank documents for one request; they are not probabilities you can compare across requests.