跳到主要内容
版本:最新版(未发布)

Hallucination checks

Vela 1.0 Halu reads the context a request carried (tool results, retrieved documents), the user's question and the model's answer, and marks the spans of the answer that the context does not support. The hallucination plugin then adds a warning header, a note in the response body, or only records the result.

It checks support, not truth: an answer can be supported by wrong context, and an answer with no context to check is reported as unverified instead.

Turn it on​

Declare the observation as a signal and enforce it with the plugin on the routes that should be checked. Vela FactCheck decides first whether a request makes claims worth checking.

routing:
signals:
fact_check:
- name: needs_fact_check
description: Requests that make factual claims.
hallucination:
- name: ungrounded_claims
description: Claims the context does not support.
decisions:
- name: grounded-answers
priority: 100
rules:
operator: AND
conditions:
- type: fact_check
name: needs_fact_check
modelRefs:
- model: answer-model
plugins:
- type: hallucination
configuration:
enabled: true
hallucination_action: header
unverified_factual_action: header
include_hallucination_details: true
global:
model_catalog:
modules:
hallucination_mitigation:
enabled: true

Without task overrides, these judgments use the default Vela 2.0 deployment. The explicit Vela 1.0 Halu binding below selects the specialist instead. That specialist reads up to 8,192 tokens of context, question, and answer together; its context preparation prioritizes retaining the answer. Input limits and complete-coverage requirements still apply; see Long inputs.

Choose where it runs​

The binding is hallucination_detector and reads text spans of the answer (token_spans.v1):

global:
model_catalog:
deployments:
vela-halu:
provider: model_runtime
artifact: vllm-sr/Vela-1.0-Encoder-307M-Halu
device: cpu
input:
max_tokens: 8192
overflow: reject
bindings:
hallucination_detector:
deployment: vela-halu
contract: token_spans.v1

On Vela 2.0​

Vela 2.0 marks unsupported claims with its router span head, as a ready-made question about the answer. Bind hallucination_detector to a Vela 2.0 deployment instead:

global:
model_catalog:
deployments:
vela2:
provider: model_runtime
artifact: vllm-sr/Vela-2.0-0.3B
device: cpu
bindings:
hallucination_detector:
deployment: vela2
contract: token_spans.v1

The model reads the whole answer and applies its own calibrated threshold, so the detector's threshold applies to Vela 1.0 Halu only. The same deployment can answer the request's decision questions and its PII question.

Check it​

These worker-level examples run inside an environment containing vllm-srun (such as the Router image). Classify, embeddings, rerank and bundle are worker APIs; the instance frontend publishes System One and decision requests.

vllm-srun serve vllm-sr/Vela-1.0-Encoder-307M-Halu --device cpu --port 8100
curl -s localhost:8100/v1/classify -H 'content-type: application/json' -d '{
"input": [{"context": "The Eiffel Tower is 330 metres tall and stands in Paris.",
"question": "How tall is the Eiffel Tower?",
"answer": "The Eiffel Tower is 450 metres tall."}]
}'

The result lists the unsupported spans of the answer, here the height, with their character offsets in the answer. Through the router, the response carries the hallucination warning headers and x-vsr-matched-hallucination.

Earlier releases could add an NLI explanation per span. That explainer is retired; see Migrate.