Skip to main content
Version: Latest

Safety models

Safety models detect risks; routing decisions and plugins determine the action. Enabling a model alone does not block or redact a request.

CheckWhat it detectsConfigure the action
Prompt guardPrompt injection and jailbreaksJailbreak signals
PIIPersonal information in textPII signals
HallucinationAnswer claims unsupported by supplied contextHallucination plugin
Fact-checkWhether a request needs factual verificationFact-check signals

Prompt guard

To enable the maintained local guard, merge this into your configuration:

global:
model_catalog:
modules:
prompt_guard:
enabled: true
variant: mmbert32k
threshold: 0.7
on_error: block

Then add the jailbreak signal and decision that should handle a match. To use a separately hosted model, follow External services. The recipe's prompt_guard binding selects the model; the threshold and routing policy remain in their existing settings.

PII

Use a complete token-classification checkpoint with the matching PII label map. Select it through the recipe's pii_classifier binding with contract: token_spans.v1. Set its deployment and adapter as in In-process models, and provide the checkpoint's mapping_path. The PII guide covers entity thresholds and redaction.

External PII services must return scored entities and valid Unicode text positions. An invalid response is an error, not an empty successful scan.

Hallucination detection

The local detector checks an answer against its context and question. An optional NLI explainer checks whether a premise supports a hypothesis. Enable the maintained models with:

global:
model_catalog:
modules:
hallucination_mitigation:
enabled: true
detector:
backend: candle
model_ref: hallucination_detector
threshold: 0.82
explainer:
model_ref: hallucination_explainer
threshold: 0.9

These local models use Candle. A remote chat service can replace the detector; NLI still requires a supported local explainer. Configure how context is supplied and how detected spans are handled in the hallucination guide.

Handle failures and missing scores

A model error produces an unknown result. The decision's rules.on_unknown chooses no_match, match, or fail_request. Without that setting, prompt guard uses on_error: allow means no match, while block means a policy match. The consuming decision still determines the resulting action.

A chat verdict or policy fallback may have no confidence score. Diagnostics show confidence: null with confidence_available: false; this is different from a model score of zero. Test both model matches and service failures when setting the policy.

For custom safety checkpoints, see the training and export guide.

For access control and rate limits, see Security hardening.