Add your own model family
The runtime is built from plugins. A model family knows a model format: how to read the package, how to turn a request into model input and how to turn the model output into an answer. An engine runs the network itself, an accelerator drives a kind of hardware, and a profile decides numerics and batching. The built-in families, engines, accelerators and profiles register exactly the way yours will, so a plugin is a first-class part of the runtime.
You need a plugin when you want to serve a model the built-in families do not understand. You do not need one for Hugging Face ModernBERT or mmBERT classifiers and embedding models; those load as they are.
The example plugin
The repository contains a complete plugin small enough to read in one
sitting:
src/model-runtime/examples/third_party_plugin.
It adds a keyword "model" family and the engine that runs it, plus an
accelerator and a profile, so it shows all four kinds of plugin. A package is
one JSON file that maps labels to keywords; the family answers /v1/classify,
/v1/embeddings and /v1/rerank from keyword counts. The runtime's test
suite builds and installs its wheel, discovers it through its entry points and
serves it next to a decision model in one process, on every endpoint, and on
the example's own accelerator and profile.
Try it:
pip install ./src/model-runtime ./src/model-runtime/examples/third_party_plugin
mkdir -p /tmp/keywords && cat > /tmp/keywords/example_model.json <<'JSON'
{"format": "vllm-sr-example/1", "labels": ["billing", "shipping", "other"],
"keywords": {"billing": ["refund", "invoice", "charge"], "shipping": ["parcel", "delivery"]}}
JSON
vllm-sr-runtime serve /tmp/keywords --engine example_counts --device cpu --port 8100
curl -s localhost:8100/v1/classify -H 'content-type: application/json' \
-d '{"input": ["Please refund the invoice", "Where is my parcel?"]}'
vllm-sr-runtime plugins
The first request returns billing for the first text and shipping for the
second. vllm-sr-runtime plugins lists example_keywords among the families,
example_counts among the engines, example_host among the accelerators and
example_one_by_one among the profiles; GET /v1/models also shows the
distribution and version each came from.
The example's accelerator offers the host CPU as a device of its own, and its profile runs every request alone, in arrival order. Name them like the built-in ones:
vllm-sr-runtime serve /tmp/keywords --engine example_counts --device example_host --profile example_one_by_one --port 8100
vllm-sr serve passes the same names to the runtime, and the runtime picks
the example's engine by itself:
vllm-sr serve /tmp/keywords --device example_host --profile example_one_by_one --port 8100
The CLI checks only that --profile is a profile name. The runtime refuses a
device or profile it has no plugin for and lists the names it has.
Write your own
A plugin is an ordinary Python distribution.
1. The family. Subclass ModelFamily and LoadedModel from
vllm_sr_runtime.plugins.base:
| Method | What it does |
|---|---|
detect(package) | Cheap check: is this directory or repository yours? Read one small file, never the weights. |
verify(package) | Read the package's files, check them and return their identity (model_sha256), limits and licence. |
describe(package) | Say what the engine must run: the backbone, the weight files and the numerics. |
load(package, spec, engine_model) | Return the loaded model with its ModelInfo: the surfaces it serves, its heads and labels, its embedding and rerank views. |
plan_surface(surface, request) | Validate one request and turn it into work items. |
run(items) | One pass of the engine over a batch of items. |
finish_surface(plan, results) | Turn the results into the response body of that endpoint. |
Declare the endpoints you serve in surfaces and describe the plugin in
descriptor(); /v1/models shows it to clients.
A family that answers questions on /v1/decisions subclasses DecisionModel
from vllm_sr_runtime.plugins.decisions instead of LoadedModel. It writes
plan (turn a request's questions into work items) and answer (one
question's answer from its result); DecisionModel serves the endpoint, its
startup self-check and its per-question metrics.
A family can also ship models of its own, pinned to a revision: name a module
whose MODELS lists them (BuiltinModel entries with the repository,
revision, identity and recorded golden answers) in builtin_table. The
runtime then serves them by repository ID, or by bare model name while no
other organisation's built-in model has that name, and checks their answers
at startup. A table that fails to import pins nothing, and the runtime logs
why. A table that pins a repository another family's table pins, or lists
another family's model, stops every model of the process from loading,
built-in ones included, until you remove the conflict. Name the module that
writes tiny test packages in
fixture_writer, and vllm-sr-runtime fixture --family <name> writes one.
The built-in families declare both the same way.
2. The engine, if you need one. Most families reuse the built-in native
(PyTorch) or onnxruntime engine by describing their network in the
ModelSpec. Write an engine (Engine and EngineModel: supports, load,
forward or encode) only for a new kind of network or a new execution
library. Set auto_priority if engine: auto should try your engine before
others (lower first; the built-in native engine is 0); without it, auto
tries it after the engines that set one, by name.
3. An accelerator or a profile, if you need one. For new hardware,
subclass Accelerator (available, devices, torch_device, kernels), and
set auto_priority if device: auto may pick it; without it, only a request
that names the device uses it. For a new way to run requests together,
subclass Profile (plan, and bind for what it reads from the model).
4. Register it in your pyproject.toml:
[project.entry-points."vllm_sr_runtime.families"]
example_keywords = "vllm_sr_example.family:KeywordFamily"
[project.entry-points."vllm_sr_runtime.engines"]
example_counts = "vllm_sr_example.engine:CountsEngine"
[project.entry-points."vllm_sr_runtime.accelerators"]
example_host = "vllm_sr_example.accelerator:HostAccelerator"
[project.entry-points."vllm_sr_runtime.profiles"]
example_one_by_one = "vllm_sr_example.profile:OneByOneProfile"
The groups are vllm_sr_runtime.families, vllm_sr_runtime.engines,
vllm_sr_runtime.accelerators and vllm_sr_runtime.profiles. A name that is
already taken is refused at startup.
5. Make it fast. Two flags turn on the runtime's shared optimizations:
- set
cache_keyon work items whose result depends only on their content, so repeated inputs are answered from the result cache; - set
fuse_bundled_jobs = Trueon the loaded model when one pass can serve several requests that arrive together in a bundle.
6. Test it the way the example is tested: install the distribution, start
a runtime on a small package and check every endpoint against the OpenAPI
contract
(tests/test_third_party_plugin.py).
Use it from the router
Install your plugin next to the runtime the router uses, then bind a feature to a deployment of your model. For a runtime you run yourself, attach to it:
global:
model_catalog:
deployments:
ticket-topics:
provider: model_runtime
endpoint: http://runtime.internal:8100
served_name: keywords
A custom classifier signal can then read its labels; see
Classify requests. The router
checks the binding against the labels your model reports in /v1/models. A
deployment names your accelerator in device and your profile in profile
the same way; the router checks only their form, and the runtime refuses a
name it has no plugin for and lists the names it has.
Rules for plugins
- Plugins run inside the runtime process. Install only plugins you trust.
- The runtime never runs code shipped inside a model package. If a format needs code, that code belongs in your plugin.
- Keep the layers apart: a family never imports an engine, and an engine never reads a package. That is what lets your family run on a built-in engine, or your engine serve a built-in family.