跳到主要内容
版本:最新版

Add your own model family

The runtime is built from plugins. A model family knows a model format: how to read the package, how to turn a request into model input and how to turn the model output into an answer. An engine runs the network itself, an accelerator drives a kind of hardware, and a profile decides numerics and batching. The built-in families, engines, accelerators and profiles register exactly the way yours will, so a plugin is a first-class part of the runtime.

You need a plugin when you want to serve a model the built-in families do not understand. You do not need one for Hugging Face ModernBERT or mmBERT classifiers and embedding models; those load as they are.

The example plugin​

The repository contains a complete plugin small enough to read in one sitting: src/model-runtime/examples/third_party_plugin. It adds a keyword "model" family and the engine that runs it, plus an accelerator and a profile, so it shows all four kinds of plugin. A package is one JSON file that maps labels to keywords; the family answers /v1/classify, /v1/embeddings and /v1/rerank from keyword counts. The runtime's test suite builds and installs its wheel, discovers it through its entry points and serves it next to a decision model in one process, on every endpoint, and on the example's own accelerator and profile.

Try it:

pip install ./src/model-runtime ./src/model-runtime/examples/third_party_plugin
mkdir -p /tmp/keywords && cat > /tmp/keywords/example_model.json <<'JSON'
{"format": "vllm-sr-example/1", "labels": ["billing", "shipping", "other"],
"keywords": {"billing": ["refund", "invoice", "charge"], "shipping": ["parcel", "delivery"]}}
JSON
vllm-sr-runtime serve /tmp/keywords --engine example_counts --device cpu --port 8100
curl -s localhost:8100/v1/classify -H 'content-type: application/json' \
-d '{"input": ["Please refund the invoice", "Where is my parcel?"]}'
vllm-sr-runtime plugins

The first request returns billing for the first text and shipping for the second. vllm-sr-runtime plugins lists example_keywords among the families, example_counts among the engines, example_host among the accelerators and example_one_by_one among the profiles; GET /v1/models also shows the distribution and version each came from.

The example's accelerator offers the host CPU as a device of its own, and its profile runs every request alone, in arrival order. Name them like the built-in ones:

vllm-sr-runtime serve /tmp/keywords --engine example_counts --device example_host --profile example_one_by_one --port 8100

vllm-sr serve passes the same names to the runtime, and the runtime picks the example's engine by itself:

vllm-sr serve /tmp/keywords --device example_host --profile example_one_by_one --port 8100

The CLI checks only that --profile is a profile name. The runtime refuses a device or profile it has no plugin for and lists the names it has.

Write your own​

A plugin is an ordinary Python distribution.

1. The family. Subclass ModelFamily and LoadedModel from vllm_sr_runtime.plugins.base:

MethodWhat it does
detect(package)Cheap check: is this directory or repository yours? Read one small file, never the weights.
verify(package)Read the package's files, check them and return their identity (model_sha256), limits and licence.
describe(package)Say what the engine must run: the backbone, the weight files and the numerics.
load(package, spec, engine_model)Return the loaded model with its ModelInfo: the surfaces it serves, its heads and labels, its embedding and rerank views.
plan_surface(surface, request)Validate one request and turn it into work items.
run(items)One pass of the engine over a batch of items.
finish_surface(plan, results)Turn the results into the response body of that endpoint.

Declare the endpoints you serve in surfaces and describe the plugin in descriptor(); /v1/models shows it to clients.

A family that answers questions on /v1/decisions subclasses DecisionModel from vllm_sr_runtime.plugins.decisions instead of LoadedModel. It writes plan (turn a request's questions into work items) and answer (one question's answer from its result); DecisionModel serves the endpoint, its startup self-check and its per-question metrics.

A family can also ship models of its own, pinned to a revision: name a module whose MODELS lists them (BuiltinModel entries with the repository, revision, identity and recorded golden answers) in builtin_table. The runtime then serves them by repository ID, or by bare model name while no other organisation's built-in model has that name, and checks their answers at startup. A table that fails to import pins nothing, and the runtime logs why. A table that pins a repository another family's table pins, or lists another family's model, stops every model of the process from loading, built-in ones included, until you remove the conflict. Name the module that writes tiny test packages in fixture_writer, and vllm-sr-runtime fixture --family <name> writes one. The built-in families declare both the same way.

2. The engine, if you need one. Most families reuse the built-in native (PyTorch) or onnxruntime engine by describing their network in the ModelSpec. Write an engine (Engine and EngineModel: supports, load, forward or encode) only for a new kind of network or a new execution library. Set auto_priority if engine: auto should try your engine before others (lower first; the built-in native engine is 0); without it, auto tries it after the engines that set one, by name.

3. An accelerator or a profile, if you need one. For new hardware, subclass Accelerator (available, devices, torch_device, kernels), and set auto_priority if device: auto may pick it; without it, only a request that names the device uses it. For a new way to run requests together, subclass Profile (plan, and bind for what it reads from the model).

4. Register it in your pyproject.toml:

[project.entry-points."vllm_sr_runtime.families"]
example_keywords = "vllm_sr_example.family:KeywordFamily"

[project.entry-points."vllm_sr_runtime.engines"]
example_counts = "vllm_sr_example.engine:CountsEngine"

[project.entry-points."vllm_sr_runtime.accelerators"]
example_host = "vllm_sr_example.accelerator:HostAccelerator"

[project.entry-points."vllm_sr_runtime.profiles"]
example_one_by_one = "vllm_sr_example.profile:OneByOneProfile"

The groups are vllm_sr_runtime.families, vllm_sr_runtime.engines, vllm_sr_runtime.accelerators and vllm_sr_runtime.profiles. A name that is already taken is refused at startup.

5. Make it fast. Two flags turn on the runtime's shared optimizations:

  • set cache_key on work items whose result depends only on their content, so repeated inputs are answered from the result cache;
  • set fuse_bundled_jobs = True on the loaded model when one pass can serve several requests that arrive together in a bundle.

6. Test it the way the example is tested: install the distribution, start a runtime on a small package and check every endpoint against the OpenAPI contract (tests/test_third_party_plugin.py).

Use it from the router​

Install your plugin next to the runtime the router uses, then bind a feature to a deployment of your model. For a runtime you run yourself, attach to it:

global:
model_catalog:
deployments:
ticket-topics:
provider: model_runtime
endpoint: http://runtime.internal:8100
served_name: keywords

A custom classifier signal can then read its labels; see Classify requests. The router checks the binding against the labels your model reports in /v1/models. A deployment names your accelerator in device and your profile in profile the same way; the router checks only their form, and the runtime refuses a name it has no plugin for and lists the names it has.

Rules for plugins​

  • Plugins run inside the runtime process. Install only plugins you trust.
  • The runtime never runs code shipped inside a model package. If a format needs code, that code belongs in your plugin.
  • Keep the layers apart: a family never imports an engine, and an engine never reads a package. That is what lets your family run on a built-in engine, or your engine serve a built-in family.