快速开始
安装 vLLM Semantic Router,启动本地栈,并发送一条请求。
系统要求
- Linux、macOS 或 WSL2
- Python 3.10 或更高版本
- Docker;Linux 可以回退到 Podman
使用 curl 安装脚本时,可传入 --runtime podman 强制使用 Podman,或传入
--runtime skip 跳过容器运行时准备。例如:
curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel stable --runtime skip
这些参数属于安装脚本。vllm-sr serve --runtime 用于选择容器运行时
(docker 或 podman);skip 不是 serve 的运行时选项。
安装
- curl
- pip
- uv
- Agent
curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel dev
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade --pre vllm-sr
uv tool install --upgrade --prerelease allow vllm-sr
将此提示词复制到你的编码 Agent 中:
Install the development channel of vLLM Semantic Router, then configure and verify it on this machine by following the official skill: https://vllm-sr.ai/install/agent/vllm-sr/SKILL.md
该提示词指向公开、自包含的 vLLM SR agent skill。 工作流和安全边界见 使用 Agent 安装。
验证 CLI:
vllm-sr --version
curl 安装程序会自动启动栈。使用 pip 或 uv 安装后,用以下命令启动:
vllm-sr serve
打开 http://localhost:8700,添加模型端点,并激活生成的配置。Agent 可以通过 CLI 和 Router 管理 API 完成同样的工作,而无需使用控制面板。
发送请求
curl http://localhost:8899/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "vllm-sr/auto",
"messages": [{"role": "user", "content": "Hello!"}]
}'