跳到主要内容
版本:v0.4

快速开始

安装 vLLM Semantic Router,启动本地栈,并发送一条请求。

系统要求​

  • Linux、macOS 或 WSL2
  • Python 3.10 或更高版本
  • Docker;Linux 可以回退到 Podman

使用 curl 安装脚本时,可传入 --runtime podman 强制使用 Podman,或传入 --runtime skip 跳过容器运行时准备。例如:

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel stable --runtime skip

这些参数属于安装脚本。vllm-sr serve --runtime 用于选择容器运行时 (docker 或 podman);skip 不是 serve 的运行时选项。

安装​

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel dev

验证 CLI:

vllm-sr --version

curl 安装程序会自动启动栈。使用 pip 或 uv 安装后,用以下命令启动:

vllm-sr serve

打开 http://localhost:8700,添加模型端点,并激活生成的配置。Agent 可以通过 CLI 和 Router 管理 API 完成同样的工作,而无需使用控制面板。

发送请求​

curl http://localhost:8899/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "vllm-sr/auto",
"messages": [{"role": "user", "content": "Hello!"}]
}'

下一步​