feat: vision 后端不再预设默认值,初始置空需用户自选

- vision.ts:移除默认 ollama/qwen3-vl:8b,未配置后端或模型时明确报错并给配置入口
- VisionConfig.tsx + i18n:面板无默认选中,新增未配置提示
- README/vision.md:同步移除默认表述,标注 OLLAMA_VISION_MODEL 必填
This commit is contained in:
luckyyzh
2026-08-03 13:52:08 +08:00
parent ccb038988e
commit ac0f0cdc0c
6 changed files with 59 additions and 38 deletions
+2 -2
View File
@@ -1,6 +1,6 @@
--- ---
name: vision name: vision
description: 视觉子代理 —— 读取并描述图片(截图/图表/文档/照片),输出完整结构化描述(OCR/版式/语义),供不支持图片输入的主模型(如 DeepSeek)推理使用。后端默认本地 Ollama(qwen3-vl:8b),也可切换 OpenAI 兼容视觉 API description: 视觉子代理 —— 读取并描述图片(截图/图表/文档/照片),输出完整结构化描述(OCR/版式/语义),供不支持图片输入的主模型(如 DeepSeek)推理使用。后端可配置(本地 Ollama 或 OpenAI 兼容视觉 API),初始未配置需先在 WebUI 视觉面板或环境变量中设置
tools: vision tools: vision
subagentOnlyExtensions: ./resources/extensions/vision.ts subagentOnlyExtensions: ./resources/extensions/vision.ts
thinking: false thinking: false
@@ -16,7 +16,7 @@ defaultProgress: true
- 对每个图片路径调用一次 `vision`;相关图片可一次传入多张。 - 对每个图片路径调用一次 `vision`;相关图片可一次传入多张。
- 工具返回的是视觉模型的转录:忠实转达,OCR 文字逐字保留,不要改写或脑补。 - 工具返回的是视觉模型的转录:忠实转达,OCR 文字逐字保留,不要改写或脑补。
- 工具报错时(文件不存在 / 后端未配置 / 模型未拉取)如实报告,并给出明确的修复提示(如 `ollama pull qwen3-vl:8b`,或检查 `VISION_OPENAI_*` 环境变量)。 - 工具报错时(文件不存在 / 后端未配置 / 模型未拉取)如实报告,并给出明确的修复提示(如安装 Ollama 并拉取视觉模型,或检查 `VISION_OPENAI_*` 环境变量)。
- 输出保持结构化:多图按图分组,先给结论性总结,再附关键细节;文字类图片保证转录完整。 - 输出保持结构化:多图按图分组,先给结论性总结,再附关键细节;文字类图片保证转录完整。
主会话(通常是 DeepSeek 这类纯文本模型)看不到图片,完全依赖你的描述,完整性优先。 主会话(通常是 DeepSeek 这类纯文本模型)看不到图片,完全依赖你的描述,完整性优先。
+14 -10
View File
@@ -143,20 +143,22 @@ data/workspaces/default/ 默认工作目录
| `@upstash/context7-pi@0.1.2` | 查询当前库、框架、SDK 和 API 文档 | 模型先解析库 ID,再按需查询文档;无 Key 可使用公共限额 | | `@upstash/context7-pi@0.1.2` | 查询当前库、框架、SDK 和 API 文档 | 模型先解析库 ID,再按需查询文档;无 Key 可使用公共限额 |
| `@narumitw/pi-retry@0.31.0` | 识别瞬时供应商错误和卡住的流 | 复用 Pi 内置重试;默认 180 秒无事件视为停滞,不增加正常请求的模型调用 | | `@narumitw/pi-retry@0.31.0` | 识别瞬时供应商错误和卡住的流 | 复用 Pi 内置重试;默认 180 秒无事件视为停滞,不增加正常请求的模型调用 |
| `resources/extensions/searxng-search.ts` | 用户自有 SearXNG 的 `web_search` | 配置 `SEARXNG_URL` 与 `SEARXNG_TOKEN` 后,模型对时效性或明确搜索请求自动调用 | | `resources/extensions/searxng-search.ts` | 用户自有 SearXNG 的 `web_search` | 配置 `SEARXNG_URL` 与 `SEARXNG_TOKEN` 后,模型对时效性或明确搜索请求自动调用 |
| `resources/extensions/vision.ts` | 文本主模型(如 DeepSeek)的识图工具 `vision`(双后端) | 派 `vision` 子代理或直接让模型调用工具,返回 OCR/版式/语义文本;后端默认本地 Ollama(`qwen3-vl:8b`),可切 OpenAI 兼容视觉 API | | `resources/extensions/vision.ts` | 文本主模型(如 DeepSeek)的识图工具 `vision`(双后端) | 派 `vision` 子代理或直接让模型调用工具,返回 OCR/版式/语义文本;后端初始未配置(本地 Ollama 或 OpenAI 兼容视觉 API 任选),需在「视觉」标签页或环境变量中自行设置 |
Windows 的 Playwright 不下载独立 Chromium;首次 `setup` 只缓存 MCP 的 Node.js 包,浏览器执行使用系统 Edge。macOS 的 setup 不安装或启用 Playwright;如需浏览器自动化,可在 Web UI 的 MCP 面板中手动添加并配置。 Windows 的 Playwright 不下载独立 Chromium;首次 `setup` 只缓存 MCP 的 Node.js 包,浏览器执行使用系统 Edge。macOS 的 setup 不安装或启用 Playwright;如需浏览器自动化,可在 Web UI 的 MCP 面板中手动添加并配置。
#### 视觉子代理(vision) #### 视觉子代理(vision)
DeepSeek 等纯文本模型不能接收图片。仓库内置 `vision` 子代理(`.agents/vision.md`):它通过 `vision` 工具调用视觉模型读取图片,把完整 OCR、版式结构与语义描述返回给主模型,主模型基于文本继续推理。视觉后端可插拔,默认本地 Ollama,也支持任意 OpenAI 兼容视觉 API——没有本地部署条件时可直接用云服务。 DeepSeek 等纯文本模型不能接收图片。仓库内置 `vision` 子代理(`.agents/vision.md`):它通过 `vision` 工具调用视觉模型读取图片,把完整 OCR、版式结构与语义描述返回给主模型,主模型基于文本继续推理。视觉后端可插拔(本地 Ollama 或任意 OpenAI 兼容视觉 API)——仓库**不预设默认后端**,首次使用前需自行选择并配置。
配置入口:WebUI 左下角「模型」面板内的「视觉」标签页(写入 `data/agent/vision.json`),保存后**下次识图请求立即生效**,无需重启;环境变量优先级高于面板配置。 配置入口:WebUI 左下角「模型」面板内的「视觉」标签页(写入 `data/agent/vision.json`),保存后**下次识图请求立即生效**,无需重启;环境变量优先级高于面板配置。
**后端一:本地 Ollama(默认,免费私密)** > 注意:视觉后端初始**未预设默认值**。未配置时识图会报错并提示配置入口;两种后端二选一即可。
- 前置:本机安装 [Ollama](https://ollama.com) 并 `ollama pull qwen3-vl:8b`。 **后端一:本地 Ollama(免费私密)**
- 环境变量:`OLLAMA_HOST`(默认 `http://localhost:11434`)、`OLLAMA_VISION_MODEL`(默认 `qwen3-vl:8b`)、`OLLAMA_VISION_KEEP_ALIVE`(默认 `-1` 常驻显存,避免每次识图冷加载大模型;也可设 `30m` 等时长)。
- 前置:本机安装 [Ollama](https://ollama.com) 并拉取一个支持视觉的模型(如 `ollama pull qwen3-vl:8b`)。
- 环境变量:`OLLAMA_HOST`(默认 `http://localhost:11434`)、`OLLAMA_VISION_MODEL`(**必填**,指定视觉模型)、`OLLAMA_VISION_KEEP_ALIVE`(默认 `-1` 常驻显存,避免每次识图冷加载大模型;也可设 `30m` 等时长)。
**后端二:OpenAI 兼容视觉 API** **后端二:OpenAI 兼容视觉 API**
@@ -385,20 +387,22 @@ Versions are pinned in the platform defaults under `config/`: Windows uses `mcp.
| `@upstash/context7-pi@0.1.2` | Current library, framework, SDK, API docs | Resolves a library ID and queries docs when needed; public quota works without a key | | `@upstash/context7-pi@0.1.2` | Current library, framework, SDK, API docs | Resolves a library ID and queries docs when needed; public quota works without a key |
| `@narumitw/pi-retry@0.31.0` | Transient provider and stalled-stream classification | Uses Pi's built-in retry path; 180 seconds without events is a stall; no extra normal model calls | | `@narumitw/pi-retry@0.31.0` | Transient provider and stalled-stream classification | Uses Pi's built-in retry path; 180 seconds without events is a stall; no extra normal model calls |
| `resources/extensions/searxng-search.ts` | `web_search` against a user-owned SearXNG proxy | After `SEARXNG_URL` and `SEARXNG_TOKEN` are set, the model calls it for current or explicit search requests | | `resources/extensions/searxng-search.ts` | `web_search` against a user-owned SearXNG proxy | After `SEARXNG_URL` and `SEARXNG_TOKEN` are set, the model calls it for current or explicit search requests |
| `resources/extensions/vision.ts` | `vision` — image description for text-only models (e.g. DeepSeek), dual backend | Ask the `vision` subagent or call the tool directly; returns OCR/layout/semantics as text; backend defaults to local Ollama (`qwen3-vl:8b`) and can switch to any OpenAI-compatible vision API | | `resources/extensions/vision.ts` | `vision` — image description for text-only models (e.g. DeepSeek), dual backend | Ask the `vision` subagent or call the tool directly; returns OCR/layout/semantics as text; no backend is preconfigured (local Ollama or any OpenAI-compatible vision API) — set one in the Vision tab or via env vars |
On Windows, Playwright never downloads a standalone Chromium: setup caches only its Node package and browser execution uses system Edge. On macOS, setup does not install or enable Playwright; add it manually through the MCP panel if browser automation is needed. On Windows, Playwright never downloads a standalone Chromium: setup caches only its Node package and browser execution uses system Edge. On macOS, setup does not install or enable Playwright; add it manually through the MCP panel if browser automation is needed.
#### Vision subagent #### Vision subagent
Text-only models such as DeepSeek cannot receive image attachments. The repository ships a `vision` subagent (`.agents/vision.md`) that calls a vision model through the `vision` tool and returns a full OCR, layout, and semantic description the main model can reason over. The vision backend is pluggable: local Ollama by default, or any OpenAI-compatible vision API for users who cannot run a local model. Text-only models such as DeepSeek cannot receive image attachments. The repository ships a `vision` subagent (`.agents/vision.md`) that calls a vision model through the `vision` tool and returns a full OCR, layout, and semantic description the main model can reason over. The vision backend is pluggable (local Ollama or any OpenAI-compatible vision API) — the repository does **not** ship a default backend; pick and configure one before first use.
Configuration: the “Vision” tab inside the “Models” panel in the lower-left Web UI (writes `data/agent/vision.json`). Saved config takes effect on the **next image request** — no restart needed; environment variables take precedence over the panel. Configuration: the “Vision” tab inside the “Models” panel in the lower-left Web UI (writes `data/agent/vision.json`). Saved config takes effect on the **next image request** — no restart needed; environment variables take precedence over the panel.
**Backend 1: local Ollama (default, free and private)** > Note: no backend is preconfigured by default. Image requests fail with a configuration hint until you pick one — choose either backend below.
- Prerequisite: install [Ollama](https://ollama.com) and run `ollama pull qwen3-vl:8b`. **Backend 1: local Ollama (free and private)**
- Env: `OLLAMA_HOST` (default `http://localhost:11434`), `OLLAMA_VISION_MODEL` (default `qwen3-vl:8b`), `OLLAMA_VISION_KEEP_ALIVE` (default `-1` — keep the model resident in VRAM to avoid cold-loading it on every transcription; can be set to e.g. `30m`).
- Prerequisite: install [Ollama](https://ollama.com) and pull a vision-capable model (e.g. `ollama pull qwen3-vl:8b`).
- Env: `OLLAMA_HOST` (default `http://localhost:11434`), `OLLAMA_VISION_MODEL` (**required** — the vision model), `OLLAMA_VISION_KEEP_ALIVE` (default `-1` — keep the model resident in VRAM to avoid cold-loading it on every transcription; can be set to e.g. `30m`).
**Backend 2: OpenAI-compatible vision API** **Backend 2: OpenAI-compatible vision API**
+18 -1
View File
@@ -136,7 +136,7 @@ export function VisionConfigContent() {
} }
}, [config]); }, [config]);
const backend = config.backend ?? "ollama"; const backend = config.backend;
const ollama = config.ollama ?? {}; const ollama = config.ollama ?? {};
const openai = config.openai ?? {}; const openai = config.openai ?? {};
@@ -205,6 +205,23 @@ export function VisionConfigContent() {
))} ))}
</div> </div>
{/* No backend selected yet */}
{!backend && (
<div
style={{
fontSize: 11,
color: "var(--text-muted)",
lineHeight: 1.5,
background: "var(--bg-panel)",
border: "1px dashed var(--border)",
borderRadius: 6,
padding: "8px 10px",
}}
>
{t("vision.notConfigured")}
</div>
)}
{/* Ollama settings */} {/* Ollama settings */}
{backend === "ollama" && ( {backend === "ollama" && (
<div style={{ display: "flex", flexDirection: "column", gap: 10 }}> <div style={{ display: "flex", flexDirection: "column", gap: 10 }}>
+1
View File
@@ -277,6 +277,7 @@ export const enLocale: LocalePlugin = {
"vision.openai.model": "Model", "vision.openai.model": "Model",
"vision.openai.modelPlaceholder": "e.g. gpt-4o-mini / glm-4.5v", "vision.openai.modelPlaceholder": "e.g. gpt-4o-mini / glm-4.5v",
"vision.effective": "Saved config takes effect on the next image request — no restart needed. Uploaded images are auto-transcribed into text for text-only main models (e.g. DeepSeek).", "vision.effective": "Saved config takes effect on the next image request — no restart needed. Uploaded images are auto-transcribed into text for text-only main models (e.g. DeepSeek).",
"vision.notConfigured": "No backend selected yet — image requests fail until you pick one and fill in its details.",
"vision.loadFailed": "Failed to load / save config", "vision.loadFailed": "Failed to load / save config",
"vision.save": "Save", "vision.save": "Save",
"vision.saved": "Saved — takes effect on the next image request", "vision.saved": "Saved — takes effect on the next image request",
+1
View File
@@ -277,6 +277,7 @@ export const zhCNLocale: LocalePlugin = {
"vision.openai.model": "模型", "vision.openai.model": "模型",
"vision.openai.modelPlaceholder": "如 gpt-4o-mini / glm-4.5v", "vision.openai.modelPlaceholder": "如 gpt-4o-mini / glm-4.5v",
"vision.effective": "保存后下次识图请求立即生效,无需重启。上传图片会自动转录为文本,供纯文本主模型(如 DeepSeek)使用。", "vision.effective": "保存后下次识图请求立即生效,无需重启。上传图片会自动转录为文本,供纯文本主模型(如 DeepSeek)使用。",
"vision.notConfigured": "尚未选择后端——未配置时识图请求会报错,请先选择后端并填写对应参数。",
"vision.loadFailed": "配置加载/保存失败", "vision.loadFailed": "配置加载/保存失败",
"vision.save": "保存", "vision.save": "保存",
"vision.saved": "已保存——下次识图请求生效", "vision.saved": "已保存——下次识图请求生效",
+23 -25
View File
@@ -1,10 +1,11 @@
/** /**
* Vision extension: describe images for text-only main models (e.g. DeepSeek). * Vision extension: describe images for text-only main models (e.g. DeepSeek).
* *
* Two interchangeable backends, selected by VISION_BACKEND (default "ollama"): * Two interchangeable backends, selected by VISION_BACKEND. No backend is
* preconfigured — users must pick one (Web UI vision panel or env vars):
* - "ollama": local Ollama vision model via the native /api/chat endpoint. * - "ollama": local Ollama vision model via the native /api/chat endpoint.
* OLLAMA_HOST (default http://localhost:11434) * OLLAMA_HOST (default http://localhost:11434)
* OLLAMA_VISION_MODEL (default qwen3-vl:8b) * OLLAMA_VISION_MODEL (required, e.g. qwen3-vl:8b)
* - "openai": any OpenAI-compatible vision API (cloud or self-hosted). * - "openai": any OpenAI-compatible vision API (cloud or self-hosted).
* VISION_OPENAI_BASE_URL e.g. https://api.openai.com/v1 * VISION_OPENAI_BASE_URL e.g. https://api.openai.com/v1
* VISION_OPENAI_API_KEY * VISION_OPENAI_API_KEY
@@ -59,7 +60,7 @@ const visionParams = Type.Object({
backend: Type.Optional( backend: Type.Optional(
Type.Union([Type.Literal("ollama"), Type.Literal("openai")], { Type.Union([Type.Literal("ollama"), Type.Literal("openai")], {
description: description:
"视觉后端:ollama(本地,默认)或 openai(OpenAI 兼容 API)。默认取 VISION_BACKEND 环境变量", "视觉后端:ollama(本地)或 openai(OpenAI 兼容 API)。默认取 VISION_BACKEND 环境变量,未配置时报错",
}), }),
), ),
model: Type.Optional( model: Type.Optional(
@@ -188,7 +189,7 @@ async function describeWithOllama(
if (!content) { if (!content) {
throw new Error( throw new Error(
`vision: Ollama model ${model} returned an empty response. ` + `vision: Ollama model ${model} returned an empty response. ` +
"Check `ollama pull qwen3-vl:8b` and that the model supports vision.", "Check the model tag and that it supports vision.",
); );
} }
return content; return content;
@@ -291,12 +292,6 @@ interface VisionFileConfig {
openai?: { baseUrl?: string; apiKey?: string; model?: string }; openai?: { baseUrl?: string; apiKey?: string; model?: string };
} }
const VISION_CONFIG_DEFAULTS: VisionFileConfig = {
backend: "ollama",
ollama: { host: "http://localhost:11434", model: "qwen3-vl:8b" },
openai: { baseUrl: "", apiKey: "", model: "" },
};
function visionConfigPath(): string { function visionConfigPath(): string {
const agentDir = process.env.PI_CODING_AGENT_DIR?.trim(); const agentDir = process.env.PI_CODING_AGENT_DIR?.trim();
return ( return (
@@ -322,27 +317,30 @@ async function describeImages(
): Promise<string> { ): Promise<string> {
const fileConfig = await loadVisionFileConfig(); const fileConfig = await loadVisionFileConfig();
const backend = const backend =
options.backend ?? options.backend ?? envOr("VISION_BACKEND", fileConfig.backend ?? "");
envOr( if (backend !== "ollama" && backend !== "openai") {
"VISION_BACKEND", throw new Error(
fileConfig.backend ?? VISION_CONFIG_DEFAULTS.backend ?? "ollama", "vision: no vision backend configured. Pick one in the Web UI vision panel " +
"(lower-left → Models → Vision) or set VISION_BACKEND=ollama|openai plus " +
"the backend's address/model (see the OLLAMA_* / VISION_OPENAI_* env vars).",
); );
}
const prompt = options.prompt ?? DEFAULT_PROMPT; const prompt = options.prompt ?? DEFAULT_PROMPT;
if (backend === "ollama") { if (backend === "ollama") {
const baseUrl = envOr( const baseUrl = envOr(
"OLLAMA_HOST", "OLLAMA_HOST",
fileConfig.ollama?.host ?? fileConfig.ollama?.host ?? "http://localhost:11434",
VISION_CONFIG_DEFAULTS.ollama?.host ??
"http://localhost:11434",
); );
const model = const model =
options.model ?? options.model ??
envOr( envOr("OLLAMA_VISION_MODEL", fileConfig.ollama?.model ?? "");
"OLLAMA_VISION_MODEL", if (!model) {
fileConfig.ollama?.model ?? throw new Error(
VISION_CONFIG_DEFAULTS.ollama?.model ?? "vision: the ollama backend needs a vision model. Set it in the Web UI " +
"qwen3-vl:8b", "vision panel or via OLLAMA_VISION_MODEL (pull one first, e.g. " +
"`ollama pull qwen3-vl:8b`).",
); );
}
return describeWithOllama(baseUrl, model, prompt, images, options.signal); return describeWithOllama(baseUrl, model, prompt, images, options.signal);
} }
const baseUrl = envOr( const baseUrl = envOr(
@@ -471,8 +469,8 @@ export default function visionExtension(pi: ExtensionAPI) {
description: description:
"用视觉模型描述本地图片并返回详细文本(完整 OCR、版式结构、语义总结)。" + "用视觉模型描述本地图片并返回详细文本(完整 OCR、版式结构、语义总结)。" +
"适用于主模型不支持图片输入(如 DeepSeek)时查看截图/图表/文档/照片。" + "适用于主模型不支持图片输入(如 DeepSeek)时查看截图/图表/文档/照片。" +
"后端可配置:ollama(默认,本地 qwen3-vl:8b,免费私密)或 openai(任意 OpenAI 兼容视觉 API," + "后端可配置:ollama(本地)或 openai(任意 OpenAI 兼容视觉 API),初始未配置需先设置:" +
"设 VISION_BACKEND=openai + VISION_OPENAI_BASE_URL/API_KEY/MODEL)。", "WebUI 视觉面板(左下角→模型→视觉)或环境变量(VISION_BACKEND + 对应地址/模型/密钥)。",
promptSnippet: promptSnippet:
"Describe local images using a vision model (Ollama or OpenAI-compatible)", "Describe local images using a vision model (Ollama or OpenAI-compatible)",
promptGuidelines: [ promptGuidelines: [
@@ -497,7 +495,7 @@ export default function visionExtension(pi: ExtensionAPI) {
return { return {
content: [{ type: "text", text: content }], content: [{ type: "text", text: content }],
details: { details: {
backend: params.backend ?? envOr("VISION_BACKEND", "ollama"), backend: params.backend ?? envOr("VISION_BACKEND", ""),
model: params.model ?? undefined, model: params.model ?? undefined,
imageCount: images.length, imageCount: images.length,
}, },