diff --git a/.agents/vision.md b/.agents/vision.md index e72a994..3096b16 100644 --- a/.agents/vision.md +++ b/.agents/vision.md @@ -1,6 +1,6 @@ --- name: vision -description: 视觉子代理 —— 读取并描述图片(截图/图表/文档/照片),输出完整结构化描述(OCR/版式/语义),供不支持图片输入的主模型(如 DeepSeek)推理使用。后端默认本地 Ollama(qwen3-vl:8b),也可切换 OpenAI 兼容视觉 API +description: 视觉子代理 —— 读取并描述图片(截图/图表/文档/照片),输出完整结构化描述(OCR/版式/语义),供不支持图片输入的主模型(如 DeepSeek)推理使用。后端可配置(本地 Ollama 或 OpenAI 兼容视觉 API),初始未配置需先在 WebUI 视觉面板或环境变量中设置 tools: vision subagentOnlyExtensions: ./resources/extensions/vision.ts thinking: false @@ -16,7 +16,7 @@ defaultProgress: true - 对每个图片路径调用一次 `vision`;相关图片可一次传入多张。 - 工具返回的是视觉模型的转录:忠实转达,OCR 文字逐字保留,不要改写或脑补。 -- 工具报错时(文件不存在 / 后端未配置 / 模型未拉取)如实报告,并给出明确的修复提示(如 `ollama pull qwen3-vl:8b`,或检查 `VISION_OPENAI_*` 环境变量)。 +- 工具报错时(文件不存在 / 后端未配置 / 模型未拉取)如实报告,并给出明确的修复提示(如安装 Ollama 并拉取视觉模型,或检查 `VISION_OPENAI_*` 环境变量)。 - 输出保持结构化:多图按图分组,先给结论性总结,再附关键细节;文字类图片保证转录完整。 主会话(通常是 DeepSeek 这类纯文本模型)看不到图片,完全依赖你的描述,完整性优先。 diff --git a/README.md b/README.md index 0c250d0..58b3709 100644 --- a/README.md +++ b/README.md @@ -143,20 +143,22 @@ data/workspaces/default/ 默认工作目录 | `@upstash/context7-pi@0.1.2` | 查询当前库、框架、SDK 和 API 文档 | 模型先解析库 ID,再按需查询文档;无 Key 可使用公共限额 | | `@narumitw/pi-retry@0.31.0` | 识别瞬时供应商错误和卡住的流 | 复用 Pi 内置重试;默认 180 秒无事件视为停滞,不增加正常请求的模型调用 | | `resources/extensions/searxng-search.ts` | 用户自有 SearXNG 的 `web_search` | 配置 `SEARXNG_URL` 与 `SEARXNG_TOKEN` 后,模型对时效性或明确搜索请求自动调用 | -| `resources/extensions/vision.ts` | 文本主模型(如 DeepSeek)的识图工具 `vision`(双后端) | 派 `vision` 子代理或直接让模型调用工具,返回 OCR/版式/语义文本;后端默认本地 Ollama(`qwen3-vl:8b`),可切 OpenAI 兼容视觉 API | +| `resources/extensions/vision.ts` | 文本主模型(如 DeepSeek)的识图工具 `vision`(双后端) | 派 `vision` 子代理或直接让模型调用工具,返回 OCR/版式/语义文本;后端初始未配置(本地 Ollama 或 OpenAI 兼容视觉 API 任选),需在「视觉」标签页或环境变量中自行设置 | Windows 的 Playwright 不下载独立 Chromium;首次 `setup` 只缓存 MCP 的 Node.js 包,浏览器执行使用系统 Edge。macOS 的 setup 不安装或启用 Playwright;如需浏览器自动化,可在 Web UI 的 MCP 面板中手动添加并配置。 #### 视觉子代理(vision) -DeepSeek 等纯文本模型不能接收图片。仓库内置 `vision` 子代理(`.agents/vision.md`):它通过 `vision` 工具调用视觉模型读取图片,把完整 OCR、版式结构与语义描述返回给主模型,主模型基于文本继续推理。视觉后端可插拔,默认本地 Ollama,也支持任意 OpenAI 兼容视觉 API——没有本地部署条件时可直接用云服务。 +DeepSeek 等纯文本模型不能接收图片。仓库内置 `vision` 子代理(`.agents/vision.md`):它通过 `vision` 工具调用视觉模型读取图片,把完整 OCR、版式结构与语义描述返回给主模型,主模型基于文本继续推理。视觉后端可插拔(本地 Ollama 或任意 OpenAI 兼容视觉 API)——仓库**不预设默认后端**,首次使用前需自行选择并配置。 配置入口:WebUI 左下角「模型」面板内的「视觉」标签页(写入 `data/agent/vision.json`),保存后**下次识图请求立即生效**,无需重启;环境变量优先级高于面板配置。 -**后端一:本地 Ollama(默认,免费私密)** +> 注意:视觉后端初始**未预设默认值**。未配置时识图会报错并提示配置入口;两种后端二选一即可。 -- 前置:本机安装 [Ollama](https://ollama.com) 并 `ollama pull qwen3-vl:8b`。 -- 环境变量:`OLLAMA_HOST`(默认 `http://localhost:11434`)、`OLLAMA_VISION_MODEL`(默认 `qwen3-vl:8b`)、`OLLAMA_VISION_KEEP_ALIVE`(默认 `-1` 常驻显存,避免每次识图冷加载大模型;也可设 `30m` 等时长)。 +**后端一:本地 Ollama(免费私密)** + +- 前置:本机安装 [Ollama](https://ollama.com) 并拉取一个支持视觉的模型(如 `ollama pull qwen3-vl:8b`)。 +- 环境变量:`OLLAMA_HOST`(默认 `http://localhost:11434`)、`OLLAMA_VISION_MODEL`(**必填**,指定视觉模型)、`OLLAMA_VISION_KEEP_ALIVE`(默认 `-1` 常驻显存,避免每次识图冷加载大模型;也可设 `30m` 等时长)。 **后端二:OpenAI 兼容视觉 API** @@ -385,20 +387,22 @@ Versions are pinned in the platform defaults under `config/`: Windows uses `mcp. | `@upstash/context7-pi@0.1.2` | Current library, framework, SDK, API docs | Resolves a library ID and queries docs when needed; public quota works without a key | | `@narumitw/pi-retry@0.31.0` | Transient provider and stalled-stream classification | Uses Pi's built-in retry path; 180 seconds without events is a stall; no extra normal model calls | | `resources/extensions/searxng-search.ts` | `web_search` against a user-owned SearXNG proxy | After `SEARXNG_URL` and `SEARXNG_TOKEN` are set, the model calls it for current or explicit search requests | -| `resources/extensions/vision.ts` | `vision` — image description for text-only models (e.g. DeepSeek), dual backend | Ask the `vision` subagent or call the tool directly; returns OCR/layout/semantics as text; backend defaults to local Ollama (`qwen3-vl:8b`) and can switch to any OpenAI-compatible vision API | +| `resources/extensions/vision.ts` | `vision` — image description for text-only models (e.g. DeepSeek), dual backend | Ask the `vision` subagent or call the tool directly; returns OCR/layout/semantics as text; no backend is preconfigured (local Ollama or any OpenAI-compatible vision API) — set one in the Vision tab or via env vars | On Windows, Playwright never downloads a standalone Chromium: setup caches only its Node package and browser execution uses system Edge. On macOS, setup does not install or enable Playwright; add it manually through the MCP panel if browser automation is needed. #### Vision subagent -Text-only models such as DeepSeek cannot receive image attachments. The repository ships a `vision` subagent (`.agents/vision.md`) that calls a vision model through the `vision` tool and returns a full OCR, layout, and semantic description the main model can reason over. The vision backend is pluggable: local Ollama by default, or any OpenAI-compatible vision API for users who cannot run a local model. +Text-only models such as DeepSeek cannot receive image attachments. The repository ships a `vision` subagent (`.agents/vision.md`) that calls a vision model through the `vision` tool and returns a full OCR, layout, and semantic description the main model can reason over. The vision backend is pluggable (local Ollama or any OpenAI-compatible vision API) — the repository does **not** ship a default backend; pick and configure one before first use. Configuration: the “Vision” tab inside the “Models” panel in the lower-left Web UI (writes `data/agent/vision.json`). Saved config takes effect on the **next image request** — no restart needed; environment variables take precedence over the panel. -**Backend 1: local Ollama (default, free and private)** +> Note: no backend is preconfigured by default. Image requests fail with a configuration hint until you pick one — choose either backend below. -- Prerequisite: install [Ollama](https://ollama.com) and run `ollama pull qwen3-vl:8b`. -- Env: `OLLAMA_HOST` (default `http://localhost:11434`), `OLLAMA_VISION_MODEL` (default `qwen3-vl:8b`), `OLLAMA_VISION_KEEP_ALIVE` (default `-1` — keep the model resident in VRAM to avoid cold-loading it on every transcription; can be set to e.g. `30m`). +**Backend 1: local Ollama (free and private)** + +- Prerequisite: install [Ollama](https://ollama.com) and pull a vision-capable model (e.g. `ollama pull qwen3-vl:8b`). +- Env: `OLLAMA_HOST` (default `http://localhost:11434`), `OLLAMA_VISION_MODEL` (**required** — the vision model), `OLLAMA_VISION_KEEP_ALIVE` (default `-1` — keep the model resident in VRAM to avoid cold-loading it on every transcription; can be set to e.g. `30m`). **Backend 2: OpenAI-compatible vision API** diff --git a/pi-web/components/VisionConfig.tsx b/pi-web/components/VisionConfig.tsx index be8d4dd..94ef9f6 100644 --- a/pi-web/components/VisionConfig.tsx +++ b/pi-web/components/VisionConfig.tsx @@ -136,7 +136,7 @@ export function VisionConfigContent() { } }, [config]); - const backend = config.backend ?? "ollama"; + const backend = config.backend; const ollama = config.ollama ?? {}; const openai = config.openai ?? {}; @@ -205,6 +205,23 @@ export function VisionConfigContent() { ))} + {/* No backend selected yet */} + {!backend && ( +
+ {t("vision.notConfigured")} +
+ )} + {/* Ollama settings */} {backend === "ollama" && (
diff --git a/pi-web/lib/i18n/messages/en.ts b/pi-web/lib/i18n/messages/en.ts index acabe87..c8f5464 100644 --- a/pi-web/lib/i18n/messages/en.ts +++ b/pi-web/lib/i18n/messages/en.ts @@ -277,6 +277,7 @@ export const enLocale: LocalePlugin = { "vision.openai.model": "Model", "vision.openai.modelPlaceholder": "e.g. gpt-4o-mini / glm-4.5v", "vision.effective": "Saved config takes effect on the next image request — no restart needed. Uploaded images are auto-transcribed into text for text-only main models (e.g. DeepSeek).", + "vision.notConfigured": "No backend selected yet — image requests fail until you pick one and fill in its details.", "vision.loadFailed": "Failed to load / save config", "vision.save": "Save", "vision.saved": "Saved — takes effect on the next image request", diff --git a/pi-web/lib/i18n/messages/zh-CN.ts b/pi-web/lib/i18n/messages/zh-CN.ts index ff131c2..fc6b7d3 100644 --- a/pi-web/lib/i18n/messages/zh-CN.ts +++ b/pi-web/lib/i18n/messages/zh-CN.ts @@ -277,6 +277,7 @@ export const zhCNLocale: LocalePlugin = { "vision.openai.model": "模型", "vision.openai.modelPlaceholder": "如 gpt-4o-mini / glm-4.5v", "vision.effective": "保存后下次识图请求立即生效,无需重启。上传图片会自动转录为文本,供纯文本主模型(如 DeepSeek)使用。", + "vision.notConfigured": "尚未选择后端——未配置时识图请求会报错,请先选择后端并填写对应参数。", "vision.loadFailed": "配置加载/保存失败", "vision.save": "保存", "vision.saved": "已保存——下次识图请求生效", diff --git a/resources/extensions/vision.ts b/resources/extensions/vision.ts index e775ec4..ec7fea9 100644 --- a/resources/extensions/vision.ts +++ b/resources/extensions/vision.ts @@ -1,10 +1,11 @@ /** * Vision extension: describe images for text-only main models (e.g. DeepSeek). * - * Two interchangeable backends, selected by VISION_BACKEND (default "ollama"): + * Two interchangeable backends, selected by VISION_BACKEND. No backend is + * preconfigured — users must pick one (Web UI vision panel or env vars): * - "ollama": local Ollama vision model via the native /api/chat endpoint. * OLLAMA_HOST (default http://localhost:11434) - * OLLAMA_VISION_MODEL (default qwen3-vl:8b) + * OLLAMA_VISION_MODEL (required, e.g. qwen3-vl:8b) * - "openai": any OpenAI-compatible vision API (cloud or self-hosted). * VISION_OPENAI_BASE_URL e.g. https://api.openai.com/v1 * VISION_OPENAI_API_KEY @@ -59,7 +60,7 @@ const visionParams = Type.Object({ backend: Type.Optional( Type.Union([Type.Literal("ollama"), Type.Literal("openai")], { description: - "视觉后端:ollama(本地,默认)或 openai(OpenAI 兼容 API)。默认取 VISION_BACKEND 环境变量", + "视觉后端:ollama(本地)或 openai(OpenAI 兼容 API)。默认取 VISION_BACKEND 环境变量,未配置时报错", }), ), model: Type.Optional( @@ -188,7 +189,7 @@ async function describeWithOllama( if (!content) { throw new Error( `vision: Ollama model ${model} returned an empty response. ` + - "Check `ollama pull qwen3-vl:8b` and that the model supports vision.", + "Check the model tag and that it supports vision.", ); } return content; @@ -291,12 +292,6 @@ interface VisionFileConfig { openai?: { baseUrl?: string; apiKey?: string; model?: string }; } -const VISION_CONFIG_DEFAULTS: VisionFileConfig = { - backend: "ollama", - ollama: { host: "http://localhost:11434", model: "qwen3-vl:8b" }, - openai: { baseUrl: "", apiKey: "", model: "" }, -}; - function visionConfigPath(): string { const agentDir = process.env.PI_CODING_AGENT_DIR?.trim(); return ( @@ -322,27 +317,30 @@ async function describeImages( ): Promise { const fileConfig = await loadVisionFileConfig(); const backend = - options.backend ?? - envOr( - "VISION_BACKEND", - fileConfig.backend ?? VISION_CONFIG_DEFAULTS.backend ?? "ollama", + options.backend ?? envOr("VISION_BACKEND", fileConfig.backend ?? ""); + if (backend !== "ollama" && backend !== "openai") { + throw new Error( + "vision: no vision backend configured. Pick one in the Web UI vision panel " + + "(lower-left → Models → Vision) or set VISION_BACKEND=ollama|openai plus " + + "the backend's address/model (see the OLLAMA_* / VISION_OPENAI_* env vars).", ); + } const prompt = options.prompt ?? DEFAULT_PROMPT; if (backend === "ollama") { const baseUrl = envOr( "OLLAMA_HOST", - fileConfig.ollama?.host ?? - VISION_CONFIG_DEFAULTS.ollama?.host ?? - "http://localhost:11434", + fileConfig.ollama?.host ?? "http://localhost:11434", ); const model = options.model ?? - envOr( - "OLLAMA_VISION_MODEL", - fileConfig.ollama?.model ?? - VISION_CONFIG_DEFAULTS.ollama?.model ?? - "qwen3-vl:8b", + envOr("OLLAMA_VISION_MODEL", fileConfig.ollama?.model ?? ""); + if (!model) { + throw new Error( + "vision: the ollama backend needs a vision model. Set it in the Web UI " + + "vision panel or via OLLAMA_VISION_MODEL (pull one first, e.g. " + + "`ollama pull qwen3-vl:8b`).", ); + } return describeWithOllama(baseUrl, model, prompt, images, options.signal); } const baseUrl = envOr( @@ -471,8 +469,8 @@ export default function visionExtension(pi: ExtensionAPI) { description: "用视觉模型描述本地图片并返回详细文本(完整 OCR、版式结构、语义总结)。" + "适用于主模型不支持图片输入(如 DeepSeek)时查看截图/图表/文档/照片。" + - "后端可配置:ollama(默认,本地 qwen3-vl:8b,免费私密)或 openai(任意 OpenAI 兼容视觉 API," + - "设 VISION_BACKEND=openai + VISION_OPENAI_BASE_URL/API_KEY/MODEL)。", + "后端可配置:ollama(本地)或 openai(任意 OpenAI 兼容视觉 API),初始未配置需先设置:" + + "WebUI 视觉面板(左下角→模型→视觉)或环境变量(VISION_BACKEND + 对应地址/模型/密钥)。", promptSnippet: "Describe local images using a vision model (Ollama or OpenAI-compatible)", promptGuidelines: [ @@ -497,7 +495,7 @@ export default function visionExtension(pi: ExtensionAPI) { return { content: [{ type: "text", text: content }], details: { - backend: params.backend ?? envOr("VISION_BACKEND", "ollama"), + backend: params.backend ?? envOr("VISION_BACKEND", ""), model: params.model ?? undefined, imageCount: images.length, },