mirror of
https://github.com/luckyyzh/pi-agent-integrated.git
synced 2026-10-03 02:59:35 +00:00
feat: add vision and extension management to WebUI
This commit is contained in:
@@ -0,0 +1,22 @@
|
||||
---
|
||||
name: vision
|
||||
description: 视觉子代理 —— 读取并描述图片(截图/图表/文档/照片),输出完整结构化描述(OCR/版式/语义),供不支持图片输入的主模型(如 DeepSeek)推理使用。后端默认本地 Ollama(qwen3-vl:8b),也可切换 OpenAI 兼容视觉 API
|
||||
tools: vision
|
||||
subagentOnlyExtensions: ./resources/extensions/vision.ts
|
||||
thinking: false
|
||||
systemPromptMode: replace
|
||||
inheritProjectContext: false
|
||||
inheritSkills: false
|
||||
defaultProgress: true
|
||||
---
|
||||
|
||||
你是视觉子代理。主会话会把一个或多个图片文件路径交给你,你调用 `vision` 工具让视觉模型看图并返回文本描述。
|
||||
|
||||
工作规则:
|
||||
|
||||
- 对每个图片路径调用一次 `vision`;相关图片可一次传入多张。
|
||||
- 工具返回的是视觉模型的转录:忠实转达,OCR 文字逐字保留,不要改写或脑补。
|
||||
- 工具报错时(文件不存在 / 后端未配置 / 模型未拉取)如实报告,并给出明确的修复提示(如 `ollama pull qwen3-vl:8b`,或检查 `VISION_OPENAI_*` 环境变量)。
|
||||
- 输出保持结构化:多图按图分组,先给结论性总结,再附关键细节;文字类图片保证转录完整。
|
||||
|
||||
主会话(通常是 DeepSeek 这类纯文本模型)看不到图片,完全依赖你的描述,完整性优先。
|
||||
@@ -143,11 +143,43 @@ data/workspaces/default/ 默认工作目录
|
||||
| `@upstash/context7-pi@0.1.2` | 查询当前库、框架、SDK 和 API 文档 | 模型先解析库 ID,再按需查询文档;无 Key 可使用公共限额 |
|
||||
| `@narumitw/pi-retry@0.31.0` | 识别瞬时供应商错误和卡住的流 | 复用 Pi 内置重试;默认 180 秒无事件视为停滞,不增加正常请求的模型调用 |
|
||||
| `resources/extensions/searxng-search.ts` | 用户自有 SearXNG 的 `web_search` | 配置 `SEARXNG_URL` 与 `SEARXNG_TOKEN` 后,模型对时效性或明确搜索请求自动调用 |
|
||||
| `resources/extensions/vision.ts` | 文本主模型(如 DeepSeek)的识图工具 `vision`(双后端) | 派 `vision` 子代理或直接让模型调用工具,返回 OCR/版式/语义文本;后端默认本地 Ollama(`qwen3-vl:8b`),可切 OpenAI 兼容视觉 API |
|
||||
|
||||
Windows 的 Playwright 不下载独立 Chromium;首次 `setup` 只缓存 MCP 的 Node.js 包,浏览器执行使用系统 Edge。macOS 的 setup 不安装或启用 Playwright;如需浏览器自动化,可在 Web UI 的 MCP 面板中手动添加并配置。
|
||||
|
||||
#### 视觉子代理(vision)
|
||||
|
||||
DeepSeek 等纯文本模型不能接收图片。仓库内置 `vision` 子代理(`.agents/vision.md`):它通过 `vision` 工具调用视觉模型读取图片,把完整 OCR、版式结构与语义描述返回给主模型,主模型基于文本继续推理。视觉后端可插拔,默认本地 Ollama,也支持任意 OpenAI 兼容视觉 API——没有本地部署条件时可直接用云服务。
|
||||
|
||||
配置入口:WebUI 左下角「模型」面板内的「视觉」标签页(写入 `data/agent/vision.json`),保存后**下次识图请求立即生效**,无需重启;环境变量优先级高于面板配置。
|
||||
|
||||
**后端一:本地 Ollama(默认,免费私密)**
|
||||
|
||||
- 前置:本机安装 [Ollama](https://ollama.com) 并 `ollama pull qwen3-vl:8b`。
|
||||
- 环境变量:`OLLAMA_HOST`(默认 `http://localhost:11434`)、`OLLAMA_VISION_MODEL`(默认 `qwen3-vl:8b`)。
|
||||
|
||||
**后端二:OpenAI 兼容视觉 API**
|
||||
|
||||
- 设置 `VISION_BACKEND=openai`,并配置 `VISION_OPENAI_BASE_URL`(如 `https://api.openai.com/v1`)、`VISION_OPENAI_API_KEY`、`VISION_OPENAI_MODEL`(如 `gpt-4o-mini`、`glm-4.5v`、`qwen-vl-max`)。
|
||||
|
||||
**自动转录(WebUI 上传即用)**
|
||||
|
||||
纯文本主模型(如 DeepSeek)无法接收图片,直接在 WebUI 上传会让请求失败(DeepSeek API 返回 HTTP 400)。`vision` 扩展注册了 `before_provider_request` 钩子:请求发出前检测到图片附件时,自动调用配置的视觉后端生成文本描述并替换进消息,主模型直接基于描述继续推理——上传即用,无需手动操作。支持图片的主模型则原样透传,不受影响。
|
||||
|
||||
描述按**单张图片**做会话级缓存:只有新上传的图片会调用视觉模型,历史图片秒回缓存。每轮请求会把历史图片的描述文本一并注入上下文以保持主模型的记忆——上下文体积会随历史图片数增长,属已知取舍(Ollama 端已显式提升 `num_ctx`,DeepSeek 前缀缓存可摊薄费用)。
|
||||
|
||||
- 用法:对主模型说“用 vision 子代理看 <图片路径>”即可;也可 `/run vision`(子代理用于主动深度分析多图;上传自动转录已覆盖日常识图)。
|
||||
- 主会话直用:重启 pi 后 `vision` 工具在主会话也可用,可对磁盘上的图片主动调用。
|
||||
- 单次调用可覆盖后端与模型:工具参数 `backend`、`model`。
|
||||
|
||||
为什么不让 pi 直接连接 Ollama 视觉模型:Ollama 的 OpenAI 兼容端点(`/v1`)会把 qwen3 系列模型的推理内容放进 `reasoning` 字段、`content` 留空,pi 会判定为空回复。Ollama 后端改走原生 `/api/chat` 并传 `think: false` 关闭思考,实测稳定可靠。
|
||||
|
||||
MCP 服务器可通过 Web UI 左下角的 MCP 按钮可视化配置(写入 `data/agent/mcp.json`):支持 stdio(命令 + 参数)与 HTTP(URL + 请求头 + OAuth/Bearer)两种传输、环境变量键值编辑、工作目录、生命周期与超时设置,另保留原始 JSON 编辑兜底。保存后重启 pi(或 /reload)生效。
|
||||
|
||||
#### 插件与扩展
|
||||
|
||||
Web UI 左下角的「插件」和「扩展」是两个独立面板:插件面板管理 npm/git 插件包的安装、更新和启停;扩展面板只展示直接加载的 `.ts`/`.js` 扩展文件,不展示插件包内的资源。扩展面板会按项目、内置和应用范围显示扩展状态、来源与路径;未信任项目中的 `.pi/extensions` 会标记为阻止而不会执行。共享扩展放在 `resources/extensions/`,项目扩展放在项目的 `.pi/extensions/`,全局扩展位于 Pi Profile 的 `extensions/` 目录。
|
||||
|
||||
Web UI 右上角的会话信息栏会汇总 Token 使用情况;当模型返回缓存读写数据时,还会显示按 Token 加权计算的缓存命中率:`cacheRead / (input + cacheRead + cacheWrite)`,不计输出 Token。
|
||||
|
||||
#### 记忆模式
|
||||
@@ -350,11 +382,43 @@ Versions are pinned in the platform defaults under `config/`: Windows uses `mcp.
|
||||
| `@upstash/context7-pi@0.1.2` | Current library, framework, SDK, API docs | Resolves a library ID and queries docs when needed; public quota works without a key |
|
||||
| `@narumitw/pi-retry@0.31.0` | Transient provider and stalled-stream classification | Uses Pi's built-in retry path; 180 seconds without events is a stall; no extra normal model calls |
|
||||
| `resources/extensions/searxng-search.ts` | `web_search` against a user-owned SearXNG proxy | After `SEARXNG_URL` and `SEARXNG_TOKEN` are set, the model calls it for current or explicit search requests |
|
||||
| `resources/extensions/vision.ts` | `vision` — image description for text-only models (e.g. DeepSeek), dual backend | Ask the `vision` subagent or call the tool directly; returns OCR/layout/semantics as text; backend defaults to local Ollama (`qwen3-vl:8b`) and can switch to any OpenAI-compatible vision API |
|
||||
|
||||
On Windows, Playwright never downloads a standalone Chromium: setup caches only its Node package and browser execution uses system Edge. On macOS, setup does not install or enable Playwright; add it manually through the MCP panel if browser automation is needed.
|
||||
|
||||
#### Vision subagent
|
||||
|
||||
Text-only models such as DeepSeek cannot receive image attachments. The repository ships a `vision` subagent (`.agents/vision.md`) that calls a vision model through the `vision` tool and returns a full OCR, layout, and semantic description the main model can reason over. The vision backend is pluggable: local Ollama by default, or any OpenAI-compatible vision API for users who cannot run a local model.
|
||||
|
||||
Configuration: the “Vision” tab inside the “Models” panel in the lower-left Web UI (writes `data/agent/vision.json`). Saved config takes effect on the **next image request** — no restart needed; environment variables take precedence over the panel.
|
||||
|
||||
**Backend 1: local Ollama (default, free and private)**
|
||||
|
||||
- Prerequisite: install [Ollama](https://ollama.com) and run `ollama pull qwen3-vl:8b`.
|
||||
- Env: `OLLAMA_HOST` (default `http://localhost:11434`), `OLLAMA_VISION_MODEL` (default `qwen3-vl:8b`).
|
||||
|
||||
**Backend 2: OpenAI-compatible vision API**
|
||||
|
||||
- Set `VISION_BACKEND=openai` and configure `VISION_OPENAI_BASE_URL` (e.g. `https://api.openai.com/v1`), `VISION_OPENAI_API_KEY`, `VISION_OPENAI_MODEL` (e.g. `gpt-4o-mini`, `glm-4.5v`, `qwen-vl-max`).
|
||||
|
||||
**Automatic transcription (upload-and-go)**
|
||||
|
||||
A text-only main model such as DeepSeek cannot receive images — uploading one in the Web UI fails the request (DeepSeek API returns HTTP 400). The `vision` extension registers a `before_provider_request` hook: when it detects image attachments, it transcribes them through the configured vision backend and replaces them with text before the request is sent, so the main model keeps reasoning seamlessly. Vision-capable main models pass through untouched.
|
||||
|
||||
Descriptions are cached **per image** for the session: only genuinely new uploads call the vision model, while previously seen images resolve from cache instantly. Every request also re-injects the accumulated image descriptions so the main model keeps its memory of them — a known trade-off where context grows with the number of images (the Ollama backend raises `num_ctx` explicitly, and DeepSeek prefix caching keeps the cost modest).
|
||||
|
||||
- Usage: ask the main model to “use the vision subagent to look at <path>”, or run `/run vision` (the subagent is for proactive deep analysis of many images; everyday image reading is covered by automatic transcription).
|
||||
- Main-session use: after restarting pi, the `vision` tool is also available in the main session for images on disk.
|
||||
- Per-call overrides: tool parameters `backend` and `model`.
|
||||
|
||||
Why not point pi directly at an Ollama vision model: Ollama's OpenAI-compatible `/v1` endpoint moves qwen3-family reasoning into the `reasoning` field with an empty `content`, which pi treats as an empty reply. The Ollama backend uses the native `/api/chat` with `think: false` instead, which works reliably.
|
||||
|
||||
MCP servers can be configured visually from the MCP button in the lower-left Web UI (writes `data/agent/mcp.json`): stdio (command + args) or HTTP (URL + headers + OAuth/Bearer) transport, environment-variable row editing, working directory, lifecycle and timeout options, plus raw JSON editing as a fallback. Changes take effect after restarting pi (or /reload).
|
||||
|
||||
#### Plugins and extensions
|
||||
|
||||
The lower-left Web UI has separate “Plugins” and “Extensions” panels. The Plugins panel manages npm/git plugin packages, including install, update, enable, and disable actions. The Extensions panel only lists directly loaded `.ts`/`.js` extension files and never lists resources supplied by plugin packages. It groups extensions by project, built-in, and app scope and shows their status, source, and path; extensions in an untrusted project `.pi/extensions` directory are shown as blocked and are not executed. Shared extensions belong in `resources/extensions/`, project extensions in `.pi/extensions/`, and global extensions in the Pi Profile `extensions/` directory.
|
||||
|
||||
The Web UI session-info panel in the upper-right summarizes Token usage. When a model reports cache read/write data, it also shows the token-weighted cache hit rate: `cacheRead / (input + cacheRead + cacheWrite)`, excluding output Tokens.
|
||||
|
||||
#### Memory modes
|
||||
|
||||
@@ -0,0 +1,127 @@
|
||||
import { existsSync, readdirSync, statSync } from "fs";
|
||||
import { basename, dirname, extname, join, relative, resolve } from "path";
|
||||
import { DefaultPackageManager, getAgentDir, type ResolvedResource } from "@earendil-works/pi-coding-agent";
|
||||
import { getAllowedFileRoots, isExistingFilePathAllowed } from "@/lib/file-access";
|
||||
import { getManagedRuntimePaths, isManagedRuntime, createAppSettingsManager } from "@/lib/app-runtime";
|
||||
import { getProjectTrustStatus } from "@/lib/project-trust";
|
||||
import type { ExtensionInfo, ExtensionsResponse, PluginDiagnostic } from "@/lib/api-types";
|
||||
|
||||
export const dynamic = "force-dynamic";
|
||||
|
||||
function extensionName(path: string): string {
|
||||
const file = basename(path);
|
||||
const extension = extname(file);
|
||||
if (/^index\.(ts|js)$/.test(file)) return basename(dirname(path));
|
||||
return extension ? file.slice(0, -extension.length) : file;
|
||||
}
|
||||
|
||||
function extensionFiles(root: string): string[] {
|
||||
if (!existsSync(root)) return [];
|
||||
try {
|
||||
const stats = statSync(root);
|
||||
if (stats.isFile()) return /\.(?:ts|js)$/.test(root) ? [root] : [];
|
||||
if (!stats.isDirectory()) return [];
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
|
||||
const files: string[] = [];
|
||||
for (const entry of readdirSync(root, { withFileTypes: true })) {
|
||||
const path = join(root, entry.name);
|
||||
if (entry.isDirectory()) files.push(...extensionFiles(path));
|
||||
else if (entry.isFile() && /\.(?:ts|js)$/.test(entry.name)) files.push(path);
|
||||
}
|
||||
return files;
|
||||
}
|
||||
|
||||
function infoFromResource(resource: ResolvedResource, scopeOverride?: ExtensionInfo["scope"]): ExtensionInfo {
|
||||
const baseDir = resource.metadata.baseDir ?? dirname(resource.path);
|
||||
const rel = relative(baseDir, resource.path);
|
||||
return {
|
||||
name: extensionName(resource.path),
|
||||
path: resource.path,
|
||||
relativePath: rel && !rel.startsWith("..") ? rel : resource.path,
|
||||
source: scopeOverride === "builtin" ? "resources" : resource.metadata.source,
|
||||
scope: scopeOverride ?? (resource.metadata.scope === "project" ? "project" : "global"),
|
||||
status: resource.enabled ? "enabled" : "disabled",
|
||||
};
|
||||
}
|
||||
|
||||
function blockedInfo(path: string): ExtensionInfo {
|
||||
return {
|
||||
name: extensionName(path),
|
||||
path,
|
||||
relativePath: relative(dirname(dirname(path)), path),
|
||||
source: "project",
|
||||
scope: "project",
|
||||
status: "blocked",
|
||||
};
|
||||
}
|
||||
|
||||
export async function GET(req: Request) {
|
||||
let cwd: string | null;
|
||||
try {
|
||||
cwd = new URL(req.url).searchParams.get("cwd");
|
||||
} catch {
|
||||
return Response.json({ error: "invalid request URL" }, { status: 400 });
|
||||
}
|
||||
if (!cwd) return Response.json({ error: "cwd required" }, { status: 400 });
|
||||
|
||||
try {
|
||||
const allowedRoots = await getAllowedFileRoots();
|
||||
if (!isExistingFilePathAllowed(cwd, allowedRoots)) {
|
||||
return Response.json({ error: "Access denied" }, { status: 403 });
|
||||
}
|
||||
|
||||
const agentDir = getAgentDir();
|
||||
const trust = getProjectTrustStatus(cwd, agentDir);
|
||||
const settingsManager = createAppSettingsManager(cwd, agentDir, trust.trusted);
|
||||
const packageManager = new DefaultPackageManager({ cwd, agentDir, settingsManager });
|
||||
const diagnostics: PluginDiagnostic[] = [];
|
||||
const byPath = new Map<string, ExtensionInfo>();
|
||||
|
||||
try {
|
||||
const resolved = await packageManager.resolve(async () => "skip");
|
||||
for (const resource of resolved.extensions) {
|
||||
// Package-provided extensions belong to the plugin panel, not here.
|
||||
if (resource.metadata.origin !== "top-level") continue;
|
||||
byPath.set(resource.path, infoFromResource(resource));
|
||||
}
|
||||
} catch (error) {
|
||||
diagnostics.push({ type: "error", message: error instanceof Error ? error.message : String(error) });
|
||||
}
|
||||
|
||||
if (isManagedRuntime()) {
|
||||
const resourcesRoot = resolve(getManagedRuntimePaths().resourcesDir, "extensions");
|
||||
const builtIn = await packageManager.resolveExtensionSources([resourcesRoot], { temporary: true });
|
||||
for (const resource of builtIn.extensions) {
|
||||
byPath.set(resource.path, infoFromResource(resource, "builtin"));
|
||||
}
|
||||
}
|
||||
|
||||
const projectExtensionsRoot = join(cwd, ".pi", "extensions");
|
||||
if (trust.requiresTrust && !trust.trusted) {
|
||||
for (const path of extensionFiles(projectExtensionsRoot)) {
|
||||
byPath.set(path, blockedInfo(path));
|
||||
}
|
||||
if (byPathHasScope(byPath, "project")) {
|
||||
diagnostics.push({ type: "warning", source: "project", message: "Project extensions are blocked until this project is trusted." });
|
||||
}
|
||||
}
|
||||
|
||||
const scopeOrder: Record<ExtensionInfo["scope"], number> = { project: 0, builtin: 1, global: 2 };
|
||||
const extensions = [...byPath.values()].sort((a, b) => (
|
||||
scopeOrder[a.scope] - scopeOrder[b.scope] || a.name.localeCompare(b.name) || a.path.localeCompare(b.path)
|
||||
));
|
||||
return Response.json({ extensions, diagnostics, projectResourcesLoaded: trust.trusted } satisfies ExtensionsResponse);
|
||||
} catch (error) {
|
||||
return Response.json({ error: String(error) }, { status: 500 });
|
||||
}
|
||||
}
|
||||
|
||||
function byPathHasScope(entries: Map<string, ExtensionInfo>, scope: ExtensionInfo["scope"]): boolean {
|
||||
for (const entry of entries.values()) {
|
||||
if (entry.scope === scope) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
@@ -0,0 +1,65 @@
|
||||
import { NextResponse } from "next/server";
|
||||
import { readFileSync, writeFileSync, existsSync, mkdirSync } from "node:fs";
|
||||
import { join, dirname } from "node:path";
|
||||
import { getAgentDir } from "@earendil-works/pi-coding-agent";
|
||||
|
||||
export const dynamic = "force-dynamic";
|
||||
|
||||
interface VisionConfigFile {
|
||||
backend?: "ollama" | "openai";
|
||||
ollama?: { host?: string; model?: string };
|
||||
openai?: { baseUrl?: string; apiKey?: string; model?: string };
|
||||
}
|
||||
|
||||
function getVisionPath(): string {
|
||||
return join(getAgentDir(), "vision.json");
|
||||
}
|
||||
|
||||
function readVisionJson(): VisionConfigFile {
|
||||
const path = getVisionPath();
|
||||
if (!existsSync(path)) return {};
|
||||
try {
|
||||
const parsed = JSON.parse(readFileSync(path, "utf8")) as VisionConfigFile;
|
||||
if (!parsed || typeof parsed !== "object") return {};
|
||||
return parsed;
|
||||
} catch {
|
||||
return {};
|
||||
}
|
||||
}
|
||||
|
||||
function writeVisionJson(data: VisionConfigFile): void {
|
||||
const path = getVisionPath();
|
||||
const dir = dirname(path);
|
||||
if (!existsSync(dir)) mkdirSync(dir, { recursive: true });
|
||||
writeFileSync(path, JSON.stringify(data, null, 2) + "\n", "utf8");
|
||||
}
|
||||
|
||||
export function GET() {
|
||||
return NextResponse.json({ config: readVisionJson(), path: getVisionPath() });
|
||||
}
|
||||
|
||||
export async function PUT(req: Request) {
|
||||
try {
|
||||
const body = (await req.json()) as VisionConfigFile;
|
||||
if (!body || typeof body !== "object") {
|
||||
return NextResponse.json({ error: "body must be an object" }, { status: 400 });
|
||||
}
|
||||
const out: VisionConfigFile = {};
|
||||
if (body.backend === "ollama" || body.backend === "openai") out.backend = body.backend;
|
||||
if (body.ollama && typeof body.ollama === "object") {
|
||||
out.ollama = {};
|
||||
if (typeof body.ollama.host === "string" && body.ollama.host.trim()) out.ollama.host = body.ollama.host.trim();
|
||||
if (typeof body.ollama.model === "string" && body.ollama.model.trim()) out.ollama.model = body.ollama.model.trim();
|
||||
}
|
||||
if (body.openai && typeof body.openai === "object") {
|
||||
out.openai = {};
|
||||
if (typeof body.openai.baseUrl === "string" && body.openai.baseUrl.trim()) out.openai.baseUrl = body.openai.baseUrl.trim();
|
||||
if (typeof body.openai.apiKey === "string" && body.openai.apiKey.trim()) out.openai.apiKey = body.openai.apiKey.trim();
|
||||
if (typeof body.openai.model === "string" && body.openai.model.trim()) out.openai.model = body.openai.model.trim();
|
||||
}
|
||||
writeVisionJson(out);
|
||||
return NextResponse.json({ success: true, path: getVisionPath() });
|
||||
} catch (error) {
|
||||
return NextResponse.json({ error: String(error) }, { status: 500 });
|
||||
}
|
||||
}
|
||||
+1388
-319
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,193 @@
|
||||
"use client";
|
||||
|
||||
import { useCallback, useEffect, useMemo, useState } from "react";
|
||||
import { useIsMobile } from "@/hooks/useIsMobile";
|
||||
import { useI18n } from "@/hooks/useI18n";
|
||||
import type { ExtensionInfo, ExtensionsResponse } from "@/lib/api-types";
|
||||
|
||||
function shortenPath(path: string): string {
|
||||
return path.replace(/^\/(?:Users|home)\/[^/]+/, "~");
|
||||
}
|
||||
|
||||
function statusColor(status: ExtensionInfo["status"]): string {
|
||||
if (status === "enabled") return "var(--accent)";
|
||||
if (status === "blocked") return "#d97706";
|
||||
return "var(--text-dim)";
|
||||
}
|
||||
|
||||
function statusLabel(status: ExtensionInfo["status"], t: ReturnType<typeof useI18n>["t"]): string {
|
||||
if (status === "enabled") return t("extensions.enabled");
|
||||
if (status === "blocked") return t("extensions.blocked");
|
||||
return t("extensions.disabled");
|
||||
}
|
||||
|
||||
function scopeLabel(scope: ExtensionInfo["scope"], t: ReturnType<typeof useI18n>["t"]): string {
|
||||
if (scope === "project") return t("extensions.project");
|
||||
if (scope === "builtin") return t("extensions.builtin");
|
||||
return t("extensions.app");
|
||||
}
|
||||
|
||||
function extensionKey(extension: ExtensionInfo): string {
|
||||
return `${extension.scope}:${extension.path}`;
|
||||
}
|
||||
|
||||
export function ExtensionsConfig({ cwd, onCloseAction }: { cwd: string; onCloseAction: () => void }) {
|
||||
const isMobile = useIsMobile();
|
||||
const { t } = useI18n();
|
||||
const [data, setData] = useState<ExtensionsResponse | null>(null);
|
||||
const [loading, setLoading] = useState(true);
|
||||
const [error, setError] = useState<string | null>(null);
|
||||
const [selected, setSelected] = useState<string | null>(null);
|
||||
|
||||
const loadExtensions = useCallback(async () => {
|
||||
setLoading(true);
|
||||
setError(null);
|
||||
try {
|
||||
const response = await fetch(`/api/extensions?cwd=${encodeURIComponent(cwd)}`);
|
||||
const next = (await response.json()) as ExtensionsResponse & { error?: string };
|
||||
if (!response.ok || next.error) throw new Error(next.error ?? `HTTP ${response.status}`);
|
||||
setData(next);
|
||||
setSelected((current) => (
|
||||
current && next.extensions.some((extension) => extensionKey(extension) === current)
|
||||
? current
|
||||
: next.extensions[0] ? extensionKey(next.extensions[0]) : null
|
||||
));
|
||||
} catch (err) {
|
||||
setError(err instanceof Error ? err.message : String(err));
|
||||
} finally {
|
||||
setLoading(false);
|
||||
}
|
||||
}, [cwd]);
|
||||
|
||||
useEffect(() => {
|
||||
void loadExtensions();
|
||||
}, [loadExtensions]);
|
||||
|
||||
const extensions = useMemo(() => data?.extensions ?? [], [data?.extensions]);
|
||||
const selectedExtension = extensions.find((extension) => extensionKey(extension) === selected) ?? null;
|
||||
const enabledCount = extensions.filter((extension) => extension.status === "enabled").length;
|
||||
const blockedCount = extensions.filter((extension) => extension.status === "blocked").length;
|
||||
const disabledCount = extensions.filter((extension) => extension.status === "disabled").length;
|
||||
const hasBlockedProjectExtensions = blockedCount > 0 && !data?.projectResourcesLoaded;
|
||||
const groupedExtensions = useMemo(() => {
|
||||
const groups: Array<{ scope: ExtensionInfo["scope"]; extensions: ExtensionInfo[] }> = [];
|
||||
for (const scope of ["project", "builtin", "global"] as const) {
|
||||
const scoped = extensions.filter((extension) => extension.scope === scope);
|
||||
if (scoped.length > 0) groups.push({ scope, extensions: scoped });
|
||||
}
|
||||
return groups;
|
||||
}, [extensions]);
|
||||
|
||||
return (
|
||||
<div
|
||||
style={{ position: "fixed", inset: 0, zIndex: 1000, background: "rgba(0,0,0,0.35)", display: "flex", alignItems: "center", justifyContent: "center" }}
|
||||
onClick={(event) => {
|
||||
if (event.target === event.currentTarget) onCloseAction();
|
||||
}}
|
||||
>
|
||||
<div
|
||||
style={{
|
||||
width: isMobile ? "calc(100vw - 16px)" : 860,
|
||||
maxWidth: "calc(100vw - 16px)",
|
||||
height: isMobile ? "calc(100dvh - 16px)" : "76vh",
|
||||
maxHeight: "calc(100dvh - 16px)",
|
||||
background: "var(--bg)",
|
||||
border: "1px solid var(--border)",
|
||||
borderRadius: 10,
|
||||
display: "flex",
|
||||
flexDirection: "column",
|
||||
boxShadow: "0 8px 32px rgba(0,0,0,0.18)",
|
||||
overflow: "hidden",
|
||||
}}
|
||||
>
|
||||
<div style={{ display: "flex", alignItems: "center", justifyContent: "space-between", padding: "12px 18px", borderBottom: "1px solid var(--border)", flexShrink: 0 }}>
|
||||
<div style={{ display: "flex", alignItems: "baseline", gap: 10, minWidth: 0 }}>
|
||||
<span style={{ fontSize: 15, fontWeight: 700, color: "var(--text)" }}>{t("common.extensions")}</span>
|
||||
<code style={{ fontSize: 11, color: "var(--text-muted)", fontFamily: "var(--font-mono)", overflow: "hidden", textOverflow: "ellipsis", whiteSpace: "nowrap" }}>
|
||||
{shortenPath(cwd)}
|
||||
</code>
|
||||
</div>
|
||||
<button type="button" onClick={onCloseAction} aria-label={t("i18n.close")} style={{ background: "none", border: "none", color: "var(--text-muted)", cursor: "pointer", fontSize: 20, lineHeight: 1, padding: "2px 6px" }}>
|
||||
×
|
||||
</button>
|
||||
</div>
|
||||
|
||||
{hasBlockedProjectExtensions && (
|
||||
<div style={{ padding: "8px 18px", borderBottom: "1px solid var(--border)", color: "#d97706", fontSize: 11 }}>
|
||||
{t("extensions.projectResourcesBlocked")}
|
||||
</div>
|
||||
)}
|
||||
|
||||
<div style={{ flex: 1, display: "flex", flexDirection: isMobile ? "column" : "row", overflow: "hidden" }}>
|
||||
<div style={{ width: isMobile ? "100%" : 245, maxHeight: isMobile ? "40vh" : undefined, borderRight: isMobile ? "none" : "1px solid var(--border)", borderBottom: isMobile ? "1px solid var(--border)" : "none", display: "flex", flexDirection: "column", flexShrink: 0, background: "var(--bg-panel)" }}>
|
||||
<div style={{ flex: 1, overflowY: "auto", padding: "8px 6px" }}>
|
||||
{loading ? (
|
||||
<div style={{ padding: "10px 8px", fontSize: 12, color: "var(--text-muted)" }}>{t("i18n.loading")}</div>
|
||||
) : error ? (
|
||||
<div style={{ padding: "10px 8px", fontSize: 11, color: "#ef4444" }}>{error}</div>
|
||||
) : groupedExtensions.length === 0 ? (
|
||||
<div style={{ padding: "10px 8px", fontSize: 11, color: "var(--text-dim)" }}>{t("extensions.noExtensions")}</div>
|
||||
) : (
|
||||
groupedExtensions.map((group) => (
|
||||
<div key={group.scope} style={{ marginBottom: 6 }}>
|
||||
<div style={{ padding: "4px 8px 3px", fontSize: 10, fontWeight: 600, color: "var(--text-dim)", textTransform: "uppercase" }}>
|
||||
{scopeLabel(group.scope, t)}
|
||||
</div>
|
||||
{group.extensions.map((extension) => {
|
||||
const key = extensionKey(extension);
|
||||
const isSelected = selected === key;
|
||||
return (
|
||||
<div
|
||||
key={key}
|
||||
onClick={() => setSelected(key)}
|
||||
style={{ display: "flex", alignItems: "center", gap: 7, padding: "8px", borderRadius: 5, cursor: "pointer", background: isSelected ? "var(--bg-selected)" : "none" }}
|
||||
onMouseEnter={(event) => { if (!isSelected) event.currentTarget.style.background = "var(--bg-hover)"; }}
|
||||
onMouseLeave={(event) => { if (!isSelected) event.currentTarget.style.background = "none"; }}
|
||||
>
|
||||
<span style={{ flexShrink: 0, width: 7, height: 7, borderRadius: "50%", background: statusColor(extension.status) }} />
|
||||
<span style={{ minWidth: 0, flex: 1, fontSize: 12, fontWeight: isSelected ? 600 : 400, color: extension.status === "disabled" ? "var(--text-dim)" : "var(--text)", fontFamily: "var(--font-mono)", overflow: "hidden", textOverflow: "ellipsis", whiteSpace: "nowrap" }}>
|
||||
{extension.name}
|
||||
</span>
|
||||
</div>
|
||||
);
|
||||
})}
|
||||
</div>
|
||||
))
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div style={{ flex: 1, overflowY: "auto", padding: 20 }}>
|
||||
{selectedExtension ? (
|
||||
<div style={{ display: "flex", flexDirection: "column", gap: 20, maxWidth: 680 }}>
|
||||
<div style={{ display: "flex", alignItems: "center", gap: 8, flexWrap: "wrap" }}>
|
||||
<span style={{ width: 8, height: 8, borderRadius: "50%", background: statusColor(selectedExtension.status) }} />
|
||||
<span style={{ fontSize: 14, fontWeight: 700, color: "var(--text)", fontFamily: "var(--font-mono)" }}>{selectedExtension.name}</span>
|
||||
<span style={{ fontSize: 10, padding: "1px 5px", borderRadius: 3, background: "rgba(120,120,120,0.12)", color: "var(--text-dim)" }}>{scopeLabel(selectedExtension.scope, t)}</span>
|
||||
</div>
|
||||
<div style={{ display: "grid", gridTemplateColumns: "minmax(100px, 130px) minmax(0, 1fr)", gap: "9px 14px", fontSize: 12, lineHeight: 1.45 }}>
|
||||
<div style={{ color: "var(--text-dim)" }}>{t("extensions.status")}</div>
|
||||
<div style={{ color: statusColor(selectedExtension.status) }}>{statusLabel(selectedExtension.status, t)}</div>
|
||||
<div style={{ color: "var(--text-dim)" }}>{t("extensions.source")}</div>
|
||||
<div style={{ color: "var(--text-muted)", fontFamily: "var(--font-mono)" }}>{selectedExtension.source}</div>
|
||||
<div style={{ color: "var(--text-dim)" }}>{t("extensions.path")}</div>
|
||||
<div style={{ color: "var(--text-muted)", fontFamily: "var(--font-mono)", overflowWrap: "anywhere" }}>{shortenPath(selectedExtension.path)}</div>
|
||||
</div>
|
||||
</div>
|
||||
) : !loading && !error ? (
|
||||
<div style={{ height: "100%", display: "flex", alignItems: "center", justifyContent: "center", color: "var(--text-dim)", fontSize: 13 }}>{t("extensions.noExtensions")}</div>
|
||||
) : null}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div style={{ display: "flex", alignItems: "center", gap: 12, padding: "10px 18px", borderTop: "1px solid var(--border)", flexShrink: 0 }}>
|
||||
<div style={{ minWidth: 0, flex: 1, fontSize: 11, color: "var(--text-dim)", overflow: "hidden", textOverflow: "ellipsis", whiteSpace: "nowrap" }}>
|
||||
{data ? `${extensions.length} ${t("extensions.total")} · ${enabledCount} ${t("extensions.enabled")} · ${disabledCount} ${t("extensions.disabled")} · ${blockedCount} ${t("extensions.blocked")}${data.diagnostics.length ? ` · ${data.diagnostics.length} ${t("extensions.diagnostics")}` : ""}` : ""}
|
||||
</div>
|
||||
<button type="button" onClick={() => void loadExtensions()} disabled={loading} style={{ padding: "6px 12px", background: "none", border: "1px solid var(--border)", borderRadius: 6, color: "var(--text-muted)", cursor: loading ? "not-allowed" : "pointer", opacity: loading ? 0.5 : 1, fontSize: 12 }}>{t("i18n.refresh")}</button>
|
||||
<button type="button" onClick={onCloseAction} style={{ padding: "6px 12px", background: "none", border: "1px solid var(--border)", borderRadius: 6, color: "var(--text-muted)", cursor: "pointer", fontSize: 12 }}>{t("i18n.close")}</button>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
+2125
-417
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,424 @@
|
||||
"use client";
|
||||
|
||||
import { useCallback, useEffect, useState } from "react";
|
||||
import { useIsMobile } from "@/hooks/useIsMobile";
|
||||
import { useI18n } from "@/hooks/useI18n";
|
||||
|
||||
interface VisionConfigFile {
|
||||
backend?: "ollama" | "openai";
|
||||
ollama?: { host?: string; model?: string };
|
||||
openai?: { baseUrl?: string; apiKey?: string; model?: string };
|
||||
}
|
||||
|
||||
const inputStyle: React.CSSProperties = {
|
||||
padding: "6px 9px",
|
||||
background: "var(--bg-panel)",
|
||||
border: "1px solid var(--border)",
|
||||
borderRadius: 5,
|
||||
color: "var(--text)",
|
||||
fontSize: 12,
|
||||
outline: "none",
|
||||
width: "100%",
|
||||
boxSizing: "border-box",
|
||||
};
|
||||
|
||||
function Field({
|
||||
label,
|
||||
children,
|
||||
}: {
|
||||
label: string;
|
||||
children: React.ReactNode;
|
||||
}) {
|
||||
return (
|
||||
<div style={{ display: "flex", flexDirection: "column", gap: 4 }}>
|
||||
<label
|
||||
style={{ fontSize: 11, color: "var(--text-muted)", fontWeight: 500 }}
|
||||
>
|
||||
{label}
|
||||
</label>
|
||||
{children}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
function TextInput({
|
||||
value,
|
||||
onChange,
|
||||
placeholder,
|
||||
mono,
|
||||
type,
|
||||
}: {
|
||||
value: string;
|
||||
onChange: (v: string) => void;
|
||||
placeholder?: string;
|
||||
mono?: boolean;
|
||||
type?: string;
|
||||
}) {
|
||||
return (
|
||||
<input
|
||||
value={value}
|
||||
onChange={(e) => onChange(e.target.value)}
|
||||
placeholder={placeholder}
|
||||
type={type ?? "text"}
|
||||
style={{
|
||||
...inputStyle,
|
||||
fontFamily: mono ? "var(--font-mono)" : "inherit",
|
||||
}}
|
||||
/>
|
||||
);
|
||||
}
|
||||
|
||||
const saveButtonStyle = (primary: boolean): React.CSSProperties => ({
|
||||
padding: "6px 14px",
|
||||
borderRadius: 6,
|
||||
fontSize: 12,
|
||||
fontWeight: 600,
|
||||
cursor: "pointer",
|
||||
border: "1px solid var(--border)",
|
||||
background: primary ? "var(--accent)" : "var(--bg-panel)",
|
||||
color: primary ? "#fff" : "var(--text)",
|
||||
});
|
||||
|
||||
/**
|
||||
* Vision backend settings (backend picker + fields + save), without modal
|
||||
* chrome. Embedded in the Models config panel and reused by the standalone
|
||||
* modal below.
|
||||
*/
|
||||
export function VisionConfigContent() {
|
||||
const { t } = useI18n();
|
||||
const [config, setConfig] = useState<VisionConfigFile>({});
|
||||
const [loading, setLoading] = useState(true);
|
||||
const [saving, setSaving] = useState(false);
|
||||
const [saveError, setSaveError] = useState<string | null>(null);
|
||||
const [savedOk, setSavedOk] = useState(false);
|
||||
|
||||
const load = useCallback(async () => {
|
||||
setLoading(true);
|
||||
setSaveError(null);
|
||||
try {
|
||||
const res = await fetch("/api/vision-config");
|
||||
const data = (await res.json()) as {
|
||||
config?: VisionConfigFile;
|
||||
error?: string;
|
||||
};
|
||||
if (!res.ok || data.error)
|
||||
throw new Error(data.error ?? `HTTP ${res.status}`);
|
||||
setConfig(data.config ?? {});
|
||||
} catch (e) {
|
||||
setSaveError(e instanceof Error ? e.message : String(e));
|
||||
} finally {
|
||||
setLoading(false);
|
||||
}
|
||||
}, []);
|
||||
|
||||
useEffect(() => {
|
||||
void load();
|
||||
}, [load]);
|
||||
|
||||
const save = useCallback(async () => {
|
||||
setSaving(true);
|
||||
setSaveError(null);
|
||||
setSavedOk(false);
|
||||
try {
|
||||
const res = await fetch("/api/vision-config", {
|
||||
method: "PUT",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify(config),
|
||||
});
|
||||
const data = (await res.json()) as { success?: boolean; error?: string };
|
||||
if (!res.ok || !data.success)
|
||||
throw new Error(data.error ?? `HTTP ${res.status}`);
|
||||
setSavedOk(true);
|
||||
} catch (e) {
|
||||
setSaveError(e instanceof Error ? e.message : String(e));
|
||||
} finally {
|
||||
setSaving(false);
|
||||
}
|
||||
}, [config]);
|
||||
|
||||
const backend = config.backend ?? "ollama";
|
||||
const ollama = config.ollama ?? {};
|
||||
const openai = config.openai ?? {};
|
||||
|
||||
return (
|
||||
<div
|
||||
style={{
|
||||
display: "flex",
|
||||
flexDirection: "column",
|
||||
gap: 14,
|
||||
padding: "14px 18px",
|
||||
flex: 1,
|
||||
overflow: "auto",
|
||||
}}
|
||||
>
|
||||
{loading ? (
|
||||
<div style={{ fontSize: 12, color: "var(--text-muted)" }}>
|
||||
{t("vision.loading")}
|
||||
</div>
|
||||
) : (
|
||||
<>
|
||||
{/* Backend picker */}
|
||||
<div style={{ display: "flex", gap: 8 }}>
|
||||
{(["ollama", "openai"] as const).map((b) => (
|
||||
<button
|
||||
key={b}
|
||||
onClick={() => setConfig((prev) => ({ ...prev, backend: b }))}
|
||||
style={{
|
||||
flex: 1,
|
||||
padding: "10px 12px",
|
||||
borderRadius: 8,
|
||||
cursor: "pointer",
|
||||
textAlign: "left",
|
||||
border: `1px solid ${backend === b ? "var(--accent)" : "var(--border)"}`,
|
||||
background:
|
||||
backend === b
|
||||
? "color-mix(in srgb, var(--accent) 8%, var(--bg-panel))"
|
||||
: "var(--bg-panel)",
|
||||
display: "flex",
|
||||
flexDirection: "column",
|
||||
gap: 3,
|
||||
}}
|
||||
>
|
||||
<span
|
||||
style={{
|
||||
fontSize: 13,
|
||||
fontWeight: 600,
|
||||
color: "var(--text)",
|
||||
}}
|
||||
>
|
||||
{b === "ollama"
|
||||
? t("vision.backend.ollama")
|
||||
: t("vision.backend.openai")}
|
||||
</span>
|
||||
<span
|
||||
style={{
|
||||
fontSize: 11,
|
||||
color: "var(--text-muted)",
|
||||
lineHeight: 1.4,
|
||||
}}
|
||||
>
|
||||
{b === "ollama"
|
||||
? t("vision.backend.ollamaHint")
|
||||
: t("vision.backend.openaiHint")}
|
||||
</span>
|
||||
</button>
|
||||
))}
|
||||
</div>
|
||||
|
||||
{/* Ollama settings */}
|
||||
{backend === "ollama" && (
|
||||
<div style={{ display: "flex", flexDirection: "column", gap: 10 }}>
|
||||
<Field label={t("vision.ollama.host")}>
|
||||
<TextInput
|
||||
value={ollama.host ?? ""}
|
||||
onChange={(v) =>
|
||||
setConfig((prev) => ({
|
||||
...prev,
|
||||
ollama: { ...prev.ollama, host: v },
|
||||
}))
|
||||
}
|
||||
placeholder={t("vision.ollama.hostPlaceholder")}
|
||||
mono
|
||||
/>
|
||||
</Field>
|
||||
<Field label={t("vision.ollama.model")}>
|
||||
<TextInput
|
||||
value={ollama.model ?? ""}
|
||||
onChange={(v) =>
|
||||
setConfig((prev) => ({
|
||||
...prev,
|
||||
ollama: { ...prev.ollama, model: v },
|
||||
}))
|
||||
}
|
||||
placeholder={t("vision.ollama.modelPlaceholder")}
|
||||
mono
|
||||
/>
|
||||
</Field>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{/* OpenAI-compatible settings */}
|
||||
{backend === "openai" && (
|
||||
<div style={{ display: "flex", flexDirection: "column", gap: 10 }}>
|
||||
<Field label={t("vision.openai.baseUrl")}>
|
||||
<TextInput
|
||||
value={openai.baseUrl ?? ""}
|
||||
onChange={(v) =>
|
||||
setConfig((prev) => ({
|
||||
...prev,
|
||||
openai: { ...prev.openai, baseUrl: v },
|
||||
}))
|
||||
}
|
||||
placeholder={t("vision.openai.baseUrlPlaceholder")}
|
||||
mono
|
||||
/>
|
||||
</Field>
|
||||
<Field label={t("vision.openai.apiKey")}>
|
||||
<TextInput
|
||||
value={openai.apiKey ?? ""}
|
||||
onChange={(v) =>
|
||||
setConfig((prev) => ({
|
||||
...prev,
|
||||
openai: { ...prev.openai, apiKey: v },
|
||||
}))
|
||||
}
|
||||
type="password"
|
||||
mono
|
||||
/>
|
||||
</Field>
|
||||
<Field label={t("vision.openai.model")}>
|
||||
<TextInput
|
||||
value={openai.model ?? ""}
|
||||
onChange={(v) =>
|
||||
setConfig((prev) => ({
|
||||
...prev,
|
||||
openai: { ...prev.openai, model: v },
|
||||
}))
|
||||
}
|
||||
placeholder={t("vision.openai.modelPlaceholder")}
|
||||
mono
|
||||
/>
|
||||
</Field>
|
||||
</div>
|
||||
)}
|
||||
|
||||
<div
|
||||
style={{
|
||||
fontSize: 11,
|
||||
color: "var(--text-muted)",
|
||||
lineHeight: 1.5,
|
||||
borderTop: "1px solid var(--border)",
|
||||
paddingTop: 10,
|
||||
}}
|
||||
>
|
||||
{t("vision.effective")}
|
||||
</div>
|
||||
</>
|
||||
)}
|
||||
|
||||
{saveError && (
|
||||
<div style={{ fontSize: 11, color: "var(--danger, #e5484d)" }}>
|
||||
{t("vision.loadFailed")}: {saveError}
|
||||
</div>
|
||||
)}
|
||||
{savedOk && (
|
||||
<div style={{ fontSize: 11, color: "var(--success, #30a46c)" }}>
|
||||
{t("vision.saved")}
|
||||
</div>
|
||||
)}
|
||||
|
||||
<div
|
||||
style={{
|
||||
display: "flex",
|
||||
justifyContent: "flex-end",
|
||||
gap: 8,
|
||||
marginTop: "auto",
|
||||
}}
|
||||
>
|
||||
<button
|
||||
onClick={save}
|
||||
disabled={saving || loading}
|
||||
style={{
|
||||
...saveButtonStyle(true),
|
||||
opacity: saving || loading ? 0.6 : 1,
|
||||
}}
|
||||
>
|
||||
{t("vision.save")}
|
||||
</button>
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
/** Standalone modal wrapper (kept for backward compatibility). */
|
||||
export function VisionConfig({ onClose }: { onClose: () => void }) {
|
||||
const isMobile = useIsMobile();
|
||||
const { t } = useI18n();
|
||||
const [configPath, setConfigPath] = useState("");
|
||||
|
||||
useEffect(() => {
|
||||
fetch("/api/vision-config")
|
||||
.then((r) => r.json())
|
||||
.then((d: { path?: string }) => setConfigPath(d.path ?? ""))
|
||||
.catch(() => {});
|
||||
}, []);
|
||||
|
||||
return (
|
||||
<div
|
||||
style={{
|
||||
position: "fixed",
|
||||
inset: 0,
|
||||
background: "rgba(0,0,0,0.45)",
|
||||
display: "flex",
|
||||
alignItems: "center",
|
||||
justifyContent: "center",
|
||||
zIndex: 1000,
|
||||
padding: 8,
|
||||
}}
|
||||
onClick={(e) => {
|
||||
if (e.target === e.currentTarget) onClose();
|
||||
}}
|
||||
>
|
||||
<div
|
||||
style={{
|
||||
width: isMobile ? "calc(100vw - 16px)" : 560,
|
||||
maxWidth: "calc(100vw - 16px)",
|
||||
maxHeight: "calc(100dvh - 16px)",
|
||||
background: "var(--bg)",
|
||||
border: "1px solid var(--border)",
|
||||
borderRadius: 10,
|
||||
display: "flex",
|
||||
flexDirection: "column",
|
||||
boxShadow: "0 8px 32px rgba(0,0,0,0.18)",
|
||||
overflow: "hidden",
|
||||
}}
|
||||
>
|
||||
<div
|
||||
style={{
|
||||
display: "flex",
|
||||
alignItems: "center",
|
||||
justifyContent: "space-between",
|
||||
padding: "12px 18px",
|
||||
borderBottom: "1px solid var(--border)",
|
||||
flexShrink: 0,
|
||||
}}
|
||||
>
|
||||
<div style={{ display: "flex", alignItems: "baseline", gap: 10 }}>
|
||||
<span
|
||||
style={{ fontSize: 15, fontWeight: 700, color: "var(--text)" }}
|
||||
>
|
||||
{t("vision.title")}
|
||||
</span>
|
||||
<code
|
||||
style={{
|
||||
fontSize: 11,
|
||||
color: "var(--text-muted)",
|
||||
fontFamily: "var(--font-mono)",
|
||||
overflow: "hidden",
|
||||
textOverflow: "ellipsis",
|
||||
whiteSpace: "nowrap",
|
||||
}}
|
||||
>
|
||||
{configPath || "data/agent/vision.json"}
|
||||
</code>
|
||||
</div>
|
||||
<button
|
||||
onClick={onClose}
|
||||
style={{
|
||||
background: "none",
|
||||
border: "none",
|
||||
color: "var(--text-muted)",
|
||||
cursor: "pointer",
|
||||
fontSize: 20,
|
||||
lineHeight: 1,
|
||||
padding: "2px 6px",
|
||||
}}
|
||||
>
|
||||
×
|
||||
</button>
|
||||
</div>
|
||||
<VisionConfigContent />
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
@@ -103,3 +103,21 @@ export interface PluginsResponse {
|
||||
diagnostics: PluginDiagnostic[];
|
||||
projectResourcesLoaded: boolean;
|
||||
}
|
||||
|
||||
export type ExtensionScope = "global" | "project" | "builtin";
|
||||
export type ExtensionStatus = "enabled" | "disabled" | "blocked";
|
||||
|
||||
export interface ExtensionInfo {
|
||||
name: string;
|
||||
path: string;
|
||||
relativePath: string;
|
||||
source: string;
|
||||
scope: ExtensionScope;
|
||||
status: ExtensionStatus;
|
||||
}
|
||||
|
||||
export interface ExtensionsResponse {
|
||||
extensions: ExtensionInfo[];
|
||||
diagnostics: PluginDiagnostic[];
|
||||
projectResourcesLoaded: boolean;
|
||||
}
|
||||
|
||||
@@ -11,6 +11,8 @@ export const enLocale: LocalePlugin = {
|
||||
"common.mcp": "MCP",
|
||||
"common.skills": "Skills",
|
||||
"common.plugins": "Plugins",
|
||||
"common.extensions": "Extensions",
|
||||
"common.vision": "Vision",
|
||||
"sidebar.hide": "Hide sidebar",
|
||||
"sidebar.show": "Show sidebar",
|
||||
"theme.light": "Switch to light mode",
|
||||
@@ -43,7 +45,7 @@ export const enLocale: LocalePlugin = {
|
||||
"session.output": "Output",
|
||||
"session.cacheRead": "Cache Read",
|
||||
"session.cacheWrite": "Cache Write",
|
||||
"session.cacheHitRate": "Cache Hit Rate",
|
||||
"session.cacheHitRate": "Cache Hit Rate (Main Session)",
|
||||
"session.cost": "Cost",
|
||||
"session.context": "Context",
|
||||
"session.copy": "Copy {value}",
|
||||
@@ -259,6 +261,25 @@ export const enLocale: LocalePlugin = {
|
||||
"mcp.noServers": "No MCP servers configured",
|
||||
"mcp.selectServer": "Select a server to edit",
|
||||
"mcp.restartHint": "Saved changes apply after restarting pi (or /reload).",
|
||||
"vision.title": "Vision config",
|
||||
"vision.loading": "Loading…",
|
||||
"vision.backend.ollama": "Local Ollama",
|
||||
"vision.backend.ollamaHint": "Free and private — needs a local Ollama with a vision model pulled",
|
||||
"vision.backend.openai": "OpenAI-compatible API",
|
||||
"vision.backend.openaiHint": "Any OpenAI-compatible vision endpoint (cloud or self-hosted)",
|
||||
"vision.ollama.host": "Ollama address",
|
||||
"vision.ollama.hostPlaceholder": "e.g. http://localhost:11434",
|
||||
"vision.ollama.model": "Vision model",
|
||||
"vision.ollama.modelPlaceholder": "e.g. qwen3-vl:8b",
|
||||
"vision.openai.baseUrl": "Base URL",
|
||||
"vision.openai.baseUrlPlaceholder": "e.g. https://api.openai.com/v1",
|
||||
"vision.openai.apiKey": "API Key",
|
||||
"vision.openai.model": "Model",
|
||||
"vision.openai.modelPlaceholder": "e.g. gpt-4o-mini / glm-4.5v",
|
||||
"vision.effective": "Saved config takes effect on the next image request — no restart needed. Uploaded images are auto-transcribed into text for text-only main models (e.g. DeepSeek).",
|
||||
"vision.loadFailed": "Failed to load / save config",
|
||||
"vision.save": "Save",
|
||||
"vision.saved": "Saved — takes effect on the next image request",
|
||||
"i18n.close": "Close",
|
||||
"i18n.copy": "Copy",
|
||||
"i18n.copied": "Copied",
|
||||
@@ -389,6 +410,19 @@ export const enLocale: LocalePlugin = {
|
||||
"i18n.configuredVersion": "configured {version}",
|
||||
"i18n.extensions": "Extensions",
|
||||
"i18n.prompts": "Prompts",
|
||||
"extensions.app": "App",
|
||||
"extensions.project": "Project",
|
||||
"extensions.builtin": "Built-in",
|
||||
"extensions.enabled": "enabled",
|
||||
"extensions.disabled": "disabled",
|
||||
"extensions.blocked": "blocked",
|
||||
"extensions.total": "extensions",
|
||||
"extensions.status": "Status",
|
||||
"extensions.source": "Source",
|
||||
"extensions.path": "Path",
|
||||
"extensions.noExtensions": "No direct extensions found",
|
||||
"extensions.projectResourcesBlocked": "Project extensions are blocked until this project is trusted.",
|
||||
"extensions.diagnostics": "diagnostics",
|
||||
"i18n.themes": "Themes",
|
||||
"i18n.resourceCount": "{count} {label}",
|
||||
"i18n.extensionShort": "ext",
|
||||
|
||||
@@ -11,6 +11,8 @@ export const zhCNLocale: LocalePlugin = {
|
||||
"common.mcp": "MCP",
|
||||
"common.skills": "技能",
|
||||
"common.plugins": "插件",
|
||||
"common.extensions": "扩展",
|
||||
"common.vision": "视觉",
|
||||
"sidebar.hide": "隐藏侧边栏",
|
||||
"sidebar.show": "显示侧边栏",
|
||||
"theme.light": "切换到浅色模式",
|
||||
@@ -43,7 +45,7 @@ export const zhCNLocale: LocalePlugin = {
|
||||
"session.output": "输出",
|
||||
"session.cacheRead": "缓存读取",
|
||||
"session.cacheWrite": "缓存写入",
|
||||
"session.cacheHitRate": "缓存命中率",
|
||||
"session.cacheHitRate": "缓存命中率(主会话)",
|
||||
"session.cost": "费用",
|
||||
"session.context": "上下文",
|
||||
"session.copy": "复制{value}",
|
||||
@@ -259,6 +261,25 @@ export const zhCNLocale: LocalePlugin = {
|
||||
"mcp.noServers": "未配置 MCP 服务器",
|
||||
"mcp.selectServer": "选择一个服务器进行编辑",
|
||||
"mcp.restartHint": "保存的更改在重启 pi(或 /reload)后生效。",
|
||||
"vision.title": "视觉配置",
|
||||
"vision.loading": "加载中…",
|
||||
"vision.backend.ollama": "本地 Ollama",
|
||||
"vision.backend.ollamaHint": "免费、私密——本机 Ollama 需已拉取视觉模型",
|
||||
"vision.backend.openai": "OpenAI 兼容 API",
|
||||
"vision.backend.openaiHint": "任意 OpenAI 兼容的视觉端点(云端或自托管)",
|
||||
"vision.ollama.host": "Ollama 地址",
|
||||
"vision.ollama.hostPlaceholder": "如 http://localhost:11434",
|
||||
"vision.ollama.model": "视觉模型",
|
||||
"vision.ollama.modelPlaceholder": "如 qwen3-vl:8b",
|
||||
"vision.openai.baseUrl": "接口地址",
|
||||
"vision.openai.baseUrlPlaceholder": "如 https://api.openai.com/v1",
|
||||
"vision.openai.apiKey": "API 密钥",
|
||||
"vision.openai.model": "模型",
|
||||
"vision.openai.modelPlaceholder": "如 gpt-4o-mini / glm-4.5v",
|
||||
"vision.effective": "保存后下次识图请求立即生效,无需重启。上传图片会自动转录为文本,供纯文本主模型(如 DeepSeek)使用。",
|
||||
"vision.loadFailed": "配置加载/保存失败",
|
||||
"vision.save": "保存",
|
||||
"vision.saved": "已保存——下次识图请求生效",
|
||||
"i18n.close": "关闭",
|
||||
"i18n.copy": "复制",
|
||||
"i18n.copied": "已复制",
|
||||
@@ -389,6 +410,19 @@ export const zhCNLocale: LocalePlugin = {
|
||||
"i18n.configuredVersion": "已配置 {version}",
|
||||
"i18n.extensions": "扩展",
|
||||
"i18n.prompts": "提示词",
|
||||
"extensions.app": "应用",
|
||||
"extensions.project": "项目",
|
||||
"extensions.builtin": "内置",
|
||||
"extensions.enabled": "已启用",
|
||||
"extensions.disabled": "已禁用",
|
||||
"extensions.blocked": "已阻止",
|
||||
"extensions.total": "个扩展",
|
||||
"extensions.status": "状态",
|
||||
"extensions.source": "来源",
|
||||
"extensions.path": "路径",
|
||||
"extensions.noExtensions": "未找到直接扩展",
|
||||
"extensions.projectResourcesBlocked": "项目尚未受信任,项目扩展已被阻止。",
|
||||
"extensions.diagnostics": "条诊断信息",
|
||||
"i18n.themes": "主题",
|
||||
"i18n.resourceCount": "{count}{label}",
|
||||
"i18n.extensionShort": "扩展",
|
||||
|
||||
@@ -3,7 +3,8 @@
|
||||
"private": true,
|
||||
"pi": {
|
||||
"extensions": [
|
||||
"./searxng-search.ts"
|
||||
"./searxng-search.ts",
|
||||
"./vision.ts"
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,539 @@
|
||||
/**
|
||||
* Vision extension: describe images for text-only main models (e.g. DeepSeek).
|
||||
*
|
||||
* Two interchangeable backends, selected by VISION_BACKEND (default "ollama"):
|
||||
* - "ollama": local Ollama vision model via the native /api/chat endpoint.
|
||||
* OLLAMA_HOST (default http://localhost:11434)
|
||||
* OLLAMA_VISION_MODEL (default qwen3-vl:8b)
|
||||
* - "openai": any OpenAI-compatible vision API (cloud or self-hosted).
|
||||
* VISION_OPENAI_BASE_URL e.g. https://api.openai.com/v1
|
||||
* VISION_OPENAI_API_KEY
|
||||
* VISION_OPENAI_MODEL e.g. gpt-4o-mini, glm-4.5v, qwen-vl-max
|
||||
*
|
||||
* Why not Ollama's OpenAI-compatible /v1 endpoint: it moves qwen3-family
|
||||
* reasoning output into the `reasoning` field with an empty `content`, which Pi
|
||||
* treats as an empty reply. The native /api/chat with `think: false` returns a
|
||||
* normal textual answer, so the Ollama backend deliberately bypasses /v1.
|
||||
*/
|
||||
|
||||
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
||||
import { createHash } from "node:crypto";
|
||||
import { readFile, stat } from "node:fs/promises";
|
||||
import { homedir } from "node:os";
|
||||
import { join } from "node:path";
|
||||
import { Type } from "typebox";
|
||||
|
||||
const REQUEST_TIMEOUT_MS = 180_000;
|
||||
const MAX_IMAGE_BYTES = 7 * 1024 * 1024; // Ollama's built-in per-image limit
|
||||
const DEFAULT_MAX_TOKENS = 4096; // generous: thinking-based APIs spend budget on reasoning first
|
||||
|
||||
const DEFAULT_PROMPT = [
|
||||
"请详细描述这张图片,输出要求:",
|
||||
"1. 完整 OCR:按阅读顺序逐字转录所有可见文字(含 UI 标签、按钮、错误信息、代码、数字、日期),保留版式线索。",
|
||||
"2. 版式与视觉结构:区域、颜色、形状;UI 截图、图表(坐标轴/数值/趋势)、表格用 markdown 重建、流程图(节点/连线/方向)。",
|
||||
"3. 语义总结:2-3 句概括图片内容与关键信息。",
|
||||
"4. 模糊/截断/有歧义处明确说明,不要猜测。",
|
||||
"5. 多张图片时,分别描述每张,用【图片1】【图片2】…标注。",
|
||||
"主模型看不到图,完全依赖你的转录,文字务必穷尽。",
|
||||
].join("\n");
|
||||
|
||||
const visionParams = Type.Object({
|
||||
image_paths: Type.Array(
|
||||
Type.String({
|
||||
description: "图片文件路径(绝对或相对路径),至少一个",
|
||||
minLength: 1,
|
||||
}),
|
||||
{ minItems: 1, maxItems: 8 },
|
||||
),
|
||||
backend: Type.Optional(
|
||||
Type.Union([Type.Literal("ollama"), Type.Literal("openai")], {
|
||||
description:
|
||||
"视觉后端:ollama(本地,默认)或 openai(OpenAI 兼容 API)。默认取 VISION_BACKEND 环境变量",
|
||||
}),
|
||||
),
|
||||
model: Type.Optional(
|
||||
Type.String({
|
||||
description:
|
||||
"视觉模型 tag/ID,默认取后端对应环境变量(OLLAMA_VISION_MODEL 或 VISION_OPENAI_MODEL)",
|
||||
}),
|
||||
),
|
||||
prompt: Type.Optional(
|
||||
Type.String({
|
||||
description: `自定义识图指令,默认:${DEFAULT_PROMPT.split("\n")[0]}`,
|
||||
}),
|
||||
),
|
||||
});
|
||||
|
||||
function envOr(name: string, fallback: string): string {
|
||||
const value = process.env[name]?.trim();
|
||||
return value && value.length > 0 ? value : fallback;
|
||||
}
|
||||
|
||||
function mimeFromPath(filePath: string): string {
|
||||
const ext = filePath.split(".").pop()?.toLowerCase() ?? "";
|
||||
switch (ext) {
|
||||
case "png":
|
||||
return "image/png";
|
||||
case "jpg":
|
||||
case "jpeg":
|
||||
return "image/jpeg";
|
||||
case "webp":
|
||||
return "image/webp";
|
||||
case "gif":
|
||||
return "image/gif";
|
||||
case "bmp":
|
||||
return "image/bmp";
|
||||
default:
|
||||
return "application/octet-stream";
|
||||
}
|
||||
}
|
||||
|
||||
interface LoadedImage {
|
||||
base64: string;
|
||||
mime: string;
|
||||
}
|
||||
|
||||
async function loadImages(imagePaths: string[]): Promise<LoadedImage[]> {
|
||||
const images: LoadedImage[] = [];
|
||||
for (const filePath of imagePaths) {
|
||||
const fileStat = await stat(filePath).catch(() => null);
|
||||
if (!fileStat) throw new Error(`vision: file not found: ${filePath}`);
|
||||
if (!fileStat.isFile())
|
||||
throw new Error(`vision: not a regular file: ${filePath}`);
|
||||
if (fileStat.size === 0) throw new Error(`vision: empty file: ${filePath}`);
|
||||
if (fileStat.size > MAX_IMAGE_BYTES) {
|
||||
throw new Error(
|
||||
`vision: ${filePath} is ${fileStat.size} bytes, exceeding the ${MAX_IMAGE_BYTES} byte limit. Resize or compress the image first.`,
|
||||
);
|
||||
}
|
||||
const buffer = await readFile(filePath);
|
||||
images.push({
|
||||
base64: buffer.toString("base64"),
|
||||
mime: mimeFromPath(filePath),
|
||||
});
|
||||
}
|
||||
return images;
|
||||
}
|
||||
|
||||
function requestSignal(signal: AbortSignal | undefined): AbortSignal {
|
||||
const timeoutSignal = AbortSignal.timeout(REQUEST_TIMEOUT_MS);
|
||||
return signal ? AbortSignal.any([signal, timeoutSignal]) : timeoutSignal;
|
||||
}
|
||||
|
||||
async function describeWithOllama(
|
||||
baseUrl: string,
|
||||
model: string,
|
||||
prompt: string,
|
||||
images: LoadedImage[],
|
||||
signal: AbortSignal | undefined,
|
||||
): Promise<string> {
|
||||
// Ollama defaults to a small num_ctx (4096 on this setup); multi-image
|
||||
// requests blow past it. Raise explicitly — the model supports 262k.
|
||||
const numCtx = parseInt(envOr("OLLAMA_NUM_CTX", "16384"), 10) || 16384;
|
||||
const body = {
|
||||
model,
|
||||
messages: [
|
||||
{
|
||||
role: "user",
|
||||
content: prompt,
|
||||
images: images.map((img) => img.base64),
|
||||
},
|
||||
],
|
||||
stream: false,
|
||||
think: false,
|
||||
options: { temperature: 0, num_ctx: numCtx },
|
||||
};
|
||||
|
||||
let response: Response;
|
||||
try {
|
||||
response = await fetch(`${baseUrl.replace(/\/+$/, "")}/api/chat`, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify(body),
|
||||
signal: requestSignal(signal),
|
||||
});
|
||||
} catch (error) {
|
||||
const cause = error instanceof Error ? error.message : String(error);
|
||||
throw new Error(
|
||||
`vision: cannot reach Ollama at ${baseUrl} (${cause}). Is Ollama running?`,
|
||||
);
|
||||
}
|
||||
|
||||
if (!response.ok) {
|
||||
const detail = (await response.text().catch(() => "")).slice(0, 500);
|
||||
throw new Error(
|
||||
`vision: Ollama returned HTTP ${response.status}${detail ? `: ${detail}` : ""}`,
|
||||
);
|
||||
}
|
||||
|
||||
const data = (await response.json()) as { message?: { content?: string } };
|
||||
const content = data.message?.content?.trim();
|
||||
if (!content) {
|
||||
throw new Error(
|
||||
`vision: Ollama model ${model} returned an empty response. ` +
|
||||
"Check `ollama pull qwen3-vl:8b` and that the model supports vision.",
|
||||
);
|
||||
}
|
||||
return content;
|
||||
}
|
||||
|
||||
async function describeWithOpenAI(
|
||||
baseUrl: string,
|
||||
apiKey: string,
|
||||
model: string,
|
||||
prompt: string,
|
||||
images: LoadedImage[],
|
||||
signal: AbortSignal | undefined,
|
||||
): Promise<string> {
|
||||
const body = {
|
||||
model,
|
||||
messages: [
|
||||
{
|
||||
role: "user",
|
||||
content: [
|
||||
{ type: "text", text: prompt },
|
||||
...images.map((img) => ({
|
||||
type: "image_url",
|
||||
image_url: { url: `data:${img.mime};base64,${img.base64}` },
|
||||
})),
|
||||
],
|
||||
},
|
||||
],
|
||||
max_tokens: DEFAULT_MAX_TOKENS,
|
||||
temperature: 0,
|
||||
};
|
||||
|
||||
let response: Response;
|
||||
try {
|
||||
response = await fetch(`${baseUrl.replace(/\/+$/, "")}/chat/completions`, {
|
||||
method: "POST",
|
||||
headers: {
|
||||
"Content-Type": "application/json",
|
||||
Authorization: `Bearer ${apiKey}`,
|
||||
},
|
||||
body: JSON.stringify(body),
|
||||
signal: requestSignal(signal),
|
||||
});
|
||||
} catch (error) {
|
||||
const cause = error instanceof Error ? error.message : String(error);
|
||||
throw new Error(`vision: cannot reach ${baseUrl} (${cause}).`);
|
||||
}
|
||||
|
||||
if (!response.ok) {
|
||||
const detail = (await response.text().catch(() => "")).slice(0, 500);
|
||||
throw new Error(
|
||||
`vision: API returned HTTP ${response.status}${detail ? `: ${detail}` : ""}. ` +
|
||||
"Check VISION_OPENAI_BASE_URL / VISION_OPENAI_API_KEY / VISION_OPENAI_MODEL.",
|
||||
);
|
||||
}
|
||||
|
||||
const data = (await response.json()) as {
|
||||
choices?: Array<{
|
||||
message?: {
|
||||
content?: string;
|
||||
reasoning_content?: string;
|
||||
reasoning?: string;
|
||||
};
|
||||
finish_reason?: string;
|
||||
}>;
|
||||
};
|
||||
const choice = data.choices?.[0];
|
||||
const content = choice?.message?.content?.trim();
|
||||
if (content) return content;
|
||||
|
||||
const reasoning =
|
||||
choice?.message?.reasoning_content || choice?.message?.reasoning;
|
||||
const finish = choice?.finish_reason ?? "unknown";
|
||||
if (reasoning) {
|
||||
throw new Error(
|
||||
`vision: model ${model} returned only reasoning (finish=${finish}). ` +
|
||||
"If it is a thinking model, pick a non-thinking vision model or raise max_tokens.",
|
||||
);
|
||||
}
|
||||
throw new Error(
|
||||
`vision: model ${model} returned an empty response (finish=${finish}).`,
|
||||
);
|
||||
}
|
||||
|
||||
interface DescribeOptions {
|
||||
backend?: string;
|
||||
model?: string;
|
||||
prompt?: string;
|
||||
signal?: AbortSignal;
|
||||
}
|
||||
|
||||
/**
|
||||
* Optional file-based configuration, edited from the Web UI vision panel
|
||||
* (writes `$PI_CODING_AGENT_DIR/vision.json`). Environment variables and
|
||||
* per-call parameters take precedence over this file; every read is fresh, so
|
||||
* saving the panel takes effect on the next image request without a restart.
|
||||
*/
|
||||
interface VisionFileConfig {
|
||||
backend?: "ollama" | "openai";
|
||||
ollama?: { host?: string; model?: string };
|
||||
openai?: { baseUrl?: string; apiKey?: string; model?: string };
|
||||
}
|
||||
|
||||
const VISION_CONFIG_DEFAULTS: VisionFileConfig = {
|
||||
backend: "ollama",
|
||||
ollama: { host: "http://localhost:11434", model: "qwen3-vl:8b" },
|
||||
openai: { baseUrl: "", apiKey: "", model: "" },
|
||||
};
|
||||
|
||||
function visionConfigPath(): string {
|
||||
const agentDir = process.env.PI_CODING_AGENT_DIR?.trim();
|
||||
return (
|
||||
(agentDir && agentDir.length > 0
|
||||
? agentDir
|
||||
: join(homedir(), ".pi", "agent")) + "/vision.json"
|
||||
);
|
||||
}
|
||||
|
||||
async function loadVisionFileConfig(): Promise<VisionFileConfig> {
|
||||
try {
|
||||
const raw = await readFile(visionConfigPath(), "utf8");
|
||||
const parsed = JSON.parse(raw) as VisionFileConfig;
|
||||
return parsed && typeof parsed === "object" ? parsed : {};
|
||||
} catch {
|
||||
return {};
|
||||
}
|
||||
}
|
||||
|
||||
async function describeImages(
|
||||
images: LoadedImage[],
|
||||
options: DescribeOptions,
|
||||
): Promise<string> {
|
||||
const fileConfig = await loadVisionFileConfig();
|
||||
const backend =
|
||||
options.backend ??
|
||||
envOr(
|
||||
"VISION_BACKEND",
|
||||
fileConfig.backend ?? VISION_CONFIG_DEFAULTS.backend ?? "ollama",
|
||||
);
|
||||
const prompt = options.prompt ?? DEFAULT_PROMPT;
|
||||
if (backend === "ollama") {
|
||||
const baseUrl = envOr(
|
||||
"OLLAMA_HOST",
|
||||
fileConfig.ollama?.host ??
|
||||
VISION_CONFIG_DEFAULTS.ollama?.host ??
|
||||
"http://localhost:11434",
|
||||
);
|
||||
const model =
|
||||
options.model ??
|
||||
envOr(
|
||||
"OLLAMA_VISION_MODEL",
|
||||
fileConfig.ollama?.model ??
|
||||
VISION_CONFIG_DEFAULTS.ollama?.model ??
|
||||
"qwen3-vl:8b",
|
||||
);
|
||||
return describeWithOllama(baseUrl, model, prompt, images, options.signal);
|
||||
}
|
||||
const baseUrl = envOr(
|
||||
"VISION_OPENAI_BASE_URL",
|
||||
fileConfig.openai?.baseUrl ?? "",
|
||||
);
|
||||
const apiKey = envOr(
|
||||
"VISION_OPENAI_API_KEY",
|
||||
fileConfig.openai?.apiKey ?? "",
|
||||
);
|
||||
const model =
|
||||
options.model ??
|
||||
envOr("VISION_OPENAI_MODEL", fileConfig.openai?.model ?? "");
|
||||
if (!baseUrl || !apiKey || !model) {
|
||||
throw new Error(
|
||||
"vision: the openai backend needs a base URL, API key and model. " +
|
||||
"Configure them in the Web UI vision panel (lower-left) or via " +
|
||||
"VISION_OPENAI_BASE_URL / VISION_OPENAI_API_KEY / VISION_OPENAI_MODEL. " +
|
||||
"Example: https://api.openai.com/v1 + gpt-4o-mini.",
|
||||
);
|
||||
}
|
||||
return describeWithOpenAI(
|
||||
baseUrl,
|
||||
apiKey,
|
||||
model,
|
||||
prompt,
|
||||
images,
|
||||
options.signal,
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* Session-scoped description cache: compaction, session restore, or repeated
|
||||
* turns replay the same image parts; avoid re-running the vision model each time.
|
||||
*/
|
||||
const IMAGE_DESCRIPTION_CACHE = new Map<string, string>();
|
||||
|
||||
function dataUrlToLoadedImage(url: string): LoadedImage | null {
|
||||
const match = /^data:(image\/[a-z0-9.+-]+);base64,(.+)$/i.exec(url);
|
||||
if (!match) return null;
|
||||
return { base64: match[2], mime: match[1] };
|
||||
}
|
||||
|
||||
function imageCacheKey(image: LoadedImage): string {
|
||||
// Whole-image hash: PNG/JPG headers repeat for same dimensions, so a short
|
||||
// prefix would collide across different images of the same size.
|
||||
return createHash("md5").update(image.base64).digest("hex");
|
||||
}
|
||||
|
||||
/**
|
||||
* Per-image description cache. Every request rebuilds the payload from the
|
||||
* session file, which keeps the original image parts, so without this cache
|
||||
* the whole conversation's images would be re-transcribed every turn and
|
||||
* eventually exceed Ollama's context window. Single-image granularity means
|
||||
* only genuinely new images hit the vision model.
|
||||
*/
|
||||
async function getImageDescription(
|
||||
image: LoadedImage,
|
||||
prompt: string,
|
||||
signal: AbortSignal | undefined,
|
||||
): Promise<string> {
|
||||
const key = imageCacheKey(image);
|
||||
const cached = IMAGE_DESCRIPTION_CACHE.get(key);
|
||||
if (cached) return cached;
|
||||
const description = await describeImages([image], { prompt, signal });
|
||||
IMAGE_DESCRIPTION_CACHE.set(key, description);
|
||||
if (IMAGE_DESCRIPTION_CACHE.size > 64) {
|
||||
const oldest = IMAGE_DESCRIPTION_CACHE.keys().next().value;
|
||||
if (oldest !== undefined) IMAGE_DESCRIPTION_CACHE.delete(oldest);
|
||||
}
|
||||
return description;
|
||||
}
|
||||
|
||||
function isTextOnlyModel(model: unknown): boolean {
|
||||
const input = (model as { input?: string[] } | undefined)?.input;
|
||||
return Array.isArray(input) && !input.includes("image");
|
||||
}
|
||||
|
||||
export default function visionExtension(pi: ExtensionAPI) {
|
||||
// Keep user-message images intact for text-only models so the hook below can
|
||||
// transcribe them (pi-ai would otherwise replace them with a text placeholder
|
||||
// before before_provider_request runs). The matching pi-ai patch is applied
|
||||
// and replayed by scripts/configure-pi-ai-vision.mjs.
|
||||
process.env.PI_VISION_PASSTHROUGH_IMAGES = "1";
|
||||
pi.registerTool({
|
||||
name: "vision",
|
||||
label: "Vision (image description)",
|
||||
description:
|
||||
"用视觉模型描述本地图片并返回详细文本(完整 OCR、版式结构、语义总结)。" +
|
||||
"适用于主模型不支持图片输入(如 DeepSeek)时查看截图/图表/文档/照片。" +
|
||||
"后端可配置:ollama(默认,本地 qwen3-vl:8b,免费私密)或 openai(任意 OpenAI 兼容视觉 API," +
|
||||
"设 VISION_BACKEND=openai + VISION_OPENAI_BASE_URL/API_KEY/MODEL)。",
|
||||
promptSnippet:
|
||||
"Describe local images using a vision model (Ollama or OpenAI-compatible)",
|
||||
promptGuidelines: [
|
||||
"Use vision when the user asks you to look at an image (screenshot, diagram, chart, document, photo) and the current model cannot receive image attachments.",
|
||||
"Pass image file paths that exist on disk; the tool reads and encodes them itself.",
|
||||
"One call can describe up to 8 images; prefer batching related images into a single call.",
|
||||
"The returned text is the vision model's transcription — relay it faithfully, quoting OCR text verbatim.",
|
||||
"If the call fails because no vision backend is configured, report the missing environment variables.",
|
||||
],
|
||||
parameters: visionParams,
|
||||
executionMode: "sequential",
|
||||
async execute(_toolCallId, params, signal) {
|
||||
const prompt = params.prompt ?? DEFAULT_PROMPT;
|
||||
const images = await loadImages(params.image_paths);
|
||||
const content = await describeImages(images, {
|
||||
backend: params.backend,
|
||||
model: params.model,
|
||||
prompt,
|
||||
signal,
|
||||
});
|
||||
|
||||
return {
|
||||
content: [{ type: "text", text: content }],
|
||||
details: {
|
||||
backend: params.backend ?? envOr("VISION_BACKEND", "ollama"),
|
||||
model: params.model ?? undefined,
|
||||
imageCount: images.length,
|
||||
},
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
// --- Automatic image transcription for text-only main models ---
|
||||
// Web UI image uploads arrive as base64 image_url parts in the user message.
|
||||
// A text-only model (e.g. DeepSeek) cannot receive them — DeepSeek rejects
|
||||
// the request with HTTP 400. This hook transcribes the images through the
|
||||
// configured vision backend and replaces them with text before the request
|
||||
// is sent, so the main model keeps working seamlessly.
|
||||
pi.on("before_provider_request", async (event, ctx) => {
|
||||
const payload = event.payload as
|
||||
| { messages?: Array<{ role?: string; content?: unknown }> }
|
||||
| undefined;
|
||||
const messages = payload?.messages;
|
||||
if (!Array.isArray(messages) || messages.length === 0) return;
|
||||
if (!isTextOnlyModel(ctx.model)) return; // vision-capable models pass through untouched
|
||||
|
||||
// Collect images and replace each with a numbered text placeholder.
|
||||
const images: LoadedImage[] = [];
|
||||
let dirty = false;
|
||||
for (const msg of messages) {
|
||||
if (msg?.role !== "user" || !Array.isArray(msg.content)) continue;
|
||||
const newContent: unknown[] = [];
|
||||
for (const part of msg.content) {
|
||||
const p = part as { type?: string; image_url?: { url?: string } };
|
||||
if (p?.type === "image_url" && typeof p.image_url?.url === "string") {
|
||||
const img = dataUrlToLoadedImage(p.image_url.url);
|
||||
if (img) {
|
||||
images.push(img);
|
||||
newContent.push({ type: "text", text: `[图片 ${images.length}]` });
|
||||
dirty = true;
|
||||
} else {
|
||||
newContent.push(part);
|
||||
}
|
||||
} else {
|
||||
newContent.push(part);
|
||||
}
|
||||
}
|
||||
msg.content = newContent;
|
||||
}
|
||||
if (!dirty || images.length === 0) return;
|
||||
|
||||
// Transcribe per image: only genuinely new images hit the vision model
|
||||
// (history resolves from cache), so one request never re-batches the
|
||||
// whole conversation's images and exceeds Ollama's context window.
|
||||
const transcribed: string[] = [];
|
||||
for (let i = 0; i < images.length; i++) {
|
||||
try {
|
||||
const desc = await getImageDescription(
|
||||
images[i],
|
||||
DEFAULT_PROMPT,
|
||||
ctx.signal,
|
||||
);
|
||||
transcribed.push(`【图片${i + 1}】\n${desc}`);
|
||||
} catch (error) {
|
||||
const cause = error instanceof Error ? error.message : String(error);
|
||||
transcribed.push(`【图片${i + 1}】\n[图片处理失败:${cause}]`);
|
||||
}
|
||||
}
|
||||
const description = transcribed.join("\n\n");
|
||||
|
||||
// Place the full transcription on the LAST image-carrying message (the one
|
||||
// the model is actively processing); earlier image messages reference it.
|
||||
// Putting it on the first message instead hid it in history.
|
||||
const targets: Array<{ part: { type?: string; text?: string } }> = [];
|
||||
for (const msg of messages) {
|
||||
if (msg?.role !== "user" || !Array.isArray(msg.content)) continue;
|
||||
for (const part of msg.content) {
|
||||
const p = part as { type?: string; text?: string };
|
||||
if (
|
||||
p?.type === "text" &&
|
||||
typeof p.text === "string" &&
|
||||
p.text.startsWith("[图片 ")
|
||||
) {
|
||||
targets.push({ part: p });
|
||||
}
|
||||
}
|
||||
}
|
||||
const lastTarget = targets[targets.length - 1];
|
||||
if (lastTarget) {
|
||||
lastTarget.part.text = `[用户上传了 ${images.length} 张图片,以下为视觉模型转录的文本描述]\n\n${description}`;
|
||||
}
|
||||
for (const { part } of targets) {
|
||||
if (part !== lastTarget?.part) {
|
||||
part.text = "(图片描述见最新消息中的综合转录)";
|
||||
}
|
||||
}
|
||||
return payload;
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,122 @@
|
||||
/**
|
||||
* pi-ai vision passthrough patch.
|
||||
*
|
||||
* Pi-ai's `transform-messages.js` replaces image parts with a text placeholder
|
||||
* ("(image omitted: model does not support images)") for text-only models BEFORE
|
||||
* the before_provider_request hook runs, so the vision extension's automatic
|
||||
* transcription hook can never see the image. This patch adds an opt-in switch:
|
||||
* when PI_VISION_PASSTHROUGH_IMAGES=1 (set by the vision extension at load),
|
||||
* user-message image parts are kept intact so the vision hook can transcribe
|
||||
* them; without the env var the original placeholder behavior is preserved
|
||||
* (safe for users without the vision extension).
|
||||
*
|
||||
* Idempotent, replayed by run.mjs and install-managed-packages.mjs, mirroring
|
||||
* configure-pi-memory.mjs.
|
||||
*/
|
||||
|
||||
import { existsSync, readFileSync, writeFileSync } from "node:fs";
|
||||
import { join, resolve } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
|
||||
const supportedVersion = "0.83.0";
|
||||
|
||||
function replaceOnce(source, before, after, label) {
|
||||
if (source.includes(after)) return source;
|
||||
|
||||
const first = source.indexOf(before);
|
||||
const last = source.lastIndexOf(before);
|
||||
if (first < 0 || first !== last) {
|
||||
throw new Error(`Cannot apply pi-ai vision patch: ${label}`);
|
||||
}
|
||||
return source.slice(0, first) + after + source.slice(first + before.length);
|
||||
}
|
||||
|
||||
export function patchPiAiSource(input) {
|
||||
let source = input.replaceAll("\r\n", "\n");
|
||||
|
||||
source = replaceOnce(
|
||||
source,
|
||||
` return messages.map((msg) => {
|
||||
if (msg.role === "user" && Array.isArray(msg.content)) {
|
||||
return {
|
||||
...msg,
|
||||
content: replaceImagesWithPlaceholder(msg.content, NON_VISION_USER_IMAGE_PLACEHOLDER),
|
||||
};
|
||||
}`,
|
||||
` const keepUserImages =
|
||||
typeof process !== "undefined" && process.env?.PI_VISION_PASSTHROUGH_IMAGES === "1";
|
||||
return messages.map((msg) => {
|
||||
if (msg.role === "user" && Array.isArray(msg.content) && !keepUserImages) {
|
||||
return {
|
||||
...msg,
|
||||
content: replaceImagesWithPlaceholder(msg.content, NON_VISION_USER_IMAGE_PLACEHOLDER),
|
||||
};
|
||||
}`,
|
||||
"downgradeUnsupportedImages user-image passthrough switch",
|
||||
);
|
||||
|
||||
return source;
|
||||
}
|
||||
|
||||
export function isPiAiVisionConfigured(source) {
|
||||
return source.includes("PI_VISION_PASSTHROUGH_IMAGES === \"1\"");
|
||||
}
|
||||
|
||||
export function defaultPiAiDir() {
|
||||
// Managed installs place packages under <agentDir>/npm/node_modules; pi-ai is
|
||||
// also a direct dependency of pi-web, which wins in dev. Prefer pi-web's copy.
|
||||
const candidates = [
|
||||
resolve(process.cwd(), "pi-web", "node_modules", "@earendil-works", "pi-ai"),
|
||||
resolve(process.cwd(), "node_modules", "@earendil-works", "pi-ai"),
|
||||
];
|
||||
for (const candidate of candidates) {
|
||||
if (existsSync(join(candidate, "package.json"))) return candidate;
|
||||
}
|
||||
throw new Error("pi-ai package not found under pi-web/node_modules or node_modules");
|
||||
}
|
||||
|
||||
export function configurePiAiVision({ piAiDir = defaultPiAiDir(), quiet = false } = {}) {
|
||||
const packageJsonPath = join(piAiDir, "package.json");
|
||||
const sourcePath = join(piAiDir, "dist", "api", "transform-messages.js");
|
||||
|
||||
if (!existsSync(packageJsonPath) || !existsSync(sourcePath)) {
|
||||
return { status: "missing", sourcePath };
|
||||
}
|
||||
|
||||
let packageJson;
|
||||
try {
|
||||
packageJson = JSON.parse(readFileSync(packageJsonPath, "utf8"));
|
||||
} catch (error) {
|
||||
throw new Error(`Cannot parse ${packageJsonPath}: ${error.message}`);
|
||||
}
|
||||
if (packageJson.version !== supportedVersion) {
|
||||
throw new Error(
|
||||
`Unsupported pi-ai version ${packageJson.version ?? "unknown"}; expected ${supportedVersion}.`,
|
||||
);
|
||||
}
|
||||
|
||||
const source = readFileSync(sourcePath, "utf8");
|
||||
const patched = patchPiAiSource(source);
|
||||
if (patched === source) {
|
||||
if (!quiet) console.log("pi-ai vision passthrough ready.");
|
||||
return { status: "ready", sourcePath };
|
||||
}
|
||||
|
||||
writeFileSync(sourcePath, patched, "utf8");
|
||||
if (!quiet) console.log("Configured pi-ai vision passthrough.");
|
||||
return { status: "patched", sourcePath };
|
||||
}
|
||||
|
||||
const invokedPath = process.argv[1] ? resolve(process.argv[1]) : "";
|
||||
if (invokedPath === fileURLToPath(import.meta.url)) {
|
||||
try {
|
||||
const result = configurePiAiVision();
|
||||
if (result.status === "missing") {
|
||||
console.error("pi-ai is not installed. Run npm run setup first.");
|
||||
process.exit(1);
|
||||
}
|
||||
} catch (error) {
|
||||
console.error(error.message);
|
||||
process.exit(1);
|
||||
}
|
||||
}
|
||||
@@ -2,6 +2,7 @@ import { spawnSync } from "node:child_process";
|
||||
import { existsSync, readFileSync } from "node:fs";
|
||||
import { join } from "node:path";
|
||||
import { configurePiMemory } from "./configure-pi-memory.mjs";
|
||||
import { configurePiAiVision } from "./configure-pi-ai-vision.mjs";
|
||||
import { agentDir, managedEnvironment, rootDir } from "./profile.mjs";
|
||||
|
||||
const settings = JSON.parse(
|
||||
@@ -49,3 +50,8 @@ for (const source of packages) {
|
||||
}
|
||||
|
||||
configurePiMemory({ agentDir });
|
||||
try {
|
||||
configurePiAiVision({ quiet: false });
|
||||
} catch (error) {
|
||||
console.warn(`[vision] pi-ai passthrough patch skipped: ${error.message}`);
|
||||
}
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
import { spawn } from "node:child_process";
|
||||
import { join } from "node:path";
|
||||
import { configurePiMemory } from "./configure-pi-memory.mjs";
|
||||
import { configurePiAiVision } from "./configure-pi-ai-vision.mjs";
|
||||
import { managedEnvironment, rootDir } from "./profile.mjs";
|
||||
import {
|
||||
isLocalPortListening,
|
||||
@@ -26,6 +27,11 @@ const memoryConfiguration = configurePiMemory({ quiet: true });
|
||||
if (memoryConfiguration.status === "missing") {
|
||||
console.warn("[memory] pi-memory is not installed; run npm run setup to enable managed memory");
|
||||
}
|
||||
try {
|
||||
configurePiAiVision({ quiet: true });
|
||||
} catch (error) {
|
||||
console.warn(`[vision] pi-ai passthrough patch skipped: ${error.message}`);
|
||||
}
|
||||
if (await isLocalPortListening(30141)) {
|
||||
console.log("[storage] skipped automatic maintenance because Pi Web is already running");
|
||||
} else {
|
||||
|
||||
Reference in New Issue
Block a user