Adds both halves of the new release as composer chips: `qwen/qwen3.8-27b` on OpenRouter and ggml-org's Q4_K_M build served by the EVOX2 LM Studio. The local one exposed a wrong assumption. Provider was inferred from the id's shape — a `/` meant OpenRouter — but LM Studio serves this model as `ggml-org/qwen3.8-27b`, prefix and all, while OpenRouter has a `qwen/qwen3.8-27b` of its own. The two differ by provider, not by shape, so the runner now resolves the provider from `~/.pi/agent/models.json` (a model must be listed there for pi to run it locally anyway) and only falls back to the shape for ids it doesn't know — which keeps the whole OpenRouter catalogue runnable without registering ids by hand. That also fixes `laguna-s-2.1`, which the bare-id rule aimed at LM Studio's port instead of its own llama.cpp one. Session pricing carried the same assumption, so a local run with a prefixed id would have been billed at OpenRouter's estimate; it now prices from the declared provider and stays $0. Qwen 3.6 stays pickable, demoted from `latest`; the pi default model follows to 3.8. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
104 lines
5.7 KiB
Markdown
104 lines
5.7 KiB
Markdown
# pi runner — the ai-agent's second harness
|
|
|
|
Runs viewer-spawned sessions on **[pi](https://pi.dev)** (`@earendil-works/
|
|
pi-coding-agent`) instead of the `claude` CLI, with open models: Qwen 3.8 (or
|
|
3.6) via **OpenRouter** (hosted) or the **EVOX2 box's LM Studio** (local, free).
|
|
|
|
## How harness switching works
|
|
|
|
The **model chip in the composer is the harness switch**. `/api/models` returns
|
|
every pickable model tagged with its `harness` (`claude` | `pi`); picking a
|
|
qwen chip routes the session through this runner, picking a Claude family runs
|
|
`claude` exactly as before. The chain:
|
|
|
|
```
|
|
composer chip → POST /api/spawn {model, harness}
|
|
→ sidecar /spawn → harness == "pi" ? RUNNER_BIN : CLAUDE_BIN
|
|
→ runner: pi --mode json --session-id <uuid> …
|
|
→ transcript mirrored to ~/.pi-runner/transcripts/<cwd>/<uuid>.jsonl
|
|
→ container watches it (PI_TRANSCRIPTS_DIR mount) → SSE → viewer
|
|
```
|
|
|
|
- The spawn stamps `harness` into the conversation metadata sidecar, so a
|
|
**resume with no model picked continues on the same harness** (backend
|
|
fallback); picking a chip from the other harness switches mid-conversation.
|
|
- The runner accepts the **same flag shape as `claude`**
|
|
(`-p --session-id/--resume --model`), so the sidecar treats both harnesses
|
|
identically — pidfiles, launch probe, SIGINT interrupt, `reused` idempotency
|
|
all work unchanged.
|
|
- **Transcript compatibility is the design**: the runner translates pi's
|
|
`message_end` events (user / assistant / toolResult) into Claude-Code-schema
|
|
JSONL, so the viewer's parser, cost-by-tool, SSE streaming, DONE detection and
|
|
lifecycle checks need no pi-specific code. Adding a third harness = another
|
|
runner that writes the same schema.
|
|
- **Built-in tool calls are remapped to Claude's shape** (`mapToolCall` in
|
|
`cli.mjs`): pi's lowercase `read/write/edit/bash/grep/find/ls` with
|
|
`path`/`edits[]` become `Read/Write/Edit|MultiEdit/Bash/Grep/Glob/LS` with
|
|
`file_path`/`old_string`/`new_string` (bash `timeout` s→ms). That's what
|
|
makes the viewer's cost-by-tool buckets, per-extension sub-breakdowns, auto
|
|
project/service tags, bash cards and Read image previews light up for pi
|
|
sessions. Unknown (extension) tools pass through and render as generic cards.
|
|
- **Real cost, not estimates**: each assistant record carries pi's own
|
|
per-message provider cost as `costUSD` (omitted when pi reports none). The
|
|
parser (`backend/conversations.py`) prefers it over its price table — so
|
|
OpenRouter runs show the billed amount and free local EVOX2 runs cost $0
|
|
(ids the picker declares `provider: "evox2"`, plus any bare qwen id, also
|
|
price at 0 in the fallback table).
|
|
|
|
## Model → provider routing
|
|
|
|
| model id | provider |
|
|
|---|---|
|
|
| any id listed under a provider in `~/.pi/agent/models.json` — `ggml-org/qwen3.8-27b`, `qwen3.6-35b-a3b`, `laguna-s-2.1` | that provider: EVOX2 LM Studio (`evox2`) or EVOX2 llama.cpp (`evox2-laguna`); the box is WoL-woken automatically |
|
|
| everything else with a vendor prefix — `openai/gpt-oss-20b`, `qwen/qwen3.8-27b`, … | OpenRouter (`OPENROUTER_API_KEY`, auto-loaded from `~/homelab/.env.claude`) |
|
|
|
|
**pi's own registry decides, and the id shape is only the fallback.** It has to:
|
|
LM Studio serves the local Qwen 3.8 as `ggml-org/qwen3.8-27b` — vendor prefix
|
|
and all — while OpenRouter has a `qwen/qwen3.8-27b` of its own, so the two
|
|
differ by provider, not by shape. A model must be listed in that file for pi to
|
|
run it locally anyway, which is what makes it the authority; anything unlisted
|
|
falls back to "has a `/` ⇒ OpenRouter", so any of the several hundred
|
|
OpenRouter models still runs here with no runner change. Registering a new
|
|
local model means adding it under its provider's `models` array there (id,
|
|
`contextWindow`, `compat.thinkingFormat`, zero `cost`) and adding a chip for it
|
|
in `backend/models.py`. Two pickers feed the composer: the
|
|
curated chips in `backend/models.py` (`PI_MODELS_DEFAULT`, override with
|
|
`PI_MODELS_JSON`) and the composer's **OpenRouter browser**, which lists the
|
|
whole catalogue with per-model prices from `backend/openrouter.py` (sourced from
|
|
APPE). Those same prices cost the session's transcript, so a run's `$` figure is
|
|
the model's real rate even before pi reports its own `costUSD`.
|
|
|
|
## Permissions
|
|
|
|
`extensions/folder-permissions.ts` blocks tool calls outside the session's
|
|
scope (`PI_PERM_SCOPE`, default = the run's cwd): `write`/`edit` only inside
|
|
the scope, reads inside `PI_PERM_READ` roots (default `$HOME`) + system
|
|
prefixes, `bash` commands may not reference paths outside the readable roots,
|
|
and a deny-always list protects `.env`/credentials/`.ssh` from writes. This is
|
|
*stricter* than the claude harness's `bypassPermissions` — pi itself has no
|
|
permission system, so the extension is the enforcement point.
|
|
|
|
## pi session state
|
|
|
|
pi sessions persist under `~/.pi-runner/sessions/` keyed by the **same UUID**
|
|
the viewer uses (`--session-dir` + `--session-id`), so resume is stateless:
|
|
the same id simply reopens the session. Skills load from `~/.claude/skills`
|
|
via `~/.pi/agent/settings.json`; pi reads the repo `CLAUDE.md` natively.
|
|
|
|
## Files
|
|
|
|
- `cli.mjs` — the runner (arg parsing, EVOX2 wake, pi spawn, transcript writer).
|
|
- `extensions/folder-permissions.ts` — the permission gate.
|
|
- Env knobs: `PI_BIN`, `PI_RUNNER_ROOT`, `RUNNER_DEFAULT_MODEL`,
|
|
`RUNNER_PERM_EXT`, `RUNNER_ENV_FILE`, `EVOX2_URL`/`EVOX2_MAC`/`EVOX2_BCAST`,
|
|
and on the sidecar side `RUNNER_BIN`, `SIDECAR_PI_MODEL`.
|
|
|
|
## Setup (already done on this host)
|
|
|
|
```bash
|
|
cd services/ai-agent/runner && npm install --ignore-scripts
|
|
# ~/.pi/agent/models.json — evox2 provider (LM Studio token inside)
|
|
# ~/.pi/agent/settings.json — {"skills": ["~/.claude/skills"]}
|
|
# OPENROUTER_API_KEY — in ~/homelab/.env.claude
|
|
```
|