From 14aca8ea094b480f9c58c9f198faea384e729158 Mon Sep 17 00:00:00 2001 From: Gabriel Vidal Date: Sun, 9 Aug 2026 23:48:52 +0200 Subject: [PATCH] feat(models): browse and price every OpenRouter model from APPE's catalogue MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The pi.dev harness could only be pointed at two hand-written qwen entries, and a pi session was costed at one flat qwen rate whatever it actually ran on. - `backend/openrouter.py` pulls APPE's published catalogue (appe.dev.gabvdl.xyz/api/models/openrouter.json — a daily models.dev sync) and serves it at `GET /api/models/openrouter`: ~340 models with $/Mtok in / out / cache-read, context window, parameter size and capability tags. Cached 6h in-process, with a `/data` disk copy so a restart with no network still lists models, and warmed at startup so nothing waits on it. - `conversations.rates_for()` prices a turn from that catalogue — per model, cache-read included (OpenRouter's cache price is per-model policy, not Anthropic's flat 10%). Unknown ids keep the old estimate; pi's own `costUSD` still wins over any of it. - The composer's model select gains `Browse OpenRouter…` on the pi harness: a FuzzyList of the whole catalogue, each row stating the three rates that bill an agent run plus size and context, with capability filters. Cheapest first. - Any vendor-prefixed id now resolves to the pi harness frontend-side (`isOpenRouterId`), so an arbitrary pick survives a resume. Co-Authored-By: Claude Opus 5 --- CLAUDE.md | 27 +++ backend/conversations.py | 71 +++++-- backend/main.py | 21 ++ backend/openrouter.py | 194 ++++++++++++++++++ backend/schemas.py | 31 +++ frontend/openapi.json | 146 +++++++++++++ frontend/src/api.ts | 11 + .../composer/components/OpenRouterBrowser.tsx | 192 +++++++++++++++++ frontend/src/generated/api.ts | 51 +++++ frontend/src/generated/model/index.ts | 2 + .../src/generated/model/openRouterModel.ts | 28 +++ .../model/openRouterModelsResponse.ts | 13 ++ frontend/src/lib/costAnalysis.ts | 17 +- frontend/src/lib/models.tsx | 127 +++++++++--- frontend/src/mock/db.ts | 3 + frontend/src/mock/handlers.ts | 1 + frontend/src/mock/seed.ts | 14 ++ runner/README.md | 11 +- 18 files changed, 907 insertions(+), 53 deletions(-) create mode 100644 backend/openrouter.py create mode 100644 frontend/src/business/composer/components/OpenRouterBrowser.tsx create mode 100644 frontend/src/generated/model/openRouterModel.ts create mode 100644 frontend/src/generated/model/openRouterModelsResponse.ts diff --git a/CLAUDE.md b/CLAUDE.md index 7a5a398..a7e805e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -143,6 +143,16 @@ store / metadata sidecar. four chips) with the CLI alias attached. A newly released model appears in the composer with no code change. Falls back to a small static list when the key is missing or the call fails. +- `openrouter.py` — the **OpenRouter catalogue** (`GET /api/models/openrouter`): + every model a pi.dev session can run on, with its $/Mtok (input / output / + cache-read), context window, parameter size and capability tags. The data is + **APPE's** (`appe.dev.gabvdl.xyz/api/models/openrouter.json`, a daily + models.dev sync), cached in-process for 6h with a disk copy under `/data` so a + restart with no network still lists models. Deliberately separate from + `/api/models`: that endpoint is the composer's short chip list, this is the few + hundred rows the model browser searches. It also **prices** pi sessions — + `conversations.rates_for()` reads a model's real rates from here instead of the + flat qwen estimate, and only an id the catalogue doesn't know falls back. - `notify_audio.py` — the sound a notification made, replayable in the browser (`POST /api/notifications/audio` → a WAV, behind the ▶ button on an expanded notification card). There is **no synthesizer here**: it POSTs to the homelab @@ -531,6 +541,23 @@ glyphs render synchronously all over the app, so they can't come from a lazy chunk, and pulling ~1500 icons into the main bundle to serve one settings page is a bad trade. +### Picking a pi model: the OpenRouter browser + +The model select lists the curated entries per harness, but on **pi** the real +catalogue is a modal: `Browse OpenRouter…` opens `OpenRouterBrowser` — a +`FuzzyList` over every OpenRouter model, each row stating the three rates that +bill an agent run (**in / out / cache-read**, $ per million tokens), the model's +parameter size and its context window, with capability chips (tools, reasoning, +vision, open weights, free) to narrow the list. Cheapest first, so an unfiltered +open lands on what costs least to try. + +Picking one sends its full id (`openai/gpt-oss-20b`) as `--model`; the runner +already routes **any** vendor-prefixed id to OpenRouter and bare ids to the local +EVOX2 provider, so no allow-list has to be kept in step. The frontend reads the +same shape: `isOpenRouterId()` (a slash — Claude ids never have one) is what +makes an arbitrary pick resolve to the pi harness on resume, without the +catalogue being loaded. + ### Effort (claude) vs. thinking (pi) How hard a run works a turn is **one knob shown two ways**, per harness — never diff --git a/backend/conversations.py b/backend/conversations.py index d78ea2e..d10e8c7 100644 --- a/backend/conversations.py +++ b/backend/conversations.py @@ -16,6 +16,8 @@ import json import pathlib import re +import openrouter + # Claude Code session ids (= main transcript file stems) are UUIDs. _UUID_RE = re.compile(r"[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}" @@ -438,12 +440,17 @@ def _worktree_events(text) -> list[dict]: # Per-model price in USD / million tokens (input, output). Cache tiers are # multiples of the input price (read ≈ 0.1×, 5m write ≈ 1.25×, 1h write ≈ 2×). -# `qwen` covers the pi-harness OpenRouter sessions; a *bare* qwen id (no vendor -# prefix) is the local EVOX2 LM Studio provider, which is free. These are the -# fallback estimates — a pi transcript record carrying the runner-mirrored -# `costUSD` (pi's own per-message provider cost) overrides them (see `feed`). +# `qwen` is the *last-resort* estimate for a pi-harness OpenRouter session whose +# model isn't in the catalogue; a *bare* qwen id (no vendor prefix) is the local +# EVOX2 LM Studio provider, which is free. These are the fallback estimates — a +# pi transcript record carrying the runner-mirrored `costUSD` (pi's own +# per-message provider cost) overrides them (see `feed`). PRICES = {"opus": (5.0, 25.0), "sonnet": (3.0, 15.0), "haiku": (1.0, 5.0), "fable": (10.0, 50.0), "qwen": (0.15, 1.0), "qwen-local": (0.0, 0.0)} +# What a cache read costs, as a multiple of the input rate — Anthropic's flat +# 10%. OpenRouter models carry their own absolute cache price instead (see +# `rates_for`), because a cache read there is per-model policy, not a fixed cut. +CACHE_READ_MULT = 0.1 MAX_BLOCK = 8000 # cap a single text/result block so payloads stay bounded @@ -456,19 +463,40 @@ def _real_model(model: str | None) -> bool: return bool(model) and not model.startswith("<") -def price_for(model: str) -> tuple[float, float]: +def rates_for(model: str) -> tuple[float, float, float]: + """(input, output, cache-read) price in USD / million tokens for a model. + + **A pi-harness session is priced from the OpenRouter catalogue** (APPE's + daily models.dev sync — see `openrouter.py`), so a run on any of the several + hundred pickable models is costed at that model's real rates, cache read + included. Only an id the catalogue doesn't know falls through to the flat + qwen estimate below. Claude models keep the built-in table and the 10% + cache-read cut. + """ m = (model or "").lower() + if "/" in m: # vendor-prefixed ⇒ an OpenRouter model on the pi harness + hit = openrouter.rates(model) + if hit: + return hit if "sonnet" in m: - return PRICES["sonnet"] - if "haiku" in m: - return PRICES["haiku"] - if "fable" in m or "mythos" in m: - return PRICES["fable"] - if "qwen" in m: - # Vendor-prefixed ids (qwen/…) run on OpenRouter; bare ids are the - # local EVOX2 LM Studio provider — free. - return PRICES["qwen"] if "/" in m else PRICES["qwen-local"] - return PRICES["opus"] + pi, po = PRICES["sonnet"] + elif "haiku" in m: + pi, po = PRICES["haiku"] + elif "fable" in m or "mythos" in m: + pi, po = PRICES["fable"] + elif "qwen" in m or "/" in m: + # Vendor-prefixed ids run on OpenRouter (priced above when known); bare + # qwen ids are the local EVOX2 LM Studio provider — free. + pi, po = PRICES["qwen"] if "/" in m else PRICES["qwen-local"] + else: + pi, po = PRICES["opus"] + return pi, po, pi * CACHE_READ_MULT + + +def price_for(model: str) -> tuple[float, float]: + """(input, output) $/Mtok — `rates_for` without the cache-read rate.""" + pi, po, _ = rates_for(model) + return pi, po def norm_usage(u: dict) -> dict: @@ -500,19 +528,20 @@ def usage_tokens(u: dict) -> int: def usage_cost(u: dict, model: str) -> float: - pi, po = price_for(model) + pi, po, pc = rates_for(model) pi /= 1e6 po /= 1e6 + pc /= 1e6 return (u["input"] * pi + u["cacheWriteUnits"] * pi - + u["cacheRead"] * 0.1 * pi + u["output"] * po) + + u["cacheRead"] * pc + u["output"] * po) def usage_cache_cost(u: dict, model: str) -> float: """The share of a turn's cost that is replayed cached prompt (billed at 10% - of the input rate). The rest is fresh tokens the model actually had to read - or write this turn.""" - pi, _ = price_for(model) - return u["cacheRead"] * 0.1 * (pi / 1e6) + of the input rate on Claude, at the model's own cache rate on OpenRouter). + The rest is fresh tokens the model actually had to read or write this turn.""" + _, _, pc = rates_for(model) + return u["cacheRead"] * (pc / 1e6) def _block_text(content) -> str: diff --git a/backend/main.py b/backend/main.py index c510c4e..39b0db7 100644 --- a/backend/main.py +++ b/backend/main.py @@ -45,6 +45,7 @@ import gitdiff import goal as goalmd import memories as memories_mod import models as models_mod +import openrouter as openrouter_mod import cron as cron_mod import ctas as ctas_mod import forms as forms_mod @@ -2053,6 +2054,22 @@ def list_models(): "default": models_mod.DEFAULT_MODEL} +@app.get("/api/models/openrouter", + responses=_r(schemas.OpenRouterModelsResponse)) +def list_openrouter_models(): + """Every model a pi.dev session can run on OpenRouter, cheapest first. + + Not part of `/api/models` on purpose: that endpoint is the composer's short + pickable list (one chip per Claude family + the curated pi entries), while + this is a few hundred models the model browser searches through. Prices + ($/Mtok in / out / cache), context window and parameter size all come from + APPE's catalogue — the homelab's model-data source of truth (openrouter.py). + """ + return {"models": openrouter_mod.catalogue(), + "source": openrouter_mod.CATALOGUE_URL, + "generatedAt": openrouter_mod.generated_at()} + + class SpawnBody(BaseModel): prompt: str model: str | None = None @@ -4027,6 +4044,10 @@ def _startup(): threading.Thread(target=_watch_sidecar_runs, daemon=True).start() # Cron jobs: fire scheduled agent sessions (see cron.py). cron_scheduler.start() + # OpenRouter catalogue (prices + sizes from APPE): pulled off the request + # path so the first model-browser open and the first priced transcript + # already have it (openrouter.py). + openrouter_mod.warm() # Conversation CTAs: the buttons a finished conversation offers (ctas.py). # The seeded Complete prompt is written into the workspace on first run; # there is no scheduler — a CTA runs when it is pressed. diff --git a/backend/openrouter.py b/backend/openrouter.py new file mode 100644 index 0000000..a81bafc --- /dev/null +++ b/backend/openrouter.py @@ -0,0 +1,194 @@ +""" +The OpenRouter catalogue — every model a pi.dev session can be spawned on. + +The composer's model select used to offer two hand-written pi entries (one +OpenRouter qwen, one local EVOX2 qwen). OpenRouter serves several hundred, and +picking between them is a *price* decision, so the picker now lists the lot with +what each one costs. + +**The prices are not ours.** They come from APPE +(https://appe.dev.gabvdl.xyz), the homelab's model-data app, which syncs +models.dev daily and publishes its catalogue as static JSON +(`/api/models/openrouter.json`, one entry per model with $/Mtok in / out / +cache-read, context window, mined parameter count, capability tags and a +measured tokens/sec). Keeping a second price table here is exactly the bug this +avoids: when OpenRouter re-prices a model, APPE's next daily sync carries it and +this module picks it up within the TTL. + +Three layers so the picker is never empty: + 1. in-process cache (``CACHE_TTL_S``); + 2. a disk copy under ``/data`` — written on every successful fetch, read when + APPE is unreachable, so a restart with no network still lists models; + 3. nothing — the picker falls back to the curated ``PI_MODELS_DEFAULT`` list + in ``models.py`` and the flat qwen price estimate in ``conversations.py``. + +The catalogue is also what **prices** a pi session: ``rates()`` hands +``conversations.py`` the per-model $/Mtok, so an OpenRouter run is costed at its +real rates instead of a one-size-fits-all qwen guess. (A pi transcript that +carries pi's own ``costUSD`` still wins over any estimate — see +``conversations.feed``.) +""" + +import json +import os +import pathlib +import threading +import time +import urllib.error +import urllib.request + +# APPE's published slice for the openrouter provider. Point at a file:// URL or +# another host with APPE_OPENROUTER_URL if APPE moves. +CATALOGUE_URL = os.environ.get( + "APPE_OPENROUTER_URL", + "https://appe.dev.gabvdl.xyz/api/models/openrouter.json") +CACHE_TTL_S = 6 * 3600 +# Survives a restart when APPE is unreachable (same writable volume as the +# conversation metadata). +CACHE_PATH = pathlib.Path( + os.environ.get("OPENROUTER_CACHE_PATH", "/data/openrouter-models.json")) +FETCH_TIMEOUT_S = 15 + +_lock = threading.Lock() +_cache: tuple[float, list[dict]] | None = None +_index: dict[str, dict] = {} # model id → entry, for the pricing lookups +_generated_at: str = "" + + +def _shape(entry: dict) -> dict | None: + """One APPE catalogue row → one pickable model. + + ``api_id`` is the id **without** the ``openrouter/`` prefix — exactly the + ``--model`` argument the pi runner passes to OpenRouter. Everything else is + renamed to the viewer's camelCase and left otherwise untouched: prices stay + USD per million tokens, ``modelSize`` is billions of parameters (null for + models that publish none), ``speedTps`` is median output tokens/sec. + """ + mid = entry.get("api_id") or "" + if not mid: + return None + return { + "id": mid, + "name": entry.get("name") or mid, + "vendor": mid.split("/")[0], + "description": entry.get("description") or "", + "inputCost": float(entry.get("input_cost") or 0.0), + "outputCost": float(entry.get("output_cost") or 0.0), + # None (not 0) when the model has no prompt cache: a missing cache price + # and a free cache are different facts and the picker shows them so. + "cacheCost": (float(entry["cache_cost"]) + if entry.get("cache_cost") is not None else None), + "contextWindow": entry.get("max_token") or None, + "modelSize": entry.get("model_size") or None, + "tags": entry.get("tags") or [], + "tier": entry.get("tier") or "", + "license": entry.get("license") or "", + "speedTps": entry.get("speed_tps") or None, + } + + +def _parse(payload: dict) -> tuple[list[dict], str]: + rows = payload.get("models") or [] + models = [m for m in (_shape(r) for r in rows) if m] + # Cheapest first (blended 3:1 in:out, the usual agent mix), so the browser's + # unfiltered list opens on what a run costs least to try. + models.sort(key=lambda m: (m["inputCost"] * 0.75 + m["outputCost"] * 0.25, + m["id"])) + return models, str(payload.get("generatedAt") or "") + + +def _from_disk() -> tuple[list[dict], str]: + try: + return _parse(json.loads(CACHE_PATH.read_text(encoding="utf-8"))) + except (OSError, ValueError): + return [], "" + + +def _fetch() -> tuple[list[dict], str]: + req = urllib.request.Request( + CATALOGUE_URL, headers={"accept": "application/json", + "user-agent": "ai-agent/openrouter-catalogue"}) + try: + with urllib.request.urlopen(req, timeout=FETCH_TIMEOUT_S) as r: + raw = r.read().decode() + payload = json.loads(raw) + models, generated = _parse(payload) + if models: + try: + CACHE_PATH.parent.mkdir(parents=True, exist_ok=True) + CACHE_PATH.write_text(raw, encoding="utf-8") + except OSError: + pass # read-only volume — the in-process cache still serves + return models, generated + except (urllib.error.URLError, OSError, ValueError): + pass + return _from_disk() + + +def _store(models: list[dict], generated: str, stamp: float | None = None) -> None: + """Publish a catalogue. `stamp=0` marks it as not-yet-refreshed (the disk + copy), so the next `catalogue()` still goes and asks APPE.""" + global _cache, _generated_at, _index + with _lock: + # A failed refresh keeps whatever we had rather than emptying the picker. + if models or not _cache: + _cache = (time.time() if stamp is None else stamp, models) + _generated_at = generated + _index = {m["id"]: m for m in models} + + +def catalogue(force: bool = False) -> list[dict]: + """Every OpenRouter model, cheapest first. Empty only if APPE was never reached. + + May block on the network (once per TTL). `warm()` calls it off the request + path at startup; the pricing lookups never do. + """ + with _lock: + if not force and _cache and time.time() - _cache[0] < CACHE_TTL_S: + return _cache[1] + _store(*_fetch()) # outside the lock: a slow call blocks nobody + return (_cache or (0.0, []))[1] + + +def warm() -> None: + """Refresh the catalogue in the background (called at startup).""" + threading.Thread(target=catalogue, kwargs={"force": True}, + name="openrouter-catalogue", daemon=True).start() + + +def generated_at() -> str: + """When APPE generated the catalogue we are serving (ISO 8601, "" if unknown).""" + catalogue() + return _generated_at + + +def find(model_id: str) -> dict | None: + """The catalogue entry for a `--model` argument, or None if it isn't OpenRouter's. + + **Never fetches.** Transcript parsing prices every turn through here, and a + parse must not stall on an HTTP call; it reads the in-process index, falling + back to the disk copy once. A cold, network-less start therefore prices at + the flat estimate rather than hanging. + """ + mid = (model_id or "").strip() + if not mid or "/" not in mid: + return None # bare ids are the local EVOX2 provider, not OpenRouter + if _cache is None: + models, generated = _from_disk() + _store(models, generated, stamp=0.0) + return _index.get(mid) + + +def rates(model_id: str) -> tuple[float, float, float] | None: + """(input, output, cache-read) $ per million tokens for an OpenRouter model. + + None when the id isn't in the catalogue, which is the caller's cue to fall + back to its own estimate. A model with no prompt cache reports its cache-read + rate as the input rate — a cache read it can't serve is billed as a fresh + read, which is what the tokens would actually cost. + """ + m = find(model_id) + if not m: + return None + cache = m["cacheCost"] if m["cacheCost"] is not None else m["inputCost"] + return m["inputCost"], m["outputCost"], cache diff --git a/backend/schemas.py b/backend/schemas.py index 218089e..1eb792c 100644 --- a/backend/schemas.py +++ b/backend/schemas.py @@ -1153,6 +1153,37 @@ class ModelsResponse(Schema): default: str # model a run uses when nothing is picked +class OpenRouterModel(Schema): + """One OpenRouter model a pi.dev session can be spawned on. + + The catalogue behind it is APPE's (models.dev, synced daily) — see + `backend/openrouter.py`. Prices are USD per **million** tokens. + """ + id: str # the `--model` argument, e.g. "openai/gpt-oss-20b" + name: str # display name ("GPT OSS 20B") + vendor: str # the id's first segment ("openai") + description: str = "" + inputCost: float + outputCost: float + # Cache-read price. None = the model has no prompt cache (a cache read is + # then billed as a fresh read), which is a different fact from a free cache. + cacheCost: Optional[float] = None + contextWindow: Optional[int] = None + # Parameters in billions, mined from the id/name by APPE. None for models + # that publish no size (most closed ones). + modelSize: Optional[float] = None + tags: list[str] = [] # vision | reasoning | tools | opensource | … + tier: str = "" # small | medium | big (blended-price heuristic) + license: str = "" + speedTps: Optional[float] = None # median output tokens/sec + + +class OpenRouterModelsResponse(Schema): + models: list[OpenRouterModel] # cheapest first (blended 3:1 in:out) + source: str # the APPE URL the catalogue came from + generatedAt: str = "" # when APPE generated it (ISO 8601) + + class SpawnResult(Schema): sessionId: str pid: Optional[int] = None diff --git a/frontend/openapi.json b/frontend/openapi.json index 4d06284..bfcc081 100644 --- a/frontend/openapi.json +++ b/frontend/openapi.json @@ -1122,6 +1122,25 @@ } } }, + "/api/models/openrouter": { + "get": { + "summary": "List Openrouter Models", + "description": "Every model a pi.dev session can run on OpenRouter, cheapest first.\n\nNot part of `/api/models` on purpose: that endpoint is the composer's short\npickable list (one chip per Claude family + the curated pi entries), while\nthis is a few hundred models the model browser searches through. Prices\n($/Mtok in / out / cache), context window and parameter size all come from\nAPPE's catalogue \u2014 the homelab's model-data source of truth (openrouter.py).", + "operationId": "list_openrouter_models_api_models_openrouter_get", + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/OpenRouterModelsResponse" + } + } + } + } + } + } + }, "/api/spawn": { "post": { "summary": "Spawn Conversation", @@ -9461,6 +9480,133 @@ ], "title": "OkResponse" }, + "OpenRouterModel": { + "properties": { + "id": { + "type": "string", + "title": "Id" + }, + "name": { + "type": "string", + "title": "Name" + }, + "vendor": { + "type": "string", + "title": "Vendor" + }, + "description": { + "type": "string", + "title": "Description", + "default": "" + }, + "inputCost": { + "type": "number", + "title": "Inputcost" + }, + "outputCost": { + "type": "number", + "title": "Outputcost" + }, + "cacheCost": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "title": "Cachecost" + }, + "contextWindow": { + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "title": "Contextwindow" + }, + "modelSize": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "title": "Modelsize" + }, + "tags": { + "items": { + "type": "string" + }, + "type": "array", + "title": "Tags", + "default": [] + }, + "tier": { + "type": "string", + "title": "Tier", + "default": "" + }, + "license": { + "type": "string", + "title": "License", + "default": "" + }, + "speedTps": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "title": "Speedtps" + } + }, + "type": "object", + "required": [ + "id", + "name", + "vendor", + "inputCost", + "outputCost" + ], + "title": "OpenRouterModel", + "description": "One OpenRouter model a pi.dev session can be spawned on.\n\nThe catalogue behind it is APPE's (models.dev, synced daily) \u2014 see\n`backend/openrouter.py`. Prices are USD per **million** tokens." + }, + "OpenRouterModelsResponse": { + "properties": { + "models": { + "items": { + "$ref": "#/components/schemas/OpenRouterModel" + }, + "type": "array", + "title": "Models" + }, + "source": { + "type": "string", + "title": "Source" + }, + "generatedAt": { + "type": "string", + "title": "Generatedat", + "default": "" + } + }, + "type": "object", + "required": [ + "models", + "source" + ], + "title": "OpenRouterModelsResponse" + }, "ParentConvRef": { "properties": { "id": { diff --git a/frontend/src/api.ts b/frontend/src/api.ts index 673b6e0..9830dc1 100644 --- a/frontend/src/api.ts +++ b/frontend/src/api.ts @@ -21,6 +21,7 @@ import type { MemoryDetail, MemorySummary, ModelsResponse, + OpenRouterModelsResponse, NotifyLogEntry, NotifyResult, EnvSaveBody, @@ -112,6 +113,16 @@ export async function fetchModels(): Promise { return r.json(); } +/** Every OpenRouter model a pi.dev session can run on, with its prices — the + * model browser's catalogue (sourced from APPE; see backend/openrouter.py). */ +export async function fetchOpenRouterModels(): Promise { + const r = await apiFetch("/api/models/openrouter", { + headers: { Accept: "application/json" }, + }); + if (!r.ok) throw new Error(`openrouter models: ${r.status}`); + return r.json(); +} + export interface ConversationsPage { conversations: ConversationSummary[]; /** Total conversations matching the filter (pre-pagination). */ diff --git a/frontend/src/business/composer/components/OpenRouterBrowser.tsx b/frontend/src/business/composer/components/OpenRouterBrowser.tsx new file mode 100644 index 0000000..f369be1 --- /dev/null +++ b/frontend/src/business/composer/components/OpenRouterBrowser.tsx @@ -0,0 +1,192 @@ +import { useMemo, useState } from "react"; +import { EmptyState, FuzzyList, Modal, TagFilter, type FuzzyRenderContext } from "@gabvdl/ui"; +import { Cpu, Pi } from "lucide-react"; +import { useOpenRouterModels } from "@/lib/models"; +import { cn } from "@/lib/utils"; +import type { OpenRouterModel } from "@/types"; + +/** + * The OpenRouter model browser — the pi.dev harness's real model picker. + * + * The composer's select lists a handful of curated entries; OpenRouter serves + * several hundred, and choosing between them is a **price** decision more than a + * name decision. So this is a fuzzy list where every row states the three rates + * that actually bill an agent run (input / output / cache-read, $ per million + * tokens), the model's parameter count and its context window — the numbers you + * would otherwise go to OpenRouter's site to compare. + * + * The catalogue comes from APPE (`/api/models/openrouter` → `backend/openrouter.py` + * → appe.dev.gabvdl.xyz), which syncs models.dev daily. Cheapest first, so an + * unfiltered open lands on what is cheapest to try; typing re-ranks by relevance. + */ + +/** `$/Mtok` as a compact figure: sub-cent rates keep their significant digits. */ +export function usd(n: number | null | undefined): string { + if (n == null) return "—"; + if (n === 0) return "free"; + if (n < 0.01) return `$${n.toFixed(4).replace(/0+$/, "")}`; + if (n < 1) return `$${n.toFixed(3).replace(/0+$/, "")}`; + return `$${n.toFixed(2)}`; +} + +/** "35B" / "1.5B" — a model's parameter count, or null when it publishes none. */ +export function paramsLabel(size: number | null | undefined): string | null { + if (!size) return null; + return size >= 1 ? `${Number(size.toFixed(1))}B` : `${Math.round(size * 1000)}M`; +} + +/** "262k" / "1M" — a context window in tokens. */ +export function ctxLabel(ctx: number | null | undefined): string | null { + if (!ctx) return null; + if (ctx >= 1_000_000) return `${Number((ctx / 1_000_000).toFixed(1))}M`; + return ctx >= 1000 ? `${Math.round(ctx / 1000)}k` : `${ctx}`; +} + +/** The capability facets worth filtering on (a row can carry several). */ +const FACETS = [ + { value: "tools", label: "Tools" }, + { value: "reasoning", label: "Reasoning" }, + { value: "vision", label: "Vision" }, + { value: "opensource", label: "Open weights" }, + { value: "free", label: "Free" }, +] as const; + +/** One price cell — label above, figure below, so three fit in a row's tail. */ +function Rate({ label, value, muted }: { label: string; value: string; muted?: boolean }) { + return ( +
+ {label} + + {value} + +
+ ); +} + +function Chip({ children }: { children: React.ReactNode }) { + return ( + + {children} + + ); +} + +export function OpenRouterBrowser({ + open, + picked, + onPick, + onClose, +}: { + open: boolean; + /** The `--model` argument currently in the composer, highlighted in the list. */ + picked?: string | null; + onPick: (id: string) => void; + onClose: () => void; +}) { + const { data, isLoading, error } = useOpenRouterModels(open); + const [facets, setFacets] = useState([]); + + const rows = useMemo(() => { + const all = data?.models ?? []; + if (!facets.length) return all; + return all.filter((m) => + facets.every((f) => + f === "free" ? m.inputCost === 0 && m.outputCost === 0 : (m.tags ?? []).includes(f), + ), + ); + }, [data, facets]); + + const renderRow = ({ item, active, highlight }: FuzzyRenderContext) => { + const params = paramsLabel(item.modelSize); + const ctx = ctxLabel(item.contextWindow); + const isPicked = picked === item.id; + return ( +
+
+
+ {highlight("name")} + {params && {params}} + {ctx && {ctx} ctx} +
+
+ {highlight("id")} +
+
+
+ + + {/* No cache price = no prompt cache: a "cache read" is billed as a + fresh read, which is why it shows the input rate, greyed. */} + +
+
+ ); + }; + + return ( + + ({ id: f.value, label: f.label }))} + value={facets} + onChange={setFacets} + /> + {isLoading ? ( +
Loading the catalogue…
+ ) : error || !data?.models.length ? ( + } + title="No catalogue" + description="APPE (appe.dev.gabvdl.xyz) couldn't be reached, so the model list is empty. The curated pi entries still work." + /> + ) : ( + + items={rows} + keys={["name", "id", "vendor", "description"]} + getItemKey={(m) => m.id} + smooth + onSelect={(m) => { + onPick(m.id); + onClose(); + }} + placeholder={'Search models… (use "quotes" for exact)'} + estimateSize={62} + overscan={8} + showCount + emptyState={ + } + title="No model matches" + description="Try another search, or clear the capability filters." + /> + } + className="min-h-0 flex-1" + renderItem={renderRow} + /> + )} +
+ ); +} diff --git a/frontend/src/generated/api.ts b/frontend/src/generated/api.ts index 58d34ad..ef13b5f 100644 --- a/frontend/src/generated/api.ts +++ b/frontend/src/generated/api.ts @@ -82,6 +82,7 @@ import type { NotifyLogResponse, NotifyResult, OkResponse, + OpenRouterModelsResponse, PlanDetail, PlanDetailApiPlanGetParams, PlansResponse, @@ -1733,6 +1734,56 @@ export const listModelsApiModelsGet = async ( options?: RequestInit): Promise
  • { + + + + + return `/api/models/openrouter` +} + +/** + * Every model a pi.dev session can run on OpenRouter, cheapest first. + * + * Not part of `/api/models` on purpose: that endpoint is the composer's short + * pickable list (one chip per Claude family + the curated pi entries), while + * this is a few hundred models the model browser searches through. Prices + * ($/Mtok in / out / cache), context window and parameter size all come from + * APPE's catalogue — the homelab's model-data source of truth (openrouter.py). + * @summary List Openrouter Models + */ +export const listOpenrouterModelsApiModelsOpenrouterGet = async ( options?: RequestInit): Promise => { + + const res = await fetch(getListOpenrouterModelsApiModelsOpenrouterGetUrl(), + { + ...options, + method: 'GET' + + + } +) + + + const body = [204, 205, 304].includes(res.status) ? null : await res.text(); + + const data: listOpenrouterModelsApiModelsOpenrouterGetResponse['data'] = body ? JSON.parse(body) : {} + return { data, status: res.status, headers: res.headers } as listOpenrouterModelsApiModelsOpenrouterGetResponse +} + + + export type spawnConversationApiSpawnPostResponse200 = { data: SpawnResult status: 200 diff --git a/frontend/src/generated/model/index.ts b/frontend/src/generated/model/index.ts index 7f1b8dc..77c5941 100644 --- a/frontend/src/generated/model/index.ts +++ b/frontend/src/generated/model/index.ts @@ -154,6 +154,8 @@ export * from './notifyResult.ts'; export * from './ogImage.ts'; export * from './ogImageKind.ts'; export * from './okResponse.ts'; +export * from './openRouterModel.ts'; +export * from './openRouterModelsResponse.ts'; export * from './parentConvRef.ts'; export * from './planActual.ts'; export * from './planConvRef.ts'; diff --git a/frontend/src/generated/model/openRouterModel.ts b/frontend/src/generated/model/openRouterModel.ts new file mode 100644 index 0000000..7ff0023 --- /dev/null +++ b/frontend/src/generated/model/openRouterModel.ts @@ -0,0 +1,28 @@ +/** + * Generated by orval v8.20.0 🍺 + * Do not edit manually. + * ai-agent + * OpenAPI spec version: 0.1.0 + */ + +/** + * One OpenRouter model a pi.dev session can be spawned on. + * + * The catalogue behind it is APPE's (models.dev, synced daily) — see + * `backend/openrouter.py`. Prices are USD per **million** tokens. + */ +export interface OpenRouterModel { + id: string; + name: string; + vendor: string; + description?: string; + inputCost: number; + outputCost: number; + cacheCost?: number | null; + contextWindow?: number | null; + modelSize?: number | null; + tags?: string[]; + tier?: string; + license?: string; + speedTps?: number | null; +} diff --git a/frontend/src/generated/model/openRouterModelsResponse.ts b/frontend/src/generated/model/openRouterModelsResponse.ts new file mode 100644 index 0000000..68d7f93 --- /dev/null +++ b/frontend/src/generated/model/openRouterModelsResponse.ts @@ -0,0 +1,13 @@ +/** + * Generated by orval v8.20.0 🍺 + * Do not edit manually. + * ai-agent + * OpenAPI spec version: 0.1.0 + */ +import type { OpenRouterModel } from './openRouterModel.ts'; + +export interface OpenRouterModelsResponse { + models: OpenRouterModel[]; + source: string; + generatedAt?: string; +} diff --git a/frontend/src/lib/costAnalysis.ts b/frontend/src/lib/costAnalysis.ts index 00a375e..4c197f5 100644 --- a/frontend/src/lib/costAnalysis.ts +++ b/frontend/src/lib/costAnalysis.ts @@ -34,8 +34,10 @@ const DAY_MS = 86_400_000; const MONTH_MS = 30.4375 * DAY_MS; // ── pricing ─────────────────────────────────────────────────────────────────── -// Mirrors the backend's `conversations.PRICES` (USD per million tokens, input / -// output). Cache reads bill at 0.1× input; cache writes are priced off their +// Mirrors the backend's `conversations.PRICES` **fallback** table (USD per +// million tokens, input / output) — the backend additionally prices OpenRouter +// models from APPE's catalogue, per model. Cache reads bill at 0.1× input (an +// OpenRouter model's own cache rate can differ); cache writes are priced off their // weighted `cacheWriteUnits` (5m ≈ 1.25×, 1h ≈ 2×) at the input rate. Kept // client-side so the dashboard can split a conversation's single `cost` figure // back into per-token-type shares. @@ -53,8 +55,15 @@ function priceFor(model: string | null | undefined): [number, number] { if (m.includes("sonnet")) return PRICES.sonnet; if (m.includes("haiku")) return PRICES.haiku; if (m.includes("fable") || m.includes("mythos")) return PRICES.fable; - if (m.includes("qwen")) - return m.includes("/") ? PRICES.qwen : PRICES["qwen-local"]; + if (m.includes("/")) { + // A vendor-prefixed id is an OpenRouter model on the pi harness. Its *real* + // rates come from the catalogue backend-side (`backend/openrouter.py` prices + // each turn, and that is what a conversation's `cost` already reflects); + // this table only splits that figure across token types, so a single open- + // model estimate is enough — and beats charging it at Opus rates. + return PRICES.qwen; + } + if (m.includes("qwen")) return PRICES["qwen-local"]; return PRICES.opus; } diff --git a/frontend/src/lib/models.tsx b/frontend/src/lib/models.tsx index 4437a4b..0ade90c 100644 --- a/frontend/src/lib/models.tsx +++ b/frontend/src/lib/models.tsx @@ -1,3 +1,4 @@ +import { useState } from "react"; import { useQuery } from "@tanstack/react-query"; import { Bot, @@ -8,11 +9,13 @@ import { Gauge, Pi, Rabbit, + Search, Sparkles, SquareTerminal, Zap, } from "lucide-react"; -import { fetchModels } from "@/api"; +import { fetchModels, fetchOpenRouterModels } from "@/api"; +import { OpenRouterBrowser } from "@/business/composer/components/OpenRouterBrowser"; import { SelectPopover, type SelectOption } from "@/technical/SelectPopover"; import { cn } from "@/lib/utils"; import type { ModelInfo } from "@/types"; @@ -41,6 +44,10 @@ import type { ModelInfo } from "@/types"; /** The agent CLI a session runs on. */ export type Harness = "claude" | "pi"; +/** Sentinel option value: "Browse OpenRouter…" opens the catalogue modal + * instead of picking a model. Can't collide with a real id (no slash). */ +const BROWSE = "__browse_openrouter__"; + /** Per-harness label + icon for the harness select. */ export const HARNESS_STYLE: Record = { claude: { @@ -65,14 +72,27 @@ const FAMILY_STYLE: Record = { qwen: { cls: "bg-sky-500/10 text-sky-500", icon: Cpu }, }; const UNKNOWN = { cls: "bg-muted text-muted-foreground", icon: Bot }; +// Any vendor-prefixed id the families above don't claim — i.e. one of the few +// hundred OpenRouter models the browser can pick (see OpenRouterBrowser). +const OPENROUTER_STYLE = { cls: "bg-sky-500/10 text-sky-500", icon: Cpu }; /** The family a model id belongs to (`claude-opus-4-8` → `opus`). */ export function modelFamily(id: string): string { return Object.keys(FAMILY_STYLE).find((f) => id.includes(f)) ?? ""; } +/** Whether a `--model` argument is an OpenRouter model (`vendor/model`). + * Claude ids never carry a slash, so the shape alone is the test — which is + * what lets a conversation resumed on an arbitrary OpenRouter model still + * resolve to the pi harness without the catalogue being loaded. */ +export function isOpenRouterId(id: string | null | undefined): boolean { + return !!id && id.includes("/"); +} + export function modelStyle(id: string) { - return FAMILY_STYLE[modelFamily(id)] ?? UNKNOWN; + const family = modelFamily(id); + if (family) return FAMILY_STYLE[family]; + return isOpenRouterId(id) ? OPENROUTER_STYLE : UNKNOWN; } /** @@ -88,7 +108,9 @@ export function modelLabel(id: string): string { const ver = tail.match(/qwen(\d+(?:\.\d+)?)/)?.[1]; return ver ? `qwen ${ver}` : tail; } - if (!family) return id.replace("claude-", ""); + // An OpenRouter pick: the vendor prefix is the least informative half of + // `openai/gpt-oss-20b`, and the chip is narrow — show the model, not the shop. + if (!family) return isOpenRouterId(id) ? (id.split("/").pop() as string) : id.replace("claude-", ""); const version = id .split(`${family}-`)[1] ?.split("-") @@ -126,6 +148,25 @@ export function useModels() { }); } +/** + * The OpenRouter catalogue — every model the **pi** harness can run, with its + * $/Mtok (input / output / cache-read), parameter size and context window. + * + * Deliberately a second query rather than more entries in `useModels()`: that + * one is the composer's short pickable list (a chip each), this is a few hundred + * rows only the browser modal renders, so it is fetched when the modal first + * opens and then cached for the app session. The data itself is APPE's — see + * `backend/openrouter.py`. + */ +export function useOpenRouterModels(enabled = true) { + return useQuery({ + queryKey: ["models", "openrouter"], + queryFn: fetchOpenRouterModels, + staleTime: 60 * 60 * 1000, + enabled, + }); +} + /** * The `--model` argument for a model entry: a family's newest model goes out as * its CLI alias (`opus`), older ones as their full id. @@ -156,7 +197,9 @@ export function harnessForModel( if (!arg) return "claude"; const hit = findModel(models, arg); if (hit?.harness) return hit.harness === "pi" ? "pi" : "claude"; - return modelFamily(arg) === "qwen" ? "pi" : "claude"; + // Not in the list: an older transcript's model, or one of the several hundred + // OpenRouter models the browser can pick (which the short list never carries). + return isOpenRouterId(arg) || modelFamily(arg) === "qwen" ? "pi" : "claude"; } /** The models a harness can run, in backend order (newest of each family first). */ @@ -242,6 +285,7 @@ export function ModelSelect({ onChange: (arg: string) => void; disabled?: boolean; }) { + const [browsing, setBrowsing] = useState(false); const list = modelsForHarness(models, harness); const options: SelectOption[] = list.map((m) => { const { icon: Icon, cls } = modelStyle(m.id); @@ -256,32 +300,65 @@ export function ModelSelect({ // What the chip shows: the pick, else the harness's first model (what the run // will actually use) — never a blank chip. const shown = model || (list[0] ? modelArg(list[0]) : ""); - if (!shown) { + // A model picked from the OpenRouter browser isn't in the short list, so it + // gets its own option — otherwise the popover would show no pick at all. + if (shown && isOpenRouterId(shown) && !options.some((o) => o.value === shown)) { + const { icon: Icon, cls } = modelStyle(shown); + options.unshift({ + value: shown, + label: modelLabel(shown), + hint: shown, + icon: , + meta: "OpenRouter", + }); + } + // The pi harness's real catalogue is the browser, not this list: several + // hundred models, picked by price. The select keeps the curated entries and + // hands off to the modal for everything else. + if (harness === "pi") { + options.push({ + value: BROWSE, + label: "Browse OpenRouter…", + hint: "every model, with its $/Mtok and size", + icon: , + }); + } + if (!shown && harness !== "pi") { return default model; } const { icon: Icon, cls } = modelStyle(shown); const entry = findModel(models, shown); return ( - - - {entry?.label || modelLabel(shown)} - - } - /> + <> + (v === BROWSE ? setBrowsing(true) : onChange(v))} + trigger={ + + + {entry?.label || modelLabel(shown) || "pick a model"} + + } + /> + {browsing && ( + setBrowsing(false)} + /> + )} + ); } diff --git a/frontend/src/mock/db.ts b/frontend/src/mock/db.ts index c7d6f1b..71440ef 100644 --- a/frontend/src/mock/db.ts +++ b/frontend/src/mock/db.ts @@ -16,6 +16,7 @@ import type { MemoryDetail, MemorySummary, ModelsResponse, + OpenRouterModelsResponse, NotifyLogEntry, PlanDetail, PlanSummary, @@ -42,6 +43,8 @@ export interface MockDb { workspace: string; pricing: Pricing; models: ModelsResponse; + /** The OpenRouter catalogue the pi-harness model browser searches. */ + openRouterModels: OpenRouterModelsResponse; files: FileEntry[]; conversations: ConversationDetail[]; /** Per-conversation commit rollup + net diff, keyed by conversation id. */ diff --git a/frontend/src/mock/handlers.ts b/frontend/src/mock/handlers.ts index a1cad2a..8e86577 100644 --- a/frontend/src/mock/handlers.ts +++ b/frontend/src/mock/handlers.ts @@ -302,6 +302,7 @@ export async function handle(url: URL, init: RequestInit | undefined, request?: if (p === "/api/deploy-status") return json(db.deployStatus); if (p === "/api/bundle") return json({ workspace: db.workspace, totals: totals(), files: db.files, pricing: db.pricing }); if (p === "/api/models") return json(db.models); + if (p === "/api/models/openrouter") return json(db.openRouterModels); if (p === "/api/dashboard") { const t = totals(); return json({ diff --git a/frontend/src/mock/seed.ts b/frontend/src/mock/seed.ts index 035da37..00b10c3 100644 --- a/frontend/src/mock/seed.ts +++ b/frontend/src/mock/seed.ts @@ -1425,6 +1425,20 @@ export function seedDb(): MockDb { { id: "qwen3-vl-8b", displayName: "Qwen3 VL 8B", family: "qwen", version: "3", alias: null, latest: false, harness: "pi", provider: "lmstudio", label: "Qwen VL", contextWindow: 32_000 }, ], }, + // A handful of the OpenRouter catalogue (the real one is ~340 models from + // APPE) — enough for the model browser to search, filter and price in mock + // mode. Cheapest first, like the API returns it. + openRouterModels: { + source: "https://appe.dev.gabvdl.xyz/api/models/openrouter.json", + generatedAt: "2026-08-01T04:30:00.000Z", + models: [ + { id: "google/gemma-4-31b-it:free", name: "Gemma 4 31B (free)", vendor: "google", description: "Open Gemma instruction model for self-hosted chat and reasoning", inputCost: 0, outputCost: 0, cacheCost: null, contextWindow: 262_144, modelSize: 31, tags: ["vision", "reasoning", "tools", "opensource"], tier: "small", license: "opensource", speedTps: 120 }, + { id: "mistralai/mistral-nemo", name: "Mistral Nemo", vendor: "mistralai", description: "Small multilingual instruct model", inputCost: 0.019, outputCost: 0.03, cacheCost: null, contextWindow: 131_072, modelSize: null, tags: ["tools", "opensource"], tier: "small", license: "opensource", speedTps: 190 }, + { id: "openai/gpt-oss-20b", name: "GPT OSS 20B", vendor: "openai", description: "Open-weight GPT model for self-hosted reasoning", inputCost: 0.03, outputCost: 0.13, cacheCost: 0.03, contextWindow: 131_072, modelSize: 20, tags: ["reasoning", "tools", "opensource"], tier: "small", license: "opensource", speedTps: 258 }, + { id: "qwen/qwen3-235b-a22b", name: "Qwen3 235B A22B", vendor: "qwen", description: "Large MoE reasoning model", inputCost: 0.13, outputCost: 0.6, cacheCost: 0.02, contextWindow: 262_144, modelSize: 235, tags: ["reasoning", "tools", "opensource"], tier: "small", license: "opensource", speedTps: 88 }, + { id: "anthropic/claude-sonnet-5", name: "Claude Sonnet 5", vendor: "anthropic", description: "Balanced frontier model", inputCost: 3, outputCost: 15, cacheCost: 0.3, contextWindow: 1_000_000, modelSize: null, tags: ["vision", "reasoning", "tools"], tier: "medium", license: "commercial", speedTps: 71 }, + ], + }, files, conversations, commits, diff --git a/runner/README.md b/runner/README.md index e246b94..c8dad59 100644 --- a/runner/README.md +++ b/runner/README.md @@ -48,11 +48,16 @@ composer chip → POST /api/spawn {model, harness} | model id | provider | |---|---| -| `qwen/qwen3.6-35b-a3b` (vendor prefix) | OpenRouter (`OPENROUTER_API_KEY`, auto-loaded from `~/homelab/.env.claude`) | +| `openai/gpt-oss-20b`, `qwen/qwen3.6-35b-a3b`, … (vendor prefix) | OpenRouter (`OPENROUTER_API_KEY`, auto-loaded from `~/homelab/.env.claude`) | | `qwen3.6-35b-a3b` (bare) | EVOX2 LM Studio (`~/.pi/agent/models.json` `evox2` provider; box is WoL-woken automatically) | -The pickable list lives in `backend/models.py` (`PI_MODELS_DEFAULT`, override -with `PI_MODELS_JSON`). +The routing is the **id shape**, not a list — so any of the several hundred +OpenRouter models runs here with no runner change. Two pickers feed it: the +curated chips in `backend/models.py` (`PI_MODELS_DEFAULT`, override with +`PI_MODELS_JSON`) and the composer's **OpenRouter browser**, which lists the +whole catalogue with per-model prices from `backend/openrouter.py` (sourced from +APPE). Those same prices cost the session's transcript, so a run's `$` figure is +the model's real rate even before pi reports its own `costUSD`. ## Permissions -- 2.49.1