Providers and models

The catalog is data, not code

Vendor model ids change faster than this package ships, so compiling them into TypeScript guarantees they are wrong within weeks. The catalog is a JSON file: packages/engine/src/providers/catalog.json.

Layers, later overriding earlier:

0. pi-ai's built-in models      whatever it knows about the catalogued vendors
   models.dev snapshot          what pi-ai does not know, for linked providers
1. the bundled catalog.json     33 providers, 6 verified models
2. ~/.oharness/catalog.json  your edits, and what discovery found
3. a catalog path in config
4. providers / models in config
5. /models refresh              asks the vendor directly

Overrides are field level, so pointing Kimi at its mainland endpoint does not mean restating everything else:

{ "providers": { "moonshot": { "baseUrl": "https://api.moonshot.cn/v1" } } }

Which providers, and where their models come from

The providers are chosen by hand: catalog.json is the list, with the label, protocol, key variable and ordering for each. International providers come first, then the leading Chinese vendors — their international endpoints only — and local servers last; the Models page lists them in that order. A mainland account overrides baseUrl in a user catalog, as below. Coverage beyond pi-ai comes from models.dev, the provider database OpenCode uses. An entry names its models.dev id under modelsDev, and packages/engine/src/providers/models-dev.json holds a snapshot of those providers' models — context length, output limit, price, image input, tool calling, reasoning tiers. Only models that answer in text, can call tools and are not deprecated are offered. Where pi-ai and models.dev both know a model, pi-ai's entry is used, for its per-vendor thinking tiers.

The snapshot is regenerated by hand and reviewed like any other change:

npm run models:sync            # rewrite the snapshot
npm run models:sync -- --check # report only

It keeps only the linked providers, one model per line so the diff reads as the models that changed, and it names any entry whose endpoint disagrees with models.dev without editing the catalog. Nothing fetches models.dev at runtime.

Plans (Z.ai GLM Coding Plan, Kimi Code, Alibaba Coding Plan) are entries of their own with their own key variable: a plan key sent to the pay-as-you-go host, or the reverse, is refused or billed per token. Their models carry no per-token price.

Provider facts are stable; model ids are not

Endpoint, environment variables and protocol are worth shipping. Model names are not — so the only entries this package states outright are the ones it could verify, and the rest come from discovery, which asks each vendor's own /models.

Between the two sits layer 0. pi-ai ships a generated catalog of the models it knows, and since the wire layer is pi-ai already, those entries are read at startup and seeded under everything else — which is why a provider you have never connected still has models to show. They carry context windows, pricing and reasoning tiers, but they are a snapshot of someone else's data and know nothing about what your account may call, so a stated entry and anything discovery returns both win over them. OHARNESS_NO_BUILTIN_MODELS=1 turns the seed off.

Two things bound it. A model is only seeded onto an entry whose protocol matches the API pi-ai catalogues it against — Mistral's models are listed against pi-ai's own conversations API, so they are left to discovery rather than sent to an OpenAI-compatible endpoint in the wrong shape. And vendors whose ids differ across the two catalogues are reconciled by naming the other one, not by guessing:

{ "providers": { "moonshot": { "piProvider": "moonshotai" } } }

A provider with no seed and no discovery yet — a local server, a gateway nobody bundles names for — still appears in the model picker, as a single row that opens the panel where you connect it.

The wire layer is pi-ai

A catalog entry's protocol names the format its endpoint speaks, and the implementation comes from [@earendil-works/pi-ai][pi-ai]:

anthropic-messages   openai-chat      openai-responses
gemini               bedrock          vertex
openai-codex

openai-codex is the odd one: it is the backend the Codex CLI talks to, and it is a different wire format from openai-responses rather than the same one on another host — store is refused, max_output_tokens is not sent, and the account id travels in its own header.

The vendor SDKs behind them are lazily imported, so a session that only talks to Anthropic never loads the AWS or Google clients.

What did not move is authentication. Credentials are resolved per request by the harness's own AuthManager — a subscription token valid at startup is not valid an hour later — so pi-ai's API implementations are called directly and handed the resolved key and headers, rather than going through its provider auth and credential store.

[pi-ai]: https://github.com/earendil-works/pi

Not included, and why

Azure OpenAIThe endpoint is per-deployment, so there is no fixed base URL to catalogue
Bedrock, VertexThe protocols work, but only with credentials already in the environment. Full SigV4 signing and gcloud ADC live in pi-ai's provider layer, which this harness deliberately does not use for auth — so no entries ship; declare one yourself if your environment is set up
A custom OpenAI-compatible serverAlready a first-class config feature: openAiCompatible

The official plan

official is the entry the pickers recommend and the one to point anyone at. It signs in with an OHarness account and asks for nothing else.

oharness auth login official    # "Sign in to your OHarness account"

Two credentials reach it, and they are the same account either way: the session from signing in, and an oh- key the account holder created on the site. The key is not a second wallet, only a second key to the same one — CI and SDKs need something that does not expire, and OHARNESS_API_KEY is how that key is carried into a process rather than a separate thing issued to a machine.

What no subscriber ever holds is an upstream credential. The OpenAI, Anthropic and Google keys stay server-side, and requests go to a metering endpoint that reads the account, checks the balance, writes the ledger entry and only then forwards. Which gateway sits behind that stays an implementation detail: replacing it is a change to one constant, not to anyone's setup.

VariableEffect
OHARNESS_API_KEYUse a key made on the site instead of signing in — how CI runs
OHARNESS_BASE_URLMove the metering endpoint (staging, self-hosted)
OHARNESS_TEAMBill this run to a team you are on, by its id — what a CI job on a team plan sets. The team key in ~/.oharness/config.json is the same choice made from the account menu
OHARNESS_SITE_URLPublic Cloud origin for browser login, exchange and refresh; Supabase configuration stays in Cloud

Credentials resolve in the order OHARNESS_API_KEY, session, stored key. GoTrue rotates its refresh token on every use, so each refresh is persisted — keeping the original would end the session at the next expiry.

Signing in

The flow is an emailed one-time code rather than a redirect: a terminal cannot receive a redirect over SSH or inside a container, and a magic link opens a page with no way to hand the session back. A code works even when the browser is on a different machine.

Sign-in requests create_user: false. Paying is the website's job, and quietly creating an account here would hand out a login that cannot call anything — an address with no account is told so instead of waiting for mail that never comes.

Both the web frontend and the desktop app run a local server, so both offer the shorter route: Sign in with the OHarness website, in the Models panel. It opens the site's own login in the browser — the page that has a password, Google, and the way to create an account — and the session comes back to a loopback callback the app is listening on.

Nothing is pasted either way. The account token is the credential for official, which is why the API key field on that card is folded away behind the sign-in and only there for OHARNESS_API_KEY in CI.

The site seals the session into a short-lived code and the app trades it back over HTTPS, so no token travels in a URL. A deploy of the site running more than one instance has to set OHARNESS_HANDOFF_SECRET, or the instance that seals a code will not be the one asked to open it. OHARNESS_SITE_URL moves the app to a different site — a staging deploy, or a local next dev.

The metering endpoint

What official points at is a service that used to live in this repository and now runs beside it, closed. Documented here anyway, because the harness speaks to it and you are entitled to know what happens to a request that leaves your machine: it attributes the request to an account, refuses it when the balance is gone, and — once the answer is in — writes what it cost to the ledger. In between it forwards, unchanged, with a master key the caller never sees. It does those three things and refuses to grow a fourth.

None of it is required. official is one provider among the ones configured in this repository, and nothing in the harness prefers it; point OTTOPORT_BASE_URL elsewhere, or use your own key with any other provider, and no part of the loop notices the difference.

Prices come from the upstream catalog's pricing field, so a model added upstream is priced the day it appears. What is ours is the pair of numbers that turn a dollar into a credit: OHARNESS_CREDIT_USD and OHARNESS_MARKUP. A model with no published price is not charged for, and says so in the log — inventing a price is the one failure that cannot be noticed later.

Streaming needs care that non-streaming does not. Usage arrives in the final SSE chunk, and only when the request asked for it, so a stream that did not ask has stream_options.include_usage set on the way out and the resulting chunk withheld on the way back. A client that never asked for usage should not start receiving it.

Subscriptions and top-ups

They are the same thing by the time anything reads them: both end as credits in the ledger. A top-up writes one purchase row; a subscription writes a grant row on every invoice it pays. So "can this account afford the call?" stays a sum, and no part of the service has to know which of the two paid.

What the subscription row is for is almost never permission but advice. An account out of credits is told to top up if it has a plan and to subscribe if it does not, and those are the opposite instruction to give the wrong person.

The exception is buying. Credits are sold to subscribers: a top-up, and the card armed to buy one unattended, are both refused for an account with no active or trialing plan. A pack covers a month that ran short rather than standing in for a subscription. The refusal has its own status and code — 403 subscription_required — so the site and the desktop paywall can offer the plans instead of reporting a failure, and both hide the buttons up front rather than letting someone pick an amount and then be turned away. Spending stays a question about the balance alone, so a plan that lapses strands nothing already bought; it only stops more being bought.

What is for sale lives in that service's source — a file called catalogue.ts — rather than in an environment variable. A price list is a product decision, and product decisions belong somewhere they can be read in a diff, reviewed and blamed; a deployment variable made the answer to "what do we sell?" something only its last editor knew, and let two environments sell different things by accident as easily as on purpose.

What legitimately differs between deployments is the Stripe price id, since the test project and the live one mint their own. Each entry carries both and the process picks by the key it was started with, so the two environments sell the same offer. Price ids are not secrets: they appear in every Checkout URL and identify nothing on their own.

Subscriptions are a ladder — one entry per tier, rendered smallest first — with a top-up pack and a pay-as-you-go plan beside them. The credits are stated rather than derived from the amount: a plan is a promise that this much money buys this much, and inferring it from Stripe would silently reprice everything the day a currency or a discount changed.

The price is the other way round. What the pricing page prints is read from the Stripe Price the plan bills against, re-read every ten minutes, so the page and the checkout cannot disagree and a price edited in the dashboard needs no deploy. Pay-as-you-go is the exception that must state amountCents and a currency itself: a recharge is an off-session PaymentIntent, which takes an amount rather than a Price. A price that cannot be read at all is published without one, and the page falls back to advertising credits alone rather than a number that might be wrong.

Five webhook events are handled — a completed payment checkout, the customer.subscription.* mirror, invoice.paid, and the two ways money goes back out, charge.refunded and charge.dispute.created. Every write is keyed on a Stripe id, because a webhook that is not retried loses money and one that is retried without a key grants it twice. Grants follow invoices rather than renewal dates: an invoice is the event that means money arrived, and a renewal that fails to charge should grant nothing.

A reversal takes back exactly what the ledger says the payment granted, read back rather than recomputed from a plan that may have been repriced since. If the credits were already spent the balance goes negative and the account settles before it calls anything again — the alternative is that refunding a top-up is a way to buy free inference. Partial refunds are logged for a human: deciding how many of a shared pool of credits a fraction of a payment bought is a policy question, not a mechanical one.

An invoice that settled at zero grants nothing. Stripe raises one for a free trial and for a fully discounted month, and an invoice is only handled here because it means money arrived — so a Stripe-native trial starts empty. That fails visibly, to whoever configured the trial, rather than invisibly to whoever reads the bill; sell a trial as its own cheap plan instead.

The dashboard renders the ledger, not just the balance. The rows carry the model and the token counts, which is what lets someone tell a runaway loop from a price change — and a balance nobody can check is a balance nobody trusts.

Cancelling, changing a card and reading past invoices are Stripe's own billing portal, reached through POST /billing/portal. Those screens are where being wrong means charging someone who cancelled, and Stripe's are already correct in every currency it sells in. The portal is offered to anyone Stripe has met, not only to subscribers: a single credit pack and a card left on file both leave something to manage.

Stripe returns the customer to the site's dashboard — /<locale>/dashboard/billing, which is where all of this is rendered. The old /<locale>/account address 308s there and keeps doing so: it is baked into shipped desktop builds and into whatever version of this service is currently deployed, so the two can be released independently.

GET /billing/invoices lists what an account was billed, read live from Stripe and never mirrored into a table here. Each row carries Stripe's own PDF and hosted links, minted per request because they expire, so the download is the original document — with the tax it charged and any credit note that reversed it — rather than a copy of ours that can disagree with it. The reply also says whether the account is a Stripe customer at all, which is what tells the site whether there is a portal to offer, and is not the same as having no invoices.

The site holds no credential that can move money. Checkout goes site → gateway, carrying the reader's own session token, and the gateway holds the Stripe key, the Supabase service role and the upstream key together.

npm run smoke:gateway stands Stripe and the account project up as fakes and checks the parts that touch money: the charge against the price list, the withheld usage chunk, the refusal of an empty account, that the caller's credential never reaches upstream, that a replayed webhook lands once, and that an invoice grants exactly one month.

Bringing your own OttoPort account

ottoport is an ordinary vendor entry, no different in kind from OpenAI or OpenRouter: you hold an OttoPort account, create an op- key in its console, and your usage is billed against your own credit balance there.

oharness auth login ottoport    # paste an op- key
VariableEffect
OTTOPORT_API_KEYUse a key without the prompt
OTTOPORT_BASE_URLMove the gateway (staging, self-hosted) — a host, /api/v1 is appended

Two things about this endpoint are the gateway's rather than any model's, and the catalog entry states both because /models mentions neither. It reserves the cost of a completion before serving it, so every request carries an Idempotency-Key — one value per call, reused only by a retry of that same call. And it caps output at its own GATEWAY_MAX_OUTPUT_TOKENS, 4096 by default, well under what the models it serves advertise; requests are clamped to that on the way out. A deployment that raises the cap should raise maxOutputTokens on the entry with it, or answers stay short for no visible reason.

It is not the official plan. That plan happens to route through this gateway today, but the subscriber never holds an OttoPort key and never sees an OttoPort account — which is what leaves the plan's billing, quota and choice of upstream ours to change.

ChatGPT subscriptions

openai-codex reaches GPT models on a ChatGPT Plus/Pro plan. It is the only catalog entry that names no env var, because its endpoint takes no API key at all — only a ChatGPT OAuth token:

codex login                        # the official CLI runs the browser flow
oharness auth login openai-codex   # choose "ChatGPT subscription via the Codex CLI"

Each request re-reads ~/.codex/auth.json (CODEX_HOME moves it) and sends that token as a bearer. Nothing here ever refreshes it. OpenAI rotates the refresh token on use, so renewing it in this process would invalidate the copy the CLI holds and sign the user out of their own codex command at random. An expired token is reported as no credential, and the fix is codex login.

This rather than hardcoding a client ID, for the reason the Anthropic entry gives: the id every other implementation uses is the one the Codex CLI registered, and logging in with it means authenticating as that application.

Two consequences of subscription: true on the entry. The seeded models carry no prices — the plan includes the usage, so quoting the vendor's list price would report spend that never happens. And the entry still requires a credential despite its empty envVars, which for any other entry would mean a local server that needs none.

excludeModels lists what the backend refuses to a ChatGPT account outright. It is plan eligibility rather than a capability, so it is data: a plan that does serve one can drop it in a user catalog.

Credentials

Resolved on every request, not once at construction — a subscription token expires, and an implementation that computes headers once breaks an hour in. Concurrent refreshes are de-duplicated, because providers that rotate the refresh token on use would otherwise invalidate each other.