Provider map
CLI agents are a distinct backend, not another/chat/completions endpoint. See CLI agent provider internals for the Companion, structured tools, gateways, and lifecycle authorization.
Core modules
The shared chat model picker (chat-model-controls.js, chat-model-preferences.js) follows the active provider and privacy mode immediately. Recommended groups are scoped to that catalog, not merged with other saved providers. reasoning-capabilities.js normalizes explicit reasoning metadata and narrow documented fallbacks; unsupported models must not get an invented effort control. The same selection is carried inside attested private transports when supported.
OpenRouter supplies per-model controls, Venice supplies reasoning metadata, and PPQ/Routstr/Custom retain supported compatible metadata. Ollama and LM Studio can expose native thinking controls. Jan, Unsloth Studio, llama.cpp, and other compatible endpoints only get controls their metadata/serving contract supports. Built-in reasoning is not evidence of configurable effort.
For non-chat feature calls use ai-feature-routing.js so CLI selection, global pause, text/image capability, and no-silent-fallback behavior remain consistent. Keep direct transport functions in the API modules; do not implement a second CLI dispatcher per feature.
Local AI adapter contract
LOCAL_AI_PROVIDER_ADAPTERS registers Ollama, LM Studio, Unsloth Studio, and generic compatible adapters with a stable id, user-facing label, normalized capabilities, and discover() method. Adapters can also implement infer(), prepareNativeRequest(), loadWithContext(), and unload().
Code outside the adapters must branch on capabilities or optional methods, not hard-coded product names. Unknown providers fall back to the compatible adapter.
Detection order
Discovery normalizes the base URL before caching or probing.- Probe LM Studio’s native API and
/v1/modelsin parallel. - If the native LM Studio API succeeds, merge its richer metadata with compatible model IDs and do not send Ollama-only probes.
- If compatible models identify
owned_by: unsloth-studio, query/api/inference/status, enrich loaded/context/vision state when available, and do not send Ollama-only probes. - Otherwise probe Ollama’s
/api/tagsand/api/psendpoints. - Prefer Ollama when it is available with models; otherwise use the identified or generic compatible result.
Location and privacy classification
getLocalAiExecutionLocation() reports:
localforlocalhost,127.0.0.1, and IPv6 loopback;lanfor.local, private IPv4, link-local, and private/link-local IPv6 hosts;remotefor other hosts; andcloudfor cloud-tagged model names, regardless of the URL.
Context planning and inference
api-local.js estimates prompt tokens, adds a safety reserve, and clamps requested output to the loaded context. If the remaining context cannot satisfy the minimum output, it fails with the current and maximum context sizes instead of sending a predictably truncated request.
Imports and meal-photo analysis opt into discovery and native context planning because a full report or structured nutrition result can exceed a default loaded context. Dense lab tables use a conservative three-characters-per-token estimate; routine chat avoids an extra discovery probe before the first token.
nutrition-ai-settings.js keeps an optional meal-specific model route separate from the main chat model while remaining on the active provider, and filters its catalog to image-capable models. Benchmarks can select models from multiple configured providers. Meal requests use recipient consent kind meal-photo, structured JSON output, no automatic retry, and their own abort signal. See Meals and Nutrition internals.
Ollama can pass num_ctx. For a current LM Studio server, prepareNativeRequest() selects a target context and loadWithContext() unloads the previous instance, loads the model through /api/v1/models/load, forces discovery, and returns the context LM Studio actually fitted. api-local.js replans output against that verified context, then generates through streaming /v1/chat/completions. A missing load route falls back to native non-streaming inference with a 15-minute ceiling. An unload failure may proceed only after forced discovery proves the old instance is gone; otherwise it fails before allocating a duplicate.
Local streaming gets a 15-minute allowance only for the first reader result because prompt prefill can be silent. Every later read keeps the shared 30-second stall guard. Provider finish reasons and native LM Studio token accounting normalize to truncated; PDF text and image imports reject that result before JSON review so a partial marker list cannot be confirmed.
All adapters return the common text/usage result plus optional diagnostics. Generic compatible endpoints and Custom API use direct browser requests to /v1/chat/completions.
Runtime handoff and VRAM
local-ai-lifecycle.js keeps only the last base URL, adapter ID, and model in module memory. Before using a different endpoint/model on the same machine, it discovers the previous runtime and unloads its loaded model when the adapter supports that operation.
Do not persist runtime-use state: it becomes stale after refresh and can expose misleading loaded-model claims. If a supported unload operation fails while changing servers in Settings, keep the new selection unsaved and return an actionable manual-unload error. An adapter without unload support can proceed without automatic release. Never attempt to unload cloud-tagged models or models on a different machine.
Model-test runtime capture is documented in AI model testing internals.
Model storage rules
Regular and private/encrypted modes may need separate model keys. Do not let a model selected for a public provider path leak into a private-mode request if the private endpoint has a smaller model list. When switching encryption/private modes:- clear or reset incompatible sessions;
- restore the last valid model for that mode when possible;
- hide unavailable controls rather than sending unsupported requests;
- update pricing/usage hints after the final model is selected.
Web search boundary
Web search can add external search results to the prompt. It must be disabled when the active transport promises private/encrypted prompt handling that would be broken by search, including PPQ Private TEE and Venice Encrypted TEE Mode. User docs should say this plainly, and UI state should match it: no visible Web toggle when search is unavailable.Custom API direct-transport boundary
api-custom.js calls /models and /chat/completions directly from the browser with credentials: 'omit'. callOpenAICompatibleAPI() receives { useProxy: false }, and connection errors explicitly state that no server retry occurred.
Remote Custom origins require recipient/origin-scoped cloud-processing consent before chat data is sent. Loopback and private-network Custom origins are treated as local for that consent gate, though the configured machine can still read the request. The endpoint must allow the exact app origin with CORS and must satisfy mixed-content rules.
Do not reintroduce an implicit relay fallback. It would change which operator can see prompts and credentials and would require a new implementation, privacy disclosure, consent route, abuse controls, and security review.
OpenRouter OAuth invariants
OpenRouter OAuth must bind state to the browser session and use PKCE/state validation. Do not accept callback tokens without matching the pending state. OAuth errors should leave existing manually-entered keys untouched unless the user explicitly clears them.Verification checklist
Before shipping provider changes:- run targeted provider tests for the touched provider;
- run
tests/api-provider-contracts.test.jsfor adapter inference and redaction contracts; - run
tests/test-provider-local-ai-runtime.jsfor discovery, context, location, and lifecycle behavior; - run chat streaming tests if transport code changed;
- run Settings panel tests if dropdowns/toggles changed;
- verify Local AI connection copy for loopback, LAN, remote, cloud, CORS, and mixed-content cases;
- verify capability-optional fields render safely for a generic compatible endpoint;
- verify switching same-machine Ollama/LM Studio runtimes unloads only when supported;
- verify Local AI imports distinguish first-token prefill from mid-stream stalls and reject truncated marker lists;
- verify web search visibility for public vs private modes;
- verify wrong/missing API key errors are clear and do not erase saved keys;
- verify private/E2EE attestation status is visible and fail-closed;
- verify no raw keys, tokens, prompts, or setup blobs appear in logs or docs.