User hit 'API error 402: ... You requested up to 131072 tokens, but can only afford 4511' on the assist alias: no provider ever sent max_tokens, so OpenRouter's credit pre-check billed the routed model's full worst-case output; user also asked to bound session history to the last 5 requests/responses. Architect decisions: - AI_MAX_TOKENS (num, default 2048): sent as max_tokens on OpenRouter and generationConfig.maxOutputTokens on Gemini — a real per-request cost ceiling. llamacpp unchanged (local/free, no pre-check). - AI_SESSION_TURNS (num, default 40 kept back-compat; messages, 2 per exchange — 10 = last 5 conversations): resolved lazily in session_push because config loads after the hardcoded line-25 default. - Both registered in the bin/pos-ai POS_CONFIG @General section, so they appear in 'pos config ai' with num: validation. Reviewer hardening (CHANGES_REQUIRED -> fixed): unguarded env input could reach jq tonumber (0/-5/010/abc all savable via config-ui's ^-?[0-9]+$) and abort the CLI; both providers and session_push now guard with ^[1-9][0-9]*$ and fall back to the default. Verified: fake-curl shim smoke (16 provider-body + 12 session-window checks incl. the 010-regression proof), make gen idempotent, make check OK, make lint 0 FAIL/0 WARN, make test 17 files / 299 checks / 0 fail (~49s), bash -n clean, git diff --check clean. Reviewer ACCEPT (twice). Tester regression round (permanent provider-body + session-pruning coverage) intentionally not run this cycle — user's call; remains a documented follow-up.
15 KiB
AI cost & session-window design — OpenRouter 402 + bounded session memory
Date: 2026-09-06
Author: Architect (big-pickle)
HEAD: 8ce5479 (clean tree)
Mode: design-only (no file edits besides this report)
TL;DR
Two tightly associated defects in pos ai / bin/pos-ai:
- 402 root cause — no provider ever sends
max_tokens, so OpenRouter's credit pre-check charges the routed model's full worst-case output (131072 onopenrouter/auto) and rejects balances under that. Fix: send an explicit cappedmax_tokens. - Unbounded-ish session —
MAX_SESSION_TURNS=40counts messages (20 exchanges); user wants "last 5 req/response" = 10 messages.
Decisions: add AI_MAX_TOKENS (num, default 2048) honored by openrouter
(body max_tokens) and gemini (generationConfig.maxOutputTokens); skip llamacpp.
Add AI_SESSION_TURNS (num, default 40 — backward compatible) honored lazily in
session_push; user sets 10 for 5 pairs. Both vars declared in the @General
section of the # POS_CONFIG: header, docs updated. No chat/alias wrapper changes.
Open items: none blocking. Tester should add provider-body + session-pruning coverage (see §Testing).
Verified fact confirmation (A–D)
All user-reported facts confirmed against source at HEAD:
A — 402 root cause: CONFIRMED.
lib/ai-providers/openrouter.sh:22-23— body is only{model,messages}; nomax_tokens.:24-29POSTs straight to OpenRouter with unchanged body.lib/ai-providers/gemini.sh:17-19— body only{contents,...}; nogenerationConfig. (:20-23adds onlysystemInstruction.)lib/ai-providers/llamacpp.sh:32-33— body{model,messages,stream:false}.- Conclusion: none carry a generation cap → OpenRouter 402 with the user's thin
balance. Fix is to send an explicit
max_tokens.
B — session window: CONFIRMED.
bin/pos-ai:25MAX_SESSION_TURNS=40.session_push()bin/pos-ai:285-290→'.messages |= .[-"$MAX_SESSION_TURNS":]'prunes to last N messages (40 msgs = 20 exchanges). User wants 5 pairs = 10 messages.
C — config surface: CONFIRMED.
# POS_CONFIG:headerbin/pos-ai:6, scopeai | ai.env.num:type already used in the same header (LLAMACPP_CTX_SIZE=num:…,LLAMACPP_GPU_LAYERS=num:…).- Validation:
lib/config-ui.sh:419*,num,*) [[ "$val" =~ ^-?[0-9]+$ ]]— integer-only on entry; empty input ="kept current value" (:468-471);-= clear (:472-477). Sonum:+ empty/unset → falls back to code default. Clean. - Env precedence:
load_env_file(lib/config-ui.sh:336-357) exports a file key only when the variable is not already set in the environment (:351-353) → env beats file beats default. Providers read config via env (AI_API_KEYpattern). PROVIDER_CONFIGheaders exist onopenrouter.sh:7-8,gemini.sh:7-8,llamacpp.sh:7for provider-specific keys.
D — provider resolution: CONFIRMED.
bin/pos-ai:679-683—--providerflag >AI_PROVIDERenv >gemini.load_config()bin/pos-ai:141-151loadsai.env(+ legacy files); called byresolve_key(line 155) and at provider resolution (line 681).- Call site
bin/pos-ai:522(ask) and:557(chat):provider_generate "$model" "$messages" "$system".
Gates/tests scan:
- No test pins
MAX_SESSION_TURNS, the provider request bodies, or header text (t-config-precedence.shtargetspos-ai-server, a separate tool). No forced test update. DOC/POS.mdhas a hand-maintained ai.env config table (:88-97) and mentions "capped at 40 turns" at:64.DOC/HOWTO.md:45andDOC/AGENT_Context_Project.md:491list ai.env vars (hand-maintained). All need doc rows/bumps for the new vars.
Decision 1: AI_MAX_TOKENS — cap generation tokens
Status: [DECIDED]
Options & trade-offs
Option 1 (chosen) — single global AI_MAX_TOKENS=num, default 2048, honored by remote providers (openrouter + gemini); skip llamacpp.
- Advantages: smallest change that fixes the 402 (OpenRouter pre-check sees a
capped cost) and is a real per-request cost ceiling; one var, one default; fits
the existing
AI_*env naming and the tool-level@Generalconfig section; no new per-provider surface. - Costs: remote providers share one ceiling (no per-provider cap without user intervention).
- Risks: a too-low cap truncates long answers — mitigated by default 2048 being
ample for terse
ask/chatCLI answers; user can raise it. - Reasoning for default 2048: conservative (user's balance affords ~4511 tokens at routed price, so 2048 passes the pre-check with margin) while being a practical, real ceiling. 2048 tokens ≈ several thousand chars — plenty for the terse, commands-first assistant role this tool plays.
Option 2 — per-provider caps via PROVIDER_CONFIG (e.g. AI_MAX_TOKENS on openrouter.sh, gemini.sh).
- Advantages: independent ceilings per provider.
- Costs: two declarations, redundant section plumbing, and the default (which is
the entire point) still has no shared home → awkward.
PROVIDER_CONFIGis for provider-specific concerns; a cost ceiling + 402 pre-check over both remote providers is tool-level, not provider-specific.
Option 3 — no gemini cap; only openrouter.
- Advantages: minimal (402 only affects OpenRouter).
- Costs: leaves Gemini without any cost ceiling while introducing the same var — inconsistent, and Gemini's own pricing can surprise. Rejected.
Option 4 — include llamacpp too (max_tokens in body; it accepts it).
- Advantages: provider parity on the OpenAI-compatible endpoint.
- Costs: local & free — no credit pre-check, no cost. Adds surface with zero user benefit. Rejected on the "smallest sufficient design" principle.
Implementation contract
- Env var name:
AI_MAX_TOKENS, typenum, default 2048. - OpenRouter body (
openrouter.sh:22-23): addmax_tokens. - Gemini body (
gemini.sh:17-19): addgenerationConfig.maxOutputTokens. - llamacpp: no change.
- Providers read
"${AI_MAX_TOKENS:-2048}"from env; value is present becauseload_configruns beforeprovider_generate(resolve path confirmed in D). - Declare in
@Generalsection of# POS_CONFIG:header (bin/pos-ai:6).
Decision 2: AI_SESSION_TURNS — bounded session window
Status: [DECIDED]
Options & trade-offs
Option 1 (chosen) — new AI_SESSION_TURNS=num, default 40 (unchanged), resolved lazily in session_push.
- Advantages: fully backward compatible — no silent memory truncation for existing
users. The user's "last 5 req/response" = setting
AI_SESSION_TURNS=10. One var, one default. - Costs: existing users must opt in (they already have the 40 behavior, so no regression).
- Why keep default 40: backward compatibility is a hard project value; silently shifting the default changes session context for every user, discards history they may rely on, and is a behavioral change not requested globally (only for this user). Keep 40.
Option 2 — change the default to 10.
- Advantages: meets the stated want out-of-the-box.
- Costs: silent behavior change for all users; discards memory; not requested globally. Rejected — keep the change opt-in via the new var.
Semantics (must be documented)
The existing prune is .messages |= .[-N:] where N counts messages — 2 messages
per exchange. So AI_SESSION_TURNS=10 ⇒ last 5 exchanges (5 user + 5 assistant).
The config description must state: "message count (2 per exchange); 10 = last 5
exchanges".
Implementation constraint — lazy resolution (important)
MAX_SESSION_TURNS is currently assigned at bin/pos-ai:25, which executes at
top-level before load_config is first called (line 681). If we wrote
MAX_SESSION_TURNS="${AI_SESSION_TURNS:-40}" at line 25, an AI_SESSION_TURNS set
only in ai.env would not yet be loaded → always 40.
Therefore:
- Keep line 25 as-is (
MAX_SESSION_TURNS=40), used for the help text (:67shows the default, accurate). - In
session_push()(lines 285-290), resolve lazily:local n="${AI_SESSION_TURNS:-$MAX_SESSION_TURNS}"and use$nin the jq prune. Becausesession_pushruns insidecmd_ask/cmd_chat— afterload_config(viaresolve_keyat:499/:538) has exportedAI_SESSION_TURNSinto the process env — the ai.env value is honored. Shell-exportedAI_SESSION_TURNSwins too (env-wins inload_env_file). - Help text
bin/pos-ai:67stays accurate ("capped at $MAX_SESSION_TURNS turns") since the default remains 40. Optionally add aConfig:help line documenting the var — recommended, and it confirms "turns = messages, 10 = 5 pairs".
Decision 3: # POS_CONFIG: header change
Status: [DECIDED]
- Single
@Generalsection addition (not per-providerPROVIDER_CONFIG):AI_MAX_TOKENS=num:Max output tokens per request (default 2048; OpenRouter/Gemini cost cap)AI_SESSION_TURNS=num:Session message cap — 2 per exchange (default 40 = 20 exchanges; 10 = last 5)
- Place both in the
@Generalsection alongsideAI_SYSTEM_PROMPT(end of the long header line,bin/pos-ai:6). - Why @General, not PROVIDER_CONFIG: the default is shared across providers
(2048, 40) and both caps are tool-level concerns.
PROVIDER_CONFIGis reserved for provider-specific keys (API keys, models). A single global declaration is the cleanest and avoids duplicating the default in two provider files. num:type confirmed suitable:lib/config-ui.sh:419enforces integer on entry; empty=keep current,-=clear (:468-477); unset → code default. No empty-parse concern.- Provider bodies consume the vars from env, so no
PROVIDER_CONFIGadditions are needed onopenrouter.sh/gemini.sh. (They could be added later if per-provider caps are ever wanted — out of scope now.)
Decision 4: Docs & gates
Status: [DECIDED]
DOC/POS.md(hand-maintained):- ai.env table (
:88-97): add rows forAI_MAX_TOKENS(no/2048/"Max output tokens per request (OpenRouter/Gemini cost cap)") andAI_SESSION_TURNS(no/40/"Session message cap — 2 per exchange; 10 = last 5 exchanges"). TheAI_SYSTEM_PROMPTrow (:93) is the placement anchor. - Line 64 text "capped at 40 turns" remains true (default unchanged) — no edit strictly needed, but a short parenthetical "(configurable via AI_SESSION_TURNS)" is recommended.
- ai.env table (
DOC/HOWTO.md:45— appendAI_MAX_TOKENS,AI_SESSION_TURNSto the listed ai.env vars.DOC/AGENT_Context_Project.md:491— append the two vars to the ai.env summary parenthetical (hand-maintained).make gen: header text change does not add commands/subcommands/flags, so the generated tree/dispatch tables are unaffected;completions/pos.bashconfig-scope table regenerates to include the new keys. Runmake gen(deterministic,LC_ALL=Cper convention), thenmake check, thenmake lint.- Line-count rows above the filetable marker in
AGENT_Context_Project.md: only bump a row if apos-*/lib/*file's length changes (it does — lib/ai-providers grow; bin/pos-ai grows). Do not touch rows for files that don't change. - tests: existing suite does not pin the session default or provider bodies, so nothing is forced. Recommend new coverage (§Testing).
Decision 5: Scope fence
Status: [DECIDED]
Approved outcome: OpenRouter 402 eliminated (explicit capped max_tokens on
remote providers) and session history bounded via configurable AI_SESSION_TURNS.
In-scope files:
bin/pos-ai—# POS_CONFIG:header (line 6, @General additions);session_pushlazyAI_SESSION_TURNSresolution (lines 285-290); optionalConfig:help lines for the two vars. Line 25 staysMAX_SESSION_TURNS=40.lib/ai-providers/openrouter.sh— addmax_tokensto body (lines 22-23).lib/ai-providers/gemini.sh— addgenerationConfig.maxOutputTokens(lines 17-19).lib/ai-providers/llamacpp.sh— no change.- Docs (hand-maintained):
DOC/POS.md,DOC/HOWTO.md,DOC/AGENT_Context_Project.md. - tests (later, Tester).
Allowed interface changes: two new config keys in scope ai; provider request
bodies gain a token cap. Provider call signature provider_generate "$model" "$messages" "$system"
is unchanged.
Explicitly out of scope: bin/pos-ai-server, chat/alias wrappers, pos-ai-alias,
HuggingFace downloader, other tools, per-provider caps, changing MAX_SESSION_TURNS
default, any llmacpp body change.
Architectural constraints:
- No new provider-side config plumbing; vars read from env (
AI_*convention). - Backward compatible: session default 40 and empty/unset values fall back to documented defaults.
- Deterministic
make gen;make check+make lintgreen (definition of done).
Open risks:
AI_MAX_TOKENStoo low truncates long--fullanswers — default 2048 mitigates; user can raise.- env-vs-file precedence: exported env var beats config file (per
load_env_file) — expected and documented.
Verification:
bash -nclean;make gen/make check/make lintgreen.- A 402 reproduction no longer triggers when
AI_MAX_TOKENSis set/at default. AI_SESSION_TURNS=10prunes the session to last 5 exchanges.
Testing budget (suggestion for Tester / Builder verification)
- Provider body cap: source each provider adapter in a sandbox with
curlstubbed, assert the request JSON containsmax_tokens(openrouter) / a truthygenerationConfig.maxOutputTokens(gemini). Pattern:t-config-precedence.shstyle curl-log capture with a fakecurlinPATH(see existing tests). - Session pruning: drive
session_pushwith a crafted messages JSON andAI_SESSION_TURNS=10, assert exactly the last 10 messages remain (5 pairs); andAI_SESSION_TURNSunset → 40 retained (default commits to backward compat). - Help/default fidelity:
pos ai --helpstill shows 40;pos config ailists the two newnum:keys and rejects a non-integer (cfg_validate). - Regression: existing
t-config-precedence.sh,run-testsfull suite green.
Handoff
Status: DECISION_READY
Problem: OpenRouter 402 (no max_tokens → full worst-case pre-check) and
unbounded session memory; user wants 5 req/response.
Decision: add AI_MAX_TOKENS=num (default 2048) honored by openrouter + gemini
(skip llamacpp); add AI_SESSION_TURNS=num (default 40, backward compatible) resolved
lazily in session_push; both in @General config section; docs updated.
Ownership: bin/pos-ai + lib/ai-providers/{openrouter,gemini}.sh (+ docs).
Interfaces: two new ai-scope config keys; provider bodies gain a token cap.
Call signature unchanged.
Approved scope / constraints / verification / out-of-scope: see Decision 5.
Risks: see Decision 5 (token-cap truncation; env-vs-file precedence — all mitigated).
Recommended next agent: Builder
Reason: the architecture and scope are fully specified with exact line-level changes (no architectural ambiguity left). Builder can implement without making architecture decisions. Tester follows for the recommended coverage.