Commit Graph
100 Commits
Author SHA1 Message Date
Yichen Jiang 2e15fc0bfe Merge branch 'stack/agent-profiles-1-seam' into stack/agent-profiles-2-configs 2026-08-07 13:34:19 +08:00
Yichen Jiang a7be2ecb0f Merge remote-tracking branch 'origin/master' into stack/agent-profiles-1-seam 2026-08-07 13:34:08 +08:00
Yichen Jiang ddacc8de45 Merge pull request #1735 from deepseek-harness/worktree/provider-credential-lifecycle
fix(web): recover provider credential lifecycle
2026-08-07 13:27:18 +08:00
Yichen Jiang abd2e2973b Merge branch 'master' into worktree/provider-credential-lifecycle 2026-08-07 11:37:41 +08:00
Yichen Jiang 6edd98e2b2 Merge branch 'master' into worktree/provider-credential-lifecycle 2026-08-07 11:28:22 +08:00
Yichen Jiang 380eaec39b Merge branch 'stack/agent-profiles-1-seam' into stack/agent-profiles-2-configs 2026-08-07 01:01:38 +08:00
Yichen Jiang 7c7cd1aa35 Merge remote-tracking branch 'origin/master' into stack/agent-profiles-1-seam 2026-08-07 00:59:28 +08:00
Yichen Jiang e27d38efd6 feat(app-boot): register cordis:group beside cordis:include
A composition that shares one `isolate` realm across rows needs a
`cordis:group` row, and a preset living outside this workspace — the
authored ones under the Harness home — cannot resolve
`@cordisjs/plugin-group` by name: Node's upward `node_modules` walk
never reaches the harness from there. Registering it as a loader builtin
beside `cordis:include` loads both through the ambient module pipeline
instead.

Record it in the preset Agent Note, which leans on the realm vocabulary
without saying where the group row comes from, and drop the preset
README's limitation claiming this builtin is unavailable — it described
the state this change ends.

The test's assertion had a vacuous escape: `provide` mints the root
symbol unconditionally, so the `rootKey === undefined` disjunct could
never hold and the comment claiming the root realm never learned the
name was wrong. Pin both halves — the symbol exists, nothing is stored
under it — and clean up the global the fixture writes.
2026-08-07 00:34:19 +08:00
Yichen Jiang 065257addb fix(agent-presets): bound the mount registry on a host that never reads it
Records are pruned by observation rather than by a disposal hook, for the
reason the module already states: three different owners can tear a
subtree down, and a cleared `uid` is what they share. That leaves the
pruning to whoever reads — and the only production reader is the
invariant companion, whose package is a development composition a
shipped host never loads.

So a live host pruned nothing: every session ever composed left a record
retaining its whole disposed subtree, since the fiber holds its config
and that config is the key its EntryTree is stored under.

Prune on the mount path too. Every session takes it, which bounds the
set at one generation of dead records instead of one per session.
2026-08-07 00:33:40 +08:00
Yichen Jiang fe8a88c96a Merge pull request #1803 from deepseek-harness/worktree/align-core-web-rl-prompt
fix(web): align core profile with RL prompt
2026-08-06 21:57:28 +08:00
Yichen Jiang 18fe174897 feat(agent-presets): compose each session's agent from a preset cordis.yml
A preset is a directory holding one `agent.cordis.yml`. Mounting it under an
agent's scope context during `setup(agentCtx)` gives that one session its own
tools and prompt sections while every other live session keeps its own.

No registry gains a tier. `dsh-tools` and `dsh-system-prompt` already file
registrations into the calling context's scope layer, and entry contexts chain
to the context a subtree was plugged into, so a composition mounted under
`agent.ctx` is that agent's alone and unwinds with it.

The mount audits itself because a directly-plugged subtree is absent from
`ctx.loader.entries()` and no boot audit covers it. It rejects an unscoped
target, a row that never became usable, and a row that published a service into
the root service realm — that last one is process-global rather than
per-session, and its collision with the next session surfaces as an unhandled
rejection `setup` never observes, leaving a half-composed agent that looks
healthy. The package invariant re-checks that rule on every service
notification, since a row publishing from a timer would escape a one-shot audit.

Raises the `packages/README.md` word ceiling from 920 to 980: the group table
must enumerate every group, and the new `preset/` row is necessary content.

Design: .agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.md
2026-08-06 21:03:18 +08:00
Yichen Jiang 53d47a8f86 docs(cli): clarify core prompt environment semantics 2026-08-06 20:45:57 +08:00
Yichen Jiang 266629f1c3 test(web): cover core profile request header 2026-08-06 20:41:19 +08:00
Yichen Jiang b5600e9376 Merge remote-tracking branch 'origin/master' into worktree/align-core-web-rl-prompt
# Conflicts:
#	apps/cli/reference/README.i18n.yaml
#	apps/cli/reference/README.md
#	apps/cli/reference/README.zh.md
#	apps/cli/src/app-cli-entry.ts
#	apps/cli/src/dump-config.ts
#	apps/cli/src/web.ts
#	apps/cli/tests/built-bin.e2e.ts
#	apps/cli/tests/web-prompt-context.spec.ts
#	apps/web/tests/scaffold.ts
2026-08-06 20:34:57 +08:00
Yichen Jiang 50c46be948 fix(cli): isolate Web runtime context from headless 2026-08-06 20:11:44 +08:00
Yichen Jiang 0cf1ba7f87 fix(web): align core profile with RL prompt 2026-08-06 19:52:16 +08:00
Yichen Jiang e43e4f187e fix(web): satisfy provider model gates 2026-08-06 17:31:25 +08:00
Yichen Jiang 66b136e6e3 Merge remote-tracking branch 'origin/master' into worktree/provider-credential-lifecycle
# Conflicts:
#	packages/client/ui-models/README.i18n.yaml
#	packages/client/ui-models/README.md
#	packages/client/ui-models/README.zh.md
#	packages/client/ui-models/src/client/ModelsSection.tsx
#	packages/client/ui-models/src/client/ProviderEditor.tsx
2026-08-06 16:54:33 +08:00
Yichen Jiang f2050bfd1e fix(web): surface provider credential status 2026-08-06 16:44:04 +08:00
Yichen Jiang b13b0e2983 Merge pull request #1368 from deepseek-harness/claude/web-pi-ai-provider-form
feat(web): declare a provider and its models from the Models page
2026-08-06 16:37:14 +08:00
Yichen Jiang 8eebf7dd40 test(web): pin the remaining markdown fixture's event times
Master stabilized two of the three hand-built markdown fixtures by spacing
their event times, so the stats line's LLM segment stops depending on how fast
the replay ran. `markdown-images` was left on the old shape and kept failing
for exactly that reason — a different test each run, always the same one-line
difference.

Pin it the same way and record the segment its golden now always shows. This
supersedes the normalizer that dropped those segments outright, reverted here:
pinning the fixture keeps a real part of the page in the goldens instead of
hiding it from every one of them, and master's two goldens already record it.
2026-08-06 16:24:20 +08:00
Yichen Jiang 5e5fd977c5 Merge branch 'claude/pi-ai-model-discovery' into claude/web-pi-ai-provider-form 2026-08-06 16:18:57 +08:00
Yichen Jiang 2a790e6120 Merge branch 'claude/web-llm-pi-ai-config-385e24' into claude/pi-ai-model-discovery 2026-08-06 16:18:15 +08:00
Yichen Jiang be4c996d97 Merge remote-tracking branch 'origin/master' into claude/web-llm-pi-ai-config-385e24 2026-08-06 16:17:33 +08:00
Yichen Jiang 3e1c63b2ea test(web): stop the stats line's wall-clock segments from deciding a golden
`StatsLine` renders its LLM, tool-call, and throughput segments only while the
matching measurement exceeds zero, and all three are wall clock taken during
the replay. A machine that finishes a step inside one millisecond drops the
segment a slower one keeps, so a golden recorded what the recording machine's
speed was rather than what the page shows. Goldens across this suite already
disagreed about the LLM segment for that reason, and CI failed on whichever
test landed on a slow enough runner — a different test each run, always the
same one-line difference.

Tokenizing the values was never enough, because presence is what moves. The
normalizer now drops those segments outright, each taking one adjacent
separator so nothing is left holding a dangling separator or a doubled space,
and the recorded goldens are normalized the same way. `TTFT avg` stays: it
gates on a step count the fixture determines.
2026-08-06 15:47:42 +08:00
Yichen Jiang dd36db5434 test(ui-models): cover reading a stored capacity back out of a row
Every capacity assertion typed into the field first, so the path that reads a
stored count — the one an already-configured route takes every time it is
opened — was never exercised, and the coverage gate said so. A row now opens
on stored counts and is expected to spell them `1M` and `256K`, the same
vocabulary the field accepts.

`capacityText` takes the row it is rendering rather than looking it up again
by index, which retires the impossible empty-row branch that lookup needed.
The list editor's pre-flight model check is unreachable from the card that
disables submit on the same failure, and says so where it stands.
2026-08-06 15:18:18 +08:00
Yichen Jiang 3a3abc2bc4 fix(ui-models): restore the add-provider row and hint an empty capacity
The two ways to gain a provider had picked up the shared button base's pill
shape and shrunk to their labels, so they read as two stray buttons of
different lengths under the list instead of its last slot. They split the row
evenly again, on the row cards' own corner and the dashed outline this page
already uses for "nothing here yet"; the rule that overrides the base now
says so in one place rather than layering a second `.addButton` block.

An empty capacity shows the adapter's route-level fallback as its placeholder,
so a blank field reads as "sized by the route" rather than as a model with no
capacity. It is a hint, not a mirror: the field counts K as 1000 while the
fallback is 262144, and a deployment may override it.

The picker's description says what the list is without promising an edit the
rows themselves already offer.
2026-08-06 15:18:18 +08:00
Yichen Jiang dc7510a902 fix(ui-models): close the media block that swallowed the fetch dialog's styles
The `@media (prefers-reduced-motion: reduce)` block never closed, so every
rule after it — the whole fetch dialog, its candidate list, and the rows
inside — applied only to viewers whose system asks for reduced motion. The
sheet still parsed and the classes still attached, so the list painted with
the browser's own bullets, indentation, and inline label, and the reported
model id ran straight into the capacity beside it.

Report the id alone: it is the string adoption writes, and the capacities the
endpoint disclosed are adopted with it and editable in the row that appears.
The candidate row's remaining font name is the one the theme declares.

The styles gate now checks that the sheet's braces balance, and reads every
theme sheet rather than the platform tokens alone, so a name declared in a
sibling is not called undeclared and a `--dsh-` typo cannot pass as a token.
2026-08-06 15:18:17 +08:00
Yichen Jiang 44484ec5f6 feat(web): declare a provider and its models from the Models page
The Models page could name a provider's key and little else. Adding an
OpenAI-compatible gateway meant opening $DSH_HOME/settings.yaml and
knowing the profile shape; correcting a stale context window meant the
same. This layer puts both on the page: a card that declares a route
pi-ai does not ship — id, endpoint, protocol, key, models — and a model
list on the pi-ai editor that can ask the provider what it serves and
adopt the answer.

It follows the DeepSeek catalog editor that landed in #1050 rather than
inventing a second look for the same job. Both editors now share the
section shell and heading, the danger-tinted delete, the add-model
button, the empty state, the per-row validator that names a bad row by
its position, and one K/M capacity vocabulary — 256K and 1M are read and
spelled back, while settings.yaml still stores plain token counts. The
row type is structurally open like that editor's, so a profile field
this card does not edit survives an edit here.

Three of that editor's decisions replaced weaker ones this branch had
made. Inheritance now reads the composition base rather than the
effective value, which would echo an override back the moment a reset
dropped it. Validation names the offending row instead of stating a
blanket problem. And emptying the list is no longer conflated with
handing the catalog back to the adapter — those are separate acts, with
separate affordances.

The create write carries the revision the card opened at, so a route
another tab declared meanwhile is a conflict rather than a silent
overwrite of its profile.
2026-08-06 15:18:17 +08:00
Yichen Jiang d97e150845 Merge branch 'claude/web-llm-pi-ai-config-385e24' into claude/pi-ai-model-discovery 2026-08-06 15:07:02 +08:00
Yichen Jiang 2abd06f21f Merge remote-tracking branch 'origin/master' into claude/web-llm-pi-ai-config-385e24 2026-08-06 15:01:19 +08:00
Yichen Jiang 2d71c928e6 Merge branch 'claude/web-llm-pi-ai-config-385e24' into claude/pi-ai-model-discovery 2026-08-06 14:26:48 +08:00
Yichen Jiang c7dbe4c6d9 Merge remote-tracking branch 'origin/master' into claude/web-llm-pi-ai-config-385e24 2026-08-06 14:24:58 +08:00
Yichen Jiang 16ad5f3f86 Merge pull request #1732 from deepseek-harness/worktree/fix-1680-cjk-markdown
fix(web): render CJK emphasis and link code URLs
2026-08-06 13:15:50 +08:00
Yichen Jiang 099b903ac6 fix(web): preserve provider credential retry checkpoint 2026-08-06 13:15:09 +08:00
Yichen Jiang 42b9ee7a53 Merge branch 'claude/web-llm-pi-ai-config-385e24' into claude/pi-ai-model-discovery
# Conflicts:
#	docs/core-data-structures/core.i18n.yaml
#	docs/event-producer-consumer.md
2026-08-06 13:05:15 +08:00
Yichen Jiang 98b894f8eb Merge remote-tracking branch 'origin/master' into claude/web-llm-pi-ai-config-385e24
# Conflicts:
#	docs/event-producer-consumer.md
2026-08-06 13:03:40 +08:00
Yichen Jiang f330b6ae79 test(web): refresh targeted provider action golden 2026-08-06 12:51:18 +08:00
Yichen Jiang af652c949f fix(web): recover provider credential lifecycle 2026-08-06 12:09:14 +08:00
Yichen Jiang 178eaf3973 fix(web): link inline-code URLs 2026-08-06 11:41:46 +08:00
Yichen Jiang 1b5f06ed0b fix(web): render CJK-adjacent strong emphasis 2026-08-06 11:41:46 +08:00
Yichen Jiang 4bd66d8e7c Merge branch 'claude/web-llm-pi-ai-config-385e24' into claude/pi-ai-model-discovery
# Conflicts:
#	docs/cordis-catalog/events.md
#	docs/core-data-structures/core.i18n.yaml
#	docs/event-producer-consumer.md
#	packages/host/apiproxy/README.i18n.yaml
#	packages/llm/llm/README.i18n.yaml
2026-08-06 10:50:20 +08:00
Yichen Jiang 17c86bbb25 Merge remote-tracking branch 'origin/master' into claude/web-llm-pi-ai-config-385e24
# Conflicts:
#	packages/llm/llm-pi-ai/tests/adapter.spec.ts
#	packages/llm/llm/README.i18n.yaml
2026-08-06 10:47:36 +08:00
Yichen Jiang 66c2cb81d3 fix(llm): let an interrogation use the credential its route already stored
A configuration surface never holds a stored secret — it edits a redacted
descriptor — so once a key is saved, the draft it sends carries the route and
the endpoint and no credential at all. The interrogation went out
unauthenticated and the endpoint's 401 came back as "check the API key",
pointing at the one thing that was fine.

A named route now supplies its own credential, resolved exactly as a request
to it would be. A key typed into the form still wins: it is the one under
test, and may be the replacement for the stored one that is failing.

Resolution is a callback the probe invokes past the catalog short-circuit and
the protocol check, so a route answered from the installed registry costs no
credential lookup — and cannot fail over a credential the question never
needed.
2026-08-05 20:55:39 +08:00
Yichen Jiang b2d0e8972f docs(host): re-record the pairings master's wording moved
Master polished two Chinese sentences this branch also edits — an
expectation is now 陈旧 rather than 过期, and the package-root sentence
spells out 包(package). Taking master's wording alongside this branch's
own additions leaves the recorded pair fingerprints stale, so they are
re-recorded against the merged text.
2026-08-05 19:53:10 +08:00
Yichen Jiang 2dd4b8e78d fix(host): pin model discovery to loopback and drop its unread wire field
llm.discoverModels was reachable from any declared trusted host. The
method takes a caller-supplied baseURL and makes the host issue a GET to
it, then reports the status or the parsed body — so on a LAN deployment
an anonymous caller had a probe for whatever the host can reach and the
browser cannot, plus a path that carries a draft credential. The
PRIVILEGED_METHODS doc already states the rule this broke: trustedHosts
is a DNS-rebinding fence, not authentication, so the configuration plane
stays loopback-same-origin. It is in that set now, asserted both against
the hand-built fence and over real HTTP beside the catalog reads that
deliberately stay reachable.

supportsDiscovery and listModelDiscoveryNamespaces are gone. The field
was required on the wire and read by nobody: its own contract said a
surface should offer the action "instead of naming an adapter family it
would have to hardcode", while the surface hardcodes llm-pi-ai in two
places and gates the button on whether there is anything to probe. Its
shape did not fit the second caller either — the create card has no row
to read a per-row field from. Keeping a required field alive for a
consumer that may never arrive costs every producer and fixture a value
nobody consults, which is exactly how the fixtures drifted. The registry
that fed it had no other production consumer, so registration and
disposal are now observed through the offer itself.

The Agent Note claimed the key is never logged, which the wire schema
beside it already contradicts, and predated both the provider field and
the catalog-answer path. The two new public types pointed at core.md
without a type-equiv block or manifest entry, so the generated service
catalog named documentation that did not exist.
2026-08-05 19:51:11 +08:00
Yichen Jiang ffd2f188f2 fix(llm): answer a catalog route's models from pi-ai's own registry
Clicking "fetch available models" on a built-in provider went to the
network. That is the wrong source: pi-ai's registry is the authoritative
list for its own providers, and it carries the context windows and output
caps a `GET /models` listing does not disclose. Asking api.deepseek.com
what DeepSeek serves is both slower and worse, and against an endpoint
that answers a different shape it failed outright.

Interrogation is still keyed by settings namespace — the provider being
added has no route — but the request may now name the route it is
editing. An adapter that already describes that route answers from what
it knows, needs no endpoint at all, and never touches the network; only a
route the catalog does not describe reaches the wire, and one naming no
endpoint is told to set one or enter its models by hand.

`ConfigurableProviderView` gained `supportsDiscovery` so a surface offers
the action where a namespace can answer instead of hardcoding an adapter
family.

Three narrower corrections ride along. Discovery no longer claims Azure
or Codex: Azure authenticates with an `api-key` header and an
`api-version` query despite its OpenAI lineage, and Codex uses OAuth, so
both reported an authentication failure as a provider with no models.
Cancellation during the body read escaped as the raw abort reason rather
than a coded ABORTED. And the schema comment claiming the probe key is
never logged overstated it: the host neither stores nor returns it, but
it rides the client's outgoing envelope like every other secret-bearing
payload, and redacting that tap is a configuration-plane-wide change.
2026-08-05 19:51:02 +08:00
Yichen Jiang ecee93ec26 feat(llm): interrogate a draft provider endpoint for its models
Once a pi-ai route became a declaration rather than a catalog lookup,
adding an OpenAI-compatible gateway meant knowing its model ids up
front. Most such endpoints publish that list at `GET /models`, but no
seam operation could ask: every one is keyed by a registered provider
route, and the provider being added has no route, no stored profile,
and no stored credential — the endpoint and key are values in a form.

Interrogation is therefore keyed by settings namespace, which a
configuration surface already holds from the configurable-provider
directory. `registerModelDiscovery` offers it per namespace,
`discoverModels` asks, and the request carries the draft itself. The
reply is candidates, not a catalog: every field but the id is optional
because most listings disclose nothing else, and adopting one is a
settings write like any other. Nothing here reads or writes settings or
credentials, so `settings.yaml` still decides what a route serves.

`llm.discoverModels` carries the same draft over the wire. Its apiKey is
the third and last payload a secret may ride, and it is never stored,
logged, or echoed; every refusal folds into `model-discovery-failed`,
naming the endpoint asked but never the credential offered.

The pi-ai side is a plain GET for OpenAI-compatible protocols only —
their listing shape is the one gateways, self-hosted servers, and the
official endpoints agree on. Others say so, sending the user to
hand-entry rather than reporting a guessed shape as an empty provider.
The reply is read under a four-megabyte ceiling held on the bytes
actually received, because the endpoint is a URL the user typed.
2026-08-05 19:50:11 +08:00
Yichen Jiang 9948a37cbc docs(llm): follow master's README hierarchy for the pi-ai adapter
A documentation rescan on master moved this README's Testing section out
of the package. This branch was still editing that section, so the
rebase asked which structure wins; master's does, and the branch keeps
only the Catalog resolution section its own change adds.

Re-records the pair fingerprint against the merged text.
2026-08-05 18:56:41 +08:00
Yichen Jiang 5d3ccdc528 test(llm): cover both names providerInfo can report for a route 2026-08-05 18:54:47 +08:00
Yichen Jiang 73fce861e5 fix(llm): let a catalog route keep the auth its provider actually declares
pi-ai resolves a request's apiKey override only through a provider that
declares an api-key method: resolveProviderAuth short-circuits to that
method when the override is present, and otherwise falls through to the
credential store and then to ambient discovery. A provider with no
api-key method at all therefore resolves to nothing, and the request
fails with "Provider is not configured" before any network I/O.

Two routes hit that. openai-codex ships OAuth alone, so moving off the
/compat dispatch broke a profile that names a key for it — the old path
handed the token straight to the provider. And a catalog route naming an
api was being rebuilt with the harness's own auth, so `openai: {api:
openai-completions}` stopped reading OPENAI_API_KEY, contradicting the
documented promise that omitting a credential keeps provider-native
discovery.

Auth is now one decision for both constructions. A catalog route keeps
its installed provider's auth, through an api override too: which
environment a provider reads belongs to the provider, not to the wire
format its models speak. A catalog provider with no api-key method gets
the harness method beside its own, but only when the profile names a
credential — a keyless codex profile keeps the honest refusal, since
this adapter holds no OAuth store to resolve through.

Materialization now spreads the installed entry instead of enumerating
the result, so a Model field this package does not model survives a
pi-ai upgrade; headers went missing from an nvidia route exactly that
way once already. providerInfo reports the configured displayName, which
also joins the registration facts so a rename re-registers rather than
leaving the old label in every selector. A refused registration swap
gets its own diagnostic naming the route, matching the directory swap
beside it.

The README documented endpoint interrogation this layer does not
implement, and still described unknown providers as kept-last-good after
they became legal declarations refused at the write point. The Agent
Note claimed per-model reasoning configurability the schema never had,
required capacities the route now defaults, and stated an apiKey
override that short-circuits unconditionally.
2026-08-05 18:54:47 +08:00
Yichen Jiang 236b1f6d97 docs(llm): record that non-reasoning catalog models lose the lone off level
The adapter omits the seam's reasoning field whenever a model carries no
reasoning metadata, which is the model's own property and says nothing
about where the model came from. Both the JSDoc and the Agent Note read
as though only hand-declared models were meant, so a reader would infer
that the 251 installed-catalog models pi-ai marks as non-reasoning still
offer their single off level. They do not, and that is the point: a
picker holding only off misrepresents a provider that thinks by default,
because off dispatches the same bytes as naming no effort at all.

Behavior is unchanged; only the prose that describes it was narrower
than the contract. adapter.spec.ts already pins the catalog case through
openai/gpt-4.1 and catalog.spec.ts pins the hand-declared one.
2026-08-05 18:54:23 +08:00
Yichen Jiang f376ee23d1 fix(llm): size unknown models and refuse a section that cannot be served
Three defects surfaced while driving the Models page.

A hand-declared model needed an explicit contextWindow and maxTokens,
but a provider listing usually returns ids and nothing else — so the
page happily wrote a profile the adapter then rejected, which took the
whole namespace down silently. Capacities now fall back to the route's
`defaultContextWindow` (262,144) and `defaultMaxTokens` (32,768). Both
are guesses by construction, which is why they are route fields a
deployment corrects once rather than constants buried in the adapter;
the fallback sizes the model and never becomes a per-request cap.

That silent failure was the second defect. A schema-valid profile the
adapter could not serve was stored and only rejected later, disabling
every route in the namespace with nothing said. `dsh-settings` gains an
optional `validate` on registration — a check for what a schema cannot
express — and `llm-pi-ai` refuses an unserviceable section at the write
that produced it. A stored section that fails keeps the namespace's last
good value, as a schema failure already did, so an externally edited
document still cannot strand the owner. The plugin's own last-good
fallback goes with it: nothing reaching it can fail any more.

Third, a model with no reasoning metadata advertised the single level
`off`, which pi-ai translates to *omitting* the reasoning option — the
same request naming no effort produces. Selecting it disabled nothing,
so a provider whose default is to think kept thinking with `off` shown
as selected. Such a model now reports no reasoning capability at all,
which is the seam's way of saying the control is unavailable, and the
per-model `reasoning` flag is gone: without a thinkingLevelMap to spell
levels it could only invent them.

The protocol table narrows to the three a hand-declared route reaches
today, most-reached first so a surface offering a choice defaults to the
one gateways actually speak.
2026-08-05 18:54:23 +08:00
Yichen Jiang 4c80cab108 fix(llm): capture an immutable snapshot per pi-ai operation
Review found four defects in the declared-provider work.

`PiAiAdapter` reused one `Models` collection and mutated it whenever the
configuration changed. `Models.streamSimple()` resolves its provider
lazily — when the stream is first consumed, which is after the adapter
awaits the route's credential — so a configuration change landing in
that window let an in-flight request finish under a configuration it
never resolved against, or fail on a provider that no longer existed.
Each resolution now produces an immutable snapshot and every operation
captures one before its first await, which is what makes the seam's
per-step freeze (`llm.prepareCall()`) hold end to end: switching models
mid-reply takes effect on the next step, never inside the one in flight.

`defaultMaxTokens` was materialized from the catalog's `Model.maxTokens`.
The two answer different questions: pi-ai requires that field as the
model's output capability, while the seam's is a cap the deployment
chose to send on requests naming none, so every request had started
carrying a number nobody picked. Only an explicitly configured cap
reaches the seam now.

The configurable-provider directory was refreshed by disposing its
registration and making a new one. A candidate set the registry refuses
— a profile keyed `deepseek-official`, which llm-deepseek declares —
left the whole directory withdrawn and the Models page empty, silently,
because the settings callback contains the failure. The seam's
registration handle now carries `replace()` with the same
validate-first atomicity `registerAdapter` has.

The protocol table offered every pi-ai streaming API, including four
whose authentication a profile cannot express: Bedrock signs with SigV4
over AWS credentials and a region, Vertex needs a project, a location,
and ADC, Azure needs provider environment plus an api-version, and Codex
uses OAuth. Offering them handed back routes that cannot authenticate.
Catalog routes still reach them through their own provider.
2026-08-05 18:54:23 +08:00
Yichen Jiang d6126c25f2 feat(llm): declare pi-ai providers instead of looking them up
A pi-ai route had to name an installed catalog provider, served that
catalog's models verbatim, and could override only the endpoint. An
OpenAI-compatible gateway, a self-hosted server, or a model newer than
the pinned pi-ai release was therefore unreachable, and a stale context
window could not be corrected without upgrading the package.

A route is now a declaration whose defaults come from the installed
catalog. `catalog.ts` merges that catalog under the profile's own model
entries, `provider.ts` builds the pi-ai Provider (reusing the catalog
provider when the route keeps its protocol, so implementations this
package cannot reconstruct keep working), and the adapter serves every
operation from one `createModels()` collection. That also retires the
`@earendil-works/pi-ai/compat` import, which pi-ai documents as a
temporary entry point it deletes with its ModelManager migration.

Credentials stay on the harness seam: the resolved key rides the request
as pi-ai's highest-priority auth override, so `Models` holds no
credential store and a named-but-missing reference still fails loud
instead of falling back to an unrelated ambient key.

A model's configured maxTokens now reaches the seam as defaultMaxTokens.
2026-08-05 18:54:23 +08:00
Yichen Jiang 96c44c83d4 test(web): stabilize and refresh the aria goldens for the speed readings
The TTFT and tok/s readings divide by measured wall time, so they are not
reproducible: the same replayed scenario yielded 69 and 70 tok/s on
consecutive local runs, and a 3 ms replayed stream reads 26333 tok/s. Baking
those into committed goldens made the Web lane flaky by construction, and the
goldens for the readings themselves were never refreshed.

Three fixes, then a refresh:

The footer's decorative dots are `aria-hidden`, so the readings concatenated
into one accessible string — `Ran for 13sTTFT 0.2s12 tok/s`. That is a real
defect on its own (a reader hears one run-on instead of three facts) and it
also denied `{{duration}}` the word boundary it matches on, so even the
previously-stable `Ran for` duration started leaking raw. The separators now
carry flanking spaces.

`normalizeAria` gains `{{throughput}}` beside `{{duration}}`, and its duration
alternation accepts the stats line's compact `2m42s` as well as the
message-chrome template's `2m 42s` — the compact form had no pattern at all,
which is why `LLM 382m39s` survived the first refresh.

Refreshed 17 goldens. They also record that the stats line's `LLM` group now
renders at all: it folds assistant `timing`, which the live transcript adapter
only began attaching in this branch, so the group was previously dead in Chat.

Verified by running the lane in replay three times after the refresh: 41/41
files green each time, goldens untouched. Before this change two consecutive
runs disagreed on both the values and the failure count.
2026-08-05 17:55:49 +08:00
Yichen Jiang 8745365547 docs(i18n): re-record the pairs the master merge left stale
The merge staged the `--ours` consistency records before the merged prose was
re-recorded, so the committed hashes still describe the pre-merge content and
`verify-translation-pairing` rejects the pair. Re-record all three against
what actually merged.
2026-08-05 17:55:36 +08:00
Yichen Jiang e43b08d7b8 Merge remote-tracking branch 'origin/master' into feat/add-session-data-preview
# Conflicts:
#	docs/core-data-structures/session.i18n.yaml
#	packages/client/ui-conversation/README.i18n.yaml
#	packages/client/ui-conversation/README.md
#	packages/llm/token-meter/README.i18n.yaml
#	packages/llm/token-meter/README.md
#	packages/llm/token-meter/README.zh.md
#	packages/llm/token-meter/src/projection.ts
2026-08-05 17:11:16 +08:00
Yichen Jiang 038699bcb4 fix(web): let the context meter see a compaction
The composer ring, percentage, and `~used / capacity` header read
`contextPressure.pressureTokens`, which moves only when a request reports
usage. Compaction reports none — compact-basic summarizes through a direct
`ctx.llm.stream()` call and appends only its own `compact/*` records plus the
replacement `user/message` — so the meter was frozen across the one action
taken to change it. Driving a real `compactNow` through the agent loop:

    BEFORE compact:  ring=4%  header=~4227/100000  rows=[18, 0, 4365]
    AFTER  compact:  ring=4%  header=~4227/100000  rows=[18, 0,  286]

The composition rows fell 93%; the ring did not move, and would not until an
entire further turn completed. The panel then contradicted itself by more than
an order of magnitude at exactly the moment a reader opens it.

`contextPressure` now also publishes `projectedTokens`: the provider sample
plus the heuristic repricing of everything the surface gained or lost since
that sample, clamped at zero, folded through the shared `surface-fold.ts`. The
sample is stamped before the same event joins the surface, so an
`assistant/message` anchors against the surface its own request carried. Only
the delta is estimated, so the figure stays provider-anchored — the estimator's
CJK and JSON-schema underpricing stays out of the occupancy number — while
reacting the moment content lands or a span is shadowed. Same run after:

    BEFORE compact:  ring=4%  header=~4323/100000  (pressure=4227, projected=4323)
    AFTER  compact:  ring=0%  header=~ 244/100000  (pressure=4227, projected= 244)

`contextOccupancy` prefers the projected figure and falls back to the bare
sample, so a projection restored from a pre-field checkpoint degrades to the
old behavior rather than disappearing. `stateVersion` moves to 3.
2026-08-05 17:00:48 +08:00
Yichen Jiang e62cbe12e4 refactor(token-meter): fold the surface once for the meter and the projection
`contextBreakdown.messageTokens` and `measure().surfaceTokens` answer the same
question in the same heuristic vocabulary, and the panel's composition rows are
only honest while they agree. Each owner carried its own copy of the positional
fold — same pricing, same `{seq, tokens}` node list, same replace-range lookup
and guard, differing only in mutable versus immutable application — so an edit
to either one would have moved the panel away from `measure()` with both sides
still green. The duplication gate caught the shared 62 tokens.

`src/surface-fold.ts` now owns `foldSurfaceTokens`: total, allocation-fresh,
returning the event's price, the next surface, and the signed total delta. The
service assigns that result where it used to prepare a commit closure, which
keeps its validate-before-mutate replay transaction intact — the fold throws
before any state is touched, so a malformed event still fails identically on
every retry. `_prepareSurfaceMutation` and `_estimateSurfaceEvent` go away with
it, and the projection's apply drops to one call.

Covers the identity with a session that appends and then compacts, asserting
the projection figure equals the service surface at each boundary; the test
fails when either side of the fold is perturbed.
2026-08-05 16:29:51 +08:00
Yichen Jiang 88a57f12b0 Merge pull request #1530 from deepseek-harness/worktree/open-settings-config-file
feat(web): open the local settings file
2026-08-05 16:10:56 +08:00
Yichen Jiang 46562b2c30 fix(web): localize the context panel headline order and drop empty bar parts
The panel header concatenated a `45%` span with a `context.used` fragment, so
Chinese rendered "45% 上下文已用" against the ring's own "上下文已用 45%". The
header now renders the one localized `context.aria` sentence split around its
`{percent}` slot: each locale owns the reading's position while the reading
keeps its primary tone, and the side a locale leaves empty collapses through
`.headline:empty` instead of spending a header gap.

The bar mapped every composition row to a segment unconditionally, and
`.segment`'s 2px min-width kept each one visible, so a 0% occupancy panel
painted an ~8px filled bar over an empty context. Parts are now computed with
their widths and zero-width parts are filtered out, which also collapses the
plain and segmented branches into one map.
2026-08-05 15:59:24 +08:00
Yichen Jiang d2aafc2a33 Merge remote-tracking branch 'origin/master' into feat/add-session-data-preview
# Conflicts:
#	docs/cordis-catalog/services.md
#	docs/core-data-structures/session.i18n.yaml
#	packages/client/ui-conversation/README.i18n.yaml
#	packages/llm/token-meter/README.i18n.yaml
2026-08-05 15:51:34 +08:00
Yichen Jiang 6c6e29a639 test(web): include config button in onboarding snapshot 2026-08-05 15:45:08 +08:00
Yichen Jiang cc6e2382cb Merge remote-tracking branch 'origin/master' into worktree/open-settings-config-file
# Conflicts:
#	.agents/notes/implemented/architecture/2026-07-30-web-config-plane.i18n.yaml
#	.agents/notes/implemented/architecture/2026-07-30-web-config-plane.md
#	.agents/notes/implemented/architecture/2026-07-30-web-config-plane.zh.md
#	packages/client/ui-settings/README.i18n.yaml
#	packages/host/apiproxy/README.i18n.yaml
#	packages/host/apiproxy/README.zh.md
2026-08-05 15:34:51 +08:00
Yichen Jiang e9b377e9c5 Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models
# Conflicts:
#	.agents/notes/implemented/architecture/2026-07-30-web-config-plane.i18n.yaml
#	.agents/notes/implemented/feature/2026-07-24-web-session-model-selector.i18n.yaml
#	packages/client/ui-model/README.i18n.yaml
#	packages/client/ui-model/README.zh.md
#	packages/client/ui-models/README.i18n.yaml
#	packages/client/ui-models/README.zh.md
#	packages/host/apiproxy/README.i18n.yaml
#	packages/llm/llm-deepseek/README.i18n.yaml
2026-08-05 12:50:58 +08:00
Yichen Jiang 086656afae Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models
# Conflicts:
#	packages/client/ui-models/src/client/ModelsSection.module.css
#	packages/client/ui-models/src/client/ModelsSection.tsx
2026-08-05 11:46:23 +08:00
Yichen Jiang a2497a0dc1 Merge remote-tracking branch 'origin/master' into worktree/open-settings-config-file
# Conflicts:
#	packages/client/connection/README.i18n.yaml
#	packages/client/connection/README.md
#	packages/client/connection/README.zh.md
2026-08-05 11:05:52 +08:00
Yichen Jiang 4f717f2da7 fix(web): address settings document review 2026-08-04 17:31:36 +08:00
Yichen Jiang e31b7221e7 feat(web): open the local settings file 2026-08-04 16:59:37 +08:00
Yichen Jiang 17b480de51 feat(web): align the model catalog with the pi-ai provider form
Both editors live in `ui-models` and list the same thing, so they now share
one row shape rather than diverging when #1368 lands: a bordered entry per
model, id and display name on the row, and the capacities behind the row's
own disclosure. The context window is joined there by the per-model output
cap the adapter just gained; both read a decimal K/M suffix.

The shared class names carry this file's token spellings, not that branch's.
`--dsw-alias-border-subtle`, `--dsw-alias-text-tertiary`, and
`--dsw-alias-text-primary` are undeclared, so they resolve to the light-mode
literals in their fallback slots — the defect this section was moved off. A
styles test now rejects any `--dsw-*` name the token sheet does not declare,
so the next editor to name one fails instead of shipping a light-only
surface.

The keystroke buffer is now per capacity field rather than per row, since a
row holds two of them.
2026-08-04 14:53:17 +08:00
Yichen Jiang 1d2ea70e9b feat(llm-deepseek): per-model output cap in the advisory catalog
A catalog entry may now carry its own `maxTokens`, matching the shape the
pi-ai adapter already exposes. Exact-model resolution prefers it over the
profile value, so capping one model no longer means capping the route.

Defaults are unchanged: an entry without a cap, and any unlisted
pass-through id, still resolve to the profile `maxTokens` (256,000), and the
shipped catalog keeps its context windows.
2026-08-04 14:53:02 +08:00
Yichen Jiang a4aace17b5 Merge remote-tracking branch 'origin/worktree/custom-deepseek-models' into worktree/custom-deepseek-models 2026-08-04 14:08:22 +08:00
Yichen Jiang adecbb71f7 Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models 2026-08-04 14:01:28 +08:00
Yichen Jiang b52b9422d4 Merge branch 'master' into worktree/custom-deepseek-models 2026-08-04 11:13:45 +08:00
Yichen Jiang 5e05838365 testing(web): re-record the markdown image golden for the model trigger
The scenario arrived with master and recorded its trigger while the Host
still synthesized an unlisted current model. This branch stopped doing that,
so the trigger reads `Select model` until the user picks an advertised one.
2026-08-03 19:10:55 +08:00
Yichen Jiang edf2cece37 Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models
# Conflicts:
#	packages/host/apiproxy/README.i18n.yaml
2026-08-03 19:02:29 +08:00
Yichen Jiang a31faf24a1 Merge pull request #927 from deepseek-harness/worktree/fix-multi-select-custom-answer
fix(user-interaction): preserve selected options with custom answers
2026-08-03 18:15:00 +08:00
Yichen Jiang 3009bc5312 Merge pull request #1003 from deepseek-harness/worktree/web-remote-markdown-images
feat(web): render remote Markdown images
2026-08-03 16:41:00 +08:00
Yichen Jiang ebe5bcea29 Merge remote-tracking branch 'origin/worktree/fix-multi-select-custom-answer' into worktree/fix-multi-select-custom-answer 2026-08-03 16:28:08 +08:00
Yichen Jiang 85b9914934 Merge remote-tracking branch 'origin/master' into worktree/fix-multi-select-custom-answer
# Conflicts:
#	packages/ui/tui/README.i18n.yaml
2026-08-03 16:26:55 +08:00
Yichen Jiang 886c39d302 test(web): refresh Markdown image snapshot 2026-08-03 16:21:02 +08:00
Yichen Jiang e0b70508a1 Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models
# Conflicts:
#	apps/web/tests/snapshots/message-actions/ui.expected.md
#	apps/web/tests/snapshots/seeded-history/command-row.expected.md
#	apps/web/tests/snapshots/seeded-history/ui.expected.md
#	packages/client/ui-model/README.i18n.yaml
#	packages/client/ui-model/README.md
#	packages/client/ui-model/README.zh.md
#	packages/client/ui-model/tests/model-select.spec.tsx
#	packages/host/apiproxy/README.i18n.yaml
2026-08-03 16:20:55 +08:00
Yichen Jiang 61cba0cda2 Merge remote-tracking branch 'origin/master' into worktree/web-remote-markdown-images
# Conflicts:
#	packages/client/ui-primitives/README.i18n.yaml
#	packages/client/ui-primitives/README.md
#	packages/client/ui-primitives/README.zh.md
#	packages/client/ui-primitives/src/markdown/MarkdownText.tsx
#	packages/client/ui-primitives/tests/markdown.spec.tsx
2026-08-03 16:09:51 +08:00
Yichen Jiang 109d5b5e1c Merge remote-tracking branch 'origin/master' into worktree/fix-multi-select-custom-answer
# Conflicts:
#	apps/web/tests/snapshots/question-composer/answered.expected.md
#	apps/web/tests/snapshots/question-composer/session.jsonl
#	docs/core-data-structures/user-interaction.i18n.yaml
#	packages/client/ui-question/README.i18n.yaml
#	packages/host/apiproxy/README.i18n.yaml
#	packages/host/apiproxy/README.md
#	packages/host/apiproxy/README.zh.md
#	packages/ui/tui/README.i18n.yaml
#	packages/ui/user-interaction/README.i18n.yaml
2026-08-03 16:09:17 +08:00
Yichen Jiang 00e08d1ec0 testing(web): open the settings dialog in the model catalog test
The test inherited an open dialog with the DeepSeek editor already expanded
from the credential test above it. That test now reloads the page to
exercise the welcome step, so nothing carries the dialog across and the
catalog test timed out looking for a fold that was no longer on screen.

The review flagged this coupling as two links deep; the merge proved it.
2026-07-31 15:19:13 +08:00
Yichen Jiang 02cd5bf733 Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models
# Conflicts:
#	apps/web/tests/onboarding-deepseek-config.e2e.ts
#	packages/client/ui-models/README.i18n.yaml
#	packages/host/apiproxy/README.i18n.yaml
2026-07-31 15:12:41 +08:00
Yichen Jiang f12d4f698e doc(web): own the Web/TUI split on an unlisted current model
The TUI still renders the unlisted current model as its own row and marks it
current; Web shows the unset trigger label and asks for a replacement. The
note recorded the Host decision but not that the two surfaces present it
differently, which reads as a missed migration rather than a choice.

Also records why `trigger.selectAria` and `trigger.fallback` hold identical
strings, so a future coalescing cleanup does not merge them.
2026-07-31 15:11:11 +08:00
Yichen Jiang d6af60f5ec testing(web): pin the row card against the editor it expands into
The previous assertion pinned the literal `bg-layer-3` fill that was just
reverted. What matters is the relationship it broke: `bg-layer-3` and
`bg-module-platform` both resolve to neutral-bluish-800 under the dark
theme, so filling the row with either erases the nested editor's boundary.
2026-07-31 15:11:02 +08:00
Yichen Jiang e42e1c3ed7 fix(web): keep every row's unreadable context-window text
The typed text was held in one active buffer, so editing a second row
displaced the first — which then fell back to rendering its stored NaN as
the literal `NaN`, losing the text the row was supposed to let the user
correct. Reset dropped the override but kept the buffer, leaving typed text
over an inherited row that no settings layer stores, and an unreadable
buffer never settles, so it stayed there.

There is now one entry per row. `remove` re-keys around the dropped row so
text travels with its own row, and reset clears them all because the rows
they annotated are gone.

Model ids are compared trimmed as well: `"   "` passed the required check,
and `"model "` slipped past the duplicate check against its own twin before
silently failing to match at the adapter. A pasted id settles on blur rather
than being trimmed per keystroke, which would stop an interior space.
2026-07-31 15:10:59 +08:00
Yichen Jiang 6056062f8f Revert "fix(web): fill the provider row card"
This reverts commit 5f96d085f46bcb712beab248bab6c1f5cd030689.
2026-07-31 14:57:15 +08:00
Yichen Jiang 267fbdbf7f Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models
# Conflicts:
#	packages/host/apiproxy/README.i18n.yaml
2026-07-31 14:48:50 +08:00
Yichen Jiang ef996473a7 fix(web): restore the default model catalog without a reload
Restore defaults dropped the user's `models` override from the draft but the
rows kept showing it, so the catalog only looked restored after closing and
reopening the card.

The inherited rows were read from the namespace's effective value, which
still carries the stored override until the unset is applied — so dropping
the override echoed it straight back. They now come from the layer beneath
the user's: what the composition entry pinned, or else the schema default
that resolution would supply.
2026-07-31 14:44:15 +08:00
Yichen Jiang aa121e409f test(web): refresh running-turn snapshots 2026-07-31 14:38:27 +08:00
Yichen Jiang 1dbf70cb8a fix(web): fill the provider row card
Raises a configured provider's row onto layer 3 so it reads as a surface
rather than an outline. No-op under the light theme, where layer 3 and the
panel are both white.
2026-07-31 14:36:23 +08:00
Yichen Jiang 935578ed98 feat(web): accept K and M suffixes in the context window field
The catalog's context window is now a text field that reads a decimal K or M
suffix — 1M is 1000K, matching how model capacities are quoted — and stores
the plain token count, so settings.yaml and the adapter are unchanged.

A stored count reads back in the shortest form that round-trips: 1000000 as
1M, 256000 as 256K, and 131072 written out, because it is not a whole number
of thousands. The field holds the typed text while its row has focus, since
re-deriving it from the parsed count on every keystroke would rewrite 1000
to 1K mid-word; text that does not parse stays on screen so the save-time
rejection names a row the user can still see and correct.
2026-07-31 14:36:16 +08:00
Yichen Jiang 2701862bf1 Merge remote-tracking branch 'origin/master' into worktree/custom-deepseek-models
# Conflicts:
#	.agents/notes/implemented/architecture/2026-07-30-web-config-plane.i18n.yaml
#	.agents/notes/implemented/architecture/2026-07-30-web-config-plane.md
#	.agents/notes/implemented/architecture/2026-07-30-web-config-plane.zh.md
#	apps/web/tests/snapshots/message-actions/ui.expected.md
#	apps/web/tests/snapshots/seeded-history/ui.expected.md
#	packages/client/ui-models/README.i18n.yaml
#	packages/client/ui-models/README.md
#	packages/client/ui-models/README.zh.md
#	packages/client/ui-models/src/client/ModelsSection.module.css
#	packages/client/ui-models/src/client/ModelsSection.tsx
#	packages/client/ui-models/tests/components.spec.tsx
#	packages/host/apiproxy/README.i18n.yaml
2026-07-31 14:36:00 +08:00
Yichen Jiang 398878e52a Merge branch 'master' into worktree/core-web-minimal-profile 2026-07-31 14:34:37 +08:00
Yichen Jiang 936b8c4033 Merge remote-tracking branch 'origin/master' into worktree/core-web-minimal-profile 2026-07-31 14:29:12 +08:00
Yichen Jiang 3d438eb329 feat(cli): add minimal Core Web profile 2026-07-31 14:28:52 +08:00