chore: verify dev before promoting to staging (WALM-460) — DO NOT MERGE until verified - #874
Merged
Merged
Conversation
… protection Replace client-generated auth messages with a server-issued one-time challenge nonce. The server generates a random nonce stored in a new walletChallenges table with a 5-minute TTL. The client signs the nonce, and the server deletes it on use — replaying a captured signature fails because the challenge no longer exists.
- Use DELETE...RETURNING in connectWallet to eliminate TOCTOU race - Clean up expired challenges in getChallenge to prevent table bloat - Add Drizzle migration for walletChallenges table
Dev removed the wallet auth UI (Enoki overhaul) but kept the connectWallet tRPC route, so the replay-protection still has a target. - UI files (use-auth, auth-button-group) take dev; wallet-button removed - package.json takes dev (Enoki/noble deps) - deathbird migration renumbered to 0003 to coexist with 0002_add_enoki_fields Server-side changes preserved. Verified e2e: happy 200, replay 404, tampered 401.
An agent connecting to an unfamiliar account had no way to learn which
namespaces hold memories. Recall is similarity-ranked and needs a namespace
to search, so the only options were guessing names or falling back to
"default" — which undercuts cross-session memory portability.
`GET /v1/owners/{owner}/namespaces` already shipped in caff567. This
surfaces it on the TS SDK, metadata-only: no blob fetch, no decryption, and
no SEAL session is built or transmitted.
Owner resolution needed solving first. The read routes take the address in
the path and reject a mismatch against the caller's credentials, but
MemWalConfig carries only the delegate key and account id — so the SDK could
not build the path at all. `resolveOwner()` learns it from POST /api/stats,
which authenticates with the same delegate scheme, needs only a namespace,
is rate-limit weight 1, and returns the owner the server resolved. Memoised
per client, single-flight guarded like the compatibility probe.
Using a stats endpoint as a whoami is indirect. It was chosen over adding a
server-side `me` alias because the SDK ships to npm and self-hosters run
their own relayer: a path-level change would 403 against every relayer older
than it, needing a compatibility branch. The indirection is contained in one
private method, no public API mentions owner, and if a self-reference lands
later `resolveOwner()` is the only thing that changes.
Result types mirror the relayer wire shape (snake_case out), matching the
existing RecallResult.dropped_count convention rather than adding a mapping
layer. The query string is part of the signed path because the server
verifies against path_and_query, not path.
MemWalMock gains the same method so the drop-in double keeps parity,
aggregating seeded memories by namespace with deterministic timestamps
derived from insertion order.
Verified live against relayer.dev.memwal.ai, not only against stubs: owner
resolution, a plain signed GET, a signed GET carrying a query string, and a
cursor round-trip whose base64 cursor contains characters that percent-encode
— all 200. Signed-GET-with-query-string had no prior call site in the SDK, so
it was the one path that unit stubs could not have validated. That run also
showed the relayer returns snapshot_version 2, so the mock now returns 2
rather than disagreeing with the server on a wire-format version.
Deferred: Python SDK, the memwal_namespaces MCP tool, and the CLI command —
all three need the SDK published first, the same blocker as WALM-428.
Without this the method would ship in a release whose CHANGELOG never mentions it, riding on an unrelated changeset's version bump. CI does not enforce changesets, so nothing would have caught the omission.
DELETE /api/document destructured the first row from getDocumentsById and read .userId off it without checking that a row came back, so an id matching no document threw TypeError and the request failed with an empty 500 body. The GET handler in the same file already guarded for this, so the two disagreed on the same missing document. Add the same guard to DELETE and cover it with a route-level test that pins both the 404 and the parity with GET.
) auth.getSession and auth.logout took a sessionId as procedure input and acted on it, so anyone who learned an id could read that session's user and Sui address, or delete the session, without ever presenting the id as a credential. getSession is a query, so the id also travelled in the request URL on every page load, into access logs and browser history. Both now read the id from ctx.sessionId, which createContext fills from the x-session-id header the way protectedProcedure and the memory REST routes already do. The input is gone, so nothing a caller sends can steer either procedure. createContext now also ignores a header that is not a uuid. The zod input schema used to reject those before they reached a uuid column, where Postgres raises instead of simply not matching. On the client the getSession query key no longer varies per session, so useAuth resets the cached lookup whenever the stored session changes. A cached null left by an expired session would otherwise be read as the answer for the session that had just been established.
…url (#778) processSource downloaded a PDF file part with a bare fetch on the url from the request body, so an authenticated chat request could make the server call anything it can reach: a loopback service, an RFC1918 neighbour, or the metadata endpoint on 169.254.169.254. Route the download through fetchPublicUrl, which parses the url, requires http or https, and rejects loopback, private, link-local, and reserved targets. Hostnames are resolved and every address returned has to be public, so a name pointing into a blocked range is refused like the literal is. Redirects are followed by hand and re-checked at each hop, since fetch's own following would skip the check on the new target. The existing denylist in extractUrlsFromText only prefix-matches a few literals, so it misses userinfo, IPv6, 127.x outside 127.0.0.1, and names that resolve inward. That path reaches its target through Jina Reader rather than from this server, so it is left alone here.
…ession (#781) session.ts and proxy.ts encoded process.env.AUTH_SECRET as it came, while enoki-challenge.ts in the same folder already refused anything under 32 characters for that same variable. Two silent failures followed. A short placeholder is brute-forceable offline from one captured cookie. Worse, an unset variable encodes to zero bytes, and jose signs and verifies HS256 with an empty key without complaint, so a deployment missing the variable issued and accepted cookies signed with a key anyone can reproduce; forging a session for any user id needed no secret. .env.example ships AUTH_SECRET= empty and the runtime image does not inherit the builder placeholder, so that is the state a missed variable actually lands in. Add getAuthSecret and getAuthSecretKey, use them in session.ts and proxy.ts, and drop the local copy in enoki-challenge.ts so one rule covers every caller. Validation stays inside the functions: doing it at module scope would run during next build, where the Dockerfile supplies a 22-character placeholder, and would fail the image build rather than a request. proxy.ts reads the secret whether or not a cookie arrived, so a misconfigured deployment fails loudly instead of serving the login page as though the caller were merely signed out.
getToken() URL-decodes Authorization before its decode try/catch, so Bearer %% threw URIError and became HTTP 500 on proxy and guest login.
* feat(mcp): expose maxDistance on memwal_recall (WALM-457) * fix(mcp): do not report maxDistance misses as decrypt failures (WALM-457) * fix(mcp): address review on maxDistance recall cutoff (WALM-457) Undecrypted matches never had a distance computed, so "All matching memories were outside maxDistance." claimed more than the code knew when dropped_count was also non-zero: those hits may well have been inside the cutoff. Report both counts instead of letting the cutoff wording imply the namespace held nothing relevant. Also: - Reject a negative cutoff at both schema boundaries (zod .nonnegative(), JSON Schema minimum: 0). maxDistance: -1 silently dropped every hit and read as an empty namespace. - Document that the cutoff runs on the `limit` nearest matches, so a tight maxDistance returns fewer than `limit` rows. - Refresh the overview line from #849, which predated printing `distance` on the result line alongside `score`. - Fold the two 0.5-boundary tests into one that pins the whole partition, drop an assertion that only restated JS float arithmetic, and cover the cutoff/undecrypted overlap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mcp): note maxDistance review copy under 0.0.12 (WALM-457) --------- Co-authored-by: Harry Phan <phanhoangvinhhien@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…alculates-negative-relevance
…ry_search-tool-calculates-negative-relevance fix(openclaw): clamp memory_search relevance when cosine distance > 1 (WALM-441)
…enylist Mapped IPv4-in-IPv6 already hit the v4 table; deprecated compatible and SIIT forms did not. Also block IPv6 multicast /8 to match v4 multicast. Leave 6to4 and local-use NAT64 out of scope.
A broken Location after 3xx threw TypeError, which the chat route turned into offline:chat. Match assertPublicUrl's 400 instead.
…-ssrf fix(researcher): check the destination before fetching a source file url (WAKM-419)
Session ids remain bearer credentials. Logout ignores a body/input id; it does not make knowledge of the id insufficient.
Route-layer tests inject a parsed sessionId, so they miss a broken malformed-header check. Assert missing/non-uuid headers skip the walletSessions lookup.
fix(noter): bind getSession and logout to the caller's own session (WALM-420)
Keep a one-liner on the URIError branch: Auth.js URL-decodes Bearer outside its JWT try/catch.
…bearer-header-crashes-the-chatbot-proxy-and fix(chatbot): treat a malformed Bearer header as no session (WALM-416)
#867) * docs: document namespace 255-byte cap and restore truncated (WALM-486) * docs: note restore limit clamp 1-100 (WALM-486)
* ci: add production-safe SEAL cross-account synthetic (COMG-715) Read-only dry-run of seal_approve so a live identity from account A cannot authorize account B's SEAL key. Missing secrets skip (exit 0). Authorization success pages via SYNTHETIC_SEAL_CROSS_ACCOUNT_FAIL. * docs: use H2 body heading in SEAL synthetic page (COMG-715) * fix(ci): parse JSON-RPC MoveAbort abort code (COMG-715) JSON-RPC returns MoveAbort as a Display string whose MoveLocation contains commas, so [^,]+ never reached abort code 100 and a healthy ENoAccess denial was classified as misconfiguration. * test(ci): drop redundant MoveAbort parser coverage (COMG-715) Keep the nested-comma JSON-RPC fixture and distinct parser/CLI branches. Remove the duplicate hex string, the old-regex assertion, and the classifyNegative helper that restated extractAbortCode results.
* feat(server): expose writes ok|paused on /health (COMG-717)
* fix(server): reject writes when WRITES_PAUSED (COMG-717)
Make WRITES_PAUSED write-path admission, not a health-only signal.
remember, remember_bulk, remember_manual, and analyze return 503
{"error":"writes are paused"} while /health stays HTTP 200 with
status ok and writes paused. Reads (recall, restore) stay available.
* fix(server): drop stub writes-paused route tests (COMG-717)
The throwaway router never mounted production remember/analyze/health
handlers, so it could not catch a dropped gate. Keep the helper and
503 mapping unit test; do not grow a handler harness.
…9) (#811) * fix(auth): don't treat Sui RPC 429 as a revoked delegate key Cache-hit re-verify used to evict on any on-chain error, including gRPC 429. That produced empty 401s that the SDK mapped to memwal_login even when the key was still registered. Keep the cache (or return 503) when Sui is unavailable; only evict on a definitive miss. WALM-429 * fix(auth): fail closed on Sui RPC 429 instead of serving a stale cache Keep the cached mapping so a later verify can succeed, but return 503 with AUTH_UPSTREAM_UNAVAILABLE and Retry-After rather than authenticating a possibly-revoked key for up to 24h. SDKs only use the credential-outage copy when that header is present. * fix(auth): expose Retry-After and stop treating object-not-found as unavailable (WALM-429) Classify GetObject failures by gRPC status code so typo'd x-account-id and NOT_FOUND evict as sign-in failures, while 429/unavailable still 503 with the cache row kept. Expose Retry-After on CORS so browsers can read it. * fix(auth): keep field-parse errors as unavailable RpcError (WALM-429) NotFound is only for absent objects / unparseable ids / NOT_FOUND. A node that returns the object but omits json/fields stays RpcError so auth 503s and keeps the cache row instead of 401+evicting a live key. * chore(auth): slim WALM-429 review follow-up Fold GetObject status tests; drop tautological field-parse cases; trim comments.
Read the key before jwtVerify's try so a short secret cannot look like an invalid cookie or a failed Enoki ownership check.
…ret-length fix(researcher): validate AUTH_SECRET before signing or verifying a session
…ay-to-discover-which-namespaces-exist-recall-is
* fix(python-sdk): warn on plaintext remote server_url (WALM-452) Fixes #748 * fix(python-sdk): redact credentials in plaintext server_url warning (WALM-452) Log only scheme/host/port so HTTPX userinfo is not written to memwal logs, while keeping the original URL for transport.
Match GET's identical guard, which has no comment.
…-nonexistent-document-returns-http-500-instead fix(chatbot): return 404 when deleting a nonexistent document (WALM-422)
… retirements (WALM-354) (#822) * fix(chatbot): replace the model ids OpenRouter retired Four of the seven ids in chatModels no longer exist upstream, and OpenRouter answers a retired id with a 404 at request time, so picking one only ever produced "Oops, an error occurred!". Two of them were also hardcoded at a call site: getTitleModel used google/gemini-2.0-flash-001, which left every chat named "New chat", and getArtifactModel used anthropic/claude-3.5-haiku, which broke document creation for everyone rather than only whoever selected it. Point the title and artifact models at ids the picker also offers, the way apps/researcher does after the same incident there, and centralise the "-thinking" handling so the marker cannot reach OpenRouter unstripped. * test(chatbot): fail CI when a curated model id is retired upstream The unit tests can only catch an id drifting away from chatModels; they cannot tell that OpenRouter has withdrawn one, which needs a live call. That gap is how the retired title model reached production: nothing in the repo changed, so there was no PR run to go red. Check the ids against the public catalog endpoint, which needs no API key. It gets its own workflow because test.yml has no schedule trigger to hang a weekly run off, and it limits pull-request runs to changes that touch the list so other PRs take no network dependency on OpenRouter. * fix(chatbot): cap the output tokens each model call asks for No call site set maxOutputTokens, so every request was quoted against the model's entire output window. OpenRouter reserves that credit up front for Anthropic and Google models, which rejected the turn outright — "requires more credits, you requested up to 65535 tokens" — however short the reply would have been. OpenAI and DeepSeek do not reserve, which is why only some of the picker appeared broken. Cap the chat turn, the background title call and all six artifact calls. The reasoning cap stays above the thinking budget the chat route requests so the answer is not truncated before it reaches the reply. * fix(chatbot): resolve a stale chat-model cookie to the default Both chat pages handed cookie.value straight to <Chat>, which sends it as selectedChatModel, and /api/chat rejects an id that is no longer in chatModels. Anyone who had a since-retired model selected therefore had every send fail as a bad request, with nothing to explain it: the picker's fallback is display-only, so it rendered a valid model name while the request body still carried the stale id. Recovery meant reopening the picker, which the error text does not hint at. Coerce the cookie where it is read, matching resolveChatModelId in apps/researcher. * fix(chatbot): make the test mock answer the ids the app actually sends myProvider registers behaviours — chat-model, chat-model-reasoning — but getLanguageModel passed the catalog id through, so every request under PLAYWRIGHT threw NoSuchModelError: No such languageModel: openai/gpt-4o-mini. Every chat turn failed in CI, and the suite stayed green only because no test asserted that a reply arrives. getResponseForPrompt had the matching problem: it searched the whole serialised prompt, and the system prompt contains "hi" inside words like "this", so it always returned the greeting and the weather and default branches were unreachable. Match the newest user turn instead. * test(chatbot): cover the reply, title and reasoning paths end to end The suite exercised the composer and the picker but never asserted that a reply comes back, which is why it stayed green while every chat turn failed under PLAYWRIGHT. These six cover the paths the recent fixes touch: a streamed reply, a reply that survives a reload, intent reaching the model, the background title landing in the sidebar, reasoning streaming for a thinking model, and a retired chat-model cookie no longer rejecting every send. * test(chatbot): make the Playwright warm-up actually compile the routes The warm-up fetched with redirect: "manual", so the proxy's guest-session redirect ended it: every request returned 307 in under 60ms without rendering anything, and the cache stayed cold. Tests then paid the compile themselves — from a cleared .next, 4 to 14 of them failed on first-hit navigations of 40s to 1.3m against a 30s navigationTimeout and a 60s test timeout. CI never reported it because by the second of its two retries the routes were warm. Follow the redirects and keep their cookies, and warm the API and chat routes a message send needs rather than the landing page alone. The same cleared-.next run now warms in 84s and passes 28/28. * chore(chatbot): restore the trailing newline models.ts had Unrelated whitespace churn from the model-id fix. * test(chatbot): pass the warm-up request method explicitly fetchFollowing inferred POST from the path containing /api/chat, which hides the one request that is not a GET behind a string match. * fix(chatbot): cap the suggestions call the earlier sweep missed requestSuggestions streams through getArtifactModel, an Anthropic id, with no maxOutputTokens, so asking for suggestions on a document still hit the credit reservation the other call sites were fixed for. Count model calls against capped calls instead of listing the files, since a named list is what let this one through: it lives under lib/ai/tools rather than beside the other artifact handlers. * test(chatbot): cover createDocument artifact path in Playwright (WALM-354) The mock now emits createDocument for essay/document prompts so CI exercises ARTIFACT_MODEL the same way a retired id used to 404.
…ject /api/analyze accepts 64 KiB but embeds the full input as its dedup query, while text-embedding-3-small caps at 8192 tokens and the embedder does not truncate. Oversized input was therefore always rejected with a 400, ~560ms after the call went out, before falling through to plain extraction. Guard the call with MAX_PRE_EXTRACT_EMBED_BYTES (8 KiB) and record pre_extract_status="skipped_oversized", so the case stops emitting a WARN that reads like an incident. Production evidence (relayer logs, 2026-08-27): a 31,782-byte input was rejected with `Invalid 'input': maximum context length is 8192` after burning 564ms. That size is pinned as a regression test. The threshold equals remember's SUMMARIZE_THRESHOLD_BYTES but is kept separate on purpose: that one decides when to pay for an LLM summary before storing a permanent embedding, this one decides when to abandon a throwaway dedup query vector. Measured bytes-per-token runs ~1.4 (base64) to ~4.5 (prose), so 8 KiB sits below even the densest content's breach point. Large inputs still extract without dedup context. Recovering it needs the pre-extraction fetch_batch timeout addressed first, which dominates this failure mode in production.
…iscover-which-namespaces-exist-recall-is feat(sdk): listNamespaces() so an agent can discover namespaces [WALM-395]
…p-the-pre-extract-embedding-call-on-inputs-known fix(analyze): skip the pre-extraction embed on inputs the API will reject [WALM-411]
…enge flow (WALM-414)
…nce-v2 feat(noter): server-issued challenge nonce for wallet auth replay protection
Collaborator
Style Guide AuditAudited 11 file(s) against the Sui Documentation Style Guide. 5 violation(s) found. All must be fixed before merge.
|
ducnmm
approved these changes
Sep 8, 2026
ducnmm
left a comment
Collaborator
There was a problem hiding this comment.
Nikola's demo-app pass on dev 3da15d36 (six merged fixes). AUTH_SECRET on staging was rotated to 64 chars. Approving dev → staging.
@harrymove-ctrl please merge when ready. Do not merge #861 (staging → main) until this is on staging — #861 still carries a leftover changeset.
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up promotion to WALM-460. The previous two promotions — #851 (WALM-460) and #865 (WALM-471) — are already merged;
devhas moved again since 2026-09-04.Delta re-checked at open time
staging←devdevHEAD3da15d36stagingHEADc3c1941efix, 5docs, 4feat, 4chore, 1test, 1ciHead is the
devbranch, not a pinned SHA — anything merged todevbefore approval joins this promotion. Re-check the counts at merge time.Dev-environment verification (2026-09-07)
devHEAD3da15d36https://relayer.dev.memwal.ai/healthstatus: ok,write_ready: true,writes: "ok"https://dev.memwal.ai6b2bd93b. The 6 commits between it anddevHEAD touchapps/noteronly — noservices/serverchange is undeployed, so the relayer running on dev is current for server codewrites: "ok"on/healthTwo things worth a second look before merge, neither blocking this PR:
"mode":"production". Expected for the Railway dev service, but confirm that is intentional.apps/noterchanges from feat(noter): server-issued challenge nonce for wallet auth replay protection #53 are ondevHEAD but ahead of the last deployed build. Noter needs a dev redeploy before the wallet-challenge flow can be verified live.The one skipped check, and why the dev env is still verified
PR checks settled at 44 pass, 0 fail, 1 skipped. The skipped one is
E2E / dev relayer— the Python SDK live E2E that actually hitsrelayer.dev.memwal.ai. It is skipped on pull requests by design, not by failure:.github/workflows/test-python-sdk.ymlgates it toworkflow_dispatch,schedule, orpushtodev, because environment secrets are withheld from fork PRs and every authenticated run writes real memories.That means the PR checks alone never touch the live dev environment. Where it does run, it is green:
6b2bd93b— the build currently deployed on devdevd64975f8dev3683d6c8The scheduled run is the one that counts: it ran against the exact commit the dev relayer is serving, about an hour before this PR opened, and passed. With
/healthalso reportingstatus: ok/write_ready: true/writes: "ok", the dev environment is verified working end to end — not merely verified as compiling.What ships
Behavior changes — need runtime verification on staging
getSuggestionsscoped to the calling userBearerheader treated as no session instead of crashinggetSessionandlogoutbound to the caller's own sessionAUTH_SECRETvalidated before signing or verifying a session/healthexposeswrites: ok|pausedlistNamespaces()so an agent can discover namespacesmaxDistanceexposed onmemwal_recallmemory_searchrelevance clamped when cosine distance > 1server_urlCI only — verified by the pipeline
Docs only — review by reading
restoretruncation documentedTest plan on staging
getSuggestionsas user A never returns user B's dataDELETEa nonexistent document → 404Authorization: Bearerwith a garbage value → treated as anonymous, no 500getSession/logoutwith another caller'ssessionId→ rejectedAUTH_SECRETmissing → fails loudly at verify/healthreportswrites: ok, flips topausedwhen writes are pausedlistNamespaces()returns the caller's namespacesmemwal_recallhonoursmaxDistancememory_searchrelevance never negativeserver_urlemits a warningstagingafter mergeKnown carry-over — not fixed here
WALM-470 (p0, Backlog) — the two regressions from #810 / #837 are already on
stagingvia #851 and are not fixed by this delta:sort=recentsilently overridden byscoring_weights—services/server/src/routes/recall.rsis untouched in this delta.embedder.rschange here is aMAX_EMBED_INPUT_BYTESvisibility bump topub(crate)for fix(analyze): skip the pre-extraction embed on inputs the API will reject [WALM-411] #828. The guard placement and the body-gated 400 mapping are both unchanged.This PR neither worsens nor resolves WALM-470. It stays open for a later promotion.
Out of scope
staging→main(release: promote staging to main #861, open).Acceptance criteria
staging, headdevdev, so it moves)