release: promote staging to main - #861
Conversation
… protection Replace client-generated auth messages with a server-issued one-time challenge nonce. The server generates a random nonce stored in a new walletChallenges table with a 5-minute TTL. The client signs the nonce, and the server deletes it on use — replaying a captured signature fails because the challenge no longer exists.
- Use DELETE...RETURNING in connectWallet to eliminate TOCTOU race - Clean up expired challenges in getChallenge to prevent table bloat - Add Drizzle migration for walletChallenges table
Dev removed the wallet auth UI (Enoki overhaul) but kept the connectWallet tRPC route, so the replay-protection still has a target. - UI files (use-auth, auth-button-group) take dev; wallet-button removed - package.json takes dev (Enoki/noble deps) - deathbird migration renumbered to 0003 to coexist with 0002_add_enoki_fields Server-side changes preserved. Verified e2e: happy 200, replay 404, tampered 401.
login-preflight.test.mjs sandboxed the home directory with HOME alone. On Windows os.homedir() reads USERPROFILE and ignores HOME, so the sandbox did nothing and the test wrote its fixture credentials over the developer's real ~/.memwal/credentials.json. The delegate private key is stored nowhere else, so running the suite destroyed it; the lost key also stays authorized on-chain, since nothing revokes it. The test is the only one that both imports the credential writer in-process and sets HOME without USERPROFILE. The other nine hand both variables to a spawned child, which is why they were never affected. Setting both variables would fix the symptom but leaves the sandbox depending on process-global mutation happening before the first import of auth.js — a requirement nothing enforces. The path was resolved once at module load, so a second test in that file with its own temp directory would silently reuse the first one's, on any platform. Resolve the directory per call instead and let MEMWAL_CREDS_DIR override it. The test now points that at a temp directory and asserts credsPath() lands there before the login flow can write anywhere, so a broken sandbox fails loudly rather than after the real file is gone. Add creds-dir-sandbox.test.mjs: one case pins the override's precedence over HOME and USERPROFILE and its per-call resolution, and a second fails if any test file sandboxes HOME without USERPROFILE. That second case is what makes this class of mistake visible on the Linux-only CI, where the Windows failure itself can never reproduce. Refs WALM-360, GH #610
…ange Review feedback on #705. The source comments named GH #610 and WALM-360. This package is published, so a tracker id in shipped code is internal process leaking out to readers who cannot resolve it. Each comment now describes the failure mode it guards against — the os.homedir()/USERPROFILE note that made them useful is kept. MEMWAL_CREDS_DIR is user-facing and was documented in the environment variable reference but recorded nowhere a user would look for changes. Add it to the 0.0.10 entry in both the package changelog and its docs counterpart, matching how other MCP env and behaviour changes are logged.
The sign-in page described a failed hand-off in listener/callback terms and showed it under a green "MCP client connected" tick, so a first-time user could not tell what had happened or what to do. The MCP side said nothing at all: memwal_login returns its URL immediately, so by the time the flow fails there is no pending response left to turn into an error. The browser now tells the two failure modes apart, says the on-chain key is already registered and needs no repeating, and links a troubleshooting section written for this case. The MCP records the reason and prepends it to the next tool call's not-signed-in error, with a notifications/message warning as a secondary channel for clients that surface those. WALM-333
…integrity-sandbox-userprofile-in-mcp-login-tests # Conflicts: # docs/mcp/changelog.mdx # packages/mcp/CHANGELOG.md # packages/mcp/src/auth.ts
…ogin-callback-failure-message-is-opaque-to-first-time # Conflicts: # apps/app/src/pages/ConnectMcp.tsx # packages/mcp/src/auth-required.ts # packages/mcp/src/bridge.ts
handleLocalLogin in bridge.ts hard-coded LOGIN_BG_TIMEOUT_MS, so MEMWAL_MCP_LOGIN_TIMEOUT_MS only controlled the auth-required listener deadline. Move the resolver into login.ts so both call sites share it.
- packages/mcp/CHANGELOG.md, docs/mcp/changelog.mdx: the #705 bullet was filed under the already-shipped 0.0.10 section; moved to Unreleased. - creds-dir-sandbox.test.mjs: the no-override assertion assumed credsPath() always falls back to the global path, but it now falls back to projectCredsPath() first (WALM-361). Sandbox cwd with a .git sentinel before asserting. - login-preflight.test.mjs: importing auth.js with its own cache-busting query created a separate module instance from the one login.js's internal ./auth.js import binds to. Import without the query so the test observes the same instance saveCreds() actually writes through. - auth.ts: trim the credsPath() JSDoc's MEMWAL_CREDS_DIR paragraph to the load-bearing behavior; move the incident narrative out of the resolver. - docs/reference/environment-variables.md: drop the misleading "without restarting" claim, which only holds for in-process env mutation (tests), not a running client picking up a shell env change.
The docs style guide requires body text between heading levels. The package CHANGELOG keeps the stacked Unreleased/Fixed pair.
…egrity-sandbox-userprofile-in-mcp-login-tests fix(mcp): keep the login test off the real credential directory
SuccessCard only renders after preflight, so opened-too-late belongs on that error path, not this card.
The unused key is already on the account and must be revoked. Signing in again registers another key, including the wallet step.
Concurrent memwal_login joins the same listener, so HTTP refusal is a stale tab or a mismatch, not a replaced listener.
The heading ternary and debug URL row already state the constraint.
loginFlow now defaults through resolveLoginTimeoutMs so CLI cannot drift from the MCP wrappers. The failure-notice suite covers a signed-in timeout warning.
Retry is still the next step, but a restart loop registers another unused key. Say a retry only helps once the client stays running, and put revoke next to it.
A concurrent memwal_login joins the in-flight session, so registering catch on every call duplicated timeout warnings. Clear lastLoginFailure when a new attempt begins.
CallbackOutcome already names those cases.
…llback-failure-message-is-opaque-to-first-time fix(app,mcp): WALM-333 explain a failed login hand-off in plain language
Move the relayer suite into test-sdk.yml so the monorepo Tests workflow no longer carries schedule guards, sdk-changed, or duplicated Node setup. Authenticated live coverage is one serial remember → recall → namespace flow; push runs wait for Railway /health to serve this SHA before writing.
…lers WALM-399: a failed remember job returned the sidecar error verbatim, so callers saw "Unable to perform gas selection due to insufficient SUI balance" alongside a relayer-owned pool-wallet address. That address is infrastructure, not the caller's, and it changes between attempts, so users read it as a funding instruction and sent SUI to an address that was never theirs to fund. PR #736 already classified this error for retry and ops alerting; this covers the client-facing half. Add WalletJobError::is_infrastructure_funding_error over the existing classifiers (gas-pool budget, low WAL balance, upload-slot congestion), and sanitize at both client read points (remember_status, remember_bulk_status): infrastructure failures collapse to one message saying it is our side, the memory was not stored, and no address shown is theirs to fund. Other errors keep their diagnostic text with 32+ hex-digit addresses redacted, so short hex like 0x2::coin and Move abort codes stay readable. The stored error_msg is unchanged, so the admin dashboard and ops alerts still see the raw text. Fixing it at the API boundary covers the SDK, MCP, and Playground without client changes. Does not address the ticket's rate-limit suggestion: the 429 is counted in middleware before a job exists, so refunding infra-failed jobs needs a new credit path into the limiter and belongs in its own change.
WALM-373. Stamp initialize.clientInfo (and x-memwal-client from the stdio bridge) as a stable agentClient id, then include it on session.identified, tool.call, and session.closed sidecar logs.
Manual bump, no changeset: every MCP/plugin manifest, both changelogs, TESTING.md, and verify-manual-sdk-release.mjs. Folds the Unreleased #705 creds-dir fix into 0.0.12 with the agent-client identity work.
Feature branches cannot match github.sha on Railway. Dispatch still polls /health until write_ready so a preview run does not hit a mid-deploy replica.
GitHub-hosted Python urllib is 403'd by Cloudflare as Python-urllib/*; the live e2e fetch path already sends a normal client UA and succeeds.
Dispatch/schedule cannot match github.sha on Railway. Empty (or GitHub's stringified false) expected-commit now waits for /health write_ready only.
fix(noter): bind getSession and logout to the caller's own session (WALM-420)
Keep a one-liner on the URIError branch: Auth.js URL-decodes Bearer outside its JWT try/catch.
…bearer-header-crashes-the-chatbot-proxy-and fix(chatbot): treat a malformed Bearer header as no session (WALM-416)
#867) * docs: document namespace 255-byte cap and restore truncated (WALM-486) * docs: note restore limit clamp 1-100 (WALM-486)
* ci: add production-safe SEAL cross-account synthetic (COMG-715) Read-only dry-run of seal_approve so a live identity from account A cannot authorize account B's SEAL key. Missing secrets skip (exit 0). Authorization success pages via SYNTHETIC_SEAL_CROSS_ACCOUNT_FAIL. * docs: use H2 body heading in SEAL synthetic page (COMG-715) * fix(ci): parse JSON-RPC MoveAbort abort code (COMG-715) JSON-RPC returns MoveAbort as a Display string whose MoveLocation contains commas, so [^,]+ never reached abort code 100 and a healthy ENoAccess denial was classified as misconfiguration. * test(ci): drop redundant MoveAbort parser coverage (COMG-715) Keep the nested-comma JSON-RPC fixture and distinct parser/CLI branches. Remove the duplicate hex string, the old-regex assertion, and the classifyNegative helper that restated extractAbortCode results.
* feat(server): expose writes ok|paused on /health (COMG-717)
* fix(server): reject writes when WRITES_PAUSED (COMG-717)
Make WRITES_PAUSED write-path admission, not a health-only signal.
remember, remember_bulk, remember_manual, and analyze return 503
{"error":"writes are paused"} while /health stays HTTP 200 with
status ok and writes paused. Reads (recall, restore) stay available.
* fix(server): drop stub writes-paused route tests (COMG-717)
The throwaway router never mounted production remember/analyze/health
handlers, so it could not catch a dropped gate. Keep the helper and
503 mapping unit test; do not grow a handler harness.
…9) (#811) * fix(auth): don't treat Sui RPC 429 as a revoked delegate key Cache-hit re-verify used to evict on any on-chain error, including gRPC 429. That produced empty 401s that the SDK mapped to memwal_login even when the key was still registered. Keep the cache (or return 503) when Sui is unavailable; only evict on a definitive miss. WALM-429 * fix(auth): fail closed on Sui RPC 429 instead of serving a stale cache Keep the cached mapping so a later verify can succeed, but return 503 with AUTH_UPSTREAM_UNAVAILABLE and Retry-After rather than authenticating a possibly-revoked key for up to 24h. SDKs only use the credential-outage copy when that header is present. * fix(auth): expose Retry-After and stop treating object-not-found as unavailable (WALM-429) Classify GetObject failures by gRPC status code so typo'd x-account-id and NOT_FOUND evict as sign-in failures, while 429/unavailable still 503 with the cache row kept. Expose Retry-After on CORS so browsers can read it. * fix(auth): keep field-parse errors as unavailable RpcError (WALM-429) NotFound is only for absent objects / unparseable ids / NOT_FOUND. A node that returns the object but omits json/fields stays RpcError so auth 503s and keeps the cache row instead of 401+evicting a live key. * chore(auth): slim WALM-429 review follow-up Fold GetObject status tests; drop tautological field-parse cases; trim comments.
Read the key before jwtVerify's try so a short secret cannot look like an invalid cookie or a failed Enoki ownership check.
…ret-length fix(researcher): validate AUTH_SECRET before signing or verifying a session
…ay-to-discover-which-namespaces-exist-recall-is
* fix(python-sdk): warn on plaintext remote server_url (WALM-452) Fixes #748 * fix(python-sdk): redact credentials in plaintext server_url warning (WALM-452) Log only scheme/host/port so HTTPX userinfo is not written to memwal logs, while keeping the original URL for transport.
Match GET's identical guard, which has no comment.
…-nonexistent-document-returns-http-500-instead fix(chatbot): return 404 when deleting a nonexistent document (WALM-422)
… retirements (WALM-354) (#822) * fix(chatbot): replace the model ids OpenRouter retired Four of the seven ids in chatModels no longer exist upstream, and OpenRouter answers a retired id with a 404 at request time, so picking one only ever produced "Oops, an error occurred!". Two of them were also hardcoded at a call site: getTitleModel used google/gemini-2.0-flash-001, which left every chat named "New chat", and getArtifactModel used anthropic/claude-3.5-haiku, which broke document creation for everyone rather than only whoever selected it. Point the title and artifact models at ids the picker also offers, the way apps/researcher does after the same incident there, and centralise the "-thinking" handling so the marker cannot reach OpenRouter unstripped. * test(chatbot): fail CI when a curated model id is retired upstream The unit tests can only catch an id drifting away from chatModels; they cannot tell that OpenRouter has withdrawn one, which needs a live call. That gap is how the retired title model reached production: nothing in the repo changed, so there was no PR run to go red. Check the ids against the public catalog endpoint, which needs no API key. It gets its own workflow because test.yml has no schedule trigger to hang a weekly run off, and it limits pull-request runs to changes that touch the list so other PRs take no network dependency on OpenRouter. * fix(chatbot): cap the output tokens each model call asks for No call site set maxOutputTokens, so every request was quoted against the model's entire output window. OpenRouter reserves that credit up front for Anthropic and Google models, which rejected the turn outright — "requires more credits, you requested up to 65535 tokens" — however short the reply would have been. OpenAI and DeepSeek do not reserve, which is why only some of the picker appeared broken. Cap the chat turn, the background title call and all six artifact calls. The reasoning cap stays above the thinking budget the chat route requests so the answer is not truncated before it reaches the reply. * fix(chatbot): resolve a stale chat-model cookie to the default Both chat pages handed cookie.value straight to <Chat>, which sends it as selectedChatModel, and /api/chat rejects an id that is no longer in chatModels. Anyone who had a since-retired model selected therefore had every send fail as a bad request, with nothing to explain it: the picker's fallback is display-only, so it rendered a valid model name while the request body still carried the stale id. Recovery meant reopening the picker, which the error text does not hint at. Coerce the cookie where it is read, matching resolveChatModelId in apps/researcher. * fix(chatbot): make the test mock answer the ids the app actually sends myProvider registers behaviours — chat-model, chat-model-reasoning — but getLanguageModel passed the catalog id through, so every request under PLAYWRIGHT threw NoSuchModelError: No such languageModel: openai/gpt-4o-mini. Every chat turn failed in CI, and the suite stayed green only because no test asserted that a reply arrives. getResponseForPrompt had the matching problem: it searched the whole serialised prompt, and the system prompt contains "hi" inside words like "this", so it always returned the greeting and the weather and default branches were unreachable. Match the newest user turn instead. * test(chatbot): cover the reply, title and reasoning paths end to end The suite exercised the composer and the picker but never asserted that a reply comes back, which is why it stayed green while every chat turn failed under PLAYWRIGHT. These six cover the paths the recent fixes touch: a streamed reply, a reply that survives a reload, intent reaching the model, the background title landing in the sidebar, reasoning streaming for a thinking model, and a retired chat-model cookie no longer rejecting every send. * test(chatbot): make the Playwright warm-up actually compile the routes The warm-up fetched with redirect: "manual", so the proxy's guest-session redirect ended it: every request returned 307 in under 60ms without rendering anything, and the cache stayed cold. Tests then paid the compile themselves — from a cleared .next, 4 to 14 of them failed on first-hit navigations of 40s to 1.3m against a 30s navigationTimeout and a 60s test timeout. CI never reported it because by the second of its two retries the routes were warm. Follow the redirects and keep their cookies, and warm the API and chat routes a message send needs rather than the landing page alone. The same cleared-.next run now warms in 84s and passes 28/28. * chore(chatbot): restore the trailing newline models.ts had Unrelated whitespace churn from the model-id fix. * test(chatbot): pass the warm-up request method explicitly fetchFollowing inferred POST from the path containing /api/chat, which hides the one request that is not a GET behind a string match. * fix(chatbot): cap the suggestions call the earlier sweep missed requestSuggestions streams through getArtifactModel, an Anthropic id, with no maxOutputTokens, so asking for suggestions on a document still hit the credit reservation the other call sites were fixed for. Count model calls against capped calls instead of listing the files, since a named list is what let this one through: it lives under lib/ai/tools rather than beside the other artifact handlers. * test(chatbot): cover createDocument artifact path in Playwright (WALM-354) The mock now emits createDocument for essay/document prompts so CI exercises ARTIFACT_MODEL the same way a retired id used to 404.
…ject /api/analyze accepts 64 KiB but embeds the full input as its dedup query, while text-embedding-3-small caps at 8192 tokens and the embedder does not truncate. Oversized input was therefore always rejected with a 400, ~560ms after the call went out, before falling through to plain extraction. Guard the call with MAX_PRE_EXTRACT_EMBED_BYTES (8 KiB) and record pre_extract_status="skipped_oversized", so the case stops emitting a WARN that reads like an incident. Production evidence (relayer logs, 2026-08-27): a 31,782-byte input was rejected with `Invalid 'input': maximum context length is 8192` after burning 564ms. That size is pinned as a regression test. The threshold equals remember's SUMMARIZE_THRESHOLD_BYTES but is kept separate on purpose: that one decides when to pay for an LLM summary before storing a permanent embedding, this one decides when to abandon a throwaway dedup query vector. Measured bytes-per-token runs ~1.4 (base64) to ~4.5 (prose), so 8 KiB sits below even the densest content's breach point. Large inputs still extract without dedup context. Recovering it needs the pre-extraction fetch_batch timeout addressed first, which dominates this failure mode in production.
…iscover-which-namespaces-exist-recall-is feat(sdk): listNamespaces() so an agent can discover namespaces [WALM-395]
…p-the-pre-extract-embedding-call-on-inputs-known fix(analyze): skip the pre-extraction embed on inputs the API will reject [WALM-411]
…enge flow (WALM-414)
…nce-v2 feat(noter): server-issued challenge nonce for wallet auth replay protection
chore: verify dev before promoting to staging (WALM-460) — DO NOT MERGE until verified
| if status == reqwest::StatusCode::BAD_REQUEST { | ||
| tracing::warn!(%status, body, "embedding API returned 400"); | ||
| return Err(AppError::BadRequest( | ||
| "embedding input exceeds the model context limit".into(), |
There was a problem hiding this comment.
[suggestion] After the 16 KiB local gate, every non-transient HTTP 400 from the embedding API is rewritten as BadRequest("embedding input exceeds the model context limit"). OpenRouter/OpenAI 400s also cover invalid model, malformed body, and similar config errors. Callers then retry by shrinking text while operators only have the tracing::warn body.
Suggestion: Map to a generic 400 (or inspect the upstream body for a context-length signal) instead of always blaming the model limit.
| /// This alone does not make "newest wins" correct: the candidate set is | ||
| /// the cosine top-`limit` from `search_similar`, so the newest row can be | ||
| /// missing from `results` entirely and no client-side sort recovers it. | ||
| /// The server-side recency mode that over-fetches is still open. |
There was a problem hiding this comment.
[nit] RecallResult.created_at still says the server-side over-fetch recency mode “is still open”. This PR adds RecallSort::Recent and candidate_limit in the same file, so the comment is already false.
Suggestion: Point at RecallSort::Recent (default recall is still cosine top-limit).
|
@harrymove-ctrl this is yours to merge (you authored it; same as #874). Daily will not merge. Leftover changeset is gone on this head ( Please confirm Railway prod |
Routine promotion. 64 commits, 83 files, 15 PRs merged into
stagingsince #796.CI on the
staginghead is green, and every test suite passes locally (1064 tests, 0 failures — breakdown at the bottom).Before you merge: one thing to fix first
Short version: merging this publishes JS SDK
0.1.6to npm, but leavesmainsaying0.1.5. After that, the JS SDK stops publishing entirely — silently, with CI still green.One-word fix, in
.changeset/config.json:Land that first, then merge this. Nothing else needs to change.
Why this happens (click to expand)
This PR carries
.changeset/recall-write-time-and-recency-weights.md, which came in with #810. The file is correct — it just tells the release workflow "bump the JS SDK by one patch."The workflow that reads it has a bug.
release-sdk.yml:.changeset/config.jsonsets"commit": true, sochangeset versioncommits the bump itself.git diff --quietonly looks at uncommitted changes, so it always sees a clean tree and reportshas_changes=false. Thegit pushstep is then skipped.But the publish step is gated only on
github.ref == 'refs/heads/main', so it runs anyway, reads the locally-bumpedpackage.json, and publishes0.1.6.The published
0.1.6is fine — correct code, correct provenance. Nobody is broken. The damage is to the pipeline: on every later push tomain, changesets recomputes0.1.5+ the still-present changeset file →0.1.6again,npm view …@0.1.6now succeeds, and publish is skipped. The JS SDK never publishes again, and CI stays green the whole time.This has been latent since March (
"commit": truein 153e373, thegit diffprobe in 71fcd0b) because MCP and the Python SDK are version-bumped by hand in their feature branches — sochangeset versionhas always been a no-op. This is the first time a changeset file reachesmain.release-mcp.ymlhas the same pattern and the same fix.Also: merging triggers a mainnet Walrus Site deploy, because
apps/app/**changed and that is the path filter ondeploy-app-walrus.yml.What ships
services/serversort=recentrecall mode +created_aton results (#810) ·OWNER_TOKEN_SECRETminimum-length check at startup (#835) · 400 instead of a provider error on oversized recall queries (#837) · clearer relayer job errors (#803)packages/mcpdocs.githubpackages/sdksortandcreated_atonrecall()packages/python-sdk-memwalremember_bulk_asyncrejects mismatchedjob_ids(#838)Versions:
memwal-mcp0.0.11 → 0.0.12 ·memwal(Python) 0.1.8 → 0.1.9. Both hand-bumped, both publish correctly. The JS SDK is the only one going through changesets — see above.Two behaviour changes worth knowing
OWNER_TOKEN_SECRETshorter than 32 bytes now panics at startup. Checked all three Railway environments before opening this: production and staging have it unset (feature disabled, boots normally), dev has 64 bytes. No environment is affected.rememberManual/remember_manualnow sendencryptedDatainstead of a pre-uploadedblobIdin both SDKs. Breaking for anyone on the manual-embed path.Test status
cargo test --bins(server)packages/mcppackages/sdkpackages/python-sdk-memwalapps/app(vitest)The 37 ignored Rust tests need Postgres and Redis. Playwright e2e was not run.
Not covered by any of the above:
sort=recentend-to-end against a live relayer — the 8 tests covering it are unit tests over the sort function and none reach/api/recall— and the MCP login/logout/concurrency fixes against a real client, since those 60 tests use a mock bridge. Worth a manual pass on dev after merging.