Skip to content

Security: cristiancmoises/keywave

SECURITY.md

Security

This document describes Keywave's threat model, cryptographic design, known limitations, and how to report a vulnerability.

Threat model

Keywave is an end-to-end encrypted chat and video tool for rooms of two (1:1) or up to six peers (group, full mesh). The signaling server is treated as untrusted: it relays the handshake but is assumed to be curious or hostile. Encryption happens entirely in the browser, per pair: every pair of participants runs its own ephemeral key agreement and derives its own key set, so compromise of one pair never touches another.

What the server (and anyone on the network path) can see:

  • IP addresses of the peers and the timing of a session.
  • That N parties joined the same room and exchanged data, including which peer addressed which (relayed signaling and chat ciphertext are directed: they carry short room-local peer ids, never broadcast).
  • Ephemeral public keys, opaque ciphertext, and WebRTC SDP/ICE during signaling.

What it cannot see:

  • Message plaintext or call audio/video.
  • The session keys (they are derived in the browser and never sent).

Keywave does not try to hide metadata (who talks to whom, when, from where). If that is part of your threat model, run it over a network layer that does.

Cryptography

All primitives use the browser's Web Crypto API. No third-party crypto code is loaded; the socket.io client and fonts are self-hosted.

Key agreement (protocol v2 — post-quantum hybrid)

Each peer generates an ephemeral, non-extractable ECDH P-256 keypair per page load. In addition, for every pair, the pair initiator generates a fresh ML-KEM-1024 keypair (FIPS 203, NIST security category 5 — the same level Evelin uses) and attaches the encapsulation key ek to its pubkey message; the responder encapsulates and returns the ciphertext in a directed kem message. Both sides then hold two independent 32-byte secrets per pair:

  • the classical ECDH P-256 shared secret (browser-native WebCrypto), and
  • the ML-KEM-1024 shared secret (vendored, hash-pinned implementation).

The pair's keys derive from both, so recorded traffic stays confidential unless an attacker breaks ECDH P-256 and ML-KEM-1024 — this is the defense against harvest-now-decrypt-later collection. There is no classical-only fallback in the code: a peer that does not present the post-quantum leg cannot complete a handshake (fail closed, no downgrade to negotiate).

Key derivation

Per pair, the two secrets are concatenated — IKM = ECDH_secret (32 B) ‖ ML-KEM_secret (32 B) — and fed to HKDF-SHA-256, whose salt is the SHA-256 of the full handshake transcript: a fixed label, both ECDH public keys in sorted order, the ML-KEM ek, and the ML-KEM ct, every field 4-byte length-prefixed so the framing is unambiguous. HKDF then expands six independent AES-256-GCM keys: chat, video, and audio, each split into a send and a receive key, with info strings keywave/v2/{chat|video|audio}:{A->B|B->A}. Send/receive direction is assigned by a lexicographic comparison of the two base64 ECDH public keys, so the peers derive mirror-image key sets. Each medium has its own key space, so a compromise of one stream's key never exposes another.

Binding the transcript into the salt means a relay that alters any handshake element — either ECDH public key, the ek, or the ct — necessarily lands the two sides on different keys and a different safety number (see below). The intermediate secrets (dk, KEM shared secret, hybrid IKM) are wiped immediately after derivation; the derived keys live only inside non-extractable WebCrypto handles.

Chat

Each message is encrypted with AES-256-GCM using a fresh random 96-bit IV. The message sequence number and timestamp are bound as additional authenticated data (AAD), so they cannot be altered without failing authentication. The receiver rejects any message whose sequence number is not strictly increasing, which gives replay and reorder detection. The sequence watermark advances only after a message authenticates: an unauthenticated or forged frame (for example one injected by a malicious relay with an inflated sequence number) is dropped and cannot advance the watermark, so it cannot censor subsequent genuine messages. Sequence counters are kept per pair. In a group room the sender encrypts the message once per recipient under that pair's key and the relay delivers each ciphertext only to its addressee — recipients cannot read (or even receive) each other's copies.

Audio and video

Media is encrypted per encoded frame with AES-256-GCM, per pair. The IV is a 4-byte random per-pipe prefix plus an 8-byte big-endian counter, so an IV is never reused within a stream. The IV is prepended to each frame.

All frame crypto runs in a dedicated worker, off the main thread, fed by one of two attachment paths (feature-detected): the standards-track RTCRtpScriptTransform (native in Safari 15.4+ and Firefox 117+) or the legacy Chromium createEncodedStreams(), whose transferable streams are handed to the same worker. The worker fails closed: until a pair's keys arrive it drops outgoing frames (nothing is ever sent in plaintext) and drops any incoming frame that does not authenticate (garbage is never fed to the decoder).

Frame encryption is negotiated per pair: each peer advertises support during the key exchange, and encoded transforms are enabled for a pair only when both of its ends support them. A pair with an unsupporting browser falls back cleanly to WebRTC's transport encryption (DTLS-SRTP) instead of breaking; in a group room other pairs keep their frame encryption. Because calls are peer-to-peer, DTLS-SRTP is already end-to-end at the transport layer; per-frame encryption is defense in depth for paths that traverse a middlebox (for example a TURN relay). The in-call badge shows the aggregate mode (frames / partial / DTLS-SRTP), and each remote tile's lock dot shows that pair's mode.

Session verification (safety number)

Public keys travel through the untrusted relay, so a malicious relay could attempt a man-in-the-middle by substituting its own keys. Keywave derives a safety number per pair with HKDF over that pair's hybrid secret, with the full handshake transcript (both ECDH public keys sorted, the ML-KEM ek, and the ML-KEM ct) bound into the derivation — independent of direction. Honest peers compute the same value; an attacker sitting in the middle of either leg — the classical ECDH exchange or the post-quantum KEM leg — necessarily produces a different value on each side (both cases are exercised by test_hybrid.node.js). In a group room, verify each pair: the Safety # control lists every peer with its verification state (1:1 skips the list).

The safety number is shown as a row of named emoji (each with its name, so the codes can be read aloud unambiguously) plus the full 64-bit value in hex. Compare it with your peer over the live call. If the codes match, no one is intercepting the keys. Until you confirm, the session is shown as unverified, and the app prompts you once to compare it.

Verifying a session

  1. Connect to a peer.
  2. Open the Safety # control in the call header (in a group room, pick the peer to verify — each pair has its own number).
  3. Read the emoji (or hex) aloud and confirm they match on both screens.
  4. If they match, mark the pair verified, and repeat for the remaining peers in a group. If they differ, end the call.

This step is what makes it safe to share a room ID over a messaging app: even if the ID is intercepted, an attacker who joins cannot reproduce your safety number.

Known limitations

  • Protocol v2 requires the same build on both ends. There is deliberately no downgrade path: a ≤1.x peer (or any peer without the post-quantum leg) cannot complete a handshake with a v2 peer. The failure is explicit, not silent.
  • The ML-KEM implementation is JavaScript and not independently audited. It is validated against 80 official NIST ACVP vectors and cross-checked against two independent implementations (RustCrypto ml-kem 0.2.3 — mirim's exact pin — and kyber-py 1.2.0), and it is hash-pinned at server boot and via SRI in the page; but constant-time execution cannot be guaranteed in JS. Exposure is bounded by the keys being ephemeral and per-pair. Full analysis: AUDIT.md §6.
  • Trust in the origin. As with any web-delivered E2E app, the server ships the client code. A compromised origin could serve a backdoored client. Self-hosted assets and a strict CSP reduce supply-chain exposure, but you ultimately trust whoever operates the origin.
  • Room IDs are bearer capabilities. Anyone who has a room ID can take a free slot (one in a 1:1 room, up to five in the largest group). The per-pair safety number defeats a racing attacker (their code will not match with anyone they impersonate). Reconnecting to an existing slot after a network drop requires a high-entropy per-session token (issued only to that peer over its own connection), so a third party who learns the short room ID cannot hijack a held slot during the reconnect grace window.
  • Verification is manual. If users skip the safety-number comparison, a relay-level MITM is not automatically detected. The app surfaces an unverified state but does not block media, because the call itself is needed to compare.
  • Per-frame media E2E requires encoded-transform support — Chromium (createEncodedStreams), Safari 15.4+, or Firefox 117+ (RTCRtpScriptTransform). It is negotiated per pair, so a pair with an older browser falls back to DTLS-SRTP rather than failing, and the in-call badge reports a partial state when a group mixes both modes.
  • Group rooms are a full mesh. Each participant uploads its media once per other participant, so uplink cost grows linearly with room size; the client steps per-sender bitrate down as the room grows, and the cap defaults to 6. There is no SFU, deliberately: no server ever touches media, even encrypted.
  • Group membership is flat. Any holder of the room link can fill a free slot; there is no creator-side admission control. Verify safety numbers.
  • No metadata protection (IPs, timing, who-talks-to-whom).
  • No in-session ratchet. Keys are ephemeral per session and discarded on hangup, which gives forward secrecy across sessions; there is no key rotation within a single session (sessions are short-lived).
  • TURN is optional and self-hosted. By default only public STUN is used, so calls across symmetric or carrier-grade NAT may fail until TURN is enabled (bundled coturn, or an external relay via env). A TURN relay only forwards encrypted media; with per-frame encryption active it cannot read call content, and even on the DTLS-SRTP path the relay sees only ciphertext.

Server hardening

  • A strict Content-Security-Policy limited to self-hosted origins with no third-party sources. Inline event handlers in the client require 'unsafe-inline' for scripts; everything else is 'self'.
  • Response headers: X-Content-Type-Options, X-Frame-Options, Referrer-Policy, Permissions-Policy, and Cross-Origin-Opener-Policy.
  • Abuse limits: room-creation rate limit, global room cap, relay rate limiting, per-event payload size validation, a socket buffer cap, and a background sweeper that reaps abandoned rooms. One room per client prevents room leaks.
  • No persistence: rooms and sessions live only in memory.
  • Container runs as a non-root user with cap_drop: ALL, no-new-privileges, a read-only root filesystem, and memory/PID limits.

Supported browsers

  • Chrome / Edge / Chromium 86+: full chat and per-frame media encryption (legacy createEncodedStreams, transferred into the crypto worker).
  • Safari 15.4+ (macOS, iPhone, iPad) and Firefox 117+: full chat and per-frame media encryption via the standards-track RTCRtpScriptTransform.
  • Older Safari / Firefox: full chat encryption; media uses DTLS-SRTP. Pairs that mix a supporting and a non-supporting browser fall back to DTLS-SRTP for that pair only.
  • A secure context is required (https:// or http://localhost); camera access and Web Crypto are unavailable on plain-HTTP origins.

Reporting a vulnerability

Please report security issues privately to the maintainer rather than opening a public issue. Include a description, affected version or commit, and steps to reproduce. Keywave is provided without warranty; see LICENSE.

There aren't any published security advisories