Server-side Node.js SDK for Twilio Video Group Rooms with raw media frame access. Built on a native C++ addon over WebRTC, it lets you push and receive decoded video and audio frames from Node.js on realtime.
Note: this is a beta release of the Twilio Media SDK for Node.js. It is provided for evaluation purposes only and should not be used with production traffic. During the beta period this SDK is not HIPAA eligible.
Install the provided .tgz file directly:
npm install ./twilio-video-node-sdk-<version>.tgzThe native binary is prebuilt and bundled — no build step required. Import it as @twilio/video-node-sdk.
Requirements:
- Node.js >= 24.0.0
- Linux x86-64, glibc >= 2.34 (Ubuntu 22.04+, Debian 12+)
Linux x86-64 is the only supported platform for the beta. Alpine/musl, arm64 and Windows are not supported, and macOS is development-only. See Platform Support for the full ABI requirements.
The connect() function takes a standard Twilio Video Access Token with a
VideoGrant, the same token format used by the JavaScript SDK. Generate one
using the twilio helper library. See
User Identity and Access Tokens
for details.
const { connect, createLocalVideoTrack } = require('@twilio/video-node-sdk');
async function main() {
const videoTrack = createLocalVideoTrack('my-camera');
const room = await connect(token, {
name: 'my-room',
videoTracks: [videoTrack],
});
console.log('Connected:', room.name, room.sid);
// Push I420 video frames. connect() resolves only once the Room is connected;
// frames written before that are dropped and counted in getWriteStats().
videoTrack.write({
format: 'I420',
width: 1280,
height: 720,
y: { data: yPlane, stride: 1280, width: 1280, height: 720 },
u: { data: uPlane, stride: 640, width: 640, height: 360 },
v: { data: vPlane, stride: 640, width: 640, height: 360 },
});
async function trackSubscribed(track) {
if (track.kind !== 'video') return;
// Awaiting each frame is the backpressure. The loop ends by itself when the
// track is unsubscribed or the Room disconnects.
for await (const frame of track.frames()) {
console.log(`${frame.width}x${frame.height} @ ${frame.timestamp}us`);
frame.close?.();
}
}
function participantConnected(participant) {
participant.on('trackSubscribed', trackSubscribed);
participant.tracks.forEach(publication => {
if (publication.isSubscribed) {
trackSubscribed(publication.track);
}
});
}
// participantConnected does not fire for participants already in the Room, and a
// track can finish subscribing before this listener is attached. Seed from
// room.participants and check isSubscribed on the publications found there.
room.participants.forEach(participantConnected);
room.on('participantConnected', participantConnected);
room.on('disconnected', () => {
room.dispose();
});
}
main().catch(err => {
console.error('Error:', err);
process.exit(1);
});Call room.dispose() when you are done with a Room. Until you do, the process
does not exit on its own. disconnect() leaves the session but does not release
the native resources behind it.
The disconnected event is the clearest place to dispose, as in the quick
start.
For what each object owns, what teardown releases, and the ordering the SDK guarantees during teardown, see LIFECYCLE.md.
This SDK shares the same Room/Participant/Track model and event names as the Twilio Video JavaScript SDK, but is designed for server-side media processing rather than browser-based conferencing. Key differences:
-
No device capture.
createLocalVideoTrack()andcreateLocalAudioTrack()return pushable tracks with no media constraints. You supply raw frames viatrack.write()instead of capturing from a camera or microphone. -
No rendering. There is no
track.attach(element). Remote media arrives as raw decoded frames (I420 video, PCM audio) throughfor await (const frame of track.frames()). -
Fixed audio input format.
LocalAudioTrack.write()accepts only 48 kHz mono S16LE PCM. Received audio may vary in sample rate and channel count. -
No adaptive simulcast or track priority. Published video tracks always use standard priority with simulcast disabled; the deprecated
TrackPriorityAPI is not exposed. Bandwidth and remote render-size hints are configurable instead viabandwidthProfile,encodingParameters, andRemoteVideoTrack.setContentPreferences(). -
Synchronous track creation.
createLocalVideoTrack()returns a track immediately (no async device permissions). Tracks can be passed toconnect()or published later vialocalParticipant.publishTrack().
| Function | Description |
|---|---|
connect(token, options?) |
Connect to a room. Returns Promise<Room>. |
createLocalVideoTrack(name?) |
Create a pushable local video track. |
createLocalAudioTrack(name?) |
Create a pushable local audio track. |
createLocalDataTrack(name | options?) |
Create a local data track. Options are name, ordered, and one of maxPacketLifeTime/maxRetransmits. |
createLocalTracks(options?) |
Create local audio and/or video tracks. With no options, returns both. If either audio or video is specified, the other defaults to false. Each key accepts true/false or a per-track options object (e.g. { name }). Returns Promise<(LocalAudioTrack | LocalVideoTrack)[]>. |
twilioErrorFromCode(code, message?) |
Build a TwilioError (or matching subclass) from a numeric error code. |
setLogLevel(level) |
Set native log level. Accepts a name ('off' | 'fatal' | 'error' | 'warning' | 'info' | 'debug' | 'trace' | 'all') or the equivalent number 0 (off) through 7 (all). |
getVersion() |
Returns the native SDK version string. |
MAX_QUEUE_CEILING |
Upper bound (1024) accepted for any maxQueue, on frames() and on source.maxQueue. |
SDK_LOCAL_CODE |
The code (0) carried by errors the SDK raises locally, which Twilio never assigns a code to. Match on the error class instead. |
| Export | Description |
|---|---|
Room |
A connected video room. Emits events, exposes participants. |
LocalParticipant |
The local participant. Publish/unpublish tracks. |
RemoteParticipant |
A remote participant. Emits trackSubscribed/trackUnsubscribed. |
LocalVideoTrack |
Pushable video track (write(frame)). |
LocalAudioTrack |
Pushable audio track (write(frame), clearBuffer()). |
LocalDataTrack |
Send arbitrary data (send). Create via createLocalDataTrack(name | options?). |
RemoteVideoTrack |
Receive video frames (frames()), getStats(), frameDropped event. |
RemoteAudioTrack |
Receive audio frames (frames()), getStats(), frameDropped event. |
RemoteDataTrack |
Receive data messages (message event). |
TrackPublication |
Base class for published tracks (trackSid, trackName, kind, isTrackEnabled). |
LocalTrackPublication |
Local publication. Exposes track and unpublish(). Subclassed per kind (LocalVideoTrackPublication, …). |
RemoteTrackPublication |
Remote publication. Exposes track and isSubscribed. Subclassed per kind (RemoteVideoTrackPublication, …). |
TwilioError |
Base error class, carrying a numeric code. One subclass per known Twilio code, plus SDK-local errors (NativeBindingLoadError, RoomConnectTimeoutError, …). |
ErrorCode |
Enum of Twilio Video error codes. |
participantConnected is not emitted for participants who were already in the Room when
connect() resolved. They are part of the Room's starting state: read them from
room.participants.
A participant who was already publishing emits trackSubscribed after connect() resolves.
Subscriptions that completed before the listener was attached are not replayed, and appear in
participant.tracks with isSubscribed set to true.
| Event | Handler Signature |
|---|---|
disconnected |
(room: Room, error?: TwilioError) => void |
connectFailure |
(error: TwilioError) => void |
reconnecting |
(error?: TwilioError) => void |
reconnected |
() => void |
participantConnected |
(participant: RemoteParticipant) => void |
participantDisconnected |
(participant: RemoteParticipant) => void |
participantReconnecting |
(participant: RemoteParticipant) => void |
participantReconnected |
(participant: RemoteParticipant) => void |
recordingStarted |
() => void |
recordingStopped |
() => void |
dominantSpeakerChanged |
(participant: RemoteParticipant | null) => void |
transcription |
(transcriptionJson: string) => void |
The Room re-emits every track event in RemoteParticipant Events,
appending the RemoteParticipant that emitted it as the last argument. Handle every
participant's tracks from one place instead of attaching a listener to each participant.
| Event | Handler Signature |
|---|---|
trackSubscribed |
(track: RemoteVideoTrack | RemoteAudioTrack | RemoteDataTrack, publication: RemoteTrackPublication) => void |
trackUnsubscribed |
(track: RemoteVideoTrack | RemoteAudioTrack | RemoteDataTrack, publication: RemoteTrackPublication) => void |
trackSubscriptionFailed |
(error: TwilioError, publication: RemoteTrackPublication) => void |
trackPublished |
(publication: RemoteTrackPublication) => void |
trackUnpublished |
(publication: RemoteTrackPublication) => void |
trackEnabled |
(publication: RemoteTrackPublication) => void |
trackDisabled |
(publication: RemoteTrackPublication) => void |
videoTrackSwitchedOff |
(track: RemoteVideoTrack) => void |
videoTrackSwitchedOn |
(track: RemoteVideoTrack) => void |
networkQualityLevelChanged |
(level: number) => void |
| Event | Handler Signature |
|---|---|
trackPublished |
(publication: LocalTrackPublication) => void |
trackPublicationFailed |
(error: TwilioError, localTrack?: LocalTrack) => void |
networkQualityLevelChanged |
(level: number) => void |
LocalTrackPublication exposes track (the local track instance) and an unpublish() method:
const pub = room.localParticipant.tracks.get(trackSid); // LocalTrackPublication
pub.unpublish(); // unpublishes the underlying trackRemoteTrackPublication exposes track (the subscribed remote track, if any) and isSubscribed.
Returns a snapshot of WebRTC stats per peer connection. Rejects if the room is disconnected.
const reports = await room.getStats();
// reports[i]: {
// peerConnectionId, localAudioTrackStats, localVideoTrackStats,
// remoteAudioTrackStats, remoteVideoTrackStats
// }Push raw I420 video frames into a room. Frames written before connect() resolves are dropped, and counted in getWriteStats().
const track = createLocalVideoTrack('camera');
track.write({
format: 'I420', // optional; only 'I420' is accepted
width,
height, // both must be positive and even
y: { data: yBuffer, stride: yStride, width, height },
u: { data: uBuffer, stride: uStride, width: width / 2, height: height / 2 },
v: { data: vBuffer, stride: vStride, width: width / 2, height: height / 2 },
timestamp, // optional microseconds; defaults to monotonic now
rotation, // optional 0 | 90 | 180 | 270
});
track.enabled = false; // muteBuffers are copied synchronously, so they can be reused as soon as write() returns.
write() returns false when the frame was dropped rather than encoded - most often because the encoder sink has not attached yet, but also when libwebrtc's adapter rate-limits or rejects the resolution. It throws TypeError/RangeError on invalid input.
Optionally pin the frame size at creation, so a mismatched frame is rejected instead of silently rescaled:
const track = createLocalVideoTrack({
name: 'camera',
source: { type: 'raw', format: 'I420', width: 1280, height: 720, fps: 30 },
});Publish-side counters:
const { framesWritten, framesDropped, sendQueueDepth, maxQueue, lastTimestamp } =
track.getWriteStats();Video publish is synchronous - write() hands the frame straight to the encoder - so there is no SDK-side send queue and sendQueueDepth/maxQueue are always 0. A framesDropped here means the frame was rejected, not shed from a queue.
Push raw PCM audio samples into a room. Format is fixed to 48kHz mono S16LE.
const track = createLocalAudioTrack('mic');
const accepted = track.write({
pcm, // Buffer of interleaved int16 samples
frames, // samples per channel, e.g. 480 for a 10ms chunk
timestamp, // optional microseconds
});Unlike video, audio publish has a real send queue, drained one 10 ms chunk at a
time. It is bounded (~500 ms by default) so a producer running faster than real
time cannot accumulate latency. A write() is accepted only if it fits whole in
the remaining space; otherwise nothing is buffered, write() returns false,
and the rejection is counted. A single write() larger than the bound can never
fit, so size maxQueue to the largest burst you intend to publish:
const track = createLocalAudioTrack({
name: 'mic',
// maxQueue is in 10ms chunks: 200 => ~2s of smoothing. It binds the
// process-wide audio device, so it applies to every local audio track.
source: { type: 'raw', format: 'PCM_S16LE', sampleRate: 48000, channels: 1, maxQueue: 200 },
});
const { framesWritten, framesDropped, sendQueueDepth, maxQueue } = track.getWriteStats();Publishing at real-time cadence should never drop. A non-zero framesDropped
means the producer is outrunning the wire, or that a single write() was larger
than maxQueue.
The timestamp on an audio frame is observability-only. Audio publish is FIFO:
the device emits queued samples on its own 10 ms cadence, so the timestamp feeds
lastTimestamp and timestampRegressions and does not change what is sent.
clearBuffer() discards whatever is still queued and not yet sent. Use it when
the queued audio has become stale rather than merely late - barge-in, where the
speaker is interrupted and the rest of the utterance should never play, is the
usual case. Writes after it resume from an empty queue.
track.clearBuffer();Send arbitrary string or binary messages. Delivery is reliable and ordered by default.
const track = createLocalDataTrack({ name: 'chat', ordered: true });
room.localParticipant.publishTrack(track);
track.send('hello');
track.send(Buffer.from([0x01, 0x02]));
// send() reports the outcome. The promise always resolves - never rejects - so
// a fire-and-forget send cannot produce an unhandled rejection.
const result = await track.send('important');
if (!result.ok) console.warn('send failed:', result.error);Messages larger than 64 KB (kMaxMessageSize) are rejected synchronously
with a RangeError and never transmitted.
Pass maxPacketLifeTime (milliseconds) or maxRetransmits (a count) to trade reliability for
latency. The two are mutually exclusive, and each must be an integer from 0 to 65535.
const telemetry = createLocalDataTrack({ name: 'telemetry', maxPacketLifeTime: 500 });
telemetry.maxPacketLifeTime; // 500
telemetry.maxRetransmits; // null
telemetry.reliable; // falsemaxPacketLifeTime and maxRetransmits are number | null, reading back as null when the
limit was not set. reliable is true only when neither is set.
Receive decoded I420 video frames from a remote participant.
for await (const frame of track.frames()) {
// frame: {
// format: 'I420',
// width, height,
// y, u, v, // I420Plane: { data: Buffer, stride, width, height }
// timestamp: number, // microseconds
// captureTimestamp?: number,
// rtpTimestamp?: number,
// frameId: number, // SDK-generated, monotonic per track
// rotation?: 0 | 90 | 180 | 270,
// close?(): void, // optional prompt release
// }
frame.close?.();
}
// The loop ends on unsubscribe or Room disconnect. `break` releases the track.
// Hint the desired render dimensions to the SFU. Width/height must be positive
// integers. Only takes effect when the room was connected with a
// bandwidthProfile that has `contentPreferencesMode: 'manual'`.
track.setContentPreferences({ renderDimensions: { width: 320, height: 240 } });
// `isSwitchedOff` is `true` when the SFU has stopped delivering this track
// (e.g. due to bandwidth-profile constraints). Pair with the
// `videoTrackSwitchedOff` / `videoTrackSwitchedOn` events on RemoteParticipant.
track.isSwitchedOff;Receive decoded PCM audio frames from a remote participant.
for await (const frame of track.frames()) {
// frame: {
// format: 'PCM_S16LE',
// sampleRate, channels, frames,
// pcm: Buffer, // interleaved int16 samples
// timestamp: number, // microseconds
// frameId: number, // SDK-generated, monotonic per track
// close?(): void,
// }
}Receive string or binary messages from a remote participant.
track.on('message', (data, track) => {
/* data is string | Buffer; track is the RemoteDataTrack it arrived on */
});maxPacketLifeTime, maxRetransmits, reliable, and ordered report how the publisher
configured delivery. A publisher's limit of 65535 reads back as null, because a subscribed
track reports it the same way it reports an unset limit; reliable still distinguishes the two.
For the precise guarantees - buffer ownership, where frames are dropped, drop policy and ordering, timestamp rules, and the publish invariants - see FRAME_CONTRACT.md.
I420 planar layout. Each plane is an I420Plane: { data: Buffer, stride, width, height }, where stride is bytes per row (≥ the plane's width, padded for alignment).
| Plane | Logical size | data size |
Description |
|---|---|---|---|
| Y | width × height |
y.stride × height |
Luminance |
| U | ⌈width/2⌉ × ⌈height/2⌉ |
u.stride × ⌈height/2⌉ |
Chrominance (Cb) |
| V | ⌈width/2⌉ × ⌈height/2⌉ |
v.stride × ⌈height/2⌉ |
Chrominance (Cr) |
Publish and receive use the same planar shape: each of y/u/v is an I420Plane ({ data, stride, width, height }). A received frame can be written straight back out without reshaping.
Timestamps are plain numbers of microseconds (timestamp). Microsecond resolution stays exact in a JS number for roughly 285 years, and the underlying engine reports microseconds natively. rotation is 0 | 90 | 180 | 270.
Interleaved 16-bit signed little-endian PCM in a single Buffer.
- Inputs to
LocalAudioTrack.write()are fixed at 48kHz mono — onlypcmandframesare accepted. - Received
AudioFrames includesampleRate,channels,frames,pcm,timestamp(microseconds), andframeId.
{
name?: string; // Room name
videoTracks?: LocalVideoTrack[]; // Tracks to publish on connect
audioTracks?: LocalAudioTrack[];
dataTracks?: LocalDataTrack[];
enableInsights?: boolean;
enableAutomaticSubscription?: boolean;
enableDominantSpeaker?: boolean;
networkQuality?: boolean | { local?: 1; remote?: 0 | 1 };
preferredAudioCodecs?: ('opus' | 'PCMU')[];
preferredVideoCodecs?: 'VP8'[];
videoEncodingMode?: 'auto';
bandwidthProfile?: BandwidthProfileOptions;
receiveTranscriptions?: boolean;
region?: string; // e.g. 'us1', 'au1'
iceOptions?: IceOptions;
encodingParameters?: EncodingParameters;
connectionTimeout?: number; // ms; default 30000, 0 waits indefinitely
}{
video?: {
mode?: 'collaboration' | 'grid' | 'presentation';
maxSubscriptionBitrate?: number; // bits per second
trackSwitchOffMode?: 'detected' | 'predicted' | 'disabled';
clientTrackSwitchOffControl?: 'auto' | 'manual';
contentPreferencesMode?: 'auto' | 'manual';
};
}{
maxAudioBitrate?: number; // bits per second
maxVideoBitrate?: number; // bits per second
}{
transportPolicy?: 'all' | 'relay'; // 'relay' forces TURN
iceServers?: IceServer[]; // { urls: string[]; username?: string; credential?: string }
}- Node.js >= 24.0.0
- OS Linux x86-64
- Distros Ubuntu 22.04+ and Debian 12+.
- CPU x86-64 only. There is no arm64 build, so on Apple Silicon Node must run under Rosetta.
The prebuilt native addon is linked against glibc and requires:
| Requirement | Minimum |
|---|---|
| glibc | 2.34 |
libstdc++ (GLIBCXX) |
3.4.30 |
C++ ABI (CXXABI) |
1.3.11 |
It also links libX11.so.6, which WebRTC requires unconditionally. Install your distro's X11 client library (libx11-6 on Debian and Ubuntu) even on headless servers.
Alpine and other musl-based distros are not supported: the addon is glibc-only. Windows is not supported. macOS x64 builds and runs for local development, but is not a supported target and is not tested as one. See DEVELOPER_GUIDE.md.
The example applications listed below demonstrate various ways to use the SDK for audio or video processing. They load credentials from a .env file at the repo root. Copy the template, fill in your credentials, and run:
cp .env.example .env
# edit .env: set TWILIO_ACCOUNT_SID / TWILIO_API_KEY / TWILIO_API_SECRET
node examples/virtual_camera.js [room-name].env is gitignored, so your real credentials are never committed.
See the examples/ directory:
| Example | Description |
|---|---|
virtual_camera.js |
Decodes an MP4 with ffmpeg and pushes I420 frames to a room. |
video_mirror.js |
Receives remote video frames and pushes them back as-is. |
audio_push.js |
Generates a sine wave tone and pushes PCM audio to a room. |
data_channel.js |
Two participants exchange string and binary messages via data tracks. |
voice_agent.js |
Bridges room audio to the OpenAI Realtime API for a spoken voice agent (requires OPENAI_API_KEY). |
cv_object_detection.js |
Runs YOLOX object detection on a participant's webcam and re-publishes the video with bounding boxes. |
cv_face_analysis.js |
Analyzes a participant's face — presence and an attention estimate (head orientation) — drawn on the video. |
The computer-vision examples (cv_*.js) run local ONNX models via
onnxruntime-node and draw with
@napi-rs/canvas. These two are
large and only these examples need them, so they live in examples/package.json
rather than the SDK's own dependencies — install them separately:
npm install --prefix examplesNo cloud service or API key is needed: each example analyzes the first participant's video and expresses its result on a re-published video track. Run them against any room you also join from a browser, publishing your webcam.
The ONNX model files are not shipped with the repo — download the ones you need
and save them to examples/.models/. The examples print these same instructions
if a model is missing.
mkdir -p examples/.models
# cv_object_detection.js — YOLOX-nano (~3.7 MB)
curl -L -o examples/.models/yolox_nano.onnx \
"https://github.com/Megvii-BaseDetection/YOLOX/releases/download/0.1.1rc0/yolox_nano.onnx"
# cv_face_analysis.js — RTMO-t (a zip containing end2end.onnx, ~27 MB extracted)
curl -L -o examples/.models/rtmo-t.zip \
"https://download.openmmlab.com/mmpose/v1/projects/rtmo/onnx_sdk/rtmo-t_8xb32-600e_body7-416x416-f48f75cb_20231219.zip"
unzip -j examples/.models/rtmo-t.zip '*end2end.onnx' -d examples/.models
mv examples/.models/end2end.onnx examples/.models/rtmo-t.onnx
# verify the downloads (the examples also check this on startup)
echo "c789161ed43c8269fcd4e67c67eeeb4e80c622da2eb296a20bc6007bd18a0b7d examples/.models/yolox_nano.onnx" | shasum -a 256 -c
echo "20aad6e2e42359cac1c5b4a0b2da00e29bfe91a72a782fdcf287d273a04c1b24 examples/.models/rtmo-t.onnx" | shasum -a 256 -cBoth models are Apache-2.0 licensed and downloaded from their projects' official channels — YOLOX by Megvii and RTMO (OpenMMLab mmpose). Each example verifies its model's SHA-256 on startup and refuses to run a file that doesn't match.
See LICENSE.md.