Realtime protocol (chat.v1)

Raw WebSocket (RFC 6455) over TLS, UTF-8 JSON text frames. The socket is the fast path; sync against per-conversation seq is the source of truth after any interruption. The byotalk SDK implements all of this for you; this page is for people writing their own client or debugging.

Connection

wss://rt.byotalk.com/v1?env=env_01J9...&sdk=js/1.0.0
Sec-WebSocket-Protocol: chat.v1
  1. TCP + TLS + HTTP upgrade. Server rejects unknown subprotocols.
  2. Client must send auth within 5 s or the server closes with 4001. The token never goes in the URL (proxy logs).
  3. Server verifies the token, binds the socket to (env_id, user_id), registers the route and presence, replies hello.
  4. Client sends sync, flushes its outbox, starts heartbeats.

Envelope

// client → server request (expects ack or error)
{ "t": "<type>", "id": "<client request id, unique per socket>", "d": { ... } }
// server → client reply
{ "t": "ack",   "re": "<id>", "d": { ... } }
{ "t": "error", "re": "<id>", "code": "<code>", "message": "<text>", "retryAfterMs": 1000 }
// server → client push
{ "t": "event", "e": "<event name>", "cid": "<conversation id>", "seq": 812, "d": { ... } }
{ "t": "<ephemeral type>", "d": { ... } }
  • Max frame size 16 KB (close 1009).
  • Unknown t or unknown fields → error invalid_frame (connection kept).
  • Every frame is validated against a JSON schema shared with the REST API.

Client → server frames

t d Ack d Notes
auth {token} hello frame instead of ack Must be first
message.send {cid, clientMsgId, text?, replyTo?, attachments?, metadata?} {clientMsgId, id, seq, createdAt} Duplicate clientMsgId → original ack
message.update {id, text?, metadata?, expectedVersion} {id, seq, version, editedAt} 409 version_conflict
message.delete {id} {id, seq, deletedAt}
typing {cid, state: "start"|"stop"} none Fire-and-forget; throttled
read {cid, seq} {cid, lastReadSeq} Monotonic
delivered {items: [{cid, seq}]} none Batched ≤ 1/s
sync {cursor | null} {conversations, removed, cursor, fullReload} Same as POST /v1/sync
presence.watch {userIds: [...]} (replaces set, ≤ 200) {presence: [{userId, online, lastSeenAt}]} Initial snapshot in ack
token.refresh {token} {expiresAt} Extends session without reconnect
hb {} none Reply to server hb

Conversation creation, membership changes, history pages and uploads use REST (they are not latency-critical and benefit from HTTP caching/retries).

Server → client frames

hello

{ "t": "hello", "d": {
    "sessionId": "ses_01J9...", "userId": "user_123", "serverTime": "2026-10-04T10:00:00.000Z",
    "heartbeatMs": 25000, "tokenExpiresAt": "2026-10-04T11:00:00Z",
    "limits": { "maxFrameBytes": 16384, "sendsPerSecond": 10 } } }

Events (in the seq stream)

e d
message.new {message}
message.updated {message} (new version)
message.deleted {id, deletedAt, hard}
member.added {userIds, actorId}
member.removed {userId, actorId, reason} — if userId is self, client drops the conversation
conversation.created {conversation} — sent to initial members; seq = conversation's first seq
conversation.updated {changes}

Ephemeral (not in the seq stream)

t d
receipt {cid, userId, lastReadSeq?, lastDeliveredSeq?}
typing {cid, userId, state}
presence {userId, online, lastSeenAt}
hb {} every 25 s
goaway {reconnectAfterMs} — node draining
resync_required {cid} — reload latest page for that conversation
resync_hint {} — node lost its fan-out subscription; run sync
token_expiring {expiresAt} — sent 5 min before expiry

Heartbeats and liveness

  • Server sends hb every 25 s; client replies hb. Client also treats any received frame as liveness.
  • Either side closes after 60 s without any frame.
  • Application-level because browser JavaScript cannot observe WebSocket ping frames, and 25 s stays under common 60 s load-balancer idle timeouts.

Close codes

Code Meaning Client action
1000 Normal (client disconnect) None
1009 Frame too large Fix client; do not resend same frame
1012 Server restart Reconnect after jitter (0–5 s)
4001 Token missing, invalid or expired Call tokenProvider once, reconnect; second failure → failed
4003 Forbidden (user deactivated, env suspended, dev token in production) failed; surface error
4008 Rate limited / plan connection ceiling Wait retryAfterMs from the close reason JSON
4009 Too many sockets for this user (10) Oldest socket is closed with this code; that client does not auto-reconnect
4400 Protocol violation Report; reconnect with backoff

Reconnection

  • Backoff: full jitter, base 500 ms, ×2, cap 30 s: delay = random(0, min(30s, 0.5s × 2^attempt)).
  • Immediate attempt on OS "online" event and app foreground (attempt counter reset).
  • goaway → reconnect after reconnectAfterMs (server sends random 0–10 s to spread load).
  • During reconnect the SDK state is reconnecting; outbox keeps accepting messages.

Sync algorithm

Client state

State Scope Persisted?
syncCursor Per device Yes, if persistence adapter enabled
lastSeq[cid] Per conversation (last contiguous event applied) Yes
Outbox entries {clientMsgId, cid, payload, createdAt} Per device Yes
Message window (≤ 500 per conversation) Per conversation Optional

Steps after hello

1. send sync{cursor}
     server: SELECT conversations where member = me AND last_activity_at > cursor - 5s
             + membership_tombstones since cursor (if cursor older than 30 d → fullReload: true)
     reply: conversations[{id,lastSeq,unreadCount}], removed[], cursor, fullReload
2. flush outbox (parallel with 1): resend each pending message.send with its clientMsgId
3. for the open conversation (and any with listeners):
     if server.lastSeq > lastSeq[cid]:
        GET /v1/conversations/{cid}/events?after=lastSeq[cid]   (repeat while hasMore)
        410 resync_required → drop window, load latest page
4. buffer live events that arrive during step 3; apply in seq order after; drop seq <= lastSeq
5. other conversations: update badges from step 1; catch up lazily when opened
6. state → connected

Gap handling while connected

  • Event with seq == lastSeq + 1 → apply.
  • seq <= lastSeq → duplicate, drop.
  • seq > lastSeq + 1 → hold, fetch events?after=lastSeq, then apply held events in order.
  • Gap fetch failing 3× → resync_required behaviour.

Presence

  • Online = at least one live socket (Redis set {env}:presence:{user}, TTL 60 s refreshed by heartbeat).
  • Last socket closes → wait 30 s grace → still none → mark offline, persist last_seen_at, notify watchers.
  • Delivered only to sockets that called presence.watch for that user. The SDK watches members of conversations currently rendered.
  • A user can watch only themself and users they share at least one conversation with; other ids in presence.watch are dropped silently (left out of the snapshot, no events), so presence can't be used to find out which user ids exist or to follow strangers.
  • Watch set ≤ 200 per socket.

Typing

  • Client sends typing start at most once per 3 s while typing, stop on send/blur.
  • Server relays to members' live sockets only (no storage, no webhooks, no push).
  • Receivers expire typing state after 6 s without refresh.

Receipts

  • delivered acks are batched by the SDK (≤ 1 frame/s) with the highest contiguous seq per conversation.
  • Server stores GREATEST(last_delivered_seq, seq); broadcasts receipt at most once per second per (conversation, user).
  • read is explicit (app calls markRead); same monotonic rule; broadcast + coalesced message.read webhook.

Scenarios

Scenario Behaviour
Network drop 3 s Send ack timeout (5 s) or socket close → reconnect at ~0.5 s → sync returns few events → outbox resends with same clientMsgId (server returns original ack if committed)
Network drop 5 min Backoff to 30 s cap, immediate retry on "online" → presence went offline after 30 s, typing expired → sync updates all badges, open conversation catches up
App backgrounds RN SDK keeps socket 30 s, then closes with 1000 → push takes over → foreground: reconnect + sync
App killed Server detects 60 s heartbeat silence → offline after grace → unsent messages resent on next launch only if persistence adapter enabled (within 24 h; dedup window 48 h)
Gateway deploy goaway with random 0–10 s → container stops accepting, drains 30 s, closes 1012 → clients reconnect to the new container and sync
Gateway crash Clients reconnect with jitter; dead node's routes expire in 60 s; lost events repaired by sync
Customer DB down No protocol effect; only REST history pages older than the buffer return 503 history_unavailable
Send during reconnection Queued sending; flushed after hello; seq assigned at commit
Duplicate request Same clientMsgId → original ack
Second device Own socket + own cursor; all events to both; read state shared per user
Token expires while connected token_expiring at T−5 min → SDK calls tokenProvider → token.refresh; if missed, server closes 4001 at expiry
User removed from conversation member.removed (self) → SDK drops conversation; server stops routing its events immediately