CRDTs in Production: Real-Time Collaboration Without Conflicts

CRDTs in Production: Real-Time Collaboration Without Conflicts

CRDTs in Production: Real-Time Collaboration Without Conflicts

Real-time collaboration looks simple from the outside: two people type, move a card, or edit a cell, and everyone sees the change. Under the hood, every keystroke is a distributed systems problem. Clients can be offline, networks can reorder messages, and multiple users can edit the same object at the same time. The hard part is not sending updates; it is merging them without losing intent or corrupting state.

Conflict-free replicated data types (CRDTs) are a family of data structures designed for this exact problem. They let replicas accept local updates independently and still converge to the same state once they exchange all updates. CRDTs are not a silver bullet, but they are the foundation of many production collaborative editors, whiteboards, and local-first apps. This guide covers the data models, architecture, scaling, security, and operational practices needed to run CRDTs in production.

Why concurrency is the core problem

A collaborative document is a replicated state machine. Each client has a copy, applies local edits immediately, and eventually syncs with others. The challenge is defining what happens when two replicas receive different updates in different orders.

The simplest strategy is last-write-wins (LWW). Each field stores a value and a timestamp, and the highest timestamp wins. LWW is easy to implement, but it silently discards concurrent edits. If two users edit different parts of the same object, or if clocks are skewed, LWW can erase valid work. It is acceptable for status fields and preferences, not for text or rich content.

Another common approach is operational transformation (OT). OT transforms operations against concurrent operations so they can be applied in a consistent order. It powers some mature editors, but it usually requires a central server to define a canonical order and a complex transformation matrix for every operation pair. As features grow, OT transformations become difficult to reason about and test.

  • Centralized ordering keeps OT simpler but creates a coordination bottleneck and weak offline support.
  • Transform complexity grows with the number of operation types, especially for rich text, tables, and comments.
  • Offline edits force OT systems to rebase long operation histories, which is fragile.

CRDTs take a different route. Instead of transforming operations at a central point, each replica stores enough metadata to merge updates deterministically. The guarantee is strong eventual consistency: if two replicas have seen the same set of updates, they have the same state. Convergence does not require a central coordinator, though production systems still use servers for persistence, discovery, and fan-out.

The CRDT toolbox: registers, counters, sets, maps, and sequences

CRDTs are composable. You rarely implement one giant CRDT for an entire document. Instead, you build a document tree from smaller types. Understanding the core types helps you choose the right model and debug convergence issues.

Registers

A register holds a single value. A last-write-wins register uses a timestamp or logical clock to pick the latest write. It is compact and fast, but concurrent writes lose all but one value. A multi-value register keeps all concurrent values until a subsequent write resolves them, which is useful for exposure to conflicts in forms or settings.

Counters

A grow-only counter (G-Counter) tracks increments per replica and sums them. A positive-negative counter (PN-Counter) combines two G-Counters, one for increments and one for decrements. Counters are ideal for likes, votes, inventory adjustments, and any metric where addition is commutative. They are not appropriate for balances that must never go negative without additional invariants.

Sets

A grow-only set only adds elements. An observed-remove set (OR-Set) supports add and remove by tagging each add with a unique identifier. A remove operation removes the tags it has observed. Concurrent add and remove can result in the element being present, which is usually the safer default. A last-write-wins element set uses timestamps per element and is simpler but less precise.

Maps

A map is a key-value store where keys are registers or nested CRDTs. Concurrent updates to different keys merge naturally. Concurrent updates to the same key depend on the value CRDT. In production, maps are the backbone of document structure: a document is often a map of metadata, a text CRDT for content, and a map of comments.

Sequences

Text, lists, and ordered collections require a sequence CRDT. Sequence CRDTs assign each element a unique position identifier that supports insertion between any two elements without renumbering. Common designs include RGA, YATA, and Fugue. Libraries such as Yjs and Automerge implement optimized sequence CRDTs for text and rich text.

Sequence CRDTs are the most complex and metadata-heavy part of the toolbox. Every character or block carries an identifier, and deleted content often becomes a tombstone. Production implementations use compression, run-length encoding, and garbage collection to keep documents manageable.

Choosing between OT, CRDTs, and server serialization

CRDTs are not the only option. The right choice depends on latency, offline requirements, conflict semantics, and team expertise.

Approach Best for Trade-offs
Server serialization Low-concurrency forms, financial ledgers, strict invariants Simple and authoritative, but poor offline support and higher latency
Operational transformation Text editors with a central server and stable operation set Good latency and compact history, but complex transformations and weak offline support
CRDTs Offline-first apps, multi-user editors, whiteboards, local-first software Automatic merge and peer-to-peer friendly, but metadata overhead and eventual consistency

A practical pattern is to use CRDTs for user-facing content and server-side invariants for critical operations. For example, a document body can be a CRDT, while billing, permission changes, and irreversible actions go through an authoritative API that validates rules before writing.

Production architecture for CRDT-backed collaboration

A production CRDT system is more than a library. It needs a sync protocol, persistence, presence, authorization, and observability. The following architecture is common for web and mobile collaboration apps.

Client model

Each client keeps a local CRDT document, usually in memory with a durable cache such as IndexedDB, SQLite, or a file store. Local edits apply immediately to the CRDT and update the UI. The client also maintains a state vector or version vector that summarizes which updates it has already seen.

When the client connects, it sends its state vector to the server. The server responds with only the missing updates. This delta sync is critical for performance; sending the full document on every reconnect does not scale.

Sync server

The sync server has three jobs: authenticate clients, relay updates, and persist the document history. It does not need to understand the application semantics of every operation, but it must enforce room membership and rate limits.

  • WebSocket is the default transport for low-latency bidirectional updates.
  • WebRTC data channels can reduce server fan-out for peer-to-peer sessions, but most production apps still use a server for reliability and persistence.
  • HTTP long polling or SSE can work for read-heavy or restricted environments, but add latency.

The server should treat CRDT updates as opaque binary blobs where possible. This keeps the server generic and lets the client library evolve independently. For authorization, the server validates the user against the room before broadcasting. For stronger guarantees, updates can be signed or grouped into authenticated batches.

Persistence and compaction

CRDT updates are append-only by nature. Persisting every update forever is simple and robust, but storage grows without bound. Production systems combine two strategies:

  1. Update log: append each update with a sequence number and document id. This supports replay and audit.
  2. Snapshots: periodically encode the full CRDT state and store it as a compact binary. New clients can load the snapshot and then apply only newer updates.

Compaction is the process of replacing a long update history with a snapshot plus a short tail. It must be done carefully. If a client is offline for a long time, it may need the full history or a snapshot plus updates since its state vector. Many libraries provide state as update and diff update APIs to generate the minimal payload for a given state vector.

Tombstones are another storage concern. Deleted characters and list items leave metadata so that concurrent inserts can still be ordered correctly. Garbage collection can remove tombstones that are older than any possible concurrent update, but this requires knowing the minimum state vector across all replicas or accepting that very old offline clients will not merge correctly.

Presence and awareness

Presence data such as cursors, selections, and online status is ephemeral. It does not need the same durability or convergence guarantees as the document. Mixing presence into the CRDT adds noise and metadata bloat.

A common pattern is an awareness channel separate from the document sync. Each client periodically broadcasts its cursor position, user id, color, and timestamp. The server relays these messages and expires stale entries after a timeout. If an awareness message is lost, the next one corrects it. This keeps the document CRDT focused on persistent content.

Auth and permissions

CRDTs solve merge, not access control. The sync server must enforce who can read and write a document. For fine-grained permissions, you can use separate rooms or document sections. For enterprise compliance, log every update with user id and timestamp, and validate permissions before accepting writes.

If you need end-to-end encryption, encrypt updates on the client before sending them to the server. The server then stores and relays ciphertext. This improves privacy, but it complicates server-side compaction, search, and moderation. You will need a key management strategy and a way for new devices to obtain the document key.

Text editing and rich content

Text is the hardest common CRDT use case. Users expect character-level merging, stable cursors, undo/redo, formatting, comments, and suggestions. A plain sequence CRDT handles basic text, but rich text requires a model for formatting spans.

There are two main approaches:

  • Mark-based formatting: formatting is stored as marks applied to ranges. Overlapping marks are merged by priority. This is flexible but can be complex to render and edit.
  • Tree-based formatting: content is a tree of blocks, inline nodes, and text leaves. Formatting is structural. This maps well to rich text editors but can produce more complex merge behavior.

Cursor positions must also be represented as CRDT references, not raw indices. If a remote user inserts text before your cursor, a raw index would point to the wrong character. A relative position or cursor CRDT keeps the cursor attached to the intended character or gap.

Comments and suggestions are usually modeled as separate CRDT maps or lists that reference stable content ids. A comment thread might anchor to a range using relative positions. When the underlying text changes, the anchor updates according to the same merge rules.

Offline-first behavior and multi-device sync

Offline-first is one of the strongest reasons to use CRDTs. A user can edit on a plane, on a train, or in a building with bad Wi-Fi, and the app keeps working. When the device reconnects, it syncs automatically.

To make this reliable:

  • Persist the CRDT document locally after every update or batch of updates.
  • Queue outgoing updates in a durable outbox.
  • Track the server state vector so you can request only missing updates.
  • Handle duplicate updates idempotently; CRDT merge should be idempotent.
  • Use exponential backoff and network status events to retry sync without draining the battery.

Multi-device sync introduces the same concurrency problems as multi-user collaboration. A user editing the same document on a laptop and phone can create concurrent changes. CRDTs merge them, but UX may still need to highlight recent changes or provide a history view.

Scaling CRDTs: performance, storage, and fan-out

CRDTs have overhead. Every character, list item, or map key carries metadata. A naive implementation can produce documents that are orders of magnitude larger than the visible content. Production performance work focuses on reducing metadata, update size, and fan-out cost.

Update size

Use binary encoding instead of JSON. Libraries like Yjs and Automerge use compact binary formats with variable-length integers and columnar encoding. Batch small edits into a single update when possible. Avoid sending a full state vector on every keystroke; send deltas.

Document size

Large documents accumulate history. Snapshots and compaction keep load times reasonable. For very large documents, consider splitting into subdocuments or blocks. A document can be a CRDT map of block ids, where each block is itself a CRDT. This enables lazy loading and partial sync.

Fan-out

In a room with many active users, the server broadcasts every update to every participant. This is simple but can become expensive. Strategies include:

  • Sharding by room so each document is handled by a single server or a small cluster.
  • Pub/sub backplane such as Redis, NATS, or Kafka to distribute updates across nodes.
  • Interest management to send updates only for visible sections in large documents.
  • Backpressure to slow down clients that produce too many updates or cannot keep up.

Memory

Servers that keep full CRDT documents in memory for fast sync need enough RAM for the largest active rooms. For cold documents, store snapshots and update logs in object storage or a database and load on demand. Monitor memory per room and set limits to avoid a single runaway document taking down a node.

Security, privacy, and compliance

Security in a CRDT system has several layers:

  • Transport security: always use TLS for WebSocket and HTTP traffic.
  • Authentication: verify user identity with tokens, sessions, or mutual TLS.
  • Authorization: enforce room-level and document-level permissions on the server.
  • Update integrity: sign updates or batches if clients cannot be trusted.
  • Rate limiting: prevent malicious clients from flooding a room with updates.

Privacy is harder. If the server stores plaintext updates, it can read document content. End-to-end encryption solves this but shifts complexity to clients. You must manage group keys, handle device revocation, and decide whether the server can compact encrypted history. Some systems use client-side compaction and upload encrypted snapshots, but this requires a trusted client to perform the merge.

Compliance adds deletion requirements. CRDT tombstones and history can conflict with the right to erasure. A practical approach is to separate personal data from document content, use per-user encryption keys, and design deletion as cryptographic erasure of the user key rather than rewriting every CRDT. This is an architectural decision, not an afterthought.

Testing, observability, and debugging

CRDT bugs are subtle. They often appear only after specific network reordering, concurrent edits, or long offline periods. Testing must go beyond unit tests.

Property-based testing

Generate random sequences of operations, apply them to multiple replicas in different orders, and assert that all replicas converge to the same state. This is the most effective way to catch merge bugs. Libraries often provide test harnesses for this.

Network simulation

Simulate partitions, delays, duplicate messages, and message reordering. Tools like Toxiproxy, network emulators, or custom test clients can help. Verify that the system recovers and converges after the network heals.

Observability

Track metrics that reveal CRDT health:

  • Update size and frequency per room.
  • Snapshot size and compaction duration.
  • Sync latency from client send to server ack.
  • Number of pending updates per client.
  • Conflict indicators, such as multi-value registers that remain unresolved.
  • Memory and CPU per room.

For debugging, store enough context to replay a session: update ids, client ids, state vectors, and snapshots. A deterministic replay tool can reproduce a convergence bug from logs.

A practical implementation recipe

If you are building a CRDT-backed app, start small and keep the server generic. Here is a production-oriented recipe:

  1. Choose a CRDT library that matches your platform. Yjs is strong for JavaScript and text. Automerge works well for JSON-like documents. Evaluate binary size, sync protocol, persistence APIs, and license.
  2. Model the document as a CRDT map with nested types. Use a sequence for text or ordered lists. Use registers for metadata. Keep ephemeral presence out of the document.
  3. Implement local-first storage so every edit is durable before or soon after applying it to the UI.
  4. Build a sync server that authenticates users, relays binary updates, stores update logs, and periodically creates snapshots.
  5. Use state vectors for delta sync. Never send the full document unless the client is new or its state is too old.
  6. Add an awareness channel for cursors, selections, and online status. Expire stale entries aggressively.
  7. Define permissions at the room or document level. Validate every write before broadcasting.
  8. Instrument metrics for update size, sync latency, room memory, and compaction time. Set alerts for runaway documents.
  9. Test with property-based and network simulation. Make convergence a testable invariant, not a hope.
const doc = new Y.Doc();
const text = doc.getText('content');
text.observe(event => {
  const update = Y.encodeStateAsUpdate(doc, remoteStateVector);
  socket.send(update);
});

The snippet above shows the core loop: apply local changes, observe the CRDT, encode only the missing state, and send it over the transport. In a real app, you would add authentication, batching, persistence, and error handling around this loop.

When CRDTs are the wrong tool

CRDTs are powerful, but they are not universal. Avoid them when:

  • You need strict global invariants. CRDTs favor availability and convergence over immediate consistency. Bank balances and inventory reservation often need a central authority.
  • Conflicts are rare and trivial. A simple form with one editor at a time can use server serialization or optimistic locking with far less complexity.
  • Document size is small and online-only. If users are always connected and the server can order operations, OT or plain server-side merging may be simpler.
  • You cannot afford metadata overhead. Very large binary files, video timelines, and huge spreadsheets may need a different architecture, such as command logs or block-level collaboration.
  • Your team lacks distributed systems experience. CRDTs reduce merge logic but introduce new failure modes around sync, compaction, presence, and debugging. Budget time for learning and observability.

Key takeaways

  • CRDTs provide automatic merge and strong eventual consistency, making them a strong fit for offline-first and real-time collaboration.
  • Use small composable types: registers, counters, sets, maps, and sequences. Do not build one monolithic CRDT.
  • Separate persistent document state from ephemeral presence. Mixing them adds noise and cost.
  • Production architecture needs delta sync, snapshots, compaction, authorization, and observability.
  • Text and rich content are the hardest cases. Use relative cursor positions and a thoughtful formatting model.
  • Scale by sharding rooms, using binary updates, compacting history, and applying backpressure.
  • Test convergence with random operations and simulated network failures. If you cannot reproduce a merge, you cannot fix it.

CRDTs turn collaboration into a data-structure problem, but production readiness comes from the systems around them. Start with a clear model, keep the server generic, measure everything, and treat convergence as a feature you continuously verify. With that discipline, you can build collaborative apps that feel instant, survive offline edits, and avoid the classic conflict nightmares that plague naive real-time systems.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *