Local-First Apps with CRDTs: Building Offline Collaboration That Just Works
Cloud software trained us to treat connectivity as a precondition. Every edit waits for a server, every screen assumes a network, and every offline moment becomes an error state. Local-first apps flip that model. They treat the device as the primary place where data lives and works, then sync opportunistically. The hard part is not storing data locally. The hard part is merging changes from many devices without a central coordinator making every decision. That is where conflict-free replicated data types, or CRDTs, come in.
Why Local-First Is a Different Architecture
Local-first software aims for immediate response, offline access, long-term data ownership, and seamless collaboration. But those user-facing benefits come from a deep architectural shift. In a traditional CRUD app, the server is the source of truth. In a local-first app, every replica can accept writes. The system must later reconcile them.
- Offline by default: Users can create, edit, and delete without waiting for a network round trip.
- Multi-device by nature: A user may edit the same document on a phone, laptop, and browser tab at the same time.
- Latency matters: Even 100 milliseconds of round-trip delay breaks the feeling of direct manipulation.
- Privacy and ownership: Data can remain on-device or be encrypted end-to-end, with the server acting as a relay rather than a reader.
- Resilience: Flaky networks, airplane mode, and server outages do not stop work.
The challenge is conflict. If two replicas edit the same field while disconnected, which edit wins? A central server can enforce order, but local-first systems need a merge strategy that works without coordination. CRDTs provide one.
What CRDTs Actually Guarantee
A CRDT is a data type that can be replicated across multiple computers and merged without conflict, as long as all replicas eventually receive the same updates. More formally, CRDTs provide strong eventual consistency. If two replicas have seen the same set of updates, they have the same state, regardless of the order in which those updates arrived.
- Commutative: Applying update A then B gives the same result as B then A.
- Associative: Grouping updates differently does not change the result.
- Idempotent: Receiving the same update twice is safe.
There are two broad families. State-based CRDTs, or CvRDTs, send state or deltas and merge them with a join operation. Operation-based CRDTs, or CmRDTs, send operations and rely on causal delivery or causal context. Hybrid approaches are common in real libraries because they balance bandwidth, latency, and complexity.
Important: CRDTs guarantee convergence, not business correctness. If two users set the same field to different values and the type uses last-writer-wins, one edit is discarded by design. That may be fine for a theme color. It is not fine for a bank balance. The merge semantics must match the domain.
Core CRDT Building Blocks
Counters
A G-Counter only increments. Each replica keeps a per-replica count, and the logical value is the sum. A PN-Counter combines increment and decrement counters. Counters are useful for likes, reactions, view counts, and additive metrics. They are not suitable for inventory when overselling is unacceptable unless you add central validation.
Sets
A G-Set supports add only. A 2P-Set supports add and remove once. An OR-Set, or observed-remove set, allows add and remove with unique tags. If an item is removed and later re-added, the re-add wins. OR-Sets are common for tags, memberships, and feature lists.
Registers
An LWW-Register stores a value with a timestamp or logical clock. Concurrent writes resolve by picking the latest. A multi-value register keeps concurrent values until the application or user resolves them. Use LWW for low-stakes metadata and MV registers when human review is necessary.
Sequences and Text
Sequence CRDTs handle ordered lists and text. Examples include RGA, Logoot, LSEQ, Yjs, and Automerge. Text editing is one of the hardest cases because users expect character-level intent. Inserting a word in the middle of a paragraph while someone else deletes nearby text can produce surprising interleavings if the model is naive. Use battle-tested libraries rather than implementing your own text CRDT.
Maps and JSON Documents
CRDT maps combine nested CRDTs. A Y.Map can hold Y.Text, Y.Array, or primitive LWW values. Automerge models JSON-like documents with history. Choose based on data shape, history requirements, language support, and performance profile.
Choosing a CRDT Model for Your Data
Start from domain semantics, not library APIs. Ask what should happen when two offline users edit the same thing. Then map each field to a merge behavior.
- Counters: likes, reactions, simple metrics. Use G-Counter or PN-Counter.
- Tags and memberships: OR-Set.
- Title or status: LWW-Register if last edit wins is acceptable, MV-Register if you need to prompt.
- Rich text: Y.Text or Automerge text. Do not use naive string concatenation.
- Ordered lists: Y.Array or a CRDT list with stable IDs.
- Relational data: CRDT document per entity plus CRDT indexes. Avoid global cross-entity invariants.
For each field, define the merge. Some fields can be last-writer-wins. Some can be additive. Some may need domain-specific resolution, such as union of tags, maximum score, or manual review. Document these decisions so future developers do not accidentally change semantics.
Designing the Sync Layer
A CRDT is only the replication algorithm. You still need transport, persistence, authentication, discovery, and lifecycle management.
Transport
Use WebSocket for realtime, WebRTC for peer-to-peer, HTTP for periodic sync, and background sync for mobile. The sync layer should be transport-agnostic. It should handle reconnects, duplicates, out-of-order delivery, and partial connectivity without corrupting state.
State Vectors and Deltas
Do not ship the entire document on every change. Each replica tracks a version vector or state vector. A client sends its state vector, and the server or peer replies with missing updates. Delta encoding keeps bandwidth and battery use reasonable. It also lets you resume after long offline periods without replaying everything.
Server Role
In local-first systems, the server is a relay, backup, and discovery service, not the sole source of truth. It can persist opaque encrypted updates, serve blobs, and enforce access. But clients must function offline and merge later. If the server becomes authoritative for every edit, you lose the local-first benefits.
Auth and Access Control
CRDT merge does not know who is allowed to edit. You need signed updates, capability tokens, or per-document encryption keys. For teams, consider end-to-end encryption with key rotation. Remember metadata: who edits what and when can leak sensitive information even if content is encrypted. Access revocation is especially hard because old replicas may still hold keys or unsynced changes.
Schema Evolution
Local-first apps have old clients. Design schema changes as additive. Never reinterpret existing fields. Use versioned migrations that create new CRDT fields and lazily copy data. Unknown fields should be preserved by libraries when possible. If you must remove a field, keep a tombstone or ignore list so old clients do not resurrect it.
Practical Stack: Yjs, Automerge, and Friends
Yjs
Yjs is a mature CRDT framework optimized for realtime collaboration. It has efficient binary encoding, shared types such as Y.Doc, Y.Text, Y.Array, and Y.Map, and awareness for presence. Providers exist for WebSocket, WebRTC, IndexedDB, and more. It is a strong default for text-heavy collaborative editors and realtime apps.
Automerge
Automerge focuses on JSON-like documents with rich history. Its Rust core compiles to WebAssembly and native targets. It provides change history, branching, and merge. It is a good fit when you need document history, offline-first data, and language bindings beyond JavaScript.
Other Options
Loro, Diamond Types, and custom state-based CRDTs are worth evaluating. The right choice depends on document size, history requirements, language, and performance profile. Benchmark with your real workload, not a demo. Test large documents, long offline periods, and many concurrent editors.
Persistence
Persist CRDT updates locally. In browsers, IndexedDB or OPFS works. In native apps, SQLite is excellent. Store updates append-only for durability, then compact periodically. Always test crash recovery, storage limits, and corruption handling. A local-first app that loses local data is worse than a cloud app.
UI Integration
Bind CRDT shared types to your framework. In React, use subscriptions or hooks. In Svelte, use stores. In Vue, use reactivity wrappers. Keep ephemeral UI state, such as open menus and scroll position, outside the CRDT. Only persist data that should sync across devices.
Handling Common Hard Problems
Rich Text and Intentions
Concurrent formatting can conflict. If one user bolds a word while another deletes it, what should happen? Rich-text CRDTs define intentions and marks. Use a library with proven rich-text support, or model formatting as attributes attached to text ranges with explicit conflict rules. Do not assume character-level merge automatically preserves meaning.
Presence and Cursors
Presence is ephemeral. It should not be persisted forever. Use awareness protocols that expire stale peers. Do not store every cursor move in the document history. Presence should be cheap, lossy, and separate from durable state.
Undo and Redo
Undo is tricky in collaborative systems. Local undo should only reverse operations from the local user, not remote edits. Implement undo stacks over operation batches with causal context. Global undo is usually a bad user experience because it can erase someone else’s work.
Garbage Collection and Compaction
Deletes leave tombstones so replicas can converge. Over time, history grows. Compaction is possible only when all replicas have seen certain updates. In practice, use library-supported snapshots and periodic compaction with careful version checks. Monitor tombstone ratios and storage growth.
Large Files
Do not put images, videos, or large binaries directly in a CRDT document. Use content-addressed storage. The CRDT stores a manifest: file ID, hash, size, and metadata. Sync blobs separately with resumable uploads. This keeps document updates small and fast.
Conflict Semantics
Convergence is not enough. Decide business rules. For calendar events, concurrent edits may need manual merge. For a shared shopping list, union and counters are natural. For payments, use a central authority and transactions. CRDTs are a tool for collaboration, not a replacement for all consistency models.
Security and Privacy
End-to-end encryption protects content, but CRDT metadata can reveal structure and timing. Use per-document keys, forward secrecy where practical, and access control lists. If the server cannot read updates, it may not be able to resolve access conflicts for you. Design the threat model before choosing the sync topology.
Case Study: Collaborative Notes App
Imagine a notes app with offline editing and realtime collaboration. The document model could be:
- Root: Y.Doc.
- Title: Y.Text for character-level collaboration.
- Blocks: Y.Array of block maps. Each block has a stable ID, type, and content.
- Block content: Y.Text for paragraphs, Y.Map for metadata.
- Tags: OR-Set or Y.Map with boolean values.
- Presence: awareness protocol with cursor and user color.
When offline, the client writes to local IndexedDB. When online, it connects to a sync server over WebSocket. The client sends its state vector; the server returns missing updates. The server stores encrypted updates in an append-only log and serves snapshots for fast bootstrap. If the server is down, peers can sync directly over WebRTC, then reconcile later. Conflicts in text merge automatically. Conflicts in tags union. Conflicts in title use Y.Text, so both edits interleave by character insertion. This is not always perfect, but it is highly usable.
Testing and Observability
CRDT bugs are often subtle and appear only after partitions, reordering, or duplication. Build a simulation harness.
- Property-based tests: Generate random operations across replicas, apply in random orders, and assert convergence.
- Network simulation: Partition, delay, duplicate, and reorder messages. Run thousands of seeds.
- Invariants: No lost acknowledged local writes, no duplicate effects, causal order preserved for dependent operations.
- Fuzzing: Feed malformed updates to ensure the decoder does not crash or corrupt state.
- Metrics: Sync latency, update size, storage growth, tombstone ratio, convergence failures, and conflict prompts.
In production, log sync sessions with document IDs, but avoid sensitive content. Track when clients are far behind. Alert on storage growth spikes and failed merges. Provide users with clear sync status so they trust the system.
When Not to Use CRDTs
CRDTs are not universal. Avoid them when strong central invariants are required.
- Unique usernames: Two offline users can both claim the same name. A central registry or reservation protocol is better.
- Payments and banking: Double-spend prevention needs authority, transactions, and consensus.
- Inventory with low stock: Concurrent sales can oversell unless you reserve inventory centrally.
- Bidding and auctions: Ordering and finality matter. Use a server-authoritative model.
- Highly dynamic access control: Revocation and permissions are hard to merge without a central authority.
Hybrids work well. Let CRDTs handle drafts, offline edits, and collaboration. Let a server validate and commit critical operations. The best architecture is often a thoughtful mix, not a purist stance.
Deployment and Scaling Patterns
Start simple. A single sync service with WebSocket and object storage can support many users. As you grow, shard by document ID. Use edge relays for low latency. Cache snapshots in CDN or regional stores. Separate blob storage from the update log. Compress updates and use binary protocols. For mobile, batch updates and respect background execution limits. For desktop, run a local daemon or embed SQLite. For peer-to-peer, use relays when NAT traversal fails. Always include a fallback to server sync.
Scaling local-first systems is often about metadata. Track document sizes, update rates, and long-lived offline clients. A single massive document with years of history can become expensive to bootstrap. Use snapshots, compaction, and lazy loading. If a document grows too large, consider splitting it into smaller CRDT documents with explicit references.
Adoption Roadmap
- Step 1: Pick one collaborative surface, such as comments or notes. Do not rewrite the whole product at once.
- Step 2: Model data with CRDT-friendly types. Identify fields that need central authority.
- Step 3: Build a transport-agnostic sync abstraction. Support reconnect, duplicate, and out-of-order updates.
- Step 4: Add local persistence. Test offline for days, not minutes.
- Step 5: Instrument convergence, storage, and sync health. Simulate partitions in CI.
- Step 6: Harden auth, encryption, and key rotation.
- Step 7: Expand to more documents and realtime features. Keep central validation for critical invariants.
Conclusion
CRDTs make local-first collaboration practical. They give users instant response, offline access, and resilient sync without a central bottleneck for every keystroke. But they are not a silver bullet. You must choose the right CRDT types, design merge semantics, secure the sync layer, and test under adversarial network conditions. Do that, and you can build software that feels fast, private, and dependable, even when the network is not.

