Local-First Software: Offline Apps That Sync Without Losing Data
Cloud-first software assumes the network is always available and the server is the source of truth. That assumption breaks on airplanes, in tunnels, on flaky mobile connections, and in privacy-sensitive environments. Local-first software flips the model: the device holds the primary copy of the data, the app works offline by default, and synchronization is a background concern. The result is software that feels instant, respects user ownership, and still supports multi-device collaboration.
This article explains the architecture, data structures, and operational practices behind local-first apps. It covers CRDTs, sync engines, conflict resolution, security, schema evolution, testing, and deployment. The goal is not to sell a single library but to give you a mental model you can apply with tools such as Automerge, Yjs, SQLite, ElectricSQL, Replicache, or a custom sync layer.
What Local-First Really Means
Local-first is not just offline caching. A local-first app treats the on-device database as the working copy and the server as a relay, backup, and collaboration point. The user can read and write without waiting for a round trip. When connectivity returns, changes merge automatically.
- Local primary copy: Every device has a full or partial replica of the data it needs. Reads and writes hit local storage first.
- Offline by default: The app remains fully functional without a network. Sync is opportunistic, not required for every action.
- Multi-device collaboration: Changes propagate between devices and users, often in real time when online.
- Automatic conflict resolution: The system uses CRDTs, operational transformation, or domain-specific merge logic to reconcile concurrent edits.
- User ownership and privacy: Data lives on the user device. End-to-end encryption can keep the server from reading content.
- Long-lived data: The local store can outlive a service. Export and import are first-class features.
The hard part is not storing data locally. The hard part is merging concurrent changes without corrupting user intent. That is where CRDTs and sync engines earn their complexity.
Why Local-First Is Having a Moment
Several trends make local-first practical today. Browsers and mobile devices have fast local databases. WebAssembly lets mature storage engines run in the browser. CRDT libraries have become more efficient and easier to embed. Users expect instant interfaces and offline access. Regulations increasingly favor data minimization and user control. Finally, edge and peer-to-peer networks reduce the cost of relaying changes.
- Powerful local runtimes: SQLite compiled to WebAssembly, IndexedDB, and native mobile databases provide durable storage.
- Mature CRDTs: Libraries such as Automerge and Yjs handle complex merge semantics for text, lists, maps, and counters.
- Sync infrastructure: Services like ElectricSQL, Replicache, and custom WebSocket relays simplify replication.
- Privacy pressure: Local-first designs reduce central data collection and support end-to-end encryption.
- Collaboration expectations: Multiplayer editing is no longer limited to documents. Users want shared tasks, notes, and dashboards.
The Core Architecture of a Local-First App
A local-first app has four main layers: local storage, a mutation or operation log, a sync engine, and a relay or identity layer. Each layer has distinct responsibilities and failure modes.
1. Local Store
The local store is the source of truth for the device. It should support transactions, indexes, and durable writes. Options include SQLite, IndexedDB, Realm, and custom embedded databases. The store must handle schema migrations and provide fast queries for the user interface.
For document-oriented data, you might store JSON blobs with indexed fields. For relational data, you might use SQL tables and translate CRDT operations into row changes. The key is that the UI reads from the local store, not from the network.
2. Mutation Log and CRDT State
Every user action becomes an operation or mutation. Examples include setting a field, inserting a list item, incrementing a counter, or deleting a record. The log records these operations with metadata: device ID, logical timestamp, causal dependencies, and sometimes a hash.
In a CRDT system, the log is applied to a replicated data type. State-based CRDTs periodically exchange full or delta states. Operation-based CRDTs exchange operations that must be delivered causally. Hybrid approaches use operation logs with state snapshots for fast catch-up.
3. Sync Engine
The sync engine moves changes between devices and the relay. It handles connection management, authentication, retries, batching, compression, and conflict detection. It also tracks what each peer has seen using version vectors or sync tokens.
A good sync engine is idempotent. Receiving the same operation twice must not duplicate effects. It is also incremental: it sends only what the peer is missing. When the network drops, the engine queues changes and resumes later.
4. Relay and Identity Layer
The relay stores encrypted operations or snapshots, authenticates users and devices, and fans out changes to connected peers. It should not need to understand application semantics if end-to-end encryption is used. Identity can be based on public keys, OAuth, or a hybrid model.
Some local-first apps use peer-to-peer transport such as WebRTC for direct device sync. Others use a simple WebSocket relay. The relay can also provide backup, search indexing over encrypted metadata, and invitation workflows.
Conflict Resolution: CRDTs, OT, and Pragmatic Merges
Conflict resolution determines whether local-first feels magical or maddening. CRDTs are a family of data structures that merge automatically and deterministically without coordination. Operational transformation (OT) is an alternative used by many collaborative editors. Domain-specific merges can also work well when the data model is simple.
- Last-write-wins (LWW): A register stores the value with the highest timestamp. Simple but can lose concurrent updates. Use only when losing one update is acceptable.
- Multi-value register: Keeps all concurrent values until a later write resolves them. Good for fields where conflicts matter and a user should choose.
- Counters: Increment and decrement operations merge by summing per-replica counts. Useful for likes, inventory, and analytics.
- Sets: Add-wins or remove-wins sets track element presence. Add-wins sets are common for tags and memberships.
- Sequences and rich text: CRDTs such as RGA, Logoot, and Yjs handle ordered lists and text. They preserve intent when multiple users insert or delete nearby characters.
CRDTs trade storage and metadata for coordination-free merging. They can grow tombstones and version vectors. Compaction, snapshots, and garbage collection are necessary for long-lived documents. OT requires a central server or a more complex peer protocol but can be efficient for text editing.
In practice, many apps combine approaches. A document might use a CRDT for text, LWW registers for settings, and a custom merge for business rules. The important thing is to define merge semantics before coding the UI.
Sync Engine Internals
A robust sync engine is more than a WebSocket that broadcasts JSON. It must handle causality, partial replication, and network partitions. Here are the core concepts.
- Causal ordering: Operations that depend on each other must be applied in order. A reply to a comment cannot arrive before the comment. Causal delivery ensures that dependencies are satisfied.
- Version vectors: Each device tracks the latest sequence number it has seen from every other device. Comparing vectors reveals missing operations and concurrent changes.
- Tombstones and compaction: Deletions often leave markers so that old operations do not resurrect data. Tombstones must eventually be garbage-collected, but only when all peers have seen the deletion.
- Snapshots and delta sync: New devices should not download the entire operation history. The relay provides a snapshot plus later deltas. Snapshots also reduce startup time and storage.
- Backpressure and batching: Mobile devices and browsers have limited memory. The sync engine should batch small operations, compress payloads, and avoid unbounded queues.
- Idempotency and deduplication: Every operation needs a unique ID. Receiving it twice should be a no-op. This protects against retries and duplicate delivery.
Logging and metrics are essential. Track sync lag, conflict rate, operation size, queue depth, and failed merges. These metrics reveal whether the data model or network layer needs work.
Storage and Library Choices
The local-first ecosystem is growing. Choose based on data model, platform, and how much control you need.
- Automerge: A JSON-like CRDT library with rich history and merge semantics. Good for documents, collaboration, and offline-first apps. It supports binary encoding and sync protocols.
- Yjs: A high-performance CRDT for shared editing. Excellent for text, rich text, and complex nested structures. It has a mature ecosystem for editors and providers.
- SQLite and SQLite WASM: A reliable relational store that runs on mobile, desktop, and the browser. You can build a sync layer on top or use it as the local cache for a CRDT engine.
- RxDB, PouchDB, WatermelonDB: Local databases with replication features. They are useful for offline-first mobile and web apps, though their conflict models may be simpler than full CRDTs.
- ElectricSQL, Replicache, TinyBase: Tools that bring local-first patterns to SQL or key-value data. ElectricSQL focuses on Postgres-to-local sync. Replicache provides a mutation and sync framework. TinyBase is a reactive local store with synchronization options.
Do not choose a library only because it is popular. Prototype the hardest conflict scenario first. If the library cannot express your merge rules, you will fight it later.
Security, Privacy, and Access Control
Local-first changes the threat model. Data is spread across devices, and the server may be untrusted. Security must cover encryption, key management, and authorization.
- End-to-end encryption: Encrypt operations or document snapshots before they leave the device. The relay stores ciphertext and cannot read content. This protects against server breaches and curious operators.
- Key management: Users and devices need keys. Options include per-document keys, per-user key pairs, and group key agreement. Key rotation and recovery are hard problems that must be designed early.
- Metadata leakage: Even with encrypted content, the relay may see document IDs, access patterns, timestamps, and sizes. Use padding, batch uploads, and anonymous credentials where needed.
- Access control and revocation: In a CRDT, removing a user does not remove the data they already have. Revocation must combine key rotation, access control lists, and policy enforcement. For sensitive data, consider forward secrecy and document re-encryption.
- Device trust: A lost device is a data breach. Support remote wipe of local keys, device approval flows, and session limits.
Privacy is a feature, not a checkbox. Local-first apps can advertise that user data stays on device and syncs encrypted. That claim must be backed by cryptographic design, not just marketing.
Schema Evolution Across Devices
In a cloud app, you migrate the database once. In a local-first app, devices may be offline for months. Old clients and new clients must coexist. Schema evolution becomes a distributed systems problem.
- Versioned documents: Include a schema version in every document or operation. The sync engine can route versions or apply lazy migrations.
- Backward and forward compatibility: New fields should be optional. Old clients should ignore unknown fields. New clients should tolerate missing fields.
- Lazy migrations: Migrate data on read or when the device comes online. Avoid blocking the UI on a full migration.
- CRDT schema constraints: Changing a CRDT type can break merge semantics. For example, converting a list to a set is not a simple migration. Plan data model changes carefully.
- Migration coordination: For breaking changes, use a feature flag, a new document version, or a server-assisted upgrade path.
Document your schema evolution strategy before launching. The cost of a bad migration grows with every offline device.
Testing, Observability, and Reliability
Local-first apps fail in ways that are hard to reproduce. Networks partition, clocks skew, and users edit the same data from multiple devices. Testing must simulate these conditions.
- Deterministic simulation: Run multiple virtual devices with controlled network delays, drops, and reordering. Assert that all replicas converge.
- Property-based testing: Generate random operations and verify invariants such as no lost updates, no duplicate effects, and eventual consistency.
- Fuzz testing: Feed malformed operations and snapshots to the sync engine. Ensure it fails safely and does not corrupt local data.
- Conflict dashboards: Track where conflicts occur and how they are resolved. High conflict rates often indicate a poor data model.
- Replay and audit logs: Store operation histories so you can replay a user session and debug merge issues.
- Observability: Measure sync latency, queue depth, error rates, and device versions. Alert on stuck syncs and growing tombstones.
Reliability also means graceful degradation. If sync is down, the app should still work. If the relay is unavailable, the user should not lose data. If a merge fails, the app should quarantine the operation and ask for help rather than silently drop it.
Deployment Patterns
Local-first does not dictate a single topology. You can choose from several patterns based on collaboration needs and privacy requirements.
- Pure peer-to-peer: Devices sync directly over WebRTC or local networks. No central server. Great for privacy, but discovery, backup, and offline delivery are harder.
- Relay server: A simple server stores and forwards encrypted operations. It provides backup and asynchronous sync. This is the most common pattern.
- Managed sync service: A vendor handles storage, identity, and conflict infrastructure. Faster to build, but you trade control and may pay per operation.
- Self-hosted sync: Run your own relay for compliance or cost reasons. You need operational maturity for scaling and backups.
- Hybrid edge: Use edge functions to reduce latency and regional data residency. The relay can run close to users while keeping encrypted blobs.
Whatever the topology, design for offline first. The network is an optimization, not a requirement.
Performance Tuning
Local-first apps often feel fast because reads are local. But sync and merge can become bottlenecks if ignored.
- Index local queries: The UI should never scan all documents. Index fields used for sorting, filtering, and search.
- Incremental sync: Send deltas, not full documents. Use binary encoding for CRDT updates when possible.
- Compression and batching: Compress operation batches and avoid one-message-per-keystroke. Yjs and Automerge updates can be merged before sending.
- Snapshot cadence: Create snapshots periodically to bound replay time. Balance snapshot size against sync speed.
- Garbage collection: Remove tombstones and old operations once all peers have seen them. Otherwise documents grow forever.
- UI virtualization: Large local datasets still need virtualized lists and pagination. Local does not mean infinite memory.
Profile on low-end devices, not just your development machine. Mobile CPUs and storage are the real constraints.
When Not to Go Local-First
Local-first is powerful, but it is not universal. Avoid it when strong consistency or central authority is non-negotiable.
- Financial ledgers and inventory: Double-spending and overselling require coordination. CRDTs can help, but you often need a central authority or consensus.
- Regulatory audit trails: Some regulations require a single immutable source of truth. Local-first can still work, but the relay must provide an authoritative log.
- Simple CRUD apps: If the app is always online and conflicts are rare, a traditional server API is simpler and cheaper.
- Large binary media: Syncing large files between devices is expensive. Use object storage and references instead of replicating blobs.
- Teams without distributed systems skills: CRDTs and sync engines add complexity. Start with a simpler offline cache if you cannot invest in the data model.
The decision is about user experience and data ownership. If offline access, instant interaction, and privacy are core, local-first is worth the complexity.
Implementation Blueprint
If you are starting a local-first app, follow a staged approach.
- Define the data model and merge semantics. Decide which fields are LWW, which are CRDTs, and which require custom merge logic. Write down conflict scenarios.
- Choose a local store and CRDT library. Prototype the hardest collaboration case. Measure document size, merge speed, and memory use.
- Build the local mutation pipeline. Every UI action writes to the local store and appends an operation. The UI reads only from local state.
- Implement the sync engine. Add authentication, version vectors, batching, retries, and idempotency. Test network partitions early.
- Add encryption and access control. Decide what the relay can see. Implement key management and device revocation before launch.
- Plan schema evolution. Version documents, support lazy migrations, and test old and new clients together.
- Instrument and simulate. Build a deterministic test harness with multiple virtual devices. Track sync lag and conflict rates in production.
- Operate and iterate. Monitor snapshots, tombstones, and storage growth. Tune compaction and batching as usage grows.
Conclusion
Local-first software is a response to the fragility of cloud-only assumptions. It puts the user device at the center, keeps the app usable offline, and treats sync as a merge problem rather than a request-response problem. The architecture is demanding: CRDTs, version vectors, encryption, schema evolution, and testing all require care. But the payoff is software that feels immediate, respects privacy, and keeps working when the network does not.
Start small. Pick one collaborative feature, model the conflicts, and prove that your sync engine converges under partition. Once that works, expand the local-first surface area. The result is not just offline support. It is a different contract with users: your data is yours, and the software works for you even when the cloud is unreachable.

