CelinQ Insights · No. 29

Backups and disaster recovery you don't have to think about

Making the loss of a laptop a non-event for the model.

A NILUS perspective on collaborative modelling for Sparx Enterprise Architect

Ask an architect what would happen if their laptop were stolen tomorrow, and watch the pause. It is not that they have no answer; it is that the answer is usually a sequence of half-measures held together by hope. There is a copy of the repository on a shared drive, though nobody is quite sure how recent it is. There was an export sent by email a few weeks ago, attached to a status update, that could probably be recovered from a sent-items folder. The client has an old version from the last formal delivery. Somewhere in this collection of fragments there is enough to reconstruct most of the model, and the reconstruction would take an unpleasant few days, and some of the recent work would simply be gone. The architect knows all of this, and they have made peace with it, because the alternative is to think about it seriously and thinking about it seriously is uncomfortable.

This is the quiet scandal of how architecture models are actually protected. The model is frequently the single most valuable intellectual asset a practice holds. It represents months or years of accumulated understanding, the distilled result of countless workshops and decisions and corrections. And it very often lives, in its most current form, on exactly one machine: the laptop of the person doing the work. Everything else is a copy of varying staleness. The protection of this asset is not a system; it is a habit, and habits fail precisely when they are most needed, which is to say on the bad day when the laptop is dropped, stolen, encrypted by ransomware, or simply fails to turn on one morning with no warning at all.

Why the ordinary approaches quietly fail

It is worth being precise about why the usual protections are weaker than they look, because the weaknesses are structural rather than accidental. Consider the manual export. An architect who periodically exports the repository to a file and stores it somewhere safe has, in principle, a backup. But the export is only as recent as the last time they remembered to do it, and the discipline of remembering degrades exactly as the work gets busy, which is the same time the work becomes most valuable. A backup regime that depends on a human remembering to perform a chore is a backup regime that is quietly out of date most of the time. The gap between the last export and the moment of disaster is pure lost work, and the busier and more productive the architect has been, the larger that gap tends to be.

Consider the shared network drive. Placing the repository file on a drive that the organisation backs up feels responsible, and it is better than nothing, but it inherits all the problems of file-based collaboration on top of its backup weaknesses. The file on the drive is only current if the architect is diligently saving to it, which they cannot do while also editing it at full speed, because the file lock forces a choice between working and sharing. In practice the drive copy is often a periodic save rather than a live one, which means it too has a staleness gap. And a backup of a single file is a backup of a single point in time; it captures the state of the model but not the sequence of changes that produced it, so a recovery restores a snapshot without the story, and often without the most recent chapters.

A backup that depends on someone remembering to make it is not a safety net. It is a promise to have been diligent, tested only on the one day diligence turns out not to have been enough.

Consider even the conscientious version, where the organisation's IT function backs up the architect's whole machine. This is genuinely useful and every serious organisation should do it, but it protects the machine, not the model specifically, and it restores the machine's state at whatever interval the backup runs. If the model changed significantly between the last machine backup and the disaster, that change is gone. More subtly, a machine-level restore gives you back a repository file, but it gives you back one architect's copy in isolation. If the practice has several people, restoring one laptop does not reconcile that person's recovered work with everything their colleagues have done since. The backup protected a device; it did not protect the shared, living, collaborative truth that the practice actually depends on.

The common thread through all of these is that they treat the model as a file to be copied, and they treat protection as an activity someone performs on that file from time to time. As long as protection is an activity, it competes with all the other activities for attention, and it loses that competition on exactly the days when it matters most. The only durable protection is one where the model is protected as a consequence of ordinary work, without anyone having to decide to protect it. That is a structural property, not a diligence property, and structural properties are the only ones that hold up on the bad day.

The shape that makes protection automatic

CelinQ's local-first design produces exactly that structural property, and it does so as a side effect of how collaboration works rather than as a bolt-on feature. Each architect edits their own local repository at full speed, and a background companion syncs each save with a shared workspace running on a server the organisation runs itself. Read that as a statement about backups rather than about collaboration and it says something quietly powerful. Every save an architect makes is reconciled, in the background, into a shared workspace that lives somewhere other than the architect's laptop. The protection is not a chore the architect performs; it is the ordinary consequence of them doing their work and saving it, which they were going to do anyway.

This inverts the relationship between work and safety. In the old world, the architect worked and then, separately and later and unreliably, tried to protect the work. In the local-first world, protecting the work is not a separate step at all. The moment a save is reconciled into the shared workspace, the model no longer lives in only one place. It lives on the architect's machine and it lives in the shared workspace on the organisation's server, and those two locations are kept in step continuously as the work proceeds. The staleness gap that haunts every manual backup regime is not reduced; it is largely designed away, because there is no interval during which the model exists only on the laptop and nowhere else waiting for someone to remember to copy it.

The consequence for the stolen-laptop scenario is the point of this whole piece. When the laptop is gone, the model is not gone. The shared workspace on the organisation's server holds the reconciled work, and a workspace can be cloned to a fresh local repository. The architect is issued a new machine, clones the workspace, and has the model back on their own machine, ready to work in at full speed, with the history intact. What used to be a multi-day reconstruction from fragments becomes an afternoon of clone-and-continue. The loss of the device is a real inconvenience and a real security event that has to be handled on its own terms, but the loss of the device is no longer the loss of the model. The two have been decoupled, which is exactly what disaster recovery is supposed to achieve and exactly what the file-on-a-laptop arrangement never could.

History is a kind of backup that snapshots cannot be

There is a dimension of recovery that snapshot backups miss entirely, and it deserves its own attention because it is where a lot of real-world pain actually lives. Most disasters are not the dramatic stolen laptop. Most disasters are quiet and internal: a restructuring that went wrong, a bulk change that turned out to be a mistake, a package deleted in haste, a well-intentioned edit that broke something three relationships away and was not noticed for a fortnight. For these, a single snapshot backup is almost useless, because by the time anyone realises the model is wrong, the snapshot has already captured the wrong state. You cannot restore your way out of a mistake you did not notice in time.

CelinQ keeps a complete, ordered revision history of the model, and this changes what recovery even means. Because every change is recorded in sequence, recovering from the quiet internal disaster is not a matter of restoring a snapshot and losing everything since; it is a matter of understanding what changed, when, and by whom, and being able to reason about the model as it was at any point along the way. The history is not a marketing feature about traceability; it is a form of backup that operates in the time dimension rather than only the space dimension. A snapshot protects you against losing the model. An ordered history protects you against the model becoming wrong, which is the more common and more insidious failure, and it does so with a granularity that a periodic file copy cannot approach.

The disaster people plan for is the building burning down. The disaster that actually happens is a bad afternoon two weeks ago that nobody caught. A complete revision history is the only kind of backup that helps with the second one.

This also matters for the peculiar recovery need that architecture practices have and that generic backup tools do not serve well: the need to answer questions about the past. When a decision made six months ago is challenged, when a regulator asks what the model showed at the time a particular commitment was made, when a colleague needs to understand why an element was structured the way it was, the answer lives in the history. A practice that has only the current state and a handful of snapshots cannot answer these questions confidently. A practice with a complete ordered history can, and it can do so because the history was accumulated automatically as a consequence of ordinary work rather than assembled painfully after the fact from exports and emails.

The server is the thing you actually back up

All of this raises the obvious question, and it is a fair one to ask of any tool that talks confidently about safety. If the model lives in a shared workspace on a server, what protects the server? Have we not simply moved the single point of failure from the laptop to the server? The answer is that we have moved it somewhere the organisation can actually protect it properly, using the disciplines the organisation already has, and that is precisely the improvement rather than a sleight of hand.

The shared workspace keeps its data in an ordinary database, SQLite for a small team or a single server, or PostgreSQL where an organisation wants the operational maturity of a proper database engine. Both are technologies that any competent infrastructure team already knows how to back up, replicate and monitor. This is a deliberately unglamorous claim and it is the whole point. There is no proprietary datastore that only the vendor understands, no opaque backup format that ties recovery to a support ticket. The organisation's own database administrators can fold the shared workspace into the same backup regime, the same replication, the same disaster-recovery plan that already protects everything else the organisation considers important. The model stops being a special case protected by an architect's good habits and becomes ordinary infrastructure protected by the organisation's ordinary discipline.

This is a far stronger position than the laptop ever offered, and the reason is not that servers are magically safer than laptops. It is that servers are things organisations already know how to protect, whereas the working repository on an architect's machine was always an orphan that fell outside the organisation's real recovery planning. By moving the authoritative shared workspace onto a server the organisation runs itself, the model comes inside the perimeter of the organisation's existing operational maturity. Nightly backups, off-site replication, tested restores, retention policies: all the things a serious infrastructure team already does for the databases it cares about now cover the architecture model too, and they cover it without the architects having to do anything except their ordinary work.

There is a redundancy here that is worth naming because it is genuinely reassuring. In a practice with several architects, the model exists in many places at once: a full local copy on each architect's machine, and the authoritative shared workspace on the server. If the server itself were lost and its backups somehow failed, the model would still exist, in near-current form, on every architect's laptop, because each of those local repositories is a real and complete copy rather than a thin client with nothing of its own. If a laptop is lost, the server has it. If the server is lost, the laptops have it. The model is protected by the same distribution that makes the collaboration work, and that distribution is not a special disaster-recovery mode; it is simply how the system operates every single day.

How small the loss window really is

It is worth being honest about the one gap that does remain, because a piece that pretended there was none would deserve to be distrusted. Between the moment an architect makes a save and the moment the background companion has reconciled that save into the shared workspace, there is a brief window in which the newest work exists only on the local machine. If disaster struck in exactly that window, the very latest edits could be lost. The point is not that this window is abolished; it is that it is measured in the ordinary rhythm of saving and syncing rather than in the days or weeks that separate one manual export from the next. The difference between a protection gap measured in seconds and a protection gap measured in fortnights is not a matter of degree. It is the difference between a rounding error and a disaster.

Smart Sync narrows this window further by adapting the cadence of reconciliation to what is actually happening. When an architect is working actively, the companion keeps their saves flowing into the shared workspace promptly, so the newest work is protected almost as fast as it is made. When little is changing, it does not hammer the network for no reason. The effect on disaster recovery is that the loss window tends to be smallest exactly when there is the most recent work to lose, which is the sensible way round. The architect does not manage any of this and does not have to think about it, which is the recurring theme of the whole arrangement: the protection tracks the work automatically, tightening when the stakes rise and relaxing when they fall, without ever asking the person doing the modelling to make a decision about their own safety.

Recovery you can actually rehearse

A backup nobody has ever restored is a rumour, not a backup, and every experienced operations person carries the scar of discovering this at the worst possible moment. One of the underrated virtues of the local-first arrangement is that its recovery path is the same as its ordinary path, which means it is exercised constantly rather than only in a crisis. Cloning a workspace to a fresh local repository is how a new architect is onboarded, how a colleague sets up a second machine, how anyone gets a clean working copy. The exact operation you would perform to recover from a lost laptop is an operation the practice performs routinely for entirely ordinary reasons. It is not a dusty procedure in a runbook that has never been run; it is a Tuesday.

This matters more than it might seem, because the failure mode of most disaster-recovery plans is not that the plan is wrong but that the plan has never been tested and turns out to have a broken step nobody noticed. When the recovery operation and the daily operation are the same operation, the plan is tested every time anyone uses the system normally, and the broken step cannot hide because someone would have hit it during ordinary work. The practice does not have to schedule a disaster-recovery drill and hold its breath. It recovers small pieces of itself all the time, as a matter of routine, and the confidence that the big recovery would work is built up quietly from the accumulated evidence that the small ones always do.

None of this means an organisation can stop thinking about backups altogether, and it would be dishonest to suggest otherwise. Somebody still has to run the server, back up its database, test the restores, and hold the disaster-recovery plan with the seriousness it deserves. What changes is which problem the organisation is solving. It is no longer trying to compensate for the fact that its most valuable model lives on a laptop protected by a habit. It is running an ordinary database with an ordinary backup regime, and the architects on top of it are protected as a consequence of the system's shape rather than their own vigilance. The responsibility does not vanish; it moves to the place best equipped to carry it, and it becomes the kind of responsibility organisations are already good at.

The honest version of the promise, then, is not that disaster becomes impossible. Laptops will still be lost, servers will still occasionally fail, and mistakes will still be made inside the model itself. The promise is that none of these events has to become a catastrophe for the model, because the model is no longer a single fragile artefact held in one place by one person's diligence. It is a distributed, continuously reconciled, fully historied thing that lives on the organisation's own infrastructure and on every architect's machine at once. The loss of any single piece is survivable because the other pieces carry the truth. That is what it means to have backups and disaster recovery you do not have to think about: not that the thinking has been abolished, but that it has been done once, structurally, so that the architect on the bad day can reach for a new machine, clone the workspace, and simply get back to work.