CelinQ Insights · No. 22

Model quality that survives many hands

Keeping a model coherent when a dozen people are shaping it.

A NILUS perspective on collaborative modelling for Sparx Enterprise Architect

There is a moment that arrives in almost every large modelling effort, usually somewhere past the halfway mark, when someone opens a diagram they have not looked at in a while and quietly loses confidence in it. The elements are all there. The relationships connect the right boxes. Nothing is broken in the way a tool would flag as broken. And yet the diagram no longer reads as though one mind made it. Two elements clearly mean the same thing under different names. A relationship type has been used one way in this corner of the model and a subtly different way three packages over. A naming convention that everyone agreed to at the start has three living dialects. The model still works, in the sense that it opens and validates, but it has stopped being trustworthy, and a model you cannot trust is a model people quietly stop using.

This is the quality problem that shows up when many hands shape a single model, and it is worth being precise about what it is and is not. It is not, for the most part, a problem of individual competence. The people contributing are usually good at their jobs. Each change, taken on its own, is defensible. The trouble is that quality in a model is not a property of any single change; it is a property of the whole, and the whole is exactly the thing that no single contributor can see while they are working. Coherence is an emergent quality, and emergent qualities are the first to erode when a team grows, because the mechanisms that hold them together were designed for a smaller room.

Why models drift even when everyone is careful

A model is a shared language before it is anything else. When an architect writes down that one capability realises another, or that a component depends on a service, they are relying on an agreement about what "realises" and "depends on" mean in this particular repository, on this particular project, for this particular audience. That agreement is rarely written down in full. It lives partly in a modelling convention document that everyone read once, partly in the examples set by whoever built the first few packages, and mostly in the heads of the people who have been on the project longest. It is a living consensus, and like all living consensus it drifts unless something actively holds it in place.

With two or three people, the consensus holds itself. Everyone sees roughly everything. When a new pattern is needed, it gets discussed over a desk and adopted by everyone at once, because everyone was in the conversation. The model stays coherent not because anyone is guarding it but because the group is small enough that the shared language never fragments. The difficulty is that this self-correcting property does not scale. Add more people and the field of view of each contributor shrinks. They see their own packages clearly and everyone else's dimly. They make locally sensible decisions in ignorance of the decision someone else made yesterday in a part of the model they never open. Nobody is being careless. The structure simply no longer lets carefulness add up to coherence.

Then there is time. A model that lives for a year or two outlives the memory of its own conventions. The person who established a pattern moves to another engagement. A new contributor joins and, finding no obvious statement of how things are done here, does what is reasonable and does it slightly differently. The old and new ways coexist, both defensible, neither dominant, and the model acquires a seam that nobody chose and nobody owns. Multiply that by every convention and every joiner and leaver, and the model develops the architectural equivalent of an accent that shifts from neighbourhood to neighbourhood.

The cost of incoherence is paid later, by someone else

The reason model drift is so corrosive is that its cost is deferred and displaced. The person who introduces a small inconsistency pays nothing for it; their change works, their diagram renders, their review passes. The cost lands later, on whoever has to read the model as a whole and act on it. It lands on the architect who is asked whether a proposed change is safe and finds they cannot answer confidently because they no longer trust that the model says what it appears to say. It lands on the newcomer who spends their first month unable to tell which of two competing patterns is the one they are supposed to follow. It lands, most expensively, on the client or the stakeholder who was told the model is the authoritative picture and then catches it contradicting itself, at which point the model's authority is spent and very hard to earn back.

A model loses its authority not in a single dramatic failure but in a slow accumulation of small contradictions, each of which was reasonable when it was made and none of which anyone owns.

What makes this worse than ordinary technical debt is that it is hard to see and harder to schedule. A broken relationship is visible; a tool will complain about it. Two elements that mean the same thing under different names will not trip any validation, because from the tool's point of view nothing is wrong. The inconsistency is semantic, and semantics are exactly what automated checks are worst at seeing. So the drift accumulates below the waterline of anything that would flag it, and it surfaces only when a human reads carefully enough to notice, which by definition happens long after the drift began and long after it would have been cheap to fix.

The two failed instincts

Faced with drift, teams reach for one of two instincts, and both fail in predictable ways. The first is to lock the model down: appoint a gatekeeper, route every change through a central authority, make coherence someone's explicit job. This does protect quality, but it does so by throttling throughput. The gatekeeper becomes the bottleneck, work queues behind their review, and the friction pushes people to make fewer and larger changes to minimise the number of times they have to pass through the gate. Large changes are harder to review well, so the gate that was supposed to protect quality ends up degrading it, because the reviewer is now looking at a month of work instead of a day's and cannot possibly hold all of it in their head.

The second instinct is the opposite: give up on active governance and rely on a periodic cleanup. Let people work freely, and every quarter someone goes through and tidies the model back into shape. This preserves throughput but treats coherence as a thing you restore rather than a thing you maintain, and restoration is always more expensive than maintenance. By the time the quarterly cleanup happens, the drift has propagated. The inconsistent pattern has been copied by three other people who took it as the norm. Fixing it now means untangling not one decision but the small ecosystem that grew up around it, and the person doing the cleanup was usually not the person who understood any of it in the first place.

The real problem with both instincts is that they treat quality and speed as a single dial you slide between. Turn it toward quality and you lose speed; turn it toward speed and you lose quality. That framing is the trap. The teams that keep a model coherent at scale are the ones that stop treating it as one dial and separate the concerns: let everyone work at full speed in their own space, and make coherence something the structure supports continuously rather than something a person has to enforce against the grain of how everyone works.

Full speed for the individual, visibility for the whole

This is the balance CelinQ is built around, and it is worth being honest that it does not make the quality problem disappear. Nothing does; coherence is human judgement and human judgement cannot be automated away. What the structure can do is remove the reasons that drift accumulates unseen, and give the people responsible for quality the raw material to act early instead of late.

The starting point is that each architect works in their own local repository, at full speed, whether online or offline. That matters for quality in a way that is easy to miss. When editing is fast and unencumbered, people make the small improvements they notice: the correction of a misnamed element, the tidying of a relationship, the alignment of a stray pattern with the convention. When editing is slow or has to be negotiated, those small improvements are exactly the ones that get skipped, because the friction outweighs the benefit of a minor fix. A great deal of model quality is nothing more than the sum of small corrections that were cheap enough to bother making. Making the individual's experience fast is therefore not merely a convenience; it is one of the quieter contributors to quality, because it keeps the cost of a small fix below the threshold at which people stop making them.

Each save is then synchronised, by a background companion, into a shared workspace that the organisation runs on its own infrastructure. This is where the whole becomes visible again. Instead of a dozen private views drifting apart and reconverging only at a quarterly cleanup, there is a continuously updated shared picture that reflects what everyone has actually done. The person responsible for coherence is no longer reading a snapshot that is weeks stale; they are looking at the model as it stands today, which means they can catch a diverging pattern while it is still one person's decision and not yet three people's habit. The economics of correction change completely when the drift is visible on the day it appears rather than the quarter after.

Drift is cheap to fix while it is one person's fresh decision and expensive to fix once it has become several people's settled habit. The difference between the two is mostly a matter of how quickly the whole team can see what changed.

Reconciliation that never quietly corrupts the model

The reconciliation of everyone's work is where quality is most at risk, because a careless merge is itself a source of drift, and a worse one than any human mistake, since nobody chose it and nobody can explain it. If reconciling two people's edits meant occasionally losing one of them, or silently picking a winner when they disagreed, the shared model would accumulate exactly the invisible contradictions that erode trust, and it would accumulate them with no human fingerprint to trace them back to. The whole approach would be self-defeating.

This is why the merge engine, CelinQ Fusion, works the way it does. It reconciles at the granularity of individual model facts rather than whole files, using a deterministic, reproducible three-way merge. Two architects editing genuinely different things both succeed, with no interaction required, which is what keeps the individual experience fast. When two people have genuinely changed the same fact in incompatible ways, the engine does not guess and it does not apply last-write-wins. It isolates the conflict, preserves both versions, and surfaces it for a human to resolve. Deletions are handled explicitly through tombstones rather than by having a thing simply vanish, so that "this was intentionally removed" is a recorded fact rather than an ambiguous absence that someone might innocently recreate.

The quality point here is subtle but central. A conflict is not a failure of the system; it is the system correctly refusing to make a semantic decision that only a person is qualified to make. When two architects have modelled the same thing two different ways, that is precisely the kind of divergence that a quarterly cleanup would have caught too late and a gatekeeper would have caught too slowly. Surfacing it at the moment of reconciliation puts the decision in front of a human while the context is fresh and the divergence is small. The engine's job is not to resolve the disagreement but to make sure the disagreement is never hidden, because a hidden disagreement is exactly how a model loses its coherence one invisible contradiction at a time.

A diverging pattern, caught on the day it appears

It helps to make this concrete, because the argument can sound abstract until you watch it play out on an ordinary working day. Suppose two architects are extending the model in adjacent areas. One of them, working through a set of business services, adopts a way of expressing how a service supports a capability that is slightly at odds with the convention the rest of the model has been using. It is not wrong. It is arguably clearer. But it is different, and if it spreads it will leave the model speaking two dialects of the same idea. In the old arrangement, where each person's work is visible to the others only at some later integration point, this divergence would sit unnoticed in a private copy for days or weeks. By the time anyone with responsibility for coherence saw it, a third architect would have looked at the new pattern, taken it for the house style, and copied it. What began as one person's local choice would have become a small competing convention with three adherents and no clear resolution.

Because synchronisation is continuous, the diverging pattern reaches the shared workspace on the day it is made. The person who watches over coherence, or simply a colleague working nearby, sees it while it is still one decision rather than a habit. The conversation that resolves it is a short one, because there is only one instance to discuss and the person who made it still has the reasoning fresh in mind. Either the convention is updated to adopt the clearer form deliberately, across the whole model, or the new instance is brought back into line, and in both cases the model ends the week speaking a single language. The divergence never got the chance to breed. This is the entire mechanism by which continuous visibility protects quality: it does not prevent people from diverging, which would be both impossible and undesirable, but it collapses the time between a divergence appearing and a human being able to decide what to do about it, and that time is where drift does all its damage.

Measuring quality honestly

There is a temptation, once a team starts taking model quality seriously, to reduce it to a set of automated checks and to treat a clean report as proof of a coherent model. It is worth resisting that temptation, because it confuses the checkable with the important. Automated validation is genuinely useful for the class of problems it can see: broken relationships, elements missing required attributes, structural violations that a rule can express. Those checks should be run and they should pass. But the qualities that make a model trustworthy to the people who rely on it are mostly semantic, and semantics resist mechanical checking. Whether two elements that carry different names actually mean the same thing, whether a relationship type is being used consistently with its intent, whether the level of abstraction is uniform enough that a reader can navigate without being jolted between the concrete and the conceptual, are all judgements that a rule cannot make.

An honest account of model quality therefore keeps the human judgement in the centre and treats the tooling as support for that judgement rather than a replacement for it. The value of a current, shared, well-reconciled model is not that it can certify its own quality, because it cannot, but that it gives the people responsible for quality the best possible conditions to exercise their judgement: a view of the whole that is close to current, a history that explains how the model came to be, and a reconciliation process that never quietly corrupts the facts they are judging. The measurement of quality stays where it belongs, with people who can read the model and say whether it hangs together, and the structure's contribution is to make sure those people are never judging a stale or corrupted picture.

History and review as the memory the team lacks

Drift feeds on forgetting. Conventions erode because the reasoning behind them is not written down and the people who held it move on. Part of what keeps a model coherent over years is simply having a reliable memory of what was decided and why, and this is where a complete, ordered revision history earns its place. When every change is recorded in sequence, the model carries its own account of how it came to be the way it is. A newcomer trying to understand which of two patterns is canonical can look at when each appeared and how each spread, rather than guessing. An architect unsure whether an inconsistency was deliberate or accidental can trace it to the change that introduced it and often to the reasoning attached to it. The history is the institutional memory that a growing team otherwise loses to turnover.

Change review and governance sit on top of that history, and the important thing is where they sit. Because everyone works at full speed locally and reconciliation is continuous, review does not have to be the gate that everything queues behind. It can be applied where it earns its cost, on the changes and the parts of the model where coherence matters most, rather than uniformly on everything as a tax on all work. This is the escape from the single-dial trap. Governance stops being the enemy of speed because it is no longer the only thing standing between a contributor and the shared model. The shared model updates continuously; review is how the team pays deliberate attention to the changes that deserve it, not how it grants permission to work at all.

The design assistant, kept in its place

There is an optional in-EA design assistant that can generate content into a package you select and analyse existing parts of the model, and it is worth being careful about how it relates to quality, because it would be easy to overclaim. Used well, it can help a contributor see a part of the model more clearly, or produce a first draft of a structure that a human then shapes to fit the conventions. It can help surface where existing parts do not hang together. But it is off by default, it is sovereign to the organisation, and the core of CelinQ works entirely without it. It is a tool that can accelerate a human's judgement about the model; it is not a substitute for that judgement, and coherence remains something people decide and defend. A model's quality is ultimately a claim its authors stand behind, and no amount of generation or analysis changes who is standing there.

What coherence actually requires

If there is a single idea worth taking away, it is that model quality at scale is not primarily a tooling problem and not primarily a discipline problem. It is a visibility-and-timing problem. Drift is not caused by bad people making bad changes; it is caused by good people making locally reasonable changes without being able to see the whole they are affecting, and by the whole becoming visible again only long after the cheap moment to correct it has passed. Everything that helps comes back to shortening the distance between a change being made and the change being seen in context by someone who can judge it.

That is why the shape matters: full-speed local work so that small corrections stay cheap enough to make; continuous synchronisation so that the whole is never weeks stale; a merge that reconciles the easy cases silently and refuses to hide the hard ones; a complete history so the model remembers its own reasoning; and review placed where it earns its keep rather than as a gate across everything. None of these on its own keeps a model coherent. Together they change the economics of coherence, so that keeping quality high stops being a heroic quarterly effort and becomes the ordinary background condition of how the team works.

A model shaped by many hands will never be the seamless product of a single mind, and it should not pretend to be. But it can be coherent, in the sense that matters: it can say one thing consistently, it can be trusted by the people who rely on it, and it can carry its own account of how it came to be. That kind of quality does not survive many hands by accident. It survives because the structure around the model makes divergence visible early, keeps correction cheap, and never lets a disagreement disappear without a person having looked at it. Get that right and the model can grow without losing the thing that made it worth building: the confidence that when it says something, it means it.