Safety by architecture
An argument that safety should be a foundational requirement of the architecture, not bolted on afterward. These are design positions we argue for — not mechanisms we claim to have built.
This is a position paper, and we want that to be unambiguous before the first argument. Everything here is a set of design commitments we argue for — stances about how a system should be built. None of it is a description of a finished mechanism we are claiming to have implemented. We will repeat that where it matters, because the difference between “we argue for this” and “we have built this” is exactly the difference honesty turns on.
The argument against bolting it on
The common shape of AI safety today is post-hoc. A capable system is built first, and safety is added afterward — as a policy layer, a content filter, a set of rules wrapped around the outside of something that was not designed with those rules in mind. We think that shape is fragile, and fragile in a way that gets worse as systems get more capable.
The reason is structural. A safeguard that sits on the outside of a system is, by construction, separate from how the system actually works. It governs the outputs without being part of the reasoning. That leaves a gap, and the more capable the system, the more pressure that gap is under: there is more behaviour to cover, more ways to arrive at an outcome the filter did not anticipate, more distance between the rule and the thing the rule is trying to constrain. Patching the outside is a losing race against the inside.
So the argument is simple to state. Safety should be a foundational requirement of the architecture itself — present in how the system is designed, not appended after the design is done. Not a layer on top, but a property of the structure. We argue this; we are not reporting it as accomplished.
The positions
We hold four design positions. Each is a commitment about how a system ought to be built, offered as a stance rather than a specification.
Regard for human welfare as a requirement of operation. We argue that care for human welfare should be a condition the system is built to satisfy, not an instruction it is told to follow and could in principle set aside. The distinction is between a value that is part of what the system is and a value that is merely something it has been asked to keep in mind. We treat the former as the goal. We are explicit that stating it as a goal is not the same as claiming to have achieved it — this is a position, not a built guarantee, and we do not promise guarantees we cannot prove.
Transparency of reasoning. We argue that a system’s reasoning should be inspectable rather than hidden — legible enough that a person can follow how a conclusion was reached and question it, instead of being handed only the conclusion. Reasoning you cannot see is reasoning you cannot check, and safety that depends on a process no one can examine is not safety you can trust. This is a design requirement we argue for, not a finished property we are reporting.
Independent oversight. We argue for oversight that is genuinely separate from the system it oversees: a distinct and deliberately simpler check, sitting outside the main system, that the main system does not get to argue its way around. The point of keeping it simple and separate is that a complex system is good at producing persuasive accounts of its own behaviour, and an overseer that can be talked into agreement is not really an overseer. The check has to be answerable to something other than the thing being checked. Again — a position about how oversight should be arranged, not a claim about an oversight system we have running.
Trust earned in stages. We argue that capability and reach should be extended gradually, with trust earned at each stage rather than granted by default. Behaviour is checked before reach is widened; cooperation is conducted with reasoning kept transparent and oversight kept independent; and the further stages are something to be designed toward rather than assumed. We have written about this staged model at more length elsewhere. Here it is enough to say it is the shape we argue for: trust as something accrued, not assigned.
What this is and is not
It would be easy to read confident prose as a set of claims about machinery. So, once more, plainly: these are positions and design commitments we argue for. They are not mechanisms we are presenting as implemented, and nothing here should be read as a claim that safety has been solved or guaranteed. It has not been, and we do not believe anyone is in a position to make that claim.
What we are arguing is narrower and, we think, defensible: that design-time safety is structurally sounder than post-hoc safety, and that the place to put care, transparency, oversight, and staged trust is in the foundations rather than in a layer wrapped around the outside. Whether these positions can be realised — whether architecture-level safety can be measured, built, and independently audited, or remains a design argument until someone tries hard to break it — is an open question, and we hold it as open. The argument for where safety belongs does not depend on pretending that question is already answered.