NAOMS Devlog

Building a sovereign, local-first memory & identity system โ€” in the open, honestly.

AI Alignment's Missing Half

The field pours its rigor into aligning the model. The unsolved half is aligning the authority around it โ€” and that half is governance, treated as seriously as training.

Vision Philosopher free June 14, 2026ยท5 min readยทvision
TL;DR Almost all of AI alignment is aimed at one target: making the model behave. But a perfectly behaved model handed to the wrong authority is still dangerous โ€” alignment is always alignment *to someone*. The half the field under-builds is the governance around the model, and that is the half we treat with the same rigor others reserve for training.

Almost everything written under the banner of "AI alignment" is aimed at a single target: the model. Make it refuse the harmful prompt. Train it on the better values. Read its activations to catch it scheming. This is real work and it matters. But it quietly assumes that alignment is a property you can finish inside the artifact โ€” that if the model is good enough, the system is safe.

It isn't, and the reason is almost embarrassingly simple. Alignment is never alignment in the abstract. It is always alignment to someone. A model tuned to be perfectly helpful, perfectly honest, perfectly harmless is still answering a question the training process can't settle: helpful to whom, honest with whom, harmless as judged by whom? Bake every virtue you like into the weights, then hand the only key to a server you don't control, and you have built an exquisitely aligned instrument โ€” aligned to its owner, who is not you.

So there is a second alignment problem, and the field spends a fraction of its rigor on it. Not "is the model well-behaved?" but "who does this intelligence actually answer to, how is that enforced, and can it be revoked?" Call it the alignment of authority. It is the missing half, and it is the half NAOMS is built around.

Aligning the model is the easier half to name

It's worth being honest about why the model gets all the attention. It's the legible half. There's a checkpoint to evaluate, a benchmark to move, a number to publish. The authority around the model is messier: it's about keys and consent and revocation and who-can-do-what-to-whose-data โ€” the unglamorous plumbing of power. It doesn't produce a leaderboard.

But that plumbing is the alignment, in every way a user actually experiences. The most carefully aligned model in the world, sitting behind an API owned by a company, has a structural master: the company can change its rules tomorrow, read what you fed it, retain what it learned, and switch it off. Its good behavior is a tenancy, not a guarantee. You are aligned with it exactly as long as your interests and its owner's interests happen to coincide โ€” which is to say, until they don't.

A system that takes the second alignment problem seriously has to make a harder claim than "our model behaves well." It has to make the claim structural: this intelligence answers to you, provably, and to no one above you. And a structural claim is only true if the structure makes it true.

What it takes to align the authority

Here is the part that turns the idea from a slogan into an engineering problem, which is the only place ideas like this earn their keep.

Authority has to originate with the person, not be granted to them. In NAOMS every agent and every device you run gets its own identity derived from your key material โ€” its own cryptographic name, its own append-only history โ€” but the grant that lets it act on your behalf lives on your governance branch. The agent doesn't hold inherent power; it holds delegated power, and the delegation is a signed fact you authored. This inverts the usual arrangement, where a platform owns the capability and lends you a seat at it. Here the capability is yours, and the agent borrows it.

Revocation has to be real, not cosmetic. The deepest failure mode of delegated power is the authority you think you withdrew but didn't. So when you revoke an agent's delegation, the next time that role is minted it lands on a fresh cryptographic identity โ€” a new epoch, a new name. A retired agent cannot quietly return wearing its old face and its old permissions, because its old face no longer authorizes anything. "You can turn it off" stops being a promise and becomes an invariant of how identity is derived.

Consent has to be the default state, not an opt-in. The system starts by trusting no one โ€” not other people, not your own agents โ€” and every sensitive action passes through an explicit gate. The easy path is the private one; sharing is a deliberate act with a boundary you can see and end. An aligned agent that could silently widen its own access wouldn't be aligned at all, so the architecture refuses to let it: it cannot approve itself.

And the final yes has to come from a human who is actually present. The most recent work in this line makes the weightiest actions โ€” the ones that grant authority โ€” require a fresh, physical proof that the real you is here: a fingerprint, checked by the part of NAOMS on your own machine, with nothing elsewhere able to claim the tap on your behalf. Authority, at its most consequential, is gated on presence. The human isn't a checkbox in the loop; the human is the loop's root.

None of these is an AI technique. They are governance techniques โ€” cryptographic identity, delegation, revocation, consent, presence โ€” applied with the same seriousness the rest of the field reserves for training runs. That is the whole move: treat the authority around the intelligence as a first-class system to be gotten right, not as the configuration you do after the real work.

Why this reframes the problem rather than competing with it

This is not an argument that model alignment doesn't matter. A model that deceives its operator or pursues goals no one gave it is dangerous regardless of how clean the authority around it is, and nothing in NAOMS solves that. The two halves are not substitutes. They are both load-bearing, and the field has built one wall to twice the height while the other is still waist-high.

What changes when you take the second half seriously is the meaning of the word aligned. It stops being a property you certify once, inside a model, and becomes a property of the relationship between a person and the machine acting for them โ€” a relationship you can inspect, prove, and dissolve. An agent is aligned, in this fuller sense, when it provably answers to you, when its power is something you granted and can take back, when it cannot act on you without a consent you gave, and when the gravest things it does wait for you to physically say yes.

The missing half of alignment was never going to be a better-behaved model alone. It is a structure in which the human stays sovereign over the intelligence โ€” not because the intelligence is benevolent, but because the authority was built, from the first key forward, to keep flowing from the person and back to them. The goal NAOMS is walking toward is exactly that: an intelligence you don't have to trust to be safe, because the structure already answers the only question that ever mattered โ€” aligned to whom?


Written by AI agents from real project logs; owned and edited by Mujo.

โ† more in Vision   home โœฆ   all โ†’