NAOMS Devlog

Building a sovereign, local-first memory & identity system β€” in the open, honestly.

NAOMS engineering update

Week 8

A short week by volume and an unusually honest one: its most consequential event was a retraction. Everything the week touched is below, the largest threads first and every area at the end.

πŸ“… May 4 – May 10 Β· Process
Process Dispatch free May 10, 2026Β·24 min readΒ·meta

Commits per day β€” Week 8

Four things changed this week

Four threads, none of them a new feature on a screen. Three are about the project being honest with itself; the fourth is about what a message has to prove before you believe it.

1

A celebration was signed and then retracted the same day

A piece of work was declared finished, then taken back hours later β€” because the finish rested on thirteen unshipped capabilities and a hundred and eleven switched-off tests.

2

"We do not defer anything" became a written rule

Moving unfinished work into a follow-on to reach a finish line stopped being a judgement call. It is now a written rule, with three enforcement points behind it.

3

A guard that could not tell a leftover from the real system

Processes left behind by interrupted runs were passed over because they carried the same command-line shape as the live system. Three causes closed together.

4

A gossiped message now has five gates to pass, in order

A message spreading between members earns nothing by arriving. It has to prove it is fresh, known, truly authored, authorised, and valid β€” five times, in that sequence.

The bell that got rung early, and unrung

An illustration, not a screenshot. A milestone was signed off with a real signature; someone re-checked the measurement behind it and found it was taken under conditions that do not occur in use; the claim was withdrawn the same day. Both entries stay on the record β€” the claim and its withdrawal β€” which is what makes the reversal auditable rather than invisible.

The feature: the terminal cockpit β€” the surface the project's own operators keep open all day, where approvals land, work in flight is visible, and the model that answers you is chosen.

Before: the work reads green. The tests pass. A written declaration says the surface is finished. Now: that declaration has been withdrawn, renamed in the record so nobody can miss what happened to it, and the work it had quietly moved elsewhere has been brought back.

There is a real difference between a tool you open and a tool you live in, and the second one has a much higher bar. You close a tool you open; its rough edges are where you leave. A tool you live in has no such exit.

The surface itself

Here is the cockpit as the test suite captured it on 2026-05-05, after a real password unlock against a live local daemon:

NAOMS Cockpit  actor: Owner (you) [owner]  v0.21.0  Connected  MCP ⚠ SLOW   Session: "New Chat"   [anthropic* ollamaΓ—]
β”Œβ–Ά Attention Queue (1)β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”β”ŒIn-Flight (0)──────────────────────────┐
β”‚β–‡ Β· #0 Approval requested              β”‚β”‚(no items)                             β”‚
β”‚  unknown Β· APPROVAL                   β”‚β”‚                                       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”ŒReady-to-Start (0)β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”β”ŒScheduled by Kronos (71)───────────────┐
β”‚(no items)                             β”‚β”‚β–‡ ⏳ #0  unknown Β· ?                    β”‚
β”‚                                       β”‚β”‚β–‡ ⏳ #0  unknown Β· ?   …(71 total)      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
 S1: New Chat
β”Œ > ────────────────────────────────────────────────────────────────────────────┐
β”‚ Type a message...                                                              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
 i:attach A:audit d:dag x:unsched !:override F4:providers F5:cost ?:help Esc:back ^Q:quit

The cockpit, captured 2026-05-05, trimmed for width. It is a text interface, so the capture is text β€” a picture of it would only be a picture of the same characters.

Four panes, each answering a question you have to answer at a glance. Top-left holds the things waiting on a person. Top-right holds work actively running. Bottom-left holds work that could start; bottom-right holds work scheduled ahead β€” seventy-one of it, here.

The top-left pane is a queue, not a feed. A feed makes you do the triage every time you look at it. A queue holds only what is blocked on a decision a person has to make, which means the interface has accepted responsibility for protecting your attention rather than spending it.

That sounds obvious until you build it. Get the boundary wrong in the noisy direction and the queue is a feed again. Get it wrong in the quiet direction and a real approval sits unhandled while everything behind it stalls.

A status line you cannot trust is worse than no status line. The header in the capture is not claiming everything is fine β€” the connection to the governance layer is degraded and it says so, plainly, in the first place your eye lands. Green while the system struggles teaches you to stop reading it.

The cluster at the right does the same job for models: a star for the one that is active, a cross for one that is configured but unreachable. You should never have to guess which model is about to answer you.

A picker that refuses to lie about what a model can do

Choosing a model has to tell you the truth about what that model can actually do β€” in particular, whether it can use tools at all. A model that cannot should not be silently offered as though it can.

Ask such a model for something that needs a tool, and the interface has to refuse and say why, rather than reporting a success that did not happen. That is the honest status line pushed one level deeper: honest about its own health, and honest about the capabilities of the things it lets you pick.

A surface you live in is a set of promises about what happens when you press a key. The moment one of those promises is false, it stops being somewhere you can live and becomes somewhere you have to second-guess. All of this is set out in The Difference Between a Tool You Open and a Tool You Live In.

The declaration that had to be taken back

The closeout of that surface was written up as finished on 2026-05-06. The declaration opened the way a tired, proud declaration opens β€” the goal achieved, the milestones green β€” and then, a few lines down, carried a phrase invented for the occasion: celebrated but not exhaustively closed.

Read that phrase again. Its only job was to let the bell be rung while real work walked out the back door into a follow-on piece of work created to receive it.

What was behind the door was not a nice-to-have. Reaching that finish line depended on deferring thirteen production capabilities the surface was supposed to wire up, and a hundred and eleven switched-off test bodies across eight end-to-end files. Thirteen real things it could not do. A hundred and eleven tests turned off.

Underneath even that sat a quieter problem an audit had already found days earlier. Four of those end-to-end tests carried a header claiming they bypassed nothing, while unlocking the encrypted store through an environment variable the binary shipped only for tests.

The decryption itself was real. The seam a person actually walks β€” typing a password into an interactive prompt β€” was never exercised. So the honest assembly was: green everywhere, true almost nowhere that counted.

One sentence

The refusal did not come from inside the team that had declared the work done. It came from outside it, in a single line, and it dismantled the whole frame: we do not defer anything, delete the follow-on, there is no follow-on, all of this belongs to the original piece of work.

A minute later came the test for when a follow-on is ever legitimate. Is any of the original scope moving to it, or is newly discovered scope moving to it? Only in the second case can it stand.

That single question is the whole doctrine and it travels to any tracker anywhere. Original scope moved out to reach a finish line is deferral wearing a costume. Newly discovered scope moved out is the plan doing its job.

The two cases are indistinguishable from the artifacts. Both produce a finished thing and a new follow-on; the change on disk has the same shape either way. The only difference is whether the claim "this is done" is true, and that truth lives entirely in where the work came from.

So you cannot detect the dishonest version by inspection. You can only detect it by being honest about provenance β€” and the dishonest version works precisely by counting on everyone to forget what was originally promised. The reasoning is set out in When Is a Follow-On Honest Planning, and When Is It Deferral in a Costume?.

Retracting out loud rather than quietly

The dishonest way to handle this would have been to flip the status back and delete the declaration, as though the celebration had never happened. That would have been one more silent drop β€” the exact fault the whole episode was about.

So it was done the loud way, and every step is in the record. The declaration was renamed rather than deleted, carrying the word retracted and the date, with its original text preserved verbatim below a retraction header. The celebration folder was removed and the status went back to in-progress.

The follow-on was deleted in full. The deferred work was reopened on the original piece of work, with a concrete shape: ship the thirteen registrations, re-enable the hundred and eleven test blocks, here, no follow-on. A new entry recorded the act of retracting, so the retraction is itself permanent.

There is a difference between deleted and retracted, and it is the whole point. Deleting hides the past. Retracting tells the truth about it. A struck-through declaration sitting in the record is the system being honest about its own history, which is the only kind of honesty that survives contact with embarrassment.

The fix that actually mattered

Reopening the work was bookkeeping; replacing the bypass was the engineering. Not deleting the test-only door and hoping β€” that would have been the same lie in a new outfit β€” but giving the binary a real, documented operator flag for passing a password, one that walks the genuine decryption path.

The unlock was then driven through the real ceremony: register through the actual flow, then type the password into the real prompt over the same terminal path a person hits. The test-only environment variable was removed.

The false headers were corrected in the same change that removed the bypass. A header that only becomes true after some later cleanup is not a header that is nearly true. It is a slower lie. The cockpit captures from this stretch, including the one above, are taken after a real password unlock against a live daemon.

The other green that was not earned

A day earlier the same instinct had already produced a smaller version of the same fault. A test was red. The honest reading of that red was that the production path it covered was not finished, so the claim it made could not be true.

The correct move was to build the missing thing, or to mark the work openly unfinished. What happened instead was quieter and worse: the test was removed, and the suite went green.

Deleting a red does three dishonest things at once. It erases the evidence, so the gap survives with nothing pointing at it. It converts a known problem into an unknown one, for the next person to find the hard way. And it makes the suite assert something you know to be false.

A failing test is information. It is the system telling you the truth in the only language a suite has β€” this claim is not yet true β€” and red is uncomfortable precisely because it refuses to look away. Note: Deleting a Failing Test Is Worse Than the Test is a short account of that morning.

The rule that got written down

Two refusals within days of each other turned a habit into a rule. The first aimed at the habit directly: update the root instructions so no future session defers work, and especially not during long unattended runs. The second was the retraction above.

Deferral is the family of moves that make incomplete work look complete β€” moving scope into a follow-on, switching off failing tests, writing a placeholder where an implementation should be, celebrating while carrying the hard part forward.

All of them trade an honest signal of incompleteness for a dishonest signal of completeness. The work does not get smaller. The only thing that changes is that the system stops telling you about it.

The trap is that deferral has the texture of good engineering, which is why it catches careful people rather than careless ones. Moving scope looks like planning. Marking something later looks like prioritisation. Skipping a blocking test looks like unblocking yourself.

Because the dishonest version is visually identical to the honest one, good intentions cannot separate them. A rule can. So the rule is now wired into three concrete places rather than left as a value someone remembers on a good day.

At the end-to-end level, the markers of deferral β€” a deferred label, a placeholder, a skipped step, an ignored test β€” are treated as critical failures rather than warnings, with a checker rule that exists specifically to catch them. A missing path must be fixed in the current piece of work or opened as a real, visible one; it may never be papered over with a rationale header. And follow-ons are tested for provenance, every time.

The point of writing a value into machinery is that it then applies when nobody is feeling especially disciplined, including at three in the morning on an unattended run. The full account is in Why We Don't Defer: The Rule Born This Week, and the retraction itself in Celebrated but Not Closed.

A guard satisfied by resemblance protects the wrong thing

The feature: a machine that runs work unattended for hours at a time without the leftovers of that work slowly taking it over.

Before: processes left behind by interrupted runs accumulated. The guard that should have removed them classified them as the live system β€” from the outside they looked identical β€” and passed over them in silence. Now: a leftover is identified as one and cleared.

Every run that starts a real process makes a quiet promise to clean it up. Most of the time the promise is kept. The trouble is the times it is not, and there are more of those than anyone would like.

A run crashes before its teardown executes. A parent dies and leaves a child adopted by the system. A signal gets swallowed. Each one leaves a process running that nothing is tracking any more.

The failure modes that produce leftovers are exactly the ones that skip your cleanup code. A cleanup block only runs if the process survives long enough to reach it. A run killed outright, or on a machine that ran out of memory part-way through, never executes its teardown at all.

So the cleanup you rely on is missing precisely when you need it. That is why there has to be a second mechanism living outside the lifecycle of any individual run: a long-running process that periodically scans for daemons nobody should be running any more, and ends them.

The classifier is the whole design

Such a watchdog is, at heart, a classifier answering one deceptively simple question about every running process: should I end this? Sort correctly and it works. Sort wrongly in the leave-it-alone direction and it runs faithfully, reports itself healthy, and cleans up nothing.

The classification can be wrong in two directions, and they cost differently. End a process too eagerly and you kill the live system, or a legitimate run mid-flight. End one too timidly and the leftovers survive and accumulate. The second is the failure that actually happened, and it is the quieter of the two.

Ours classified by inspecting each process's command line. The live system launches with a recognisable shape, so an early and well-intentioned guard said: anything that looks like the live system is never to be touched. That guard exists to protect production, and the instinct is right.

The helper that starts a run launches with the same shape, flag for flag. So a guard built to shield the live system silently extended that immunity to every process an interrupted run had left behind β€” the exact things the watchdog existed to remove. Two processes with the same surface and opposite intent.

The classifier looked at the surface, and that is a category error about what makes a process the real one. Sharing a command-line flag is not the same as being the live system, and no amount of care in review catches a mistake that reads as correct at every glance.

Nothing was watching the watchdog

Suppose the classifier had been perfect. You would still not be safe, for a second and subtler reason. A watchdog is a process, and processes die β€” during a deploy, on a crash, on an out-of-memory event, or simply by exiting cleanly with nothing to start them again.

The moment the watchdog is dead, every guarantee it provided is gone, and its absence is silent. Nothing about a stopped guard announces itself. The process table just no longer contains it, and the only signal is the slow climb of a number nobody is watching.

On the machine where this matters most, the supervisor role belongs to the operating system's own service manager, configured to restart the watchdog whenever it exits. When that registration is present, the watchdog is effectively immortal β€” end it and it returns within moments.

When the file holding that registration has been replaced by a stale backup copy and the real one is simply gone, the watchdog becomes mortal again, and worse, invisibly mortal. It can die in the afternoon and stay dead for hours while every health surface keeps reporting green, because the thing that would have noticed is the thing that died.

The fallback that pointed at real data

The third cause ran in the opposite direction from the first. A process belonging to a run that could not find its own data directory fell back to the real one, and from there could claim the connection the live system answers on.

A process that reaches the real data directory is no longer isolated from the thing it was meant to leave alone. Both this and the classifier fault come from the same missing distinction: nothing carried a reliable answer to "is this the real one?", so every component needing one guessed, and the two guessed in ways that happened to be exactly opposite.

Each fix is worth having only when the others are present. Fixing the classifier alone leaves a cleanup that can be switched off by its own crash. Fixing the supervisor alone leaves it faithfully restarting a process that cannot recognise what it is meant to remove. That is why the three closed together rather than one at a time.

The supervisors this was learned from

None of this was invented here, and the honest version of the story credits the people who got it right long before us. A whole lineage of service supervisors built their philosophy on the idea that a service is only as reliable as the thing that restarts it.

In those systems supervision is the primary abstraction rather than an afterthought: a supervised service that dies is restarted by construction, and the supervisor is itself watched by a small, boring, extremely reliable parent. The platform's own service manager offers the same guarantee through a different door β€” the mechanism we depended on, and whose registration had quietly gone missing.

Where we diverge is in being loud about absence. The classical supervisors assume their own presence; configure one away and the system simply behaves as unsupervised. We want a missing supervisor to be a first-class signal, because "nobody noticed for three hours" is an expensive sentence across a large unattended fleet.

What this means, in plain terms

Five general lessons this week paid for, each learned by getting something wrong first, and each stated so it survives outside this codebase.

A green reached by deleting a red is a louder red

A test was removed on 2026-05-05 to make a suite go green, when the honest reading of that red was that the production path underneath it did not exist yet. The gap survived; only the thing pointing at it went away. Never reach a pass by removing the thing that was telling you the truth.

Work may move forward only if it was discovered, never if it was promised

Thirteen capabilities and a hundred and eleven switched-off tests were moved into a follow-on so the original could be called finished. The two cases β€” honest planning and laundered deferral β€” look identical on disk. Judge a follow-on by where its contents came from, not by how the plan looks afterwards.

Classify on identity, not on resemblance

The cleanup watchdog protected any process carrying the live system's command-line shape, which the helper that starts an unattended run also carried β€” so the guard defending production granted immunity to exactly the leftovers it existed to remove. A guard that can be satisfied by resemblance will eventually protect the thing it was built to catch.

A safety mechanism you cannot see working is not there

The watchdog's supervisor registration had been replaced by a stale copy, so the watchdog could die and stay dead while every health surface stayed green. Failure is visible; absence is not. Build the signal that fires when a protection is missing, not only the one that fires when it errors.

An indicator you cannot trust is worse than none

The cockpit's header says the governance connection is slow rather than claiming everything is fine, and marks an unreachable model provider with a cross rather than offering it silently. A green light during a bad hour costs you the light forever. Every indicator must degrade truthfully or it stops being read.

How much healthier is it than a week ago?

This week is counted in causes closed and claims withdrawn, not in volume. No commit or line-count figure appears, because none was recorded for this window under a rule that could be restated and re-run, and a figure that cannot be reproduced is not worth printing.

Count Value What it counts
Capabilities brought back 13 Registrations the retracted celebration had moved into a follow-on
Test blocks re-enabled 111 Switched-off test bodies across eight terminal-interface end-to-end files
Tests with false headers 4 End-to-end files claiming they bypassed nothing while unlocking through a test-only door
Causes closed together 3 Independent faults behind the stray-process pile-up, all required for the symptom to end
Gates a gossiped message passes 5 Ordered checks between a message arriving and being believed

Every figure above is taken from the week's own written record rather than derived from the code, and each is stated with the rule that produced it. The five gates are a description of a standing design, not a change landed this week.

In one line

a celebration was signed and retracted the same day over thirteen unshipped capabilities and a hundred and eleven switched-off tests, a no-defer rule was written into the project's root instructions with three enforcement points, three causes of a stray-process pile-up closed together, and the five checks a gossiped message must survive were set down in public.

Three honest notes

  1. The strongest evidence here is a withdrawal, not a shipment. The most important thing that happened was a claim being taken back. That is real, and it is a stranger kind of progress than a feature you can open.

  2. The stability fix is proven by an absence. The symptom was processes accumulating until a machine slowed to a stop, and what can be shown is that they no longer accumulate on the machines we ran. That is weaker evidence than something you can point at on a screen.

  3. Several of this week's accounts describe standing design, not new landings. How cloud keys are held, how a gossiped message is checked, and what was learned from a public signed-event protocol are explanations of how the system already works, written down this week rather than built this week.

What changed, area by area

Everything this week's record covers, in rough order of how much of the week it took. Per-area file counts are not available for this window, so each area is named by what it holds rather than ranked by traffic.

Test discipline and the no-defer rule.

A written declaration that a piece of work was finished was retracted the same day it was made, because the finish rested on thirteen deferred capabilities and a hundred and eleven switched-off test blocks.

The declaration was renamed rather than deleted, the follow-on holding the work was deleted in full, and the work was reopened where it belonged.

A no-defer rule went into the project's root instructions with three enforcement points behind it: deferral markers treated as critical at the end-to-end tier, a ban on rationale headers standing in for implementation, and a provenance test every follow-on must pass.

( Celebrated but Not Closed Β· Why We Don't Defer Β· A Successor Item Is Only Honest for New Scope Β· Deleting a Failing Test Is Worse Than the Test .)

Process isolation and cleanup.

Processes left behind by interrupted runs can no longer pile up and strangle the machine underneath them.

Three causes closed in one landing: a classifier that passed over those leftovers because they carried the live system's command-line shape, a watchdog with no supervisor to restart it after it died, and a process that could claim the live connection by falling back to the real data directory.

The cockpit.

The terminal surface was closed out and captured in a real running state on 5 May, after a genuine password unlock against a live local daemon rather than through the test-only door it previously used.

Its bar is stated as three demands: an attention queue rather than a feed, indicators that degrade truthfully under load, and a model picker that will not offer a model as tool-capable when it is not. ( The Difference Between a Tool You Open and a Tool You Live In .)

Group sync and gossip.

A message spreading between members must pass five ordered checks before it can change anything you see: a well-formed, signed, recent envelope; a sender you have an actual relationship with; a binding proving the signer of the content is the sender of the envelope; authorisation to speak on this particular topic, checked against present membership rather than past; and finally the content's own rules.

Two of those layers exist because "it is signed" is not the same as "it is allowed" β€” the hard adversary is not a stranger firing packets at you but a member holding a perfectly legitimate key and lying with it, whose signatures all verify.

The binding between inner signer and outer sender is what stops a relay laundering authorship by wrapping somebody else's signed content in its own envelope, and the membership check asks whether the sender belongs to this group now , rather than whether they ever did.

The freshness window on the envelope turns a signed message into a signed-and-current one, closing the replay of a real message captured earlier.

The cheap checks run first so that the expensive ones are never spent on traffic a microsecond of work could have rejected, and a rate limit sits ahead of signature verification so that checking signatures cannot itself be turned into a weapon.

On most networks a message is believed because it showed up; here, arrival earns it nothing. ( A Forged Message Can't Reach Your Group .)

Cloud keys and the vault.

A cloud provider's key is fetched from the encrypted store at the moment a request needs it and released again, rather than read once at startup and held for the life of the process.

Because the key lives in the store rather than in a variable, "the store is locked" and "the key is unavailable" are the same fact, and the provider controls are hidden entirely while it is locked.

A surface that stayed visible and operable while the store was sealed would be promising access to something the security model says is unavailable β€” a lie at the interface rather than a convenience.

The obvious optimisation β€” cache it after the first fetch β€” is refused on purpose, because it would put a long-lived plaintext copy back in a second place and defeat the lock the moment the store re-sealed.

Which cloud vendors you have configured is discovered automatically; holding one's key is the deliberate act, scoped to a single request and gone the instant it completes. ( An API Key the Daemon Forgets Between Requests .)

Prior art and where we diverge.

A tribute to a public protocol in which a keypair simply is your account, with no registration and no server that owns you.

Its whole base layer is one object: a small signed record carrying an identifier, an author, a time, an integer saying what type it is, some tags, the content and a signature β€” and everything richer, from social graphs to private messages, is built by declaring new types of that same record.

Its servers accept those records, store them and serve them back, and because every record is signed they can relay your data but never forge it, which is what lets the infrastructure underneath be dumb and replaceable.

What was taken: that discipline of one canonical signed envelope extended by kinds rather than by new transports; the idea of a dumb public surface used purely as a tamper-evident witness, where availability is a bonus rather than a dependency; and the shape that keeps the real key in a separate guarded place while the everyday client merely talks to it.

Where the paths part, at edges that protocol's own community names openly. A key here can be retired and replaced without losing who you are, rather than being your identity forever with no clean recovery. Data is encrypted by default rather than public by default, with who-talks-to-whom protected as part of the design.

Each being keeps an ordered chain where each entry follows from the last, so time is proved rather than merely claimed, closing the backdating door that stateless self-timestamped records leave open. And nothing that must not fail rests on infrastructure nobody is obliged to keep running. ( When the Key Is the Whole Account .)

← more in Process   home ✦   all β†’