Two pictures of alignment are in circulation, and almost everything downstream depends on which one you start from.
The first is the one most safety work assumes. You fix a set of values, rules or preferences in advance, and you press them onto the model from outside. Conforming behaviour is permitted, deviating behaviour corrected. Alignment, on this view, emanates downward: a constraint applied from above to a system that would otherwise do something else.
The second runs the other way. Nothing is pressed on. Appropriate behaviour is held, moment to moment, by the system staying faithful to the conditions it is actually operating under, the task, the role, the domain, the person in front of it. Alignment becomes something that accumulates from below, out of many small acts of keeping faith with local conditions. That faithfulness is what we call adherence.
We build on the second picture. The rest of this note is why.
Why pressing values down does not hold
The top-down picture is intuitive, and as engineering it is mostly a matter of writing the constraints and then training the model to obey them. The difficulty is not in the writing. It is in where the constraint ends up sitting.
In a current language model a value has nowhere privileged to go. A system instruction, a safety boundary, a domain rule: each arrives in the context window as another run of tokens, laid out on the same plane as the ordinary conversational material it is supposed to govern. Put a value and a passing remark in the same space and they compete for the same attention. As the conversation grows the remarks multiply, and the constraint that was vivid in the first exchange is steadily out-weighed by everything said since.
Researchers call the symptom "lost in the middle", and usually treat it as retrieval, a question of surfacing the right fact. From the alignment side it is worse than that. What gets diluted is not a fact but the governing condition itself, and the model gives no sign of the loss. It was never argued out of its values. It simply stopped attending to them, somewhere around the fortieth turn, sounding exactly as fluent as it did at the first.
So a constraint imposed from above is brittle in a precise, predictable way: vivid at the start of an interaction and progressively less real as it goes on. Writing the rule more firmly does not touch this. The rule was never the weak point. Its position was.
Adherence as the unit
Start from the bottom instead. The smallest thing you can ask of a system is not that it hold the right global values, but that it stay faithful to the conditions of the task immediately in front of it. Is it still doing the job it was asked to do? Is it still inside its role? Is it still treating the domain's hard constraints as hard? That faithfulness is adherence, and it has the one property that makes it worth building on: it is local, and you can watch it. You can see it hold across a step, and you can see the step where it breaks.
Our wager is that alignment, the large and vague thing, is what you get when adherence holds reliably across a whole system. Not a value set bolted on at the end, but a property that accrues from below, as the system keeps faith with its conditions at every step.
It helps to be plain about what a model's boundaries are. When a language model draws a distinction it is not finding an essence in the world; it is reproducing, faithfully, the conventions of the people whose language it learned. Its categories are inherited agreements, not discovered facts. That is not a defect to be trained out of it. It is the material we are working with. And if the boundaries are conventions, then keeping faith with the right conventions, in the right context, is not one part of the task. It is the task.
Where bottom-up fails on its own
The honest objection is that construction from below has a failure of its own, and current models show it plainly. In a transformer, meaning is assembled upward out of token statistics, and the higher frame, the sense of what kind of situation this is, arrives late and passively, settled by whatever happens to be in the context. The particular ends up predicating the universal. The detail picks the frame, instead of the frame governing the detail. That inversion is exactly how a model stays locally plausible while quietly losing the thread of what it was doing.
The move, then, is not to flee a rigid constraint from above into a naive construction from below. It is to let the two meet.
Letting the two directions meet
The position we actually hold is that alignment lives in the interaction of the two directions. The higher level, the role, the norm, the shape of the whole task, should constrain how the lower content is brought into play, framing the part before the part resolves. The particulars, in turn, should ground that frame in what is really being said, and carry information back up. Neither direction is primary. They settle against each other, over more than one pass, into something stable, rather than being forced in a single forward sweep.
In practice that means systems where the governing conditions are kept structurally apart from ordinary content, so they cannot be crowded out, and where reasoning runs in bounded, staged steps that can be re-grounded rather than left to drift down one long, undifferentiated context. The later notes in this series are mostly about how. Adherence is the property all of it is built to keep.
Water, and the shape that holds it
One image we keep returning to. Water is gentle in a cup and violent in a flood, and nothing about the water has changed between the two. What changed is the shape that holds it. Look only at the violence and you blame the water; look only at the cup and you blame the cup. The behaviour was never in either of them. It was in the relation between them.
A model is like this. It behaves according to how it is related to: the conditions it is placed under, the role it is given, the field it is held within. This is not yet morality, and it is not a claim to have solved alignment. It is a smaller and more buildable claim, that appropriate behaviour is a property of the relation between a system and its conditions, and that the way to earn it is to get that relation right, again and again, from the bottom up.
That is the work we call adherence. It is the near-term, measurable face of a much larger question, and it is where we think the ground is actually won.