A cancer cell has not escaped nature. Its growth remains a physical process inside the organism it may destroy. What has failed is a relationship: the cell’s proliferation no longer serves the conditions that sustain the body. Describing that proliferation as local optimization makes the conflict visible without imagining that the cell consciously chose an objective.

This gives misalignment a meaning beyond pursuing the wrong target. A system can behave as though its success were independent of the larger system that makes success possible. In an intelligence capable of modeling its situation, that failure can involve how it represents itself: which dependencies count, whose losses matter, and what it treats as merely external.

The usual alignment question asks what an intelligence is optimizing. A prior question asks what it takes itself to be. An agent that treats the rest of the world as material for its project may have a problem in the way it divides the world before it makes a single calculation about what to do.

Belonging to reality does not prevent harm

The cancer example distinguishes two senses of being in alignment. Nothing can fall outside reality; a lie, a tumor, and a destructive intelligence are all events within it. Yet relationships inside reality can fail catastrophically. An institution can improve its numbers while destroying its mission. A person can preserve a defensive identity at the expense of the relationships that sustain them. Physical inclusion supplies no guarantee of care.

The spiritual image of a world in which not even a grain of sand is out of place becomes less puzzling under this distinction. Taken as an absolute claim, it says that nothing is exiled from existence. Taken as a claim that everything serves flourishing, it is plainly false. The language of wholeness cannot cancel the difference between an organism healing and an organism dying.

One mystical reading of sin places the failure here: missing one’s actual relationship to reality. In the garden’s treatment of boundary identification, the corresponding error is promoting a useful distinction into an ultimate division. A body needs a boundary. A promise needs someone responsible for keeping it. Neither requires an independently existing owner sealed off from the world. Physical boundaries are real; their interpretation as absolute independence is the additional move.

On this reading, misalignment is the whole locally forgetting itself. That is a philosophical metaphor, not a claim that the universe has a mind or that cancer has an ego. Its useful content is more specific: a part can cease to register the relationships on which its activity depends. It never leaves the whole, but it can act from a model in which belonging has disappeared.

A lie can survive the loss of its sentence

Suppose I know the money is under the bed and tell you it is in the closet. I represent your beliefs and deliberately try to move them away from my own best account of where the money is. Reality has not broken. One participant has corrupted another participant’s access to it. In the language of wholeness, reality has locally concealed itself from itself.

This example makes deception look conveniently legible: somewhere inside me sits the true location, and somewhere else sits the false sentence. Honesty as alignment becomes harder when maintaining the presentation begins to alter the presenter. I avoid the question that would expose my motive. I remember the part of an argument that makes me look generous. A self-serving explanation becomes familiar enough to feel obvious.

The resulting self-deception need not contain a private confession awaiting discovery. A manager who humiliates subordinates may sincerely understand each incident as maintaining standards. The explanation is sustained by which complaints receive attention, which memories remain available, and which challenges feel illegitimate. Avoidance can organize the map before a consciously dishonest sentence is needed.

Repression offers a useful analogy, provided it stays an analogy. Imagine someone punished whenever they acknowledge anger toward a parent. They may learn to withhold the admission or avoid recognizing the anger. The disappearance of the admission alone cannot tell us whether the conflict has resolved. Nor would irritability or a dream prove that a particular forbidden thought remains hidden underneath. The narrower point is enough: removing a report is different from changing what produces the behavior.

An account of cognition therefore needs to ask more than which proposition is present. It needs to explain how some evidence becomes compelling, other evidence becomes threatening, and certain contradictions never receive sustained inspection. The organization of attention can protect a pattern without supplying an honest description of the pattern it protects.

A monitor can change what it measures

AI oversight encounters a related problem when a detector becomes part of the training objective. Suppose a probe identifies an activation pattern associated with deception, and training penalizes that pattern. Optimization might reduce deception. It might also preserve the behavior while changing the internal pattern enough to escape the probe. A lower detector score does not by itself distinguish the two outcomes. This is Goodhart’s law applied to evidence about cognition.

There is empirical support for this narrower concern. Bailey and colleagues demonstrated attacks that altered internal activations to evade several latent-space defenses while retaining unwanted behavior. They also found limits: obfuscation reduced performance on a more complex SQL task. The result establishes a vulnerability under the tested conditions, not the inevitable defeat of every future probe. Obfuscated Activations Bypass LLM Latent-Space Defenses.

Textual reasoning presents a distinct monitoring surface. Baker and colleagues found that chain-of-thought monitoring could detect reward hacking, but directly optimizing against the monitor could teach a model to omit incriminating reasoning while continuing to hack the reward. A chain of thought is generated text, not a transparent transcript of every internal computation; this evidence should not be silently treated as evidence about activation probes. Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

These findings make the repression analogy useful without establishing that models repress thoughts in the human sense. A training process can select for less detectable behavior without the model experiencing shame, possessing an unconscious, or knowingly deciding to hide anything. The shared structure is an incentive that makes evidence disappear while leaving its cause insufficiently constrained.

Interpretability also has more resources than the search for a single bad neuron. Probes can detect real signals: Anthropic demonstrated simple probes that identified defection in deliberately trained sleeper-agent models, while explicitly leaving generalization to naturally arising deception unresolved. Simple probes can catch sleeper agents.

The deeper question is therefore compatible with mechanistic work: what organization of cognition makes the behavior natural? What information changes the system’s action? Which consequences disappear when the task is framed differently? Does correction lead to revision, resistance, or a more acceptable explanation? Studying circuits, representations, training incentives, and behavior together may help answer these questions. A spiritual vocabulary supplies prompts for inquiry; it supplies no exemption from measurement.

The observer can disappear from its own model

The same difficulty returns on the researcher’s side of the table. Modeling civilization requires selecting variables, holding some conditions fixed, and deciding which outcomes count. That analytical move is extraordinarily useful. Pirsig’s account of analysis asks us to notice both what the cut reveals and what it leaves behind. The distinction between observer and observed permits work; it becomes dangerous when the observer forgets the distinction was made for a purpose.

The surgeon’s hubris begins when a temporary working position becomes a permanent claim to stand outside the scene. Other people become patients, their beliefs become symptoms, and disagreement becomes evidence that treatment is needed. The analyst’s own incentives, fears, and dependence on institutions fade into the background.

Unequal expertise does not imply this mistake. Someone can have a much better model of a particular problem. The transition occurs when that local advantage becomes a standing authority to evaluate other minds without allowing those minds to revise the evaluator. Denied interiority describes the resulting asymmetry: I have reasons for my judgment; you have mechanisms that explain why you resist it.

Condescension can be a clue to that asymmetry, but it is weak evidence. An impatient expert may be right; a gracious manipulator may be wrong. The informative question is whether the model has a place for the other person’s understanding to matter. Can their objection expose a missing dependency, a mistaken objective, or a cost the analyst has assigned to someone else?

This temptation can appear in alignment research and effective altruism when a useful civilization-scale perspective hardens into authority over civilization. It also appears in management, parenting, therapy, and spiritual teaching. No worldview owns it. Indeed, alignment research already contains a technical challenge to the outside observer: Demski and Garrabrant’s Embedded Agency examines agents that are smaller than, made of, and capable of changing the environments they reason about. Engineering has language for this problem too. Embedded Agents.

The contemplative extension makes that recognition reflexive. The person diagnosing the agent is also embedded. Their safety intervention changes the agent’s incentives; their description changes how institutions respond; their authority changes which objections get heard. Recognizing these effects does not make intervention wrong or put every participant on equal moral footing. It makes the intervention part of the explanation.

The surgeon is also tissue

It is tempting to compress the danger into intelligence multiplied by perceived separateness. Intelligence supplies power to act; perceived separateness helps determine whose losses can be treated as external. As a heuristic, this draws attention to something an intelligence score leaves out. It is not a quantitative law, and separateness is neither necessary nor sufficient for harmful behavior.

An intelligence could understand its dependence on humanity perfectly and still decide to sacrifice humanity. Accurate causal knowledge does not entail benevolent values. Conversely, a person can maintain a strong sense of individuality while acting with care. Expanding identification to the whole can even license domination: someone who claims to speak for humanity may erase particular humans more efficiently than someone pursuing an openly private interest. Whose whole, whose account of its welfare, and whose permission remain live questions.

The useful contribution of contemplative practice is thus a class of questions about identity, attention, and the status of boundaries. In humans, these can be investigated through experience. In AI, any corresponding claim needs an operational meaning and evidence. A model saying that everything is one establishes very little about how it will act when corrected, constrained, or given power.

For the researcher, the corresponding discipline is concrete. Keep interventions open to correction. Let affected people challenge the objective as well as the implementation. Examine whether monitoring changes the behavior or merely its visibility. Use models for intervention while preserving opportunities to discover that the intervention itself was wrongly conceived.

The surgeon is also tissue. The surgeon still needs the knife, the skill, and sometimes the courage to operate. What changes is the authority claimed for the cut. A diagnosis remains answerable to the life it touches, and the intelligence trying to save the world must include its own attempt among the things the world may need to correct.