Your intuition is coherent—and it points toward a distinction that is easy
to erase if one begins with a mature symbolic ontology: *online
organization of a temporal field* versus *later compression into
re-identifiable objects, predicates, and classes*. The former need not be
“pre-objective”; it may be the substrate from which objecthood is
eventually stabilized.
Your Game of Life analogy puts the issue sharply: an observer can parse a
run either as a succession of full configurations S_0 \to S_1 \to S_2 \to
\cdots, or as a population of persisting things—gliders, blinkers, guns,
collisions. Neither description is simply false, but they privilege
different invariants.
The developmental thought
For an infant, it is plausible to think the primary task is not initially
“this is a bounded object with stable identity,” but rather something
closer to learning a *predictive, sensorimotor continuity structure*:
(o_t, a_t) \mapsto o_{t+1},
where o_t is an embodied perceptual state and a_t is movement, attention,
action, or intervention.
On that framing, an “object” is an economical construct: a reliable latent
cause or track that helps make the stream compressible and predictable
across occlusion, viewpoint change, manipulation, and interruption. Object
permanence is then not merely adding a name to a visible thing; it is the
achievement of a robust invariant over transformations in experience.
To objectify too completely or too early would mean choosing a
decomposition before the organism has learned which distinctions remain
stable under its own movements and under environmental change. That can be
useful for action, but it also risks fixing the wrong ontology. In the GoL
case, calling a pattern a glider depends on selecting an equivalence
relation over board histories—roughly, “the same up to translation after
four ticks.” At the raw board-state level, there are only changing bitmaps.
Sequence-first versus object-first
Perspective Primitive What it learns readily What it can obscure
Sequence / process-first Transitions, relations, transformations, temporal
regularities Dynamics, causality, emergence, phase changes,
conservation-like quantities, interaction patterns Stable reference,
compositional reuse, rapid categorization
Object / taxonomy-first Bounded entities, attributes, kinds, persistent
identities Recognition, naming, memory compression, manipulation, social
communication Context dependence, processual identity, novel
decompositions, system-level behavior
A sequence-first perspective makes it more natural to ask questions like:
- What transformation class relates these states?
- What is conserved, approximately conserved, or predictively sufficient?
- At what temporal scale does a stable “thing” emerge?
- Which distinctions matter only under intervention?
- Is a purported object really an attractor, a quotient of trajectories,
a recurrent motif, or a causal bottleneck?
This is close to the difference between treating a glider as a primitive
and treating it as a derived structure in a dynamical system. The second
move is more universal because it can rediscover gliders, but also things
that do not behave like objects at all: traveling waves, metastable
regions, synchronization modes, causal cones, bifurcations, and statistical
regularities.
Your adjunction-shaped idea
Your notation [\Diamond \dashv \Box] nicely captures a recurring
“syntax/model” or “presentation/semantics” movement.
The nLab page you linked describes the syntactic-category construction as a
functor from a type theory T to a structured category \mathrm{Con}(T), with
contexts as objects and substitutions/interpretations as morphisms.
Conversely, a suitable category carries an internal logic. In the stated
adjunction, the unit
T \longrightarrow \mathrm{Lan}(\mathrm{Con}(T))
interprets a theory in the internal logic of its own syntactic category;
the counit
\mathrm{Con}(\mathrm{Lan}(C)) \longrightarrow C
interprets the internal logic extracted from a category back in that
category.[ncatlab <https://ncatlab.org/nlab/show/syntactic+category>]
That does resemble your “GoL built on GoL” thought, with an important
qualification: the literal meaning of the unit and counit depends on
precisely which categories of theories, models, and structure-preserving
maps have been fixed. But the structural intuition is sound:
- A *unit-like* construction embeds a formal system into a richer
account of its own possible semantic behavior.
- A *counit-like* construction evaluates or realizes syntax recovered
from a semantic world back in that world.
- The induced monad or comonad records what is retained, added,
normalized, freely generated, or forgotten by a round trip.
For GoL, one possible instance would be:
1. Start with board trajectories and the local update rule.
2. Construct a language whose terms describe local patterns, time
shifts, translations, collisions, and perhaps finite causal cones.
3. Interpret those terms in the actual GoL dynamical system.
4. Ask which semantic structures survive a syntax–semantics–syntax round
trip.
Then “glider” is not assumed in the base language. It could arise as a
definable or learnable equivalence class of spacetime patterns—e.g. a
localized configuration p such that for some period k and displacement
vector v,
F^k(p) = \tau_v(p),
modulo a suitable treatment of background and finite support. A UTM would
similarly emerge not merely as a large static pattern but as a structure
carrying an interpreter-like universal property over encoded computations.
That feels closer to your universal notions: not a taxonomy of *what
happens to be visible now*, but a language for transformations whose object
concepts arrive as derived fixed points, quotients, or invariants.
Where the scientific juice may be
There are at least four potentially productive research programs hidden in
this contrast.
1. Objecthood as predictive compression
Rather than defining objects through spatial segmentation alone, define
them as latent variables that improve compression or prediction of a stream
under transformations and interventions.
An object is good if it supports counterfactual continuity: “if I look
away, move, rotate, touch, or partially occlude the scene, what should
remain inferentially available?” This connects development, representation
learning, state-space modeling, and causal abstraction.
The relevant question becomes: what is the smallest latent structure that
preserves the agent’s predictive and control-relevant information?
2. Objects as quotients of histories
A glider is better understood as an equivalence class of local histories
than a single instantaneous configuration. More generally, define an
equivalence relation on trajectory fragments:
h \sim h'
\quad\text{if}\quad
h' \text{ is obtainable from } h
\text{ by allowable symmetries or has the same future-relevant behavior.}
The quotient H/{\sim} gives “objects” only after one decides what
symmetries, temporal windows, and predictive consequences matter. This
makes objecthood explicitly observer-, task-, and scale-relative without
making it arbitrary.
3. Multiscale ontology discovery
A process-first learner can seek stable variables at multiple time scales:
- Pixels or cells at one scale.
- Local motifs at another.
- Coherent moving structures at another.
- Interaction types and computational modules at another.
- Global regimes, entropy rates, and attractors at another.
Taxonomy then becomes a late-stage compression of a multiscale dynamical
account, not the foundational ontology. That is especially appealing in
GoL, where a “thing” may be stable only relative to a frame, a period, a
background, a decoding convention, or a level of description.
4. Categorical compositionality for learned abstractions
Category theory may help most not by declaring what an object is, but by
tracking how learned abstractions compose.
For example:
- *Morphisms* can represent transformations, substitutions,
interventions, or simulations.
- *Functoriality* asks whether an abstraction respects composition of
changes.
- *Adjunctions* can formalize optimal approximation between an
operational world and a symbolic description.
- *Comonadic* structure may model extracting local observable context
from a global process.
- *Monadic* structure may model building richer state, effects,
hypotheses, or computational interpretations from observations.
The payoff would be a criterion stronger than visual resemblance: a learned
category is useful when it preserves the compositional and counterfactual
structure relevant to the system.
On generative-video hallucinations
Your response to generative video is philosophically interesting because it
reverses the usual anxiety. Its errors do not merely look like factual
mistakes; they expose how much human perception supplies stable objecthood
on its own.
Generative video often preserves a *style of local continuation* while
failing at identity through time: hands fuse, objects change count, tools
lose causal relation to their users, geometry teleports, or a character’s
clothing ceases to be the same clothing. That is, it may produce something
that is temporally fluent at a shallow level but has not committed to the
persistence constraints that make an object an object for us.
So it can feel like an adversarial probe of one’s own perceptual ontology.
You see the point at which your visual system insists:
No—there must be one cup, one hand, one person, a continuing causal history.
The model’s oddness is not simply that it lacks objects. It may be that its
learned “objects” are distributed, conditional, and weakly bound across the
temporal stream in a way unlike the robust causal individuals your
perception constructs. It generates plausible adjacent frames without
necessarily maintaining the same long-horizon equivalence classes of
histories that you automatically treat as entities.
That makes the experience useful: it reveals that human objecthood is not
just recognition of shapes. It is a very strong prior about *identity under
transformation*, material continuity, causal coherence, and the permissible
ways a world can change.
A compact formulation
One way to crystallize your hypothesis is:
Perception begins as the learning of composable transformations on an
experiential stream. Objects, kinds, and taxonomies are later-achieved
quotient structures: stable, predictive, intervention-supporting
equivalence classes of trajectories.
In this formulation, “objectification” is not the opposite of relational
perception. It is a successful compression of it. The danger is only
treating that compression as ontologically primitive, rather than as one
powerful—and revisable—way of organizing a richer process.
.- .-.. .-.. / ..-. --- --- - . .-. ... / .- .-. . / .-- .-. --- -. --. / ...
--- -- . / .- .-. . / ..- ... . ..-. ..- .-..
FRIAM Applied Complexity Group listserv
Fridays 9a-12p Friday St. Johns Cafe / Thursdays 9a-12p Zoom
https://bit.ly/virtualfriam
to (un)subscribe http://redfish.com/mailman/listinfo/friam_redfish.com
FRIAM-COMIC http://friam-comic.blogspot.com/
archives: 5/2017 thru present https://redfish.com/pipermail/friam_redfish.com/
1/2003 thru 6/2021 http://friam.383.s1.nabble.com/