LLM Research Ignores the Actual Structure
Why LLM Research Ignores the Actual Structure of Semantic Clouds
Table of Contents
The Missing Object
1.1 The strange success of LLMs without a theory of what they learned
1.2 Objective, architecture, computation, representation, and object are different things
1.3NEXT-TOKEN PREDICTION != INTELLIGENCE
1.4TRAINING OBJECTIVE != LEARNED STRUCTURE
1.5 Why describing input-output behavior does not identify the generative object
1.6 The semantic cloud as the missing level of description
1.7MODEL = learned possibility field, not database, lookup table, or bag of facts
1.8 Why semantic clouds are invisible to research organized around measurable projectionsWhat Is a Semantic Cloud?
2.1 From stored items to relational possibility fields
2.2 Alternatives as the primitive of decision
2.3CURRENT STATE + POSSIBILITY FIELD + CONSTRAINT -> RESPONSE
2.4 Relations, transformations, generators, accessibility, and constraint propagation
2.5 Semantic possibility versus explicit symbolic representation
2.6 Why the cloud is distributed rather than addressable
2.7 Why facts need not have locations
2.8 Why knowledge can be reconstructible without being directionally stored
2.9 Local realization versus global semantic organization
2.10 The cloud as a changing topology rather than a static memoryTraining Creates Structure, Not Merely Prediction Ability
3.1 Training examples as repeated pressures on representation
3.2 Recurrence across variation
3.3EVENT -> INVARIANT -> RELATION -> TRANSFORMATION -> GENERATOR
3.4 Why repeated combinations matter more than repeated examples
3.5 Detaching invariants from individual observations
3.6 Compression as transfer of ownership
3.7 From copying states to copying generators
3.8 Why generalization is structural reorganization
3.9 Why scale changes representational organization
3.10 Semantic-cloud formation as accumulated relational compressionDouble Descent as a Carrier Phase Transition
4.1 Why the usual statistical description is downstream
4.2 Small-data regime: datapoints become representational objects
4.3 Large-data regime: reusable structure takes ownership
4.4DATAPOINT CARRIER -> STRUCTURAL CARRIER
4.5 Capacity competition between event memory and reusable generators
4.6 Destruction of obsolete representations
4.7 Why the loss bump is a readout, not the phenomenon
4.8 Combinatorial recurrence as the transition driver
4.9 Hysteresis and representational irreversibility
4.10 TSCT: learning as ownership transfer between carriersCopy as the Primitive Behind Learning, Evolution, and Intelligence
5.1 Why copying is more fundamental than “computation”
5.2 Stable state, stochasticity, variation, and differential retention
5.3COPY STATE -> VARY -> SELECT -> RETAIN
5.4 Evolution as externalized branching search
5.5 Intelligence as internalized branching search
5.6 Counterfactual copying
5.7 Restart, branching, simulation, and discard
5.8 Why discard is as important as generation
5.9COPY EVENT -> COPY FEATURE -> COPY RELATION -> COPY GENERATOR
5.10 Intelligence as recursive replication of generative structurePrompting Is Constraint, Not Retrieval
6.1 Why the retrieval metaphor fails
6.2 Prompt as accessibility transformation
6.3CONSTRAINT change -> ACCESSIBILITY change -> RESPONSE FAMILY change
6.4 Generated prefixes as endogenous constraints
6.5 Why autoregression creates path dependence
6.6 Restart as removal of self-generated constraint
6.7 Why one prompt reveals almost nothing about total capability
6.8 Response surfaces instead of single-response evaluation
6.9 Prompt engineering as accidental discovery of semantic-cloud control
6.10 Context compilation as topology engineeringThe First Research Error: Treating Readout as Mechanism
7.1 Tokens are outputs, not semantic objects
7.2 Activations are states, not necessarily mechanisms
7.3 Features are decompositions, not necessarily native objects
7.4 Circuits are causal slices, not necessarily global organization
7.5 Probes recover correlations under imposed concepts
7.6 Sparse autoencoders manufacture coordinate systems
7.7 Jacobians capture local response, not complete generative structure
7.8PROJECTION != SOURCE
7.9REPRESENTATION != OBJECT
7.10 How measurement convenience becomes ontologyMechanistic Interpretability Searches for Parts Before Establishing the Whole
8.1 The decomposition assumption
8.2 “Carving at the joints” presupposes joints
8.3 Neurons fail as natural units
8.4 Attention heads fail as natural units
8.5 Layers fail as natural units
8.6 Features inherit the same problem at a new coordinate system
8.7 Global geometry versus bags of features
8.8 Cross-layer and distributed representation
8.9 Why native computation may occur in superposition
8.10 The missing question: what is the native causal object?Mechanistic Unfaithfulness Exposes the Problem
9.1 A faithful output can arise from an unfaithful mechanism
9.2 The absolute-value transcoder example
9.3 Repeated datapoints create fictitious mechanisms
9.4 Sparsity rewards explanations that the source model never used
9.5LOW ERROR + CLEAN FEATURES != TRUE MECHANISM
9.6 Interpretability systems can hallucinate ontology
9.7 Jacobian matching as a stronger intervention constraint
9.8 Why Jacobian matching still stops at first-order response
9.9 Higher-order response and off-manifold divergence
9.10 Mechanistic equivalence as preserved intervention structureThe Second Research Error: Human Concepts Own the Experiment
10.1 Task definition precedes circuit discovery
10.2 Human labels constrain what can be found
10.3 Concept probes recover what researchers ask for
10.4 Dataset construction silently defines ontology
10.5PROMPT FRAME != SOURCE ONTOLOGY
10.6 Human-readable does not mean model-native
10.7 Interpretability illusions
10.8 Why increasingly sophisticated tools can deepen the same error
10.9 Unknown concepts cannot be discovered by fixed concept vocabularies
10.10 Successor representations require freedom from current grammarThe Third Research Error: Static Decomposition of a Developmental Process
11.1 A trained network is the residue of a history
11.2 Features form, merge, rotate, disappear, and transfer ownership
11.3 Capability emergence as structural transition
11.4 Stagewise learning
11.5 Parameter trajectory changes
11.6 Double descent as developmental evidence
11.7 Why endpoint interpretation loses formation history
11.8 Tracking transformations through training
11.9 Identifying surviving invariants across representational rewrites
11.10 Developmental interpretability versus post-hoc decompositionThe Fourth Research Error: Locality
12.1 Why local circuits are experimentally convenient
12.2 Why semantic meaning is relational
12.3 Local feature identity depends on surrounding geometry
12.4 Global constraint propagation
12.5 Distributed ownership
12.6 Semantic accessibility as a global property
12.7LOCAL CAUSAL PATH != GLOBAL GENERATIVE ORGANIZATION
12.8 Why local closure cannot establish global closure
12.9 Semantic clouds as nonlocal structured fields
12.10 The geometry of constraint as the central unsolved problemThe Fifth Research Error: Confusing Accessibility with Storage
13.1 “Known,” “encoded,” “stored,” and “recoverable” are different claims
13.2 One-shot failure does not establish absence
13.3 Priming changes accessibility
13.4 Reasoning changes accessibility
13.5 Context changes accessibility
13.6 Finetuning can expose or mask existing capability
13.7NOT RECOVERED != NOT PRESENT
13.8 Behavioral readout cannot establish internal content ontology
13.9 Knowledge as reconstructibility across interventions
13.10 Mapping accessibility topology instead of cataloguing stored factsSemantic Clouds and the Geometry of Constraint
14.1 The real object of inference
14.2 Regions, boundaries, transitions, and accessibility
14.3 Constraint propagation across relations
14.4 Competitive continuations
14.5 Constraint-induced phase transitions
14.6 Semantic attractors without assuming literal dynamical attractors
14.7 Path dependence under generated prefixes
14.8 Branch suppression and branch amplification
14.9 Constraint composition
14.10WHAT IS THE GEOMETRY OF CONSTRAINT OVER A SEMANTIC CLOUD?From Features to Generators
15.1 Features are intermediate compressions
15.2 Relations explain reuse
15.3 Transformations explain families of relations
15.4 Generators explain families of states
15.5 Why intelligence prefers generative compression
15.6 Novel reconstruction from old structure
15.7 Compositional novelty without explicit storage
15.8 Rule formation as stable compression of response structure
15.9 Abstractions as reusable generative operators
15.10 Learning as migration toward increasingly generative ownershipWhy Scaling Laws See the Shadow but Miss the Object
16.1 Loss as a scalar projection
16.2 Smooth aggregate curves from structural transitions
16.3 Quanta and discrete acquisition
16.4 Why independent quanta are insufficient
16.5 Hierarchical dependence among learned structures
16.6 Polygenic prediction
16.7 Scaling as expansion of relational reconstructibility
16.8 Capability thresholds as accessibility transitions
16.9 Why parameter count cannot describe semantic organization
16.10 Scaling laws without semantic-cloud theoryWhy Benchmarks Reinforce the Wrong Ontology
17.1 Benchmark = fixed task + fixed readout
17.2 Scores collapse response surfaces
17.3 Capability becomes identified with elicited behavior
17.4 One-shot interfaces hide recursive search
17.5 Instruction following hides latent alternatives
17.6 Benchmark optimization reshapes accessibility without explaining it
17.7 Better score can mean better access, not richer structure
17.8 System capability versus model capability
17.9 Why benchmark progress does not reveal what changed
17.10 Replace benchmark points with intervention mapsWhy LLM Research Keeps Returning to the Wrong Objects
18.1 Measurability selects research objects
18.2 Tokens are easy to count
18.3 Activations are easy to extract
18.4 Features are easy to name
18.5 Circuits are easy to diagram
18.6 Benchmarks are easy to rank
18.7 Semantic clouds are relational, distributed, developmental, and intervention-dependent
18.8 Scientific tooling favors projections over sources
18.9 Publication incentives reward local tractable questions
18.10 The research program recursively optimizes what its instruments can already seeA Research Program for Semantic Clouds
19.1 Stop asking where concepts are stored
19.2 Measure response fields across controlled constraints
19.3 Identify minimal interventions that change accessibility
19.4 Track ownership migration during training
19.5 Follow invariants across representational rewrites
19.6 Discover generators rather than merely features
19.7 Measure combinatorial recurrence thresholds
19.8 Test hysteresis after structural transitions
19.9 Compare models through intervention-equivalence
19.10 Study global constraint propagation
19.11 Separate representation from mechanism
19.12 Separate mechanism from semantic organization
19.13 Let new causal objects emerge rather than predeclaring themTSCT as a Developmental Theory of Semantic Structure
20.1 Event ownership
20.2 Carrier competition
20.3 Recurrence pressure
20.4 Invariant detachment
20.5 Ownership transfer
20.6 Generator formation
20.7 Carrier destruction and reconstruction
20.8 Recursive reuse
20.9 Semantic-cloud densification
20.10 Intelligence as repeated transfer from outcomes toward generatorsFrom LLMs to General Intelligence
21.1 Semantic clouds need not be linguistic
21.2 Any decision process requires alternatives
21.3 Alternatives require a possibility field
21.4 Constraints select realizations
21.5 Memory allows accumulation
21.6 Copy allows branching
21.7 Stochasticity creates variation
21.8 Selection preserves consequential differences
21.9 Recursive internal branching creates intelligence
21.10 Intelligence as controlled evolution in possibility spaceThe Final Inversion
22.1 LLM research starts from outputs and works inward
22.2 Semantic-cloud science starts from the generative field and explains outputs outward
22.3TOKENis not the object
22.4FEATUREis not the object
22.5CIRCUITis not the object
22.6BENCHMARK CAPABILITYis not the object
22.7 The object is structured possibility under constraint
22.8 Learning is transformation of that possibility structure
22.9 Inference is traversal and realization within it
22.10 Intelligence is recursive manipulation of the field itselfConclusion: The Science LLMs Actually Require
23.1 From output prediction to possibility structure
23.2 From decomposition to formation
23.3 From features to generators
23.4 From retrieval to accessibility
23.5 From local circuits to global relational organization
23.6 From static representation to developmental transformation
23.7 From one-shot behavior to intervention surfaces
23.8 From interpretability to semantic-cloud science
23.9 The missing object was not hidden inside a neuron
23.10 It was the structured field making all intelligent continuation possible
Comments
Post a Comment