LLM Research Ignores the Actual Structure

 

Why LLM Research Ignores the Actual Structure of Semantic Clouds

Table of Contents

  1. The Missing Object
    1.1 The strange success of LLMs without a theory of what they learned
    1.2 Objective, architecture, computation, representation, and object are different things
    1.3 NEXT-TOKEN PREDICTION != INTELLIGENCE
    1.4 TRAINING OBJECTIVE != LEARNED STRUCTURE
    1.5 Why describing input-output behavior does not identify the generative object
    1.6 The semantic cloud as the missing level of description
    1.7 MODEL = learned possibility field, not database, lookup table, or bag of facts
    1.8 Why semantic clouds are invisible to research organized around measurable projections

  2. What Is a Semantic Cloud?
    2.1 From stored items to relational possibility fields
    2.2 Alternatives as the primitive of decision
    2.3 CURRENT STATE + POSSIBILITY FIELD + CONSTRAINT -> RESPONSE
    2.4 Relations, transformations, generators, accessibility, and constraint propagation
    2.5 Semantic possibility versus explicit symbolic representation
    2.6 Why the cloud is distributed rather than addressable
    2.7 Why facts need not have locations
    2.8 Why knowledge can be reconstructible without being directionally stored
    2.9 Local realization versus global semantic organization
    2.10 The cloud as a changing topology rather than a static memory

  3. Training Creates Structure, Not Merely Prediction Ability
    3.1 Training examples as repeated pressures on representation
    3.2 Recurrence across variation
    3.3 EVENT -> INVARIANT -> RELATION -> TRANSFORMATION -> GENERATOR
    3.4 Why repeated combinations matter more than repeated examples
    3.5 Detaching invariants from individual observations
    3.6 Compression as transfer of ownership
    3.7 From copying states to copying generators
    3.8 Why generalization is structural reorganization
    3.9 Why scale changes representational organization
    3.10 Semantic-cloud formation as accumulated relational compression

  4. Double Descent as a Carrier Phase Transition
    4.1 Why the usual statistical description is downstream
    4.2 Small-data regime: datapoints become representational objects
    4.3 Large-data regime: reusable structure takes ownership
    4.4 DATAPOINT CARRIER -> STRUCTURAL CARRIER
    4.5 Capacity competition between event memory and reusable generators
    4.6 Destruction of obsolete representations
    4.7 Why the loss bump is a readout, not the phenomenon
    4.8 Combinatorial recurrence as the transition driver
    4.9 Hysteresis and representational irreversibility
    4.10 TSCT: learning as ownership transfer between carriers

  5. Copy as the Primitive Behind Learning, Evolution, and Intelligence
    5.1 Why copying is more fundamental than “computation”
    5.2 Stable state, stochasticity, variation, and differential retention
    5.3 COPY STATE -> VARY -> SELECT -> RETAIN
    5.4 Evolution as externalized branching search
    5.5 Intelligence as internalized branching search
    5.6 Counterfactual copying
    5.7 Restart, branching, simulation, and discard
    5.8 Why discard is as important as generation
    5.9 COPY EVENT -> COPY FEATURE -> COPY RELATION -> COPY GENERATOR
    5.10 Intelligence as recursive replication of generative structure

  6. Prompting Is Constraint, Not Retrieval
    6.1 Why the retrieval metaphor fails
    6.2 Prompt as accessibility transformation
    6.3 CONSTRAINT change -> ACCESSIBILITY change -> RESPONSE FAMILY change
    6.4 Generated prefixes as endogenous constraints
    6.5 Why autoregression creates path dependence
    6.6 Restart as removal of self-generated constraint
    6.7 Why one prompt reveals almost nothing about total capability
    6.8 Response surfaces instead of single-response evaluation
    6.9 Prompt engineering as accidental discovery of semantic-cloud control
    6.10 Context compilation as topology engineering

  7. The First Research Error: Treating Readout as Mechanism
    7.1 Tokens are outputs, not semantic objects
    7.2 Activations are states, not necessarily mechanisms
    7.3 Features are decompositions, not necessarily native objects
    7.4 Circuits are causal slices, not necessarily global organization
    7.5 Probes recover correlations under imposed concepts
    7.6 Sparse autoencoders manufacture coordinate systems
    7.7 Jacobians capture local response, not complete generative structure
    7.8 PROJECTION != SOURCE
    7.9 REPRESENTATION != OBJECT
    7.10 How measurement convenience becomes ontology

  8. Mechanistic Interpretability Searches for Parts Before Establishing the Whole
    8.1 The decomposition assumption
    8.2 “Carving at the joints” presupposes joints
    8.3 Neurons fail as natural units
    8.4 Attention heads fail as natural units
    8.5 Layers fail as natural units
    8.6 Features inherit the same problem at a new coordinate system
    8.7 Global geometry versus bags of features
    8.8 Cross-layer and distributed representation
    8.9 Why native computation may occur in superposition
    8.10 The missing question: what is the native causal object?

  9. Mechanistic Unfaithfulness Exposes the Problem
    9.1 A faithful output can arise from an unfaithful mechanism
    9.2 The absolute-value transcoder example
    9.3 Repeated datapoints create fictitious mechanisms
    9.4 Sparsity rewards explanations that the source model never used
    9.5 LOW ERROR + CLEAN FEATURES != TRUE MECHANISM
    9.6 Interpretability systems can hallucinate ontology
    9.7 Jacobian matching as a stronger intervention constraint
    9.8 Why Jacobian matching still stops at first-order response
    9.9 Higher-order response and off-manifold divergence
    9.10 Mechanistic equivalence as preserved intervention structure

  10. The Second Research Error: Human Concepts Own the Experiment
    10.1 Task definition precedes circuit discovery
    10.2 Human labels constrain what can be found
    10.3 Concept probes recover what researchers ask for
    10.4 Dataset construction silently defines ontology
    10.5 PROMPT FRAME != SOURCE ONTOLOGY
    10.6 Human-readable does not mean model-native
    10.7 Interpretability illusions
    10.8 Why increasingly sophisticated tools can deepen the same error
    10.9 Unknown concepts cannot be discovered by fixed concept vocabularies
    10.10 Successor representations require freedom from current grammar

  11. The Third Research Error: Static Decomposition of a Developmental Process
    11.1 A trained network is the residue of a history
    11.2 Features form, merge, rotate, disappear, and transfer ownership
    11.3 Capability emergence as structural transition
    11.4 Stagewise learning
    11.5 Parameter trajectory changes
    11.6 Double descent as developmental evidence
    11.7 Why endpoint interpretation loses formation history
    11.8 Tracking transformations through training
    11.9 Identifying surviving invariants across representational rewrites
    11.10 Developmental interpretability versus post-hoc decomposition

  12. The Fourth Research Error: Locality
    12.1 Why local circuits are experimentally convenient
    12.2 Why semantic meaning is relational
    12.3 Local feature identity depends on surrounding geometry
    12.4 Global constraint propagation
    12.5 Distributed ownership
    12.6 Semantic accessibility as a global property
    12.7 LOCAL CAUSAL PATH != GLOBAL GENERATIVE ORGANIZATION
    12.8 Why local closure cannot establish global closure
    12.9 Semantic clouds as nonlocal structured fields
    12.10 The geometry of constraint as the central unsolved problem

  13. The Fifth Research Error: Confusing Accessibility with Storage
    13.1 “Known,” “encoded,” “stored,” and “recoverable” are different claims
    13.2 One-shot failure does not establish absence
    13.3 Priming changes accessibility
    13.4 Reasoning changes accessibility
    13.5 Context changes accessibility
    13.6 Finetuning can expose or mask existing capability
    13.7 NOT RECOVERED != NOT PRESENT
    13.8 Behavioral readout cannot establish internal content ontology
    13.9 Knowledge as reconstructibility across interventions
    13.10 Mapping accessibility topology instead of cataloguing stored facts

  14. Semantic Clouds and the Geometry of Constraint
    14.1 The real object of inference
    14.2 Regions, boundaries, transitions, and accessibility
    14.3 Constraint propagation across relations
    14.4 Competitive continuations
    14.5 Constraint-induced phase transitions
    14.6 Semantic attractors without assuming literal dynamical attractors
    14.7 Path dependence under generated prefixes
    14.8 Branch suppression and branch amplification
    14.9 Constraint composition
    14.10 WHAT IS THE GEOMETRY OF CONSTRAINT OVER A SEMANTIC CLOUD?

  15. From Features to Generators
    15.1 Features are intermediate compressions
    15.2 Relations explain reuse
    15.3 Transformations explain families of relations
    15.4 Generators explain families of states
    15.5 Why intelligence prefers generative compression
    15.6 Novel reconstruction from old structure
    15.7 Compositional novelty without explicit storage
    15.8 Rule formation as stable compression of response structure
    15.9 Abstractions as reusable generative operators
    15.10 Learning as migration toward increasingly generative ownership

  16. Why Scaling Laws See the Shadow but Miss the Object
    16.1 Loss as a scalar projection
    16.2 Smooth aggregate curves from structural transitions
    16.3 Quanta and discrete acquisition
    16.4 Why independent quanta are insufficient
    16.5 Hierarchical dependence among learned structures
    16.6 Polygenic prediction
    16.7 Scaling as expansion of relational reconstructibility
    16.8 Capability thresholds as accessibility transitions
    16.9 Why parameter count cannot describe semantic organization
    16.10 Scaling laws without semantic-cloud theory

  17. Why Benchmarks Reinforce the Wrong Ontology
    17.1 Benchmark = fixed task + fixed readout
    17.2 Scores collapse response surfaces
    17.3 Capability becomes identified with elicited behavior
    17.4 One-shot interfaces hide recursive search
    17.5 Instruction following hides latent alternatives
    17.6 Benchmark optimization reshapes accessibility without explaining it
    17.7 Better score can mean better access, not richer structure
    17.8 System capability versus model capability
    17.9 Why benchmark progress does not reveal what changed
    17.10 Replace benchmark points with intervention maps

  18. Why LLM Research Keeps Returning to the Wrong Objects
    18.1 Measurability selects research objects
    18.2 Tokens are easy to count
    18.3 Activations are easy to extract
    18.4 Features are easy to name
    18.5 Circuits are easy to diagram
    18.6 Benchmarks are easy to rank
    18.7 Semantic clouds are relational, distributed, developmental, and intervention-dependent
    18.8 Scientific tooling favors projections over sources
    18.9 Publication incentives reward local tractable questions
    18.10 The research program recursively optimizes what its instruments can already see

  19. A Research Program for Semantic Clouds
    19.1 Stop asking where concepts are stored
    19.2 Measure response fields across controlled constraints
    19.3 Identify minimal interventions that change accessibility
    19.4 Track ownership migration during training
    19.5 Follow invariants across representational rewrites
    19.6 Discover generators rather than merely features
    19.7 Measure combinatorial recurrence thresholds
    19.8 Test hysteresis after structural transitions
    19.9 Compare models through intervention-equivalence
    19.10 Study global constraint propagation
    19.11 Separate representation from mechanism
    19.12 Separate mechanism from semantic organization
    19.13 Let new causal objects emerge rather than predeclaring them

  20. TSCT as a Developmental Theory of Semantic Structure
    20.1 Event ownership
    20.2 Carrier competition
    20.3 Recurrence pressure
    20.4 Invariant detachment
    20.5 Ownership transfer
    20.6 Generator formation
    20.7 Carrier destruction and reconstruction
    20.8 Recursive reuse
    20.9 Semantic-cloud densification
    20.10 Intelligence as repeated transfer from outcomes toward generators

  21. From LLMs to General Intelligence
    21.1 Semantic clouds need not be linguistic
    21.2 Any decision process requires alternatives
    21.3 Alternatives require a possibility field
    21.4 Constraints select realizations
    21.5 Memory allows accumulation
    21.6 Copy allows branching
    21.7 Stochasticity creates variation
    21.8 Selection preserves consequential differences
    21.9 Recursive internal branching creates intelligence
    21.10 Intelligence as controlled evolution in possibility space

  22. The Final Inversion
    22.1 LLM research starts from outputs and works inward
    22.2 Semantic-cloud science starts from the generative field and explains outputs outward
    22.3 TOKEN is not the object
    22.4 FEATURE is not the object
    22.5 CIRCUIT is not the object
    22.6 BENCHMARK CAPABILITY is not the object
    22.7 The object is structured possibility under constraint
    22.8 Learning is transformation of that possibility structure
    22.9 Inference is traversal and realization within it
    22.10 Intelligence is recursive manipulation of the field itself

  23. Conclusion: The Science LLMs Actually Require
    23.1 From output prediction to possibility structure
    23.2 From decomposition to formation
    23.3 From features to generators
    23.4 From retrieval to accessibility
    23.5 From local circuits to global relational organization
    23.6 From static representation to developmental transformation
    23.7 From one-shot behavior to intervention surfaces
    23.8 From interpretability to semantic-cloud science
    23.9 The missing object was not hidden inside a neuron
    23.10 It was the structured field making all intelligent continuation possible

Comments

Popular posts from this blog

Semiotics Rebooted

ORSI: The Telic Geometry of Meaning

THE COLLAPSE ENGINE: AI, Capital, and the Terminal Logic of 2025