Kullback–Leibler Divergence: From Relative Entropy to a Theory of Distinguishability

 

Kullback–Leibler Divergence: Basic Geometry and Dynamics of Probability Distributions Under Inference, Transport, Coarse-Graining, and Model Change

Prologue — From Relative Entropy to a Theory of Distinguishability

0.1 Probability distributions as states of uncertainty
0.2 Comparison before distance
0.3 Why KL divergence is directional
0.4 D(P‖Q) as expected log-likelihood ratio
0.5 Distinguishability rather than geometric distance
0.6 Probability law ≠ density ≠ parameterization ≠ statistical model
0.7 Transformation of distributions as the central dynamical operation
0.8 Four basic transformation regimes: inference, transport, coarse-graining, model change
0.9 Preservation, contraction, amplification, and destruction of distinguishability
0.10 Local geometry versus global divergence
0.11 Empirical fluctuations versus dynamical trajectories
0.12 Relative entropy as the canonical empirical-measure large-deviation cost
0.13 Beyond large deviations: path ensembles, thermodynamic structure, information flow, renormalization, and model evolution
0.14 Central thesis: probability states → distinguishability → constrained transformation → information residue → reconstruction

PART I — THE BASIC PROBABILITY FIELD

1. Probability Laws Before KL

1.1 Measurable spaces
1.2 Probability measures
1.3 Densities as representations
1.4 Dominating measures
1.5 Absolute continuity
1.6 Singular measures
1.7 Support structure
1.8 Product distributions
1.9 Marginals
1.10 Conditionals
1.11 Pushforwards
1.12 Markov kernels
1.13 Statistical models as families of probability laws
1.14 Parameters as coordinates rather than ontology
1.15 Distribution space as the primary comparison carrier

2. The Likelihood-Ratio Field

2.1 Radon–Nikodym derivative dP/dQ
2.2 Pointwise likelihood ratio
2.3 Log-likelihood ratio
2.4 Evidence accumulation
2.5 Sequential log evidence
2.6 Additivity under independent composition
2.7 Conditional likelihood ratios
2.8 Change of reference measure
2.9 Support mismatch
2.10 Infinite evidence
2.11 Likelihood ratio as the precursor of statistical distinguishability
2.12 From pointwise discrimination to expected discrimination

PART II — FORMATION OF RELATIVE ENTROPY

3. Kullback–Leibler Divergence

3.1 Definition of D(P‖Q)
3.2 Discrete distributions
3.3 Continuous distributions
3.4 General measure-theoretic form
3.5 Non-negativity
3.6 Equality condition
3.7 Directionality
3.8 Lack of triangle inequality
3.9 Finiteness conditions
3.10 Joint convexity
3.11 Lower semicontinuity
3.12 Product additivity
3.13 Chain rule
3.14 Conditional KL
3.15 Marginal KL
3.16 Mutual information as relative entropy
3.17 Multi-information
3.18 Relative entropy rate

4. KL as Basic Distinguishability

4.1 Expected evidence against Q when P generates data
4.2 Excess logarithmic prediction loss
4.3 Excess coding length
4.4 Model mismatch cost
4.5 Information gain
4.6 Projection cost
4.7 Fluctuation cost
4.8 Thermodynamic excess
4.9 Statistical discrimination rate
4.10 Why these interpretations share one mathematical architecture

PART III — BASIC GEOMETRY OF PROBABILITY DISTRIBUTIONS

5. Infinitesimal KL Geometry

5.1 Nearby distributions
5.2 Expansion around a reference law
5.3 Vanishing first-order contribution
5.4 Quadratic term
5.5 Fisher information
5.6 Fisher information matrix
5.7 Statistical line element
5.8 Local distinguishability ellipsoids
5.9 Coordinate invariance
5.10 Fisher–Rao metric
5.11 Local equivalence of smooth statistical divergences
5.12 Why local equivalence does not imply global equivalence

6. Fisher–Rao Geometry

6.1 Statistical manifolds
6.2 Tangent vectors as infinitesimal probability changes
6.3 Fisher metric
6.4 Statistical geodesics
6.5 Geodesic distance
6.6 Curvature
6.7 Isometry groups
6.8 Geodesic completeness
6.9 Boundary geometry
6.10 Product manifolds
6.11 Location families
6.12 Scale families
6.13 Location-scale families

7. Higher-Order Geometry

7.1 KL beyond second order
7.2 Third-order asymmetry
7.3 Amari–Chentsov tensor
7.4 Statistical skewness
7.5 Alpha-connections
7.6 Exponential connection
7.7 Mixture connection
7.8 Levi-Civita connection
7.9 Dual affine structure
7.10 Curvature of statistical connections
7.11 Local orientation retained beyond Fisher geometry

PART IV — CONVEX GEOMETRY AND INFORMATION PROJECTION

8. KL as a Bregman-Type Geometry

8.1 Convex potentials
8.2 Legendre duality
8.3 Bregman divergences
8.4 KL in exponential families
8.5 Natural parameters
8.6 Expectation parameters
8.7 Cumulant potential
8.8 Convex conjugates
8.9 Canonical divergence
8.10 Generalized Pythagorean relations

9. Projection Geometry

9.1 Information projection
9.2 Reverse information projection
9.3 Convex constraint sets
9.4 Moment constraints
9.5 Maximum entropy
9.6 Maximum likelihood
9.7 Projection onto exponential families
9.8 Projection onto mixture families
9.9 Alternating projection
9.10 Iterative proportional fitting
9.11 Projection residue as model inadequacy
9.12 Projection failure as pressure for model change

PART V — TRANSFORMATION GEOMETRY

10. Invertible Transport

10.1 Pushforward by bijections
10.2 Change of variables
10.3 Jacobian cancellation
10.4 KL invariance under common invertible transport
10.5 Group actions on statistical families
10.6 Orbit structure
10.7 Stabilizers
10.8 Pair-orbit invariants
10.9 Source-preserving transformations
10.10 Divergences as orbit readouts

11. Stochastic Transport

11.1 Markov kernels
11.2 Noisy channels
11.3 Transition operators
11.4 Markov semigroups
11.5 Relative entropy contraction
11.6 Data-processing inequality
11.7 Equality conditions
11.8 Sufficient statistics
11.9 Recoverability
11.10 Approximate recoverability
11.11 Information loss under stochastic maps

12. Distinguishability Under Transformation

12.1 Invertible transport → preserve
12.2 Sufficient compression → preserve relevant distinctions
12.3 Markov noise → contract
12.4 Coarse-graining → destroy selected distinctions
12.5 Model projection → retain only representable distinctions
12.6 Model expansion → recover previously inaccessible distinctions
12.7 Transformation families as orders on statistical experiments
12.8 Blackwell informativeness
12.9 Statistical experiment comparison
12.10 Distinguishability as a transformation-sensitive resource

PART VI — COARSE-GRAINING

13. Coarse-Graining as Information Loss

13.1 Many-to-one maps
13.2 Marginalization
13.3 Binning
13.4 Quantization
13.5 Hidden variables
13.6 Partial observation
13.7 Macroscopic observables
13.8 Projection onto sufficient statistics
13.9 KL contraction
13.10 Irrecoverable distinctions

14. Recoverability and Sufficiency

14.1 Exact sufficient statistics
14.2 Minimal sufficient statistics
14.3 Reverse reconstruction maps
14.4 Equality in data processing
14.5 Approximate sufficiency
14.6 Information bottlenecks
14.7 Relevant versus discarded information
14.8 Reconstruction error
14.9 Coarse-grained equivalence classes
14.10 Residual distinguishability hidden below the coarse grain

PART VII — INFERENCE AS DISTRIBUTIONAL DYNAMICS

15. Bayesian Updating

15.1 Prior as initial probability state
15.2 Likelihood as interaction
15.3 Posterior as transformed state
15.4 Bayes rule as probability transport
15.5 KL information gain
15.6 Expected information gain
15.7 Mutual information
15.8 Sequential Bayesian dynamics
15.9 Filtering
15.10 Bayesian experimental design
15.11 Active learning
15.12 Posterior contraction

16. Variational Inference

16.1 Target distribution
16.2 Restricted approximation family
16.3 Forward KL
16.4 Reverse KL
16.5 Mode-seeking versus mass-covering behavior
16.6 Evidence lower bound
16.7 Mean-field approximation
16.8 Structured approximation
16.9 Amortized inference
16.10 Approximation residue
16.11 Support mismatch
16.12 When residual KL signals wrong model class

17. Learning Dynamics

17.1 Empirical risk under logarithmic scoring
17.2 Cross-entropy minimization
17.3 Gradient descent in parameter space
17.4 Natural gradient
17.5 Fisher-preconditioned dynamics
17.6 Mirror descent
17.7 KL-proximal updates
17.8 Online learning
17.9 Replicator dynamics
17.10 Information-geometric learning flows

PART VIII — LARGE DEVIATIONS: THE FIRST MAJOR LEVEL ABOVE LOCAL GEOMETRY

18. Empirical Distributions

18.1 Independent sampling
18.2 Empirical measure
18.3 Typical distributions
18.4 Atypical empirical distributions
18.5 Method of types
18.6 Exponential probability scales
18.7 Concentration near the generating law
18.8 Statistical fluctuations as geometry in distribution space

19. Sanov Theory

19.1 Sanov’s theorem
19.2 Relative entropy as empirical-measure rate function
19.3 P(empirical law ≈ Q) ≈ exp(-n D(Q‖P))
19.4 Why directionality matters
19.5 Constraint sets
19.6 Most likely atypical empirical distribution
19.7 Information projection as rare-event selection
19.8 Gibbs conditioning
19.9 Contraction principle
19.10 Large deviations as global probability geometry

20. Large-Deviation Geometry

20.1 Rate functions as landscape geometry
20.2 Zero-cost states
20.3 Rare states
20.4 Exponential cost contours
20.5 Constrained minimizers
20.6 Metastable basins
20.7 Rare-event pathways
20.8 Relative entropy versus dynamical action
20.9 Static versus pathwise large deviations
20.10 From empirical distributions to trajectory ensembles

PART IX — ABOVE LARGE DEVIATIONS I: PATH-SPACE GEOMETRY

21. Probability Measures on Trajectories

21.1 States versus paths
21.2 Path-space distributions
21.3 Reference dynamics
21.4 Alternative dynamics
21.5 Path-space likelihood ratio
21.6 Path-space relative entropy
21.7 Dynamical distinguishability
21.8 Finite-time trajectory comparison
21.9 Entropy rate
21.10 Relative entropy rate

22. Dynamical Large Deviations

22.1 Time-averaged observables
22.2 Empirical flows
22.3 Empirical currents
22.4 Donsker–Varadhan theory
22.5 Level-1 large deviations
22.6 Level-2 empirical measures
22.7 Level-2.5 empirical measure-current theory
22.8 Dynamical rate functions
22.9 Optimal fluctuation trajectories
22.10 Rare dynamical phases

23. Action Functionals

23.1 Path likelihood
23.2 Onsager–Machlup structure
23.3 Freidlin–Wentzell action
23.4 Minimum-action trajectories
23.5 Instantons
23.6 Escape events
23.7 Transition pathways
23.8 Large deviations as mechanics on probability paths
23.9 KL between processes versus KL between states
23.10 Dynamic information geometry

PART X — ABOVE LARGE DEVIATIONS II: STOCHASTIC THERMODYNAMICS

24. Relative Entropy and Free Energy

24.1 Gibbs distributions
24.2 Equilibrium reference law
24.3 Free-energy excess
24.4 Relative entropy representation of nonequilibrium free energy
24.5 Equilibrium as information projection
24.6 Available work
24.7 Dissipation
24.8 Free-energy landscapes
24.9 Thermodynamic constraints as probability constraints
24.10 Statistical mechanics as constrained distribution geometry

25. Entropy Production

25.1 Forward trajectory ensemble
25.2 Time-reversed trajectory ensemble
25.3 Path-space KL between forward and reverse dynamics
25.4 Entropy production
25.5 Detailed balance
25.6 Broken detailed balance
25.7 Nonequilibrium steady states
25.8 Probability currents
25.9 Irreversibility as distinguishability of temporal orientation
25.10 Thermodynamic arrow from trajectory asymmetry

26. Fluctuation Relations

26.1 Crooks relation
26.2 Jarzynski equality
26.3 Gallavotti–Cohen symmetry
26.4 Rare negative entropy production
26.5 Exponential weighting
26.6 Path-space likelihood ratios
26.7 Work distributions
26.8 Dissipation bounds
26.9 Information-theoretic formulations of the second law
26.10 Fluctuation theorems as symmetries of dynamical distinguishability

PART XI — ABOVE LARGE DEVIATIONS III: INFORMATION FLOW

27. Information as a Dynamical Resource

27.1 Information stored in correlations
27.2 Information transferred between subsystems
27.3 Mutual-information dynamics
27.4 Directed information
27.5 Transfer entropy
27.6 Predictive information
27.7 Information currents
27.8 Information production
27.9 Information dissipation
27.10 Information conservation versus contraction

28. Open Systems

28.1 System-environment distributions
28.2 Marginalization over environment
28.3 Effective dynamics
28.4 Memory
28.5 Non-Markovianity
28.6 Relative entropy backflow
28.7 Information recovery
28.8 Hidden degrees of freedom
28.9 Effective coarse-grained laws
28.10 Distinguishability flow between visible and hidden sectors

PART XII — ABOVE LARGE DEVIATIONS IV: RENORMALIZATION AND SCALE

29. Distributional Coarse-Graining Across Scale

29.1 Microscopic probability law
29.2 Block variables
29.3 Mesoscopic law
29.4 Macroscopic law
29.5 Repeated coarse-graining
29.6 Information contraction across scale
29.7 Effective theories
29.8 Emergent variables
29.9 Scale-dependent distinguishability
29.10 Information retained by universality

30. Renormalization Flow

30.1 Coarse-grain
30.2 Rescale
30.3 Iterate
30.4 Distributional fixed points
30.5 Relevant perturbations
30.6 Irrelevant perturbations
30.7 Marginal perturbations
30.8 Universality classes
30.9 Relative entropy between nearby RG trajectories
30.10 Information geometry of coupling space

31. Emergence

31.1 Microscopic distinctions
31.2 Macroscopic equivalence
31.3 Stable collective variables
31.4 Information lost under scale reduction
31.5 Information retained as relevant structure
31.6 Effective state spaces
31.7 Emergent laws
31.8 Closure
31.9 Universality as quotient geometry
31.10 Emergence as constrained distinguishability preservation

PART XIII — ABOVE LARGE DEVIATIONS V: MODEL CHANGE

32. Parameter Change Versus Model Change

32.1 Updating parameters inside a fixed family
32.2 Enlarging the family
32.3 Restricting the family
32.4 Introducing latent variables
32.5 Changing dependency structure
32.6 Changing symmetry assumptions
32.7 Changing dimensionality
32.8 Changing observation model
32.9 Changing causal structure
32.10 Model change as motion between statistical manifolds

33. Persistent KL Residue

33.1 Best-fit model
33.2 Residual divergence
33.3 Finite-sample residue
33.4 Structural residue
33.5 Tail residue
33.6 Dependence residue
33.7 Multimodal residue
33.8 Latent-variable residue
33.9 Support residue
33.10 Persistent residue as evidence of model inadequacy

34. Successor Models

34.1 Model extension
34.2 Mixture formation
34.3 Hierarchical models
34.4 Nonparametric models
34.5 State-space enlargement
34.6 New latent carriers
34.7 New symmetry classes
34.8 New causal factorizations
34.9 Comparison against simpler rivals
34.10 Liftback to held-out observables

35. Statistical Theory Evolution

35.1 Fixed inference problem
35.2 Model failure
35.3 Residue localization
35.4 Successor hypothesis
35.5 Re-estimation
35.6 External discrimination
35.7 Model replacement
35.8 Preservation of valid prior structure
35.9 Theory change without target leakage
35.10 Scientific discovery as model-space dynamics

PART XIV — ABOVE LARGE DEVIATIONS VI: DISTINGUISHABILITY ORDER

36. Statistical Experiments

36.1 Experiment as family of distributions
36.2 Comparing experiments
36.3 Informativeness
36.4 Garbling
36.5 Blackwell order
36.6 Sufficient channels
36.7 Deficiency
36.8 Decision-theoretic comparison
36.9 Distinguishability as resource
36.10 Transformation preorder on experiments

37. Divergence Monotonicity

37.1 f-divergences
37.2 KL as distinguished f-divergence
37.3 Data processing
37.4 Monotone statistical quantities
37.5 Contractive maps
37.6 Equality under sufficient transformations
37.7 Information monotonicity as a structural law
37.8 Chentsov uniqueness of Fisher metric
37.9 Local monotonicity versus global divergence families
37.10 Geometry generated by allowed transformations

38. Resource Theory of Distinguishability

38.1 States
38.2 Allowed transformations
38.3 Free transformations
38.4 Information-degrading maps
38.5 Monotones
38.6 Relative entropy monotones
38.7 Convertibility
38.8 Catalytic statistical resources
38.9 Asymptotic conversion rates
38.10 Distinguishability as a generalized resource theory

PART XV — ABOVE LARGE DEVIATIONS VII: GEOMETRY OF MODEL SPACES

39. Spaces of Statistical Manifolds

39.1 A model as a manifold of probability laws
39.2 A family of models as a higher-order state space
39.3 Parameter changes within one manifold
39.4 Model changes between manifolds
39.5 Inclusion relations
39.6 Quotient relations
39.7 Embeddings
39.8 Singular statistical models
39.9 Stratified model spaces
39.10 Geometry of model families rather than individual distributions

40. Model-Space Distinguishability

40.1 Distribution-to-distribution KL
40.2 Model-to-distribution KL
40.3 Model-to-model comparison
40.4 Best achievable KL between families
40.5 Model-family overlap
40.6 Approximation frontiers
40.7 Singular intersection geometry
40.8 Phase transitions in model preference
40.9 Structural versus parametric uncertainty
40.10 Higher-order information geometry of changing model classes

PART XVI — SPECIAL FAMILY LABORATORIES

41. Cauchy Family

41.1 Location-scale Cauchy laws
41.2 Closed KL formula
41.3 Symmetry of KL
41.4 Pair invariant q
41.5 Möbius source symmetry
41.6 Hyperbolic upper half-plane
41.7 Fisher metric
41.8 Pair-orbit collapse
41.9 KL as a readout of hyperbolic separation
41.10 Why the Cauchy family exhibits unusually strong global collapse

42. Gaussian Family

42.1 Location-scale Gaussian laws
42.2 Directional KL
42.3 Relative mean and scale invariants
42.4 Fisher hyperbolic geometry
42.5 Local KL/Fisher agreement
42.6 Global directional fracture
42.7 Exponential-family duality
42.8 Multivariate covariance geometry
42.9 Affine group action
42.10 Source symmetry versus Fisher isometry symmetry

43. Exponential Families

43.1 Natural coordinates
43.2 Expectation coordinates
43.3 Log-partition function
43.4 Convex potential
43.5 Dual flatness
43.6 KL as Bregman divergence
43.7 Maximum entropy
43.8 Maximum likelihood
43.9 Projection geometry
43.10 Why exponential families provide the cleanest inference geometry

PART XVII — DISTINGUISHABILITY UNDER CONSTRAINED TRANSFORMATION

44. The Fundamental Transformation Problem

44.1 Start with probability states P and Q
44.2 Specify admissible transformation class T
44.3 Ask which distinctions survive every allowed transformation
44.4 Identify transformations preserving KL
44.5 Identify transformations contracting KL
44.6 Identify transformations destroying recoverability
44.7 Identify transformations creating apparent equivalence
44.8 Separate representation invariance from source equivalence
44.9 Define transformation-relative distinguishability
44.10 Distinguishability is meaningful only relative to allowed operations

45. Transformation Classes

45.1 Coordinate transformation
45.2 Deterministic invertible transport
45.3 Deterministic noninvertible transport
45.4 Stochastic transport
45.5 Marginalization
45.6 Conditionalization
45.7 Bayesian update
45.8 Projection
45.9 Coarse-graining
45.10 Model replacement

46. Preservation and Loss

46.1 Exact preservation
46.2 Monotone contraction
46.3 Partial recoverability
46.4 Approximate recoverability
46.5 Irreversible information loss
46.6 Newly exposed distinctions
46.7 Hidden distinctions
46.8 Representation-created distinctions
46.9 Distinguishability residue
46.10 Reconstruction after loss

PART XVIII — A HIERARCHY OF INFORMATION-GEOMETRIC LEVELS

47. Level 0 — Probability State

47.1 Distribution
47.2 Support
47.3 Conditional structure
47.4 Transformation law

48. Level 1 — Pairwise Distinguishability

48.1 Likelihood ratio
48.2 KL divergence
48.3 f-divergence family
48.4 Hypothesis discrimination

49. Level 2 — Local Geometry

49.1 Fisher metric
49.2 Geodesics
49.3 Curvature
49.4 Alpha-connections

50. Level 3 — Inference Geometry

50.1 Exponential families
50.2 Mixture families
50.3 Bregman projection
50.4 Bayesian and variational dynamics

51. Level 4 — Transformation Geometry

51.1 Markov kernels
51.2 Data processing
51.3 Sufficiency
51.4 Coarse-graining

52. Level 5 — Large-Deviation Geometry

52.1 Empirical measures
52.2 Sanov rate
52.3 Rare-event cost
52.4 Constraint minimization

53. Level 6 — Path-Space Geometry

53.1 Process distributions
53.2 Relative entropy rate
53.3 Dynamical large deviations
53.4 Action functionals

54. Level 7 — Nonequilibrium Thermodynamic Geometry

54.1 Forward/reverse path distinguishability
54.2 Entropy production
54.3 Fluctuation relations
54.4 Free-energy dissipation

55. Level 8 — Multiscale Information Geometry

55.1 Coarse-graining flows
55.2 Renormalization
55.3 Emergence
55.4 Universality

56. Level 9 — Model-Space Dynamics

56.1 Model selection
56.2 Model inadequacy
56.3 Successor-model generation
56.4 Statistical theory change

57. Level 10 — Distinguishability Under Constrained Transformation

57.1 States
57.2 Allowed operations
57.3 Monotones
57.4 Equivalence classes
57.5 Conversion order
57.6 Reconstruction limits
57.7 Information as transformation-relative structure

PART XIX — GRM BASE-T RECONSTRUCTION

58. Source-First Formation

58.1 Probability law before density
58.2 Pair before divergence
58.3 Comparison before geometry
58.4 Likelihood ratio before KL
58.5 KL before Fisher metric
58.6 Transformation before invariance claim
58.7 Coarse-graining before information-loss claim
58.8 Persistent residue before model-change claim

59. GRM Spine

59.1 Σ_probability
59.2 → pair state (P,Q)
59.3 → comparison carrier
59.4 → likelihood-ratio field
59.5 → KL
59.6 → transformation court
59.7 → {preserved ∥ contracted ∥ destroyed}
59.8 → residue
59.9 → reconstruction / successor model
59.10 → liftback / replay

60. First Noninvertible Arrows

60.1 Marginalization
60.2 Quantization
60.3 Compression
60.4 Projection
60.5 Latent-variable removal
60.6 Finite sampling
60.7 Markov noise
60.8 Support truncation
60.9 Model-family restriction
60.10 Which lost distinctions can be reconstructed?

61. Distinguishability Residue

61.1 Residue after inference
61.2 Residue after compression
61.3 Residue after coarse-graining
61.4 Residue after approximation
61.5 Residue after model selection
61.6 Persistent residue
61.7 Counterkernel construction
61.8 Successor-model pressure
61.9 Replay under alternative carrier
61.10 Model change as repair of unresolved distinguishability

PART XX — FINAL SYNTHESIS

62. KL as Basic Geometry

62.1 Relative entropy is not a metric
62.2 Yet it generates local Fisher geometry
62.3 Its asymmetry carries higher-order structure
62.4 Its invariance reveals transformation structure
62.5 Its contraction reveals information loss

63. KL as Basic Dynamics

63.1 Updating changes probability states
63.2 Transport moves probability mass
63.3 Noise contracts distinguishability
63.4 Coarse-graining destroys microscopic distinctions
63.5 Projection selects representable structure
63.6 Model change alters the available state space

64. Large Deviations as the First Higher Dynamical Theory

64.1 Distribution fluctuations become geometric trajectories
64.2 Relative entropy becomes exponential rarity cost
64.3 Projection becomes selection of the most likely atypical state
64.4 Constraint geometry becomes fluctuation geometry
64.5 Static inference connects to statistical mechanics

65. The Levels Above Large Deviations

65.1 Empirical distributions → path distributions
65.2 Path distributions → dynamical actions
65.3 Dynamical actions → nonequilibrium thermodynamics
65.4 Nonequilibrium thermodynamics → information currents
65.5 Information currents → multiscale coarse-graining
65.6 Multiscale coarse-graining → renormalization
65.7 Renormalization → emergent effective laws
65.8 Persistent residuals → model-space evolution
65.9 Model-space evolution → transformation-relative distinguishability theory

66. The Larger Theory

66.1 Probability as a space of relational states
66.2 KL as the basic directional distinguishability functional
66.3 Fisher information as infinitesimal distinguishability
66.4 Large deviations as fluctuation distinguishability
66.5 Path-space KL as dynamical distinguishability
66.6 Entropy production as time-direction distinguishability
66.7 Data processing as distinguishability contraction
66.8 Renormalization as scale-dependent distinguishability loss
66.9 Model change as reconstruction after persistent distinguishability residue
66.10 Scientific inference as dynamics on a hierarchy of probability-state spaces

67. Final Compression

probability states
likelihood-ratio field
KL distinguishability
{Fisher geometry ∥ convex projection ∥ f-divergence contraction}
inference ∥ transport ∥ coarse-graining
empirical fluctuations
large-deviation rate
path-space action
entropy production / nonequilibrium dynamics
information flow
renormalization / emergence
model-space evolution
theory of distinguishability under constrained transformation
reconstruction ↺

The key escalation is: KL begins as pairwise comparison, becomes geometry locally, becomes dynamics under transformation, becomes fluctuation cost at the large-deviation level, becomes trajectory action at path-space level, becomes irreversibility in stochastic thermodynamics, becomes information flow in open systems, becomes scale-selection under renormalization, and ultimately becomes one component of a more general theory describing which distinctions survive, contract, disappear, or force model change under constrained transformations.

Comments

Popular posts from this blog

Semiotics Rebooted

ORSI: The Telic Geometry of Meaning

THE COLLAPSE ENGINE: AI, Capital, and the Terminal Logic of 2025