Kullback–Leibler Divergence: From Relative Entropy to a Theory of Distinguishability
Kullback–Leibler Divergence: Basic Geometry and Dynamics of Probability Distributions Under Inference, Transport, Coarse-Graining, and Model Change
Prologue — From Relative Entropy to a Theory of Distinguishability
0.1 Probability distributions as states of uncertainty
0.2 Comparison before distance
0.3 Why KL divergence is directional
0.4 D(P‖Q) as expected log-likelihood ratio
0.5 Distinguishability rather than geometric distance
0.6 Probability law ≠ density ≠ parameterization ≠ statistical model
0.7 Transformation of distributions as the central dynamical operation
0.8 Four basic transformation regimes: inference, transport, coarse-graining, model change
0.9 Preservation, contraction, amplification, and destruction of distinguishability
0.10 Local geometry versus global divergence
0.11 Empirical fluctuations versus dynamical trajectories
0.12 Relative entropy as the canonical empirical-measure large-deviation cost
0.13 Beyond large deviations: path ensembles, thermodynamic structure, information flow, renormalization, and model evolution
0.14 Central thesis: probability states → distinguishability → constrained transformation → information residue → reconstruction
PART I — THE BASIC PROBABILITY FIELD
1. Probability Laws Before KL
1.1 Measurable spaces
1.2 Probability measures
1.3 Densities as representations
1.4 Dominating measures
1.5 Absolute continuity
1.6 Singular measures
1.7 Support structure
1.8 Product distributions
1.9 Marginals
1.10 Conditionals
1.11 Pushforwards
1.12 Markov kernels
1.13 Statistical models as families of probability laws
1.14 Parameters as coordinates rather than ontology
1.15 Distribution space as the primary comparison carrier
2. The Likelihood-Ratio Field
2.1 Radon–Nikodym derivative dP/dQ
2.2 Pointwise likelihood ratio
2.3 Log-likelihood ratio
2.4 Evidence accumulation
2.5 Sequential log evidence
2.6 Additivity under independent composition
2.7 Conditional likelihood ratios
2.8 Change of reference measure
2.9 Support mismatch
2.10 Infinite evidence
2.11 Likelihood ratio as the precursor of statistical distinguishability
2.12 From pointwise discrimination to expected discrimination
PART II — FORMATION OF RELATIVE ENTROPY
3. Kullback–Leibler Divergence
3.1 Definition of D(P‖Q)
3.2 Discrete distributions
3.3 Continuous distributions
3.4 General measure-theoretic form
3.5 Non-negativity
3.6 Equality condition
3.7 Directionality
3.8 Lack of triangle inequality
3.9 Finiteness conditions
3.10 Joint convexity
3.11 Lower semicontinuity
3.12 Product additivity
3.13 Chain rule
3.14 Conditional KL
3.15 Marginal KL
3.16 Mutual information as relative entropy
3.17 Multi-information
3.18 Relative entropy rate
4. KL as Basic Distinguishability
4.1 Expected evidence against Q when P generates data
4.2 Excess logarithmic prediction loss
4.3 Excess coding length
4.4 Model mismatch cost
4.5 Information gain
4.6 Projection cost
4.7 Fluctuation cost
4.8 Thermodynamic excess
4.9 Statistical discrimination rate
4.10 Why these interpretations share one mathematical architecture
PART III — BASIC GEOMETRY OF PROBABILITY DISTRIBUTIONS
5. Infinitesimal KL Geometry
5.1 Nearby distributions
5.2 Expansion around a reference law
5.3 Vanishing first-order contribution
5.4 Quadratic term
5.5 Fisher information
5.6 Fisher information matrix
5.7 Statistical line element
5.8 Local distinguishability ellipsoids
5.9 Coordinate invariance
5.10 Fisher–Rao metric
5.11 Local equivalence of smooth statistical divergences
5.12 Why local equivalence does not imply global equivalence
6. Fisher–Rao Geometry
6.1 Statistical manifolds
6.2 Tangent vectors as infinitesimal probability changes
6.3 Fisher metric
6.4 Statistical geodesics
6.5 Geodesic distance
6.6 Curvature
6.7 Isometry groups
6.8 Geodesic completeness
6.9 Boundary geometry
6.10 Product manifolds
6.11 Location families
6.12 Scale families
6.13 Location-scale families
7. Higher-Order Geometry
7.1 KL beyond second order
7.2 Third-order asymmetry
7.3 Amari–Chentsov tensor
7.4 Statistical skewness
7.5 Alpha-connections
7.6 Exponential connection
7.7 Mixture connection
7.8 Levi-Civita connection
7.9 Dual affine structure
7.10 Curvature of statistical connections
7.11 Local orientation retained beyond Fisher geometry
PART IV — CONVEX GEOMETRY AND INFORMATION PROJECTION
8. KL as a Bregman-Type Geometry
8.1 Convex potentials
8.2 Legendre duality
8.3 Bregman divergences
8.4 KL in exponential families
8.5 Natural parameters
8.6 Expectation parameters
8.7 Cumulant potential
8.8 Convex conjugates
8.9 Canonical divergence
8.10 Generalized Pythagorean relations
9. Projection Geometry
9.1 Information projection
9.2 Reverse information projection
9.3 Convex constraint sets
9.4 Moment constraints
9.5 Maximum entropy
9.6 Maximum likelihood
9.7 Projection onto exponential families
9.8 Projection onto mixture families
9.9 Alternating projection
9.10 Iterative proportional fitting
9.11 Projection residue as model inadequacy
9.12 Projection failure as pressure for model change
PART V — TRANSFORMATION GEOMETRY
10. Invertible Transport
10.1 Pushforward by bijections
10.2 Change of variables
10.3 Jacobian cancellation
10.4 KL invariance under common invertible transport
10.5 Group actions on statistical families
10.6 Orbit structure
10.7 Stabilizers
10.8 Pair-orbit invariants
10.9 Source-preserving transformations
10.10 Divergences as orbit readouts
11. Stochastic Transport
11.1 Markov kernels
11.2 Noisy channels
11.3 Transition operators
11.4 Markov semigroups
11.5 Relative entropy contraction
11.6 Data-processing inequality
11.7 Equality conditions
11.8 Sufficient statistics
11.9 Recoverability
11.10 Approximate recoverability
11.11 Information loss under stochastic maps
12. Distinguishability Under Transformation
12.1 Invertible transport → preserve
12.2 Sufficient compression → preserve relevant distinctions
12.3 Markov noise → contract
12.4 Coarse-graining → destroy selected distinctions
12.5 Model projection → retain only representable distinctions
12.6 Model expansion → recover previously inaccessible distinctions
12.7 Transformation families as orders on statistical experiments
12.8 Blackwell informativeness
12.9 Statistical experiment comparison
12.10 Distinguishability as a transformation-sensitive resource
PART VI — COARSE-GRAINING
13. Coarse-Graining as Information Loss
13.1 Many-to-one maps
13.2 Marginalization
13.3 Binning
13.4 Quantization
13.5 Hidden variables
13.6 Partial observation
13.7 Macroscopic observables
13.8 Projection onto sufficient statistics
13.9 KL contraction
13.10 Irrecoverable distinctions
14. Recoverability and Sufficiency
14.1 Exact sufficient statistics
14.2 Minimal sufficient statistics
14.3 Reverse reconstruction maps
14.4 Equality in data processing
14.5 Approximate sufficiency
14.6 Information bottlenecks
14.7 Relevant versus discarded information
14.8 Reconstruction error
14.9 Coarse-grained equivalence classes
14.10 Residual distinguishability hidden below the coarse grain
PART VII — INFERENCE AS DISTRIBUTIONAL DYNAMICS
15. Bayesian Updating
15.1 Prior as initial probability state
15.2 Likelihood as interaction
15.3 Posterior as transformed state
15.4 Bayes rule as probability transport
15.5 KL information gain
15.6 Expected information gain
15.7 Mutual information
15.8 Sequential Bayesian dynamics
15.9 Filtering
15.10 Bayesian experimental design
15.11 Active learning
15.12 Posterior contraction
16. Variational Inference
16.1 Target distribution
16.2 Restricted approximation family
16.3 Forward KL
16.4 Reverse KL
16.5 Mode-seeking versus mass-covering behavior
16.6 Evidence lower bound
16.7 Mean-field approximation
16.8 Structured approximation
16.9 Amortized inference
16.10 Approximation residue
16.11 Support mismatch
16.12 When residual KL signals wrong model class
17. Learning Dynamics
17.1 Empirical risk under logarithmic scoring
17.2 Cross-entropy minimization
17.3 Gradient descent in parameter space
17.4 Natural gradient
17.5 Fisher-preconditioned dynamics
17.6 Mirror descent
17.7 KL-proximal updates
17.8 Online learning
17.9 Replicator dynamics
17.10 Information-geometric learning flows
PART VIII — LARGE DEVIATIONS: THE FIRST MAJOR LEVEL ABOVE LOCAL GEOMETRY
18. Empirical Distributions
18.1 Independent sampling
18.2 Empirical measure
18.3 Typical distributions
18.4 Atypical empirical distributions
18.5 Method of types
18.6 Exponential probability scales
18.7 Concentration near the generating law
18.8 Statistical fluctuations as geometry in distribution space
19. Sanov Theory
19.1 Sanov’s theorem
19.2 Relative entropy as empirical-measure rate function
19.3 P(empirical law ≈ Q) ≈ exp(-n D(Q‖P))
19.4 Why directionality matters
19.5 Constraint sets
19.6 Most likely atypical empirical distribution
19.7 Information projection as rare-event selection
19.8 Gibbs conditioning
19.9 Contraction principle
19.10 Large deviations as global probability geometry
20. Large-Deviation Geometry
20.1 Rate functions as landscape geometry
20.2 Zero-cost states
20.3 Rare states
20.4 Exponential cost contours
20.5 Constrained minimizers
20.6 Metastable basins
20.7 Rare-event pathways
20.8 Relative entropy versus dynamical action
20.9 Static versus pathwise large deviations
20.10 From empirical distributions to trajectory ensembles
PART IX — ABOVE LARGE DEVIATIONS I: PATH-SPACE GEOMETRY
21. Probability Measures on Trajectories
21.1 States versus paths
21.2 Path-space distributions
21.3 Reference dynamics
21.4 Alternative dynamics
21.5 Path-space likelihood ratio
21.6 Path-space relative entropy
21.7 Dynamical distinguishability
21.8 Finite-time trajectory comparison
21.9 Entropy rate
21.10 Relative entropy rate
22. Dynamical Large Deviations
22.1 Time-averaged observables
22.2 Empirical flows
22.3 Empirical currents
22.4 Donsker–Varadhan theory
22.5 Level-1 large deviations
22.6 Level-2 empirical measures
22.7 Level-2.5 empirical measure-current theory
22.8 Dynamical rate functions
22.9 Optimal fluctuation trajectories
22.10 Rare dynamical phases
23. Action Functionals
23.1 Path likelihood
23.2 Onsager–Machlup structure
23.3 Freidlin–Wentzell action
23.4 Minimum-action trajectories
23.5 Instantons
23.6 Escape events
23.7 Transition pathways
23.8 Large deviations as mechanics on probability paths
23.9 KL between processes versus KL between states
23.10 Dynamic information geometry
PART X — ABOVE LARGE DEVIATIONS II: STOCHASTIC THERMODYNAMICS
24. Relative Entropy and Free Energy
24.1 Gibbs distributions
24.2 Equilibrium reference law
24.3 Free-energy excess
24.4 Relative entropy representation of nonequilibrium free energy
24.5 Equilibrium as information projection
24.6 Available work
24.7 Dissipation
24.8 Free-energy landscapes
24.9 Thermodynamic constraints as probability constraints
24.10 Statistical mechanics as constrained distribution geometry
25. Entropy Production
25.1 Forward trajectory ensemble
25.2 Time-reversed trajectory ensemble
25.3 Path-space KL between forward and reverse dynamics
25.4 Entropy production
25.5 Detailed balance
25.6 Broken detailed balance
25.7 Nonequilibrium steady states
25.8 Probability currents
25.9 Irreversibility as distinguishability of temporal orientation
25.10 Thermodynamic arrow from trajectory asymmetry
26. Fluctuation Relations
26.1 Crooks relation
26.2 Jarzynski equality
26.3 Gallavotti–Cohen symmetry
26.4 Rare negative entropy production
26.5 Exponential weighting
26.6 Path-space likelihood ratios
26.7 Work distributions
26.8 Dissipation bounds
26.9 Information-theoretic formulations of the second law
26.10 Fluctuation theorems as symmetries of dynamical distinguishability
PART XI — ABOVE LARGE DEVIATIONS III: INFORMATION FLOW
27. Information as a Dynamical Resource
27.1 Information stored in correlations
27.2 Information transferred between subsystems
27.3 Mutual-information dynamics
27.4 Directed information
27.5 Transfer entropy
27.6 Predictive information
27.7 Information currents
27.8 Information production
27.9 Information dissipation
27.10 Information conservation versus contraction
28. Open Systems
28.1 System-environment distributions
28.2 Marginalization over environment
28.3 Effective dynamics
28.4 Memory
28.5 Non-Markovianity
28.6 Relative entropy backflow
28.7 Information recovery
28.8 Hidden degrees of freedom
28.9 Effective coarse-grained laws
28.10 Distinguishability flow between visible and hidden sectors
PART XII — ABOVE LARGE DEVIATIONS IV: RENORMALIZATION AND SCALE
29. Distributional Coarse-Graining Across Scale
29.1 Microscopic probability law
29.2 Block variables
29.3 Mesoscopic law
29.4 Macroscopic law
29.5 Repeated coarse-graining
29.6 Information contraction across scale
29.7 Effective theories
29.8 Emergent variables
29.9 Scale-dependent distinguishability
29.10 Information retained by universality
30. Renormalization Flow
30.1 Coarse-grain
30.2 Rescale
30.3 Iterate
30.4 Distributional fixed points
30.5 Relevant perturbations
30.6 Irrelevant perturbations
30.7 Marginal perturbations
30.8 Universality classes
30.9 Relative entropy between nearby RG trajectories
30.10 Information geometry of coupling space
31. Emergence
31.1 Microscopic distinctions
31.2 Macroscopic equivalence
31.3 Stable collective variables
31.4 Information lost under scale reduction
31.5 Information retained as relevant structure
31.6 Effective state spaces
31.7 Emergent laws
31.8 Closure
31.9 Universality as quotient geometry
31.10 Emergence as constrained distinguishability preservation
PART XIII — ABOVE LARGE DEVIATIONS V: MODEL CHANGE
32. Parameter Change Versus Model Change
32.1 Updating parameters inside a fixed family
32.2 Enlarging the family
32.3 Restricting the family
32.4 Introducing latent variables
32.5 Changing dependency structure
32.6 Changing symmetry assumptions
32.7 Changing dimensionality
32.8 Changing observation model
32.9 Changing causal structure
32.10 Model change as motion between statistical manifolds
33. Persistent KL Residue
33.1 Best-fit model
33.2 Residual divergence
33.3 Finite-sample residue
33.4 Structural residue
33.5 Tail residue
33.6 Dependence residue
33.7 Multimodal residue
33.8 Latent-variable residue
33.9 Support residue
33.10 Persistent residue as evidence of model inadequacy
34. Successor Models
34.1 Model extension
34.2 Mixture formation
34.3 Hierarchical models
34.4 Nonparametric models
34.5 State-space enlargement
34.6 New latent carriers
34.7 New symmetry classes
34.8 New causal factorizations
34.9 Comparison against simpler rivals
34.10 Liftback to held-out observables
35. Statistical Theory Evolution
35.1 Fixed inference problem
35.2 Model failure
35.3 Residue localization
35.4 Successor hypothesis
35.5 Re-estimation
35.6 External discrimination
35.7 Model replacement
35.8 Preservation of valid prior structure
35.9 Theory change without target leakage
35.10 Scientific discovery as model-space dynamics
PART XIV — ABOVE LARGE DEVIATIONS VI: DISTINGUISHABILITY ORDER
36. Statistical Experiments
36.1 Experiment as family of distributions
36.2 Comparing experiments
36.3 Informativeness
36.4 Garbling
36.5 Blackwell order
36.6 Sufficient channels
36.7 Deficiency
36.8 Decision-theoretic comparison
36.9 Distinguishability as resource
36.10 Transformation preorder on experiments
37. Divergence Monotonicity
37.1 f-divergences
37.2 KL as distinguished f-divergence
37.3 Data processing
37.4 Monotone statistical quantities
37.5 Contractive maps
37.6 Equality under sufficient transformations
37.7 Information monotonicity as a structural law
37.8 Chentsov uniqueness of Fisher metric
37.9 Local monotonicity versus global divergence families
37.10 Geometry generated by allowed transformations
38. Resource Theory of Distinguishability
38.1 States
38.2 Allowed transformations
38.3 Free transformations
38.4 Information-degrading maps
38.5 Monotones
38.6 Relative entropy monotones
38.7 Convertibility
38.8 Catalytic statistical resources
38.9 Asymptotic conversion rates
38.10 Distinguishability as a generalized resource theory
PART XV — ABOVE LARGE DEVIATIONS VII: GEOMETRY OF MODEL SPACES
39. Spaces of Statistical Manifolds
39.1 A model as a manifold of probability laws
39.2 A family of models as a higher-order state space
39.3 Parameter changes within one manifold
39.4 Model changes between manifolds
39.5 Inclusion relations
39.6 Quotient relations
39.7 Embeddings
39.8 Singular statistical models
39.9 Stratified model spaces
39.10 Geometry of model families rather than individual distributions
40. Model-Space Distinguishability
40.1 Distribution-to-distribution KL
40.2 Model-to-distribution KL
40.3 Model-to-model comparison
40.4 Best achievable KL between families
40.5 Model-family overlap
40.6 Approximation frontiers
40.7 Singular intersection geometry
40.8 Phase transitions in model preference
40.9 Structural versus parametric uncertainty
40.10 Higher-order information geometry of changing model classes
PART XVI — SPECIAL FAMILY LABORATORIES
41. Cauchy Family
41.1 Location-scale Cauchy laws
41.2 Closed KL formula
41.3 Symmetry of KL
41.4 Pair invariant q
41.5 Möbius source symmetry
41.6 Hyperbolic upper half-plane
41.7 Fisher metric
41.8 Pair-orbit collapse
41.9 KL as a readout of hyperbolic separation
41.10 Why the Cauchy family exhibits unusually strong global collapse
42. Gaussian Family
42.1 Location-scale Gaussian laws
42.2 Directional KL
42.3 Relative mean and scale invariants
42.4 Fisher hyperbolic geometry
42.5 Local KL/Fisher agreement
42.6 Global directional fracture
42.7 Exponential-family duality
42.8 Multivariate covariance geometry
42.9 Affine group action
42.10 Source symmetry versus Fisher isometry symmetry
43. Exponential Families
43.1 Natural coordinates
43.2 Expectation coordinates
43.3 Log-partition function
43.4 Convex potential
43.5 Dual flatness
43.6 KL as Bregman divergence
43.7 Maximum entropy
43.8 Maximum likelihood
43.9 Projection geometry
43.10 Why exponential families provide the cleanest inference geometry
PART XVII — DISTINGUISHABILITY UNDER CONSTRAINED TRANSFORMATION
44. The Fundamental Transformation Problem
44.1 Start with probability states P and Q
44.2 Specify admissible transformation class T
44.3 Ask which distinctions survive every allowed transformation
44.4 Identify transformations preserving KL
44.5 Identify transformations contracting KL
44.6 Identify transformations destroying recoverability
44.7 Identify transformations creating apparent equivalence
44.8 Separate representation invariance from source equivalence
44.9 Define transformation-relative distinguishability
44.10 Distinguishability is meaningful only relative to allowed operations
45. Transformation Classes
45.1 Coordinate transformation
45.2 Deterministic invertible transport
45.3 Deterministic noninvertible transport
45.4 Stochastic transport
45.5 Marginalization
45.6 Conditionalization
45.7 Bayesian update
45.8 Projection
45.9 Coarse-graining
45.10 Model replacement
46. Preservation and Loss
46.1 Exact preservation
46.2 Monotone contraction
46.3 Partial recoverability
46.4 Approximate recoverability
46.5 Irreversible information loss
46.6 Newly exposed distinctions
46.7 Hidden distinctions
46.8 Representation-created distinctions
46.9 Distinguishability residue
46.10 Reconstruction after loss
PART XVIII — A HIERARCHY OF INFORMATION-GEOMETRIC LEVELS
47. Level 0 — Probability State
47.1 Distribution
47.2 Support
47.3 Conditional structure
47.4 Transformation law
48. Level 1 — Pairwise Distinguishability
48.1 Likelihood ratio
48.2 KL divergence
48.3 f-divergence family
48.4 Hypothesis discrimination
49. Level 2 — Local Geometry
49.1 Fisher metric
49.2 Geodesics
49.3 Curvature
49.4 Alpha-connections
50. Level 3 — Inference Geometry
50.1 Exponential families
50.2 Mixture families
50.3 Bregman projection
50.4 Bayesian and variational dynamics
51. Level 4 — Transformation Geometry
51.1 Markov kernels
51.2 Data processing
51.3 Sufficiency
51.4 Coarse-graining
52. Level 5 — Large-Deviation Geometry
52.1 Empirical measures
52.2 Sanov rate
52.3 Rare-event cost
52.4 Constraint minimization
53. Level 6 — Path-Space Geometry
53.1 Process distributions
53.2 Relative entropy rate
53.3 Dynamical large deviations
53.4 Action functionals
54. Level 7 — Nonequilibrium Thermodynamic Geometry
54.1 Forward/reverse path distinguishability
54.2 Entropy production
54.3 Fluctuation relations
54.4 Free-energy dissipation
55. Level 8 — Multiscale Information Geometry
55.1 Coarse-graining flows
55.2 Renormalization
55.3 Emergence
55.4 Universality
56. Level 9 — Model-Space Dynamics
56.1 Model selection
56.2 Model inadequacy
56.3 Successor-model generation
56.4 Statistical theory change
57. Level 10 — Distinguishability Under Constrained Transformation
57.1 States
57.2 Allowed operations
57.3 Monotones
57.4 Equivalence classes
57.5 Conversion order
57.6 Reconstruction limits
57.7 Information as transformation-relative structure
PART XIX — GRM BASE-T RECONSTRUCTION
58. Source-First Formation
58.1 Probability law before density
58.2 Pair before divergence
58.3 Comparison before geometry
58.4 Likelihood ratio before KL
58.5 KL before Fisher metric
58.6 Transformation before invariance claim
58.7 Coarse-graining before information-loss claim
58.8 Persistent residue before model-change claim
59. GRM Spine
59.1 Σ_probability
59.2 → pair state (P,Q)
59.3 → comparison carrier
59.4 → likelihood-ratio field
59.5 → KL
59.6 → transformation court
59.7 → {preserved ∥ contracted ∥ destroyed}
59.8 → residue
59.9 → reconstruction / successor model
59.10 → liftback / replay
60. First Noninvertible Arrows
60.1 Marginalization
60.2 Quantization
60.3 Compression
60.4 Projection
60.5 Latent-variable removal
60.6 Finite sampling
60.7 Markov noise
60.8 Support truncation
60.9 Model-family restriction
60.10 Which lost distinctions can be reconstructed?
61. Distinguishability Residue
61.1 Residue after inference
61.2 Residue after compression
61.3 Residue after coarse-graining
61.4 Residue after approximation
61.5 Residue after model selection
61.6 Persistent residue
61.7 Counterkernel construction
61.8 Successor-model pressure
61.9 Replay under alternative carrier
61.10 Model change as repair of unresolved distinguishability
PART XX — FINAL SYNTHESIS
62. KL as Basic Geometry
62.1 Relative entropy is not a metric
62.2 Yet it generates local Fisher geometry
62.3 Its asymmetry carries higher-order structure
62.4 Its invariance reveals transformation structure
62.5 Its contraction reveals information loss
63. KL as Basic Dynamics
63.1 Updating changes probability states
63.2 Transport moves probability mass
63.3 Noise contracts distinguishability
63.4 Coarse-graining destroys microscopic distinctions
63.5 Projection selects representable structure
63.6 Model change alters the available state space
64. Large Deviations as the First Higher Dynamical Theory
64.1 Distribution fluctuations become geometric trajectories
64.2 Relative entropy becomes exponential rarity cost
64.3 Projection becomes selection of the most likely atypical state
64.4 Constraint geometry becomes fluctuation geometry
64.5 Static inference connects to statistical mechanics
65. The Levels Above Large Deviations
65.1 Empirical distributions → path distributions
65.2 Path distributions → dynamical actions
65.3 Dynamical actions → nonequilibrium thermodynamics
65.4 Nonequilibrium thermodynamics → information currents
65.5 Information currents → multiscale coarse-graining
65.6 Multiscale coarse-graining → renormalization
65.7 Renormalization → emergent effective laws
65.8 Persistent residuals → model-space evolution
65.9 Model-space evolution → transformation-relative distinguishability theory
66. The Larger Theory
66.1 Probability as a space of relational states
66.2 KL as the basic directional distinguishability functional
66.3 Fisher information as infinitesimal distinguishability
66.4 Large deviations as fluctuation distinguishability
66.5 Path-space KL as dynamical distinguishability
66.6 Entropy production as time-direction distinguishability
66.7 Data processing as distinguishability contraction
66.8 Renormalization as scale-dependent distinguishability loss
66.9 Model change as reconstruction after persistent distinguishability residue
66.10 Scientific inference as dynamics on a hierarchy of probability-state spaces
67. Final Compression
probability states
→ likelihood-ratio field
→ KL distinguishability
→ {Fisher geometry ∥ convex projection ∥ f-divergence contraction}
→ inference ∥ transport ∥ coarse-graining
→ empirical fluctuations
→ large-deviation rate
→ path-space action
→ entropy production / nonequilibrium dynamics
→ information flow
→ renormalization / emergence
→ model-space evolution
→ theory of distinguishability under constrained transformation
→ reconstruction ↺
The key escalation is: KL begins as pairwise comparison, becomes geometry locally, becomes dynamics under transformation, becomes fluctuation cost at the large-deviation level, becomes trajectory action at path-space level, becomes irreversibility in stochastic thermodynamics, becomes information flow in open systems, becomes scale-selection under renormalization, and ultimately becomes one component of a more general theory describing which distinctions survive, contract, disappear, or force model change under constrained transformations.
Comments
Post a Comment