Kullback–Leibler Divergence: Geometry and Dynamics of Probability Distributions
Kullback–Leibler Divergence: Geometry and Dynamics of Probability Distributions Under Inference, Transport, Coarse-Graining, and Model Change
Prologue — KL Divergence as a Structural Object
0.1 Why KL divergence is more than a dissimilarity measure
0.2 Probability distributions as states rather than formulas
0.3 Distinguishability as the central organizing principle
0.4 Density representation ≠ probability law ≠ statistical model
0.5 Directionality: why D(P‖Q) ≠ D(Q‖P)
0.6 Local geometry versus global divergence
0.7 Inference, transport, coarse-graining, and model change as four transformation classes
0.8 KL as logarithmic discrimination cost
0.9 KL as excess coding cost
0.10 KL as large-deviation rate
0.11 KL as free-energy excess
0.12 KL as convex projection cost
0.13 KL as a bridge between probability, geometry, dynamics, and thermodynamics
0.14 When KL becomes symmetric
0.15 When KL becomes infinite
0.16 What KL forgets
0.17 What KL preserves
0.18 The central architecture: probability state → comparison → transformation → residue → projection → reconstruction
PART I — PROBABILITY STATES BEFORE DIVERGENCE
1. Probability Measures as the Source Carrier
1.1 Sample space and measurable structure
1.2 Probability measures before densities
1.3 Dominating measures and density representations
1.4 Absolute continuity
1.5 Singular components
1.6 Radon–Nikodym derivatives
1.7 Support and effective support
1.8 Probability laws versus parameter coordinates
1.9 Pushforwards under measurable transformations
1.10 Markov kernels as probabilistic transports
1.11 Product measures and independent composition
1.12 Conditional distributions
1.13 Marginalization
1.14 Mixture formation
1.15 Model families as submanifolds of probability space
2. Log-Likelihood Ratio as the Primitive Comparison Field
2.1 Likelihood ratio dP/dQ
2.2 Log-likelihood ratio L = log(dP/dQ)
2.3 Pointwise evidence versus expected evidence
2.4 Change of reference measure
2.5 Likelihood ratios under common transformations
2.6 Additivity under independent products
2.7 Conditional decomposition
2.8 Sequential accumulation of log evidence
2.9 Martingale structure of likelihood ratios
2.10 Hypothesis testing interpretation
2.11 Support mismatch and divergent evidence
2.12 The log-ratio field as the carrier from which KL descends
PART II — FORMATION OF KL DIVERGENCE
3. Relative Entropy
3.1 Definition of D(P‖Q)
3.2 Discrete form
3.3 Continuous form
3.4 General measure-theoretic form
3.5 Conditions for finiteness
3.6 Non-negativity
3.7 Equality condition
3.8 Asymmetry
3.9 Non-metric character
3.10 Convexity in the second argument
3.11 Joint convexity
3.12 Lower semicontinuity
3.13 Tensorization
3.14 Conditional relative entropy
3.15 Chain rules
3.16 KL under mixtures
3.17 KL under products
3.18 KL under marginalization
4. Operational Meanings of KL
4.1 Excess log-loss
4.2 Expected evidence against a model
4.3 Coding redundancy
4.4 Model mismatch penalty
4.5 Hypothesis discrimination rate
4.6 Rare-event cost
4.7 Free-energy penalty
4.8 Bayesian updating cost
4.9 Projection onto constrained families
4.10 Irreversibility and entropy production
4.11 Why all these interpretations converge on the logarithm
PART III — KL AS LOCAL GEOMETRY
5. Second-Order Expansion and Fisher Information
5.1 Nearby probability distributions
5.2 Taylor expansion of KL
5.3 Vanishing first-order term
5.4 Fisher information as the quadratic coefficient
5.5 Fisher matrix
5.6 Coordinate covariance
5.7 Local statistical distinguishability
5.8 Fisher–Rao line element
5.9 Infinitesimal KL spheres
5.10 Local equivalence of statistical divergences
5.11 Why global KL cannot be recovered from the Fisher metric alone
5.12 Higher-order terms as hidden structure
6. Fisher–Rao Geometry
6.1 Statistical manifolds
6.2 Riemannian metric from Fisher information
6.3 Geodesics
6.4 Geodesic distance
6.5 Curvature
6.6 Isometries induced by statistical transformations
6.7 Location families
6.8 Scale families
6.9 Location-scale families
6.10 Product families
6.11 Fisher geometry of Bernoulli distributions
6.12 Fisher geometry of categorical distributions
6.13 Fisher geometry of Gaussian families
6.14 Fisher geometry of Cauchy families
6.15 Geodesic completeness and boundary behavior
PART IV — THIRD-ORDER STRUCTURE AND AMARI DUALITY
7. The Amari–Chentsov Tensor
7.1 Third-order score moments
7.2 Definition of the cubic tensor
7.3 Why Fisher geometry is insufficient
7.4 Directionality beyond quadratic order
7.5 Skewness of statistical structure
7.6 Tensor transformation properties
7.7 Vanishing versus non-vanishing cubic tensor
7.8 What T = 0 structurally means
7.9 Cauchy as connection-collapse example
7.10 Gaussian as nontrivial dual-connection example
8. Alpha-Connections
8.1 Levi-Civita connection
8.2 Definition of alpha-connections
8.3 Duality of +alpha and −alpha
8.4 Exponential connection
8.5 Mixture connection
8.6 Parallel transport
8.7 Connection curvature
8.8 Connection torsion
8.9 E-geodesics
8.10 M-geodesics
8.11 Ambient density interpolation versus induced manifold geodesics
8.12 Why mixture interpolation usually leaves a restricted model family
9. Dually Flat Geometry
9.1 Exponential families
9.2 Natural coordinates
9.3 Expectation coordinates
9.4 Cumulant generating potential
9.5 Convex conjugacy
9.6 Legendre transform
9.7 Dual affine coordinates
9.8 Bregman divergence
9.9 KL as canonical divergence
9.10 Generalized Pythagorean theorem
9.11 Orthogonal projection
9.12 Maximum-likelihood projection
9.13 Maximum-entropy projection
9.14 Gaussian dual flatness
9.15 Why Cauchy is not dually flat in the same sense
PART V — GLOBAL GEOMETRY OF SPECIAL FAMILIES
10. Cauchy Geometry
10.1 Location-scale Cauchy family
10.2 Closed-form KL divergence
10.3 Symmetry of Cauchy KL
10.4 Finiteness despite heavy tails
10.5 The pair invariant q
10.6 Density-integral derivation
10.7 Parameter-algebra derivation
10.8 Hyperbolic derivation
10.9 Fisher–Rao derivation
10.10 Upper-half-plane representation
10.11 Möbius transformations
10.12 PSL(2,R) action
10.13 Pair-orbit dimension collapse
10.14 q as complete two-point orbit invariant
10.15 KL as a monotone readout of q
10.16 Fisher distance as a monotone readout of q
10.17 Chi-square divergence as another readout
10.18 Why many f-divergences reduce to q
10.19 Connection collapse from T = 0
10.20 Global symmetry collapse versus local connection collapse
11. Gaussian Geometry
11.1 Univariate Gaussian family
11.2 Directional Gaussian KL
11.3 Affine orbit invariants r and δ
11.4 Reverse KL
11.5 Jeffreys divergence
11.6 Fisher metric
11.7 Hyperbolic Fisher manifold
11.8 Fisher quotient q_G
11.9 Why q_G does not reconstruct directional KL
11.10 Lossless ordered-pair invariant versus lossy symmetric quotient
11.11 Local KL/Fisher agreement
11.12 Global KL/Fisher fracture
11.13 Exponential-family structure
11.14 Natural and expectation coordinates
11.15 Multivariate Gaussian extension
11.16 Covariance congruence invariants
11.17 Mahalanobis structure
11.18 Relative orientation between mean and covariance eigenspaces
11.19 Affine invariance
11.20 Why Gaussian source symmetry is smaller than Fisher isometry symmetry
12. Comparative Statistical Geometry
12.1 Cauchy versus Gaussian
12.2 One pair invariant versus multiple pair invariants
12.3 Symmetric versus directional KL
12.4 Source-preserving group versus intrinsic metric isometry group
12.5 Connection collapse versus dual flatness
12.6 Heavy-tail geometry
12.7 Exponential-family geometry
12.8 Orbit-space dimension as predictor of divergence complexity
12.9 When a divergence becomes a function of geodesic distance
12.10 When Fisher geometry loses global information
PART VI — KL AS AN f-DIVERGENCE
13. The f-Divergence Family
13.1 General definition
13.2 Convex generating functions
13.3 KL as an f-divergence
13.4 Reverse KL
13.5 Pearson chi-square
13.6 Hellinger divergence
13.7 Total variation
13.8 Jensen–Shannon divergence
13.9 Rényi-related constructions
13.10 Common invariance laws
13.11 Data processing inequality
13.12 Contraction under Markov kernels
13.13 Equality cases
13.14 Sufficient statistics
13.15 Blackwell comparison
13.16 Information monotonicity
14. Coarse-Graining
14.1 Coarse-graining as probabilistic many-to-one transport
14.2 Markov morphisms
14.3 Marginalization as coarse-graining
14.4 Quantization and binning
14.5 Hidden variables
14.6 Projection onto observables
14.7 Loss of distinguishability
14.8 Data-processing inequality as geometric contraction
14.9 Irrecoverable information loss
14.10 Equality and sufficient statistics
14.11 Recoverability maps
14.12 Approximate sufficiency
14.13 Coarse-graining semigroups
14.14 Renormalization-style interpretation
14.15 KL decay under repeated coarse-graining
PART VII — TRANSPORT OF PROBABILITY DISTRIBUTIONS
15. Deterministic Transport
15.1 Pushforward measures
15.2 Invertible change of variables
15.3 Jacobian cancellation
15.4 KL invariance under common bijections
15.5 Non-invertible maps
15.6 Quotienting and information loss
15.7 Sufficient transformations
15.8 Group actions on model families
15.9 Orbit invariants
15.10 Stabilizer structure
15.11 Transport-equivalent probability states
16. Stochastic Transport
16.1 Markov kernels
16.2 Channels
16.3 Transition semigroups
16.4 Relative entropy contraction
16.5 Ergodic evolution
16.6 Diffusions
16.7 Fokker–Planck dynamics
16.8 Birth-death processes
16.9 Filtering
16.10 Hidden Markov models
16.11 Information decay through noisy channels
16.12 Recoverability after stochastic transport
17. Optimal Transport versus Information Geometry
17.1 Wasserstein distance
17.2 KL divergence
17.3 Geometry of displacement versus geometry of distinguishability
17.4 Probability flows
17.5 Wasserstein gradient flows
17.6 Entropy as a functional on distribution space
17.7 Fokker–Planck as entropy-gradient dynamics
17.8 JKO scheme
17.9 Entropic regularization
17.10 Schrödinger bridge
17.11 Sinkhorn divergence
17.12 Fisher–Rao versus Wasserstein geometry
17.13 Wasserstein–Fisher–Rao hybrid geometry
17.14 When mass moves
17.15 When mass appears or disappears
PART VIII — KL AS DYNAMICS OF INFERENCE
18. Bayesian Updating
18.1 Prior distribution
18.2 Likelihood
18.3 Posterior
18.4 Bayes rule as multiplicative transport
18.5 Posterior KL from prior
18.6 Information gain
18.7 Expected information gain
18.8 Mutual information
18.9 Sequential updating
18.10 Bayesian experimental design
18.11 Active learning
18.12 Information gain as acquisition criterion
19. Variational Inference
19.1 Intractable posterior
19.2 Approximation family
19.3 Forward versus reverse KL
19.4 Mode seeking
19.5 Mass covering
19.6 Evidence lower bound
19.7 KL projection
19.8 Mean-field approximation
19.9 Structured variational families
19.10 Amortized inference
19.11 Variational autoencoders
19.12 Geometry of approximation error
19.13 Model-family boundary effects
19.14 Failure under support mismatch
20. Information Projection
20.1 I-projection
20.2 Reverse I-projection
20.3 Convex constraint sets
20.4 Moment constraints
20.5 Maximum entropy
20.6 Iterative proportional fitting
20.7 Alternating projections
20.8 Generalized Pythagorean decomposition
20.9 Projection onto exponential families
20.10 Projection onto mixture families
20.11 Projection instability under model misspecification
PART IX — KL AND STATISTICAL DECISION
21. Hypothesis Testing
21.1 Binary testing
21.2 Likelihood-ratio test
21.3 Neyman–Pearson structure
21.4 Chernoff bounds
21.5 Stein’s lemma
21.6 Error exponents
21.7 Asymmetric testing and directional KL
21.8 Composite hypotheses
21.9 Sequential probability ratio tests
21.10 Multiple hypotheses
21.11 Distinguishability geometry of experiments
22. Coding and Compression
22.1 Shannon coding
22.2 Cross-entropy
22.3 Redundancy
22.4 Mismatched coding
22.5 KL as expected excess codelength
22.6 Universal coding
22.7 Minimum description length
22.8 Model complexity
22.9 Compression as inference
22.10 Compression versus coarse-graining
PART X — KL AND LARGE DEVIATIONS
23. Empirical Measures
23.1 Law of large numbers
23.2 Empirical distributions
23.3 Fluctuations around the generating law
23.4 Method of types
23.5 Exponential probability scales
23.6 Sanov’s theorem
23.7 KL as the empirical-measure rate function
23.8 Contraction principle
23.9 Conditional large deviations
23.10 Gibbs conditioning principle
24. Rare Events and Dynamical Fluctuations
24.1 Path-space probability measures
24.2 Relative entropy on trajectories
24.3 Action functionals
24.4 Freidlin–Wentzell theory
24.5 Entropy production
24.6 Nonequilibrium fluctuations
24.7 Large-deviation geometry
24.8 Most probable rare path
24.9 Importance sampling
24.10 KL-optimal change of measure
PART XI — THERMODYNAMICS AND STATISTICAL MECHANICS
25. Entropy, Relative Entropy, and Free Energy
25.1 Shannon entropy
25.2 Gibbs entropy
25.3 Cross-entropy
25.4 Relative entropy
25.5 Canonical ensembles
25.6 Gibbs distributions
25.7 Free-energy variational principle
25.8 Free-energy difference as KL
25.9 Equilibrium as KL projection
25.10 Nonequilibrium free energy
25.11 Available work
25.12 Entropy production
26. Irreversible Dynamics
26.1 Detailed balance
26.2 Markov semigroups
26.3 KL as Lyapunov functional
26.4 H-theorem
26.5 Dissipation of relative entropy
26.6 Log-Sobolev inequalities
26.7 Convergence to equilibrium
26.8 Spectral gaps
26.9 Fisher information dissipation
26.10 De Bruijn identity
26.11 Entropy power and diffusion
PART XII — MODEL CHANGE
27. Model Families as Dynamic Objects
27.1 Fixed-model inference
27.2 Model misspecification
27.3 Structural model change
27.4 Parameter update versus model-class update
27.5 Enlargement of model families
27.6 Restriction and pruning
27.7 Nested models
27.8 Non-nested models
27.9 Mixture expansion
27.10 Latent-variable introduction
27.11 Representation change versus ontology change
27.12 KL under model refinement
28. Projection Failure and Residual Structure
28.1 Best approximation does not imply adequate model
28.2 Irreducible KL residue
28.3 Support mismatch
28.4 Tail mismatch
28.5 Multimodality mismatch
28.6 Dependency mismatch
28.7 Wrong latent dimension
28.8 Wrong symmetry class
28.9 Wrong causal factorization
28.10 Residual diagnostics
28.11 When model error demands a successor family
29. Model Selection
29.1 Expected log-likelihood
29.2 KL risk
29.3 AIC as asymptotic KL-risk correction
29.4 Cross-validation
29.5 Predictive selection
29.6 Bayesian evidence
29.7 Minimum description length
29.8 Complexity penalties
29.9 Out-of-distribution model comparison
29.10 Model selection versus model discovery
PART XIII — COARSE-GRAINING, RENORMALIZATION, AND EMERGENCE
30. Hierarchies of Description
30.1 Microstate distributions
30.2 Mesostates
30.3 Macrostates
30.4 Sufficient macroscopic variables
30.5 Projection operators
30.6 Information lost across scale
30.7 Effective models
30.8 Closure approximations
30.9 Memory generated by coarse-graining
30.10 Mori–Zwanzig perspective
31. Renormalization as Distributional Flow
31.1 Scale transformations
31.2 Coarse-graining maps
31.3 Rescaling
31.4 Fixed points
31.5 Relevant directions
31.6 Irrelevant directions
31.7 Information loss along RG flow
31.8 Relative entropy between scales
31.9 Universality classes
31.10 Emergent invariants
32. Emergence and Information Bottlenecks
32.1 Compression while preserving predictive structure
32.2 Information bottleneck
32.3 Relevant information
32.4 Minimal sufficient representation
32.5 Predictive state compression
32.6 Rate-distortion theory
32.7 Coarse-grained causal states
32.8 Emergent macroscopic coordinates
32.9 Lost microscopic distinguishability
32.10 Recoverability boundaries
PART XIV — MACHINE LEARNING AND REPRESENTATION CHANGE
33. Cross-Entropy Training
33.1 Empirical risk
33.2 Log-loss
33.3 Cross-entropy
33.4 KL decomposition
33.5 Classification
33.6 Density estimation
33.7 Calibration
33.8 Label smoothing
33.9 Distribution shift
33.10 Overconfidence and support failure
34. Generative Models
34.1 Maximum likelihood
34.2 Autoregressive models
34.3 Normalizing flows
34.4 Variational autoencoders
34.5 Diffusion models
34.6 Energy-based models
34.7 Forward and reverse KL objectives
34.8 Mode collapse
34.9 Tail coverage
34.10 Model geometry under latent-variable maps
35. Representation Learning
35.1 Data-processing inequality
35.2 Compression of representations
35.3 Mutual information
35.4 Sufficient representation
35.5 Invariance
35.6 Equivariance
35.7 Nuisance-variable removal
35.8 Information bottleneck objectives
35.9 Representation collapse
35.10 Recoverability from latent carriers
PART XV — CAUSAL AND CONDITIONAL INFORMATION GEOMETRY
36. Conditional KL
36.1 Conditional distributions
36.2 Conditional relative entropy
36.3 Chain rule
36.4 Factorization of joint models
36.5 Dependency structure
36.6 Conditional independence
36.7 Graphical models
36.8 KL decomposition over Bayesian networks
36.9 Local versus global model mismatch
36.10 Intervention-conditioned divergence
37. Information Under Intervention
37.1 Observation versus intervention
37.2 Causal model distributions
37.3 Distribution shift under do-operations
37.4 Intervention distinguishability
37.5 Causal model comparison
37.6 Counterfactual distributions
37.7 KL and mechanism change
37.8 Environment-indexed models
37.9 Invariant mechanisms
37.10 Detecting structural change versus parameter drift
PART XVI — DYNAMICS ON DISTRIBUTION SPACE
38. Probability Flows
38.1 Distribution-valued trajectories P_t
38.2 Velocity in probability space
38.3 Continuity equations
38.4 Generator operators
38.5 Markov semigroups
38.6 Diffusion semigroups
38.7 Replicator equations
38.8 Bayesian filtering flows
38.9 Gradient flows
38.10 Information-geometric flows
39. KL Production and Dissipation
39.1 Time derivative of KL
39.2 Relative entropy production
39.3 Fisher-information production
39.4 Contractive dynamics
39.5 Expansive dynamics
39.6 Detailed-balance systems
39.7 Nonequilibrium steady states
39.8 Information currents
39.9 Entropy-production decomposition
39.10 Path-space KL
PART XVII — FAILURES, BOUNDARIES, AND NON-EQUIVALENCES
40. Failure Modes of KL
40.1 Infinite divergence
40.2 Support mismatch
40.3 Zero-density boundaries
40.4 Heavy tails
40.5 Singular measures
40.6 Degenerate distributions
40.7 Numerical instability
40.8 Monte Carlo estimation error
40.9 High-dimensional concentration
40.10 Misspecified reference distributions
41. What KL Does Not Measure
41.1 Geometric displacement of support
41.2 Semantic similarity
41.3 Causal equivalence
41.4 Wasserstein transport cost
41.5 Topological similarity
41.6 Temporal alignment
41.7 Pointwise error
41.8 Symmetric metric distance
41.9 Physical energy unless a model identifies the two
41.10 Ontological identity
42. Divergence Fractures
42.1 Same Fisher metric, different global divergences
42.2 Same KL, different distributions
42.3 Same symmetric divergence, different directional structure
42.4 Same coarse-grained KL, different microscopic structure
42.5 Same endpoint, different inference path
42.6 Same model family, different group orbit
42.7 Same geometry, different source-preserving symmetry
42.8 Local equivalence versus global non-equivalence
PART XVIII — GRM/TSCT REBUILD OF KL
43. Source-First Reconstruction
43.1 Probability law before density representation
43.2 Comparison field before scalar divergence
43.3 Likelihood ratio as earned relational structure
43.4 Expectation as aggregation operator
43.5 Logarithm as compositional normalization
43.6 KL as derived readout
43.7 SOURCE ≠ DENSITY ≠ PARAMETERIZATION ≠ DIVERGENCE
43.8 Carrier-change audit
43.9 Support ownership
43.10 Directionality ownership
44. Carrier Invariance
44.1 Density carrier
44.2 Parameter carrier
44.3 Group-action carrier
44.4 Fisher carrier
44.5 Convex carrier
44.6 Large-deviation carrier
44.7 Thermodynamic carrier
44.8 Coding carrier
44.9 Independent descent versus reformulation
44.10 Common-bottom test
45. Structural Descent
45.1 density ratio → log ratio → expectation → KL
45.2 KL → second-order residue → Fisher metric
45.3 KL → third-order residue → Amari tensor
45.4 KL → Markov contraction → coarse-graining order
45.5 KL → rate function → large deviations
45.6 KL → convex potential → Bregman geometry
45.7 KL → free-energy excess → thermodynamic geometry
45.8 KL → code redundancy → communication geometry
45.9 Determine which branches share genuine ancestry
45.10 Separate equivalence from analogy
46. FNA and Information Loss
46.1 First noninvertible transformation in probability pipelines
46.2 Marginalization
46.3 Quantization
46.4 Sufficient statistics
46.5 Latent compression
46.6 Model projection
46.7 Coarse-graining
46.8 Support truncation
46.9 Approximation-family restriction
46.10 Reconstruction from compressed statistics
47. Residue and Successor Models
47.1 KL residual after best fit
47.2 Residual structure not captured by the model
47.3 Counterkernel construction
47.4 Detecting wrong parameterization
47.5 Detecting wrong family
47.6 Detecting missing latent variables
47.7 Detecting missing dependencies
47.8 Successor-model generation
47.9 Liftback to held-out consequences
47.10 Replay under changed carrier
PART XIX — A UNIFIED THEORY OF DISTINGUISHABILITY
48. Pairwise Probability Structure
48.1 Ordered pair (P,Q) as the primitive comparison object
48.2 Symmetry group of the family
48.3 Pair orbit space
48.4 Dimension of pair invariants
48.5 Complete pair invariants
48.6 Directional invariants
48.7 Symmetric quotients
48.8 Divergences as readouts of orbit structure
48.9 Cauchy one-dimensional orbit bottom
48.10 Gaussian higher-dimensional orbit bottom
49. Local, Global, and Dynamical Distinguishability
49.1 Local distinguishability → Fisher metric
49.2 Global distinguishability → divergence
49.3 Sequential distinguishability → accumulated log evidence
49.4 Dynamical distinguishability → path-space KL
49.5 Asymptotic distinguishability → large-deviation rate
49.6 Operational distinguishability → hypothesis-testing exponent
49.7 Thermodynamic distinguishability → free-energy excess
49.8 Computational distinguishability → coding redundancy
50. Probability Geometry Under Transformation
50.1 Invertible transport preserves information
50.2 Stochastic transport contracts information
50.3 Coarse-graining removes distinctions
50.4 Inference changes the reference state
50.5 Model projection discards unresolved directions
50.6 Model expansion restores representational capacity
50.7 Dynamics redistribute distinguishability
50.8 Equilibrium minimizes distinguishability to the reference law
50.9 Learning as controlled motion through distribution space
50.10 Scientific model change as reconstruction after persistent divergence residue
PART XX — FINAL SYNTHESIS
51. KL as a Central Transition Object
51.1 Not a distance, but a directional comparison law
51.2 Not merely entropy, but a relation between probability states
51.3 Fisher geometry as its infinitesimal shadow
51.4 Amari duality as its higher-order affine structure
51.5 Large deviations as its asymptotic dynamics
51.6 Free energy as its thermodynamic realization
51.7 Coding redundancy as its operational realization
51.8 Bayesian projection as its inferential realization
51.9 Data processing as its coarse-graining law
51.10 Group invariance as its transport law
52. The Larger Theory
52.1 Probability distributions as relational states
52.2 Distinguishability as conserved or dissipated structure
52.3 Geometry from infinitesimal discrimination
52.4 Dynamics from changing distributions
52.5 Inference as constrained distributional motion
52.6 Transport as relocation without necessary information loss
52.7 Coarse-graining as irreversible distinguishability loss
52.8 Model change as response to persistent reconstruction residue
52.9 Statistical physics, machine learning, coding, and inference as realizations of one architecture
52.10 STATE → DISTINGUISH → TRANSFORM → CONTRACT/PRESERVE → PROJECT → RESIDUE → RECONSTRUCT
53. Open Structural Questions
53.1 Which statistical families collapse all natural divergences to one pair invariant?
53.2 Which source-preserving symmetry groups force such collapse?
53.3 When does global divergence become a monotone function of Fisher distance?
53.4 Which higher-order tensors determine global divergence structure?
53.5 Can divergence geometry classify model-family expressivity?
53.6 Can coarse-graining loss be decomposed into reconstructible and irrecoverable components?
53.7 What is the minimal invariant underlying families of f-divergences?
53.8 How should Fisher, Wasserstein, and divergence geometries be combined?
53.9 Can model-change dynamics be formulated intrinsically on a space of statistical manifolds?
53.10 Can persistent KL residue identify when the statistical ontology itself must change?
The compressed architecture of the entire work is:
probability state → likelihood-ratio field → KL
→ {Fisher metric ∥ Amari duality ∥ f-divergence contraction ∥ Bregman projection ∥ large-deviation rate ∥ free-energy cost}
→ {inference ∥ transport ∥ coarse-graining ∥ dynamics ∥ model change}
→ residue
→ reconstruction of a richer probability geometry.
Comments
Post a Comment