Statistical Estimation and Prediction

 

Statistical Estimation and Prediction

Table of Contents

Part I — The pre-statistical situation: what must exist before statistics is possible

1. Why statistical reasoning begins with multiple possible states of the world

1.1 How uncertainty requires more than one possible state
1.2 How a realized state differs from the alternatives that could have occurred
1.3 Why the set of imagined possibilities can itself be incomplete
1.4 How impossible, possible, plausible, and observed states differ
1.5 Why probability cannot be introduced before possibilities have been distinguished
1.6 What happens when the assumed possibility space excludes the actual state

2. How one realized world becomes only partially accessible to an observer

2.1 The difference between what exists and what becomes observable
2.2 How observation exposes some properties while hiding others
2.3 Why every observation system performs selection and transformation
2.4 How measurement resolution determines which distinctions survive observation
2.5 Why repeated measurement does not necessarily recover hidden structure
2.6 How information can be permanently destroyed before statistical analysis begins

3. How distinctions turn an observed world into representable differences

3.1 What it means to distinguish one observed state from another
3.2 How equality, ordering, magnitude, and distance create different representational structures
3.3 Why some distinctions exist in the source but not in the representation
3.4 How discretization creates categories that do not exist naturally
3.5 How measurement resolution collapses distinct source states into identical records
3.6 Why indistinguishability in data does not imply identity in the source

4. How multiplicity creates the possibility of statistical comparison

4.1 Why statistics requires more than an isolated realization
4.2 How variation can arise across units, occasions, locations, and conditions
4.3 Why repeated measurements and repeated systems are different forms of multiplicity
4.4 How temporal, spatial, hierarchical, and network multiplicity differ
4.5 Why multiple observations do not automatically constitute repeated instances of one problem
4.6 How apparent repetition can conceal changes in the generating system

5. Why comparability is the foundational inferential commitment

5.1 What it means to let one observation inform another
5.2 Which properties must remain sufficiently stable for comparison to be meaningful
5.3 How comparability can hold globally, locally, conditionally, or only partially
5.4 Why heterogeneous observations may still be inferentially related
5.5 How comparison fails when relevant generating conditions change
5.6 Why comparability is asserted before it is formalized probabilistically
5.7 How every statistical inference depends on an implicit relation among cases


Part II — Defining the question before constructing the mathematics

6. How an investigator converts incomplete observation into an inferential problem

6.1 The difference between what is observed and what someone wants to know
6.2 How scientific, operational, and commercial purposes create different targets
6.3 Why the same dataset can support many incompatible inferential questions
6.4 How an inferential problem is created by pairing observations with a target
6.5 Why the target is imposed by inquiry rather than supplied automatically by data

7. How to distinguish the true target from the recorded quantity used in its place

7.1 What makes a target directly observable, indirectly observable, or latent
7.2 How labels become proxies for quantities that actually matter
7.3 Why an available outcome variable may be easier to measure but conceptually wrong
7.4 How target definitions determine what counts as success
7.5 How target drift occurs when institutional objectives change
7.6 Why a perfectly predicted proxy can still answer the wrong question

8. How consequences determine what kinds of prediction errors matter

8.1 Why numerical error and practical consequence are different objects
8.2 How asymmetric consequences change the preferred prediction
8.3 Why ranking, classification, forecasting, and probability estimation require different objectives
8.4 How different users can rationally prefer different predictions from the same information
8.5 Why a prediction problem cannot be fully specified without stating its intended use
8.6 Where statistical prediction ends and decision-making begins


Part III — Constitutional distinctions that prevent statistical categories from collapsing

9. Why the source system must never be identified with its observations

9.1 The source as the system that generates observable consequences
9.2 The observation as a partial trace of the source
9.3 Why measurement is not recovery of the underlying state
9.4 How unobserved mechanisms can dominate observed outcomes
9.5 Why more records cannot compensate for systematically missing dimensions

10. Why an observation must not be confused with its recorded representation

10.1 How observations become encoded values
10.2 How coding schemes introduce distinctions and collapse distinctions
10.3 How aggregation changes the object being analyzed
10.4 How transformations preserve some information and destroy other information
10.5 Why the stored dataset is already several transformations removed from the source

11. Why the observed sample must not be identified with the intended population

11.1 How observed cases differ from the domain of intended inference
11.2 Why sampling frames determine who can possibly enter the data
11.3 How coverage error differs from sampling variability
11.4 Why convenience data can be statistically precise and inferentially irrelevant
11.5 How transport beyond the observed sample requires additional assumptions

12. Why estimand, estimator, and estimate are fundamentally different objects

12.1 The estimand as the quantity the analysis seeks
12.2 The estimator as the rule that maps observations to a claim
12.3 The estimate as the realized output of that rule
12.4 Why many estimators can target the same estimand
12.5 Why the same estimator can behave differently under different source structures
12.6 How estimator performance depends on assumptions about the observation process

13. Why goodness of fit must not be mistaken for truth

13.1 How multiple incompatible structures can fit the same observations
13.2 Why perfect interpolation does not establish correct representation
13.3 How underdetermination arises from finite observations
13.4 Why model compatibility is weaker than source correspondence
13.5 How predictive adequacy and explanatory adequacy can diverge

14. Why dependence, association, mechanism, and intervention must remain distinct

14.1 How variables can move together without one generating the other
14.2 How conditional dependence differs from causal dependence
14.3 How common causes create predictive relationships without direct mechanisms
14.4 How reverse direction produces misleading interpretations
14.5 Why observing a variable differs from deliberately changing it
14.6 Why predictive importance does not imply intervention value

15. Why model-conditional uncertainty is not total uncertainty

15.1 Uncertainty about parameters inside a fixed model
15.2 Uncertainty about future observations inside a fixed distribution
15.3 Uncertainty about which model should be used
15.4 Uncertainty created by representation choice
15.5 Uncertainty created by missing mechanisms
15.6 Uncertainty created by future regime changes
15.7 Why conventional confidence statements usually cover only part of the epistemic problem


Part IV — How observation systems manufacture the data available to statistics

16. How different observation systems expose different fragments of reality

16.1 Sensors that transform physical states into recorded signals
16.2 Surveys that transform responses into coded variables
16.3 Administrative systems that record activity for institutional purposes
16.4 Transaction systems that observe only events passing through a platform
16.5 Human classification systems that embed judgment into labels
16.6 Automated logging systems that record what engineers chose to instrument
16.7 Experimental systems that deliberately create informative variation

17. How measurement quality constrains everything downstream

17.1 Defining measurement rules before analyzing their outputs
17.2 Distinguishing resolution, precision, accuracy, and reproducibility
17.3 How calibration links an instrument to an external reference
17.4 How instrument drift changes measurements over time
17.5 How systematic measurement error differs from random variation
17.6 Why measurement error can attenuate, distort, or manufacture relationships

18. How coverage determines which parts of the source can ever become evidence

18.1 Geographic regions that never enter the observation system
18.2 Time periods that are absent or poorly observed
18.3 Populations excluded by the sampling frame
18.4 Rare events absent because the observation period is too short
18.5 Extreme states absent because instruments saturate or systems fail
18.6 Why missing coverage cannot be repaired merely by increasing sample size

19. How selection determines which realizations become visible

19.1 Random selection and its inferential advantages
19.2 Convenience selection and hidden population distortions
19.3 Self-selection by observed units
19.4 Selection created by institutional procedures
19.5 Survivorship as conditioning on continued existence
19.6 Selection created by previous predictions or decisions
19.7 Why selection can reverse apparently stable relationships

20. How missingness can occur at several structurally different levels

20.1 Missing values within otherwise observed cases
20.2 Missing variables that were never measured
20.3 Missing units that never entered the observation system
20.4 Missing regions of the state space
20.5 Informative missingness caused by the hidden quantity itself
20.6 Censoring in which the target is only partially revealed
20.7 Truncation in which some cases cannot appear at all
20.8 When no statistical adjustment can recover the missing information

21. How incentives alter what gets reported, measured, and recorded

21.1 Reporting incentives that distort measured outcomes
21.2 Administrative targets that change classification behavior
21.3 Compensation systems that alter recorded performance
21.4 Strategic actors who respond to measurement rules
21.5 Goodhart effects when a measure becomes an objective
21.6 Adversarial behavior against scoring and screening systems
21.7 Why the observation mechanism may change after deployment


Part V — Constructing the represented statistical world

22. How observed distinctions become statistical variables

22.1 Numerical variables representing magnitude
22.2 Ordered variables representing rank without known distance
22.3 Categorical variables representing discrete distinctions
22.4 Binary variables representing deliberately collapsed alternatives
22.5 Derived variables created from other measurements
22.6 Proxy variables standing in for inaccessible targets
22.7 Latent variables representing unobserved structure

23. How units of analysis determine what counts as one observation

23.1 Individuals, organizations, events, and transactions as possible units
23.2 Repeated measurements nested within the same unit
23.3 Units nested within groups, institutions, or regions
23.4 Temporal units whose dependence changes with separation
23.5 Spatial units whose relationships depend on proximity
23.6 Network units whose dependence follows connections rather than distance
23.7 Why choosing the wrong unit can manufacture pseudo-replication

24. How representation choices determine what relationships can be expressed

24.1 Encoding categorical states into analyzable quantities
24.2 Scaling variables without changing their substantive meaning
24.3 Transforming skewed or multiplicative quantities
24.4 Constructing interactions between represented properties
24.5 Creating basis expansions for nonlinear relationships
24.6 Aggregating observations across time or groups
24.7 How representation can manufacture apparent simplicity
24.8 How representation can make real structure impossible to recover

25. How domains define where a statistical claim is intended to hold

25.1 The source domain that generates possible realizations
25.2 The observed domain that actually enters the data
25.3 The training domain used for constructing a model
25.4 The validation domain used for selecting among alternatives
25.5 The test domain used for final evaluation
25.6 The deployment domain where predictions will actually be used
25.7 The target population about which conclusions are intended
25.8 How domain mismatch becomes a generalization problem


Part VI — Probability as a formal calculus over represented uncertainty

26. Why probability is introduced only after representation and comparability exist

26.1 The need to represent uncertainty across possible outcomes
26.2 The need to combine information coherently across related observations
26.3 The difference between uncertainty in the source and uncertainty in the representation
26.4 Why probabilities operate on represented events rather than directly on reality
26.5 What probability adds that descriptive counting alone cannot provide

27. How events define the propositions to which probability can be assigned

27.1 Events as subsets of represented possibilities
27.2 Simple events and compound events
27.3 Mutually exclusive and overlapping events
27.4 Exhaustive collections of events
27.5 Event spaces generated by observable distinctions
27.6 Why changing representation changes the event space

28. How probability assignments obey a coherent mathematical structure

28.1 Nonnegative weighting of possible events
28.2 Normalization across the represented possibility space
28.3 Additivity across mutually exclusive events
28.4 Conditional probability after information becomes available
28.5 Competing interpretations of probability
28.6 Why probabilistic coherence does not guarantee empirical adequacy

29. How random variables connect represented quantities to probability

29.1 Random variables as mappings from possible outcomes to values
29.2 Discrete random variables and countable outcomes
29.3 Continuous random variables and numerical ranges
29.4 Vector-valued random variables representing multiple measurements
29.5 Indicator variables representing events numerically
29.6 Why the same source can support many different random-variable representations

30. How joint, marginal, and conditional distributions encode relationships

30.1 Joint distributions over several represented quantities
30.2 Marginalization by ignoring parts of the joint structure
30.3 Conditional distributions after observing additional information
30.4 Dependence as failure of probabilistic factorization
30.5 Independence as a strong structural simplification
30.6 Conditional independence as dependence mediated through other variables

31. How expectation summarizes probabilistic quantities

31.1 Expected value as probability-weighted averaging
31.2 Conditional expectation after information becomes available
31.3 Expectation of transformed random quantities
31.4 Linearity of expectation and why it matters
31.5 Variance as expected squared deviation
31.6 Covariance as expected joint deviation
31.7 Why expectation is the bridge from probability to many estimators and predictors


Part VII — Formalizing inferential relationships among observations

32. How exchangeability formalizes one important kind of comparability

32.1 Why order may sometimes carry no inferential information
32.2 Finite exchangeability across observations
32.3 Infinite exchangeability and mixture representations
32.4 Conditional exchangeability after accounting for known structure
32.5 Why exchangeability is weaker than assuming identical mechanisms
32.6 Where exchangeability fails in real systems

33. How independence simplifies inference and where it becomes unrealistic

33.1 Independent observations under repeated sampling
33.2 Conditional independence after controlling relevant information
33.3 Dependence induced by shared environments
33.4 Dependence created by repeated measurements
33.5 Dependence created by networks and spatial proximity
33.6 Consequences of pretending dependent observations are independent

34. How stationarity formalizes stability across time

34.1 Stable marginal behavior through time
34.2 Stable joint behavior under time translation
34.3 Local stationarity over restricted periods
34.4 Seasonal structure that violates simple stationarity
34.5 Structural breaks that destroy historical comparability
34.6 Why future prediction requires some form of temporal persistence

35. How hierarchical structure creates partial rather than complete comparability

35.1 Observations nested within shared groups
35.2 Group-specific variation around common structure
35.3 Partial pooling across related groups
35.4 Random effects as probabilistic representations of group heterogeneity
35.5 Why complete pooling and complete separation are limiting cases
35.6 How multilevel models encode graded inferential relationships


Part VIII — From repeated evidence to stable statistical quantities

36. Why empirical quantities can stabilize as comparable observations accumulate

36.1 Sample averages as functions of repeated observations
36.2 The law of large numbers as stabilization under assumptions
36.3 Weak and strong forms of convergence
36.4 Why convergence does not repair systematic observation bias
36.5 Why larger samples improve precision without necessarily improving validity

37. How concentration describes the probability of large deviations from typical behavior

37.1 Deviations of empirical averages from expectations
37.2 The role of boundedness and variance assumptions
37.3 Concentration inequalities as finite-sample guarantees
37.4 Why high-dimensional settings make concentration especially important
37.5 How dependence weakens familiar concentration results

38. How central-limit behavior creates approximate distributions for estimation error

38.1 Standardized sums of many contributions
38.2 Conditions supporting normal approximation
38.3 Why convergence in distribution differs from convergence in probability
38.4 How central-limit approximations produce standard errors
38.5 Where heavy tails, dependence, or small samples break the approximation


Part IX — Defining the estimation problem precisely

39. How population properties become explicit inferential targets

39.1 Means, variances, quantiles, and probabilities as distributional functionals
39.2 Regression functions as conditional population quantities
39.3 Risk measures as summaries of tail behavior
39.4 Latent quantities that require additional assumptions
39.5 Structural quantities that cannot be identified from association alone

40. How estimands specify exactly what is being asked of the data

40.1 Specifying a quantity before choosing an estimator
40.2 Distinguishing descriptive estimands from predictive estimands
40.3 Population-wide estimands versus subgroup-specific estimands
40.4 Conditional estimands indexed by observed information
40.5 Why changing the estimand changes the scientific question

41. Why identification must be solved before estimation

41.1 What it means for observations to distinguish competing parameter values
41.2 Global identification across the entire parameter space
41.3 Local identification near the true parameter value
41.4 Partial identification when data determine only a set of possibilities
41.5 Non-identification when observationally equivalent explanations remain
41.6 How additional assumptions can create identification
41.7 Why stronger assumptions can manufacture apparent certainty

42. How estimators map finite observations into claims about unknown quantities

42.1 Statistics as functions of the observed data
42.2 Estimators as rules defined before observing the realized sample
42.3 Estimates as sample-specific numerical outputs
42.4 Deterministic and randomized estimators
42.5 Plug-in estimation using estimated distributions
42.6 Why estimator choice encodes assumptions and priorities


Part X — Constructing and comparing estimators

43. How empirical averages generate the simplest estimators

43.1 Estimating means from sample averages
43.2 Estimating probabilities from empirical frequencies
43.3 Estimating moments from empirical moments
43.4 Estimating conditional quantities through stratification
43.5 Why empirical estimation becomes difficult in high dimensions

44. How estimating equations generalize moment-based estimation

44.1 Matching theoretical and empirical moments
44.2 Solving estimating equations for unknown parameters
44.3 Overidentified systems with more equations than parameters
44.4 Generalized method-of-moments reasoning
44.5 Sensitivity to poorly chosen moment conditions

45. How likelihood converts a probabilistic model into an estimation criterion

45.1 Likelihood as compatibility of parameters with observed data
45.2 Log-likelihood as an additive optimization objective
45.3 Maximum likelihood estimation
45.4 Score functions and information
45.5 Likelihood under model misspecification
45.6 Why maximum likelihood optimizes within the assumed family rather than proving it correct

46. How Bayesian estimation combines a prior representation with observed evidence

46.1 Prior distributions over uncertain quantities
46.2 Likelihood contributions from observed data
46.3 Posterior distributions after conditioning
46.4 Posterior expectations, medians, and modes as estimators
46.5 Posterior predictive distributions
46.6 Sensitivity to prior and likelihood assumptions
46.7 Why Bayesian coherence does not eliminate structural uncertainty

47. How estimator quality is judged under repeated hypothetical samples

47.1 Bias as systematic displacement from the estimand
47.2 Variance as sensitivity to the realized sample
47.3 Mean squared error combining bias and variance
47.4 Consistency as convergence toward the estimand
47.5 Efficiency as relative use of available information
47.6 Robustness against contamination and model deviation
47.7 Stability under small perturbations of the observed data


Part XI — Quantifying uncertainty about estimates

48. How sampling distributions describe estimator variability

48.1 Repeated-sample thought experiments
48.2 Exact sampling distributions when derivable
48.3 Approximate sampling distributions when exact forms are unavailable
48.4 Standard errors as scales of estimator variability
48.5 Why sampling uncertainty is conditional on the assumed data-generating structure

49. How confidence intervals convert estimator variability into coverage statements

49.1 Constructing intervals from pivotal quantities
49.2 Interpreting repeated-sample coverage correctly
49.3 Why confidence level is not posterior probability
49.4 Approximate intervals based on asymptotic normality
49.5 Bootstrap confidence intervals
49.6 How model misspecification destroys nominal coverage

50. How Bayesian credible intervals answer a different uncertainty question

50.1 Posterior probability assigned to parameter regions
50.2 Equal-tail credible intervals
50.3 Highest-posterior-density regions
50.4 Prior sensitivity of credible regions
50.5 Why credible intervals and confidence intervals can coincide numerically but differ conceptually

51. How hypothesis tests turn claims into procedures for detecting disagreement with data

51.1 Null hypotheses as restricted statistical worlds
51.2 Alternative hypotheses representing departures of interest
51.3 Test statistics designed to expose those departures
51.4 Reference distributions under the null
51.5 p-values as tail probabilities under an assumed reference model
51.6 Type I and Type II errors
51.7 Power as the probability of detecting specified departures
51.8 Why statistical significance does not establish substantive importance

52. How repeated searching manufactures apparently strong evidence

52.1 Multiple testing across many hypotheses
52.2 Data-dependent hypothesis generation
52.3 Selective reporting of favorable results
52.4 Model search followed by naive inference
52.5 Winner's curse among selected effects
52.6 False discovery control
52.7 Selective inference after data-dependent selection


Part XII — Defining the prediction problem

53. How prediction differs from estimating a population property

53.1 Predicting an unobserved value for a particular case
53.2 Predicting a probability distribution rather than a single number
53.3 Predicting class membership
53.4 Predicting relative rank
53.5 Predicting event timing
53.6 Predicting sequences and trajectories
53.7 Why estimation and prediction can prefer different models

54. How information available at prediction time constrains legitimate predictors

54.1 Predictors known before the target occurs
54.2 Predictors observed only after the target occurs
54.3 Leakage through future or post-outcome information
54.4 Prediction from partial histories
54.5 Real-time prediction under delayed measurements
54.6 Why deployment-time information must match evaluation-time information

55. How loss functions translate prediction errors into consequences

55.1 Squared loss emphasizing large numerical errors
55.2 Absolute loss emphasizing typical absolute deviation
55.3 Classification loss penalizing incorrect categories
55.4 Logarithmic loss evaluating full probability assignments
55.5 Asymmetric loss for unequal consequences
55.6 Cost-sensitive prediction under heterogeneous stakes
55.7 Why the optimal predictor changes when the loss changes

56. How optimal prediction follows from the conditional distribution under stated assumptions

56.1 Conditional expectation under squared-error loss
56.2 Conditional median under absolute-error loss
56.3 Conditional mode under simple classification loss
56.4 Probability thresholds under asymmetric classification costs
56.5 Bayes risk as minimum expected loss within the assumed representation
56.6 Irreducible uncertainty remaining even under the optimal predictor
56.7 Why Bayes-optimal does not mean optimal outside the assumed probabilistic world


Part XIII — Constructing predictive functions from finite observations

57. How candidate function classes restrict which predictive relationships can be learned

57.1 Linear function classes
57.2 Additive function classes
57.3 Local function classes
57.4 Partition-based function classes
57.5 Kernel-based function classes
57.6 Compositional neural function classes
57.7 Why every function class encodes representational exclusions

58. How empirical risk minimization converts prediction into an optimization problem

58.1 Empirical loss calculated over observed training cases
58.2 Selecting a function that minimizes observed loss
58.3 Population risk that remains unobserved
58.4 Optimization error versus statistical error
58.5 Local minima, saddle points, and numerical approximation
58.6 Why successful optimization does not imply successful generalization

59. How approximation, estimation, and optimization errors arise from different sources

59.1 Approximation error from an inadequate function class
59.2 Estimation error from finite observations
59.3 Optimization error from incomplete numerical search
59.4 Representation error from missing predictive information
59.5 Observation error inherited from the data-production system
59.6 Why reducing one error source can expose another

60. How regularization deliberately restricts fitting to improve future performance

60.1 Complexity penalties added to empirical objectives
60.2 Ridge penalties shrinking parameter magnitudes
60.3 Lasso penalties encouraging sparse solutions
60.4 Elastic-net penalties combining shrinkage and sparsity
60.5 Early stopping as optimization-dependent regularization
60.6 Architectural restrictions as implicit regularization
60.7 Why regularization trades flexibility for stability rather than creating truth


Part XIV — Evaluating whether predictive performance generalizes

61. Why training performance cannot estimate future performance without correction

61.1 Optimism created by evaluating on fitted observations
61.2 Training error as a biased estimate of future error
61.3 Validation data for model and hyperparameter selection
61.4 Test data reserved for final evaluation
61.5 How repeated test-set use converts test data into training information

62. How resampling estimates performance when observations are limited

62.1 Simple holdout evaluation
62.2 K-fold cross-validation
62.3 Leave-one-out cross-validation
62.4 Repeated cross-validation
62.5 Bootstrap estimation of prediction error
62.6 Nested cross-validation for model-selection pipelines
62.7 Why resampling still depends on comparability between folds and deployment

63. How evaluation design must reproduce the structure of deployment

63.1 Random splitting for exchangeable observations
63.2 Grouped splitting for clustered observations
63.3 Temporal splitting for future prediction
63.4 Spatial splitting for geographic generalization
63.5 Leave-domain-out validation for transport problems
63.6 External validation on independently collected data
63.7 Why the wrong split can manufacture convincing but useless performance

64. How predictive metrics answer different operational questions

64.1 Mean squared error for numerical prediction
64.2 Mean absolute error for robust numerical prediction
64.3 Accuracy for symmetric categorical decisions
64.4 Precision and recall under imbalanced outcomes
64.5 Sensitivity and specificity under diagnostic framing
64.6 ROC curves and threshold-independent ranking behavior
64.7 Precision-recall curves for rare outcomes
64.8 Ranking metrics when ordering matters more than classification

65. How calibration evaluates whether predicted probabilities mean what they claim

65.1 Marginal calibration across all predictions
65.2 Calibration curves across probability ranges
65.3 Conditional calibration within relevant subgroups
65.4 Calibration under changing prevalence
65.5 Calibration drift after deployment
65.6 Why discrimination can remain strong while probability estimates become wrong

66. How aggregate evaluation can conceal catastrophic local failure

66.1 Average performance hiding subgroup errors
66.2 Tail performance hidden by mean metrics
66.3 Rare catastrophic events overwhelmed by common easy cases
66.4 Unequal consequences across populations
66.5 Performance degradation concentrated near regime boundaries
66.6 Why operational evaluation often needs a failure distribution rather than one score


Part XV — Representation, dimension, and latent structure

67. How high dimensionality changes the geometry of statistical estimation

67.1 Growth of possible configurations with additional variables
67.2 Sparsity of observations in high-dimensional spaces
67.3 Distance concentration and failure of intuitive neighborhood concepts
67.4 Parameter proliferation relative to available observations
67.5 Why structural assumptions become more important as dimension increases

68. How dimension reduction trades retained information for simpler representation

68.1 Principal components as directions of maximal variation
68.2 Low-rank approximations of high-dimensional observations
68.3 Factor models representing shared variation
68.4 Sparse representations preserving selected dimensions
68.5 Nonlinear embeddings
68.6 Why useful compression is not evidence that compressed dimensions are ontologically real

69. How latent-variable models posit hidden structure to explain observed relationships

69.1 Hidden common factors generating observable dependence
69.2 Mixture components representing unobserved subpopulations
69.3 Hidden states evolving through time
69.4 Latent classes explaining categorical heterogeneity
69.5 Identification problems in latent-variable models
69.6 Why multiple latent structures can explain the same observations


Part XVI — Major predictive model families as implementations of earlier principles

70. How linear and generalized linear models impose global parametric structure

70.1 Ordinary linear regression for conditional means
70.2 Multiple regression with several represented predictors
70.3 Interaction terms representing conditional effects
70.4 Polynomial and basis expansions representing nonlinear trends
70.5 Logistic regression for conditional class probabilities
70.6 Generalized linear models for non-Gaussian responses
70.7 Diagnostics for systematic residual structure

71. How local prediction methods infer from nearby observations

71.1 Nearest-neighbor prediction
71.2 Kernel-weighted local averaging
71.3 Local polynomial regression
71.4 Bandwidth choice controlling neighborhood size
71.5 Failure of locality in high-dimensional spaces
71.6 Why local similarity depends entirely on representation

72. How partitioning methods divide represented space into predictive regions

72.1 Recursive binary partitioning
72.2 Split criteria based on predictive improvement
72.3 Tree depth as model complexity
72.4 Pruning to reduce unstable partitions
72.5 Interpretability of tree-based decision rules
72.6 Instability of individual trees under small data changes

73. How ensemble methods combine unstable predictors into more reliable systems

73.1 Bagging through repeated resampling and averaging
73.2 Random forests through feature-randomized trees
73.3 Boosting through sequential correction of predictive residuals
73.4 Gradient boosting as functional optimization
73.5 Diversity among component predictors
73.6 Why ensemble accuracy can reduce interpretability

74. How margin and kernel methods construct separating boundaries

74.1 Linear separating hyperplanes
74.2 Maximum-margin classification
74.3 Soft margins allowing classification violations
74.4 Kernel functions representing implicit feature transformations
74.5 Support vectors determining the separating boundary
74.6 Sensitivity to kernel and regularization choices

75. How neural predictors construct highly flexible compositional functions

75.1 Layered function composition
75.2 Nonlinear activation functions
75.3 Gradient-based parameter fitting
75.4 Representation learning inside predictive optimization
75.5 Overparameterization beyond classical parameter-count intuition
75.6 Explicit and implicit regularization
75.7 Scaling behavior with data, parameters, and computation
75.8 Why high predictive capacity does not eliminate observation or target failure


Part XVII — Estimation under incomplete and structured observation

76. How missing-data methods depend on assumptions about why values are missing

76.1 Missing completely at random
76.2 Missing at random conditional on observed information
76.3 Missing not at random
76.4 Complete-case analysis
76.5 Weighting adjustments
76.6 Multiple imputation
76.7 Sensitivity analysis for unverifiable missingness assumptions

77. How censoring and truncation alter what can be learned about outcomes

77.1 Right censoring of incomplete event times
77.2 Left censoring below detection thresholds
77.3 Interval censoring between observation occasions
77.4 Truncation excluding entire cases from observation
77.5 Survival functions and hazard functions
77.6 Kaplan-Meier estimation under censoring assumptions
77.7 Regression with censored outcomes

78. How complex sampling designs alter estimation and uncertainty

78.1 Unequal-probability sampling
78.2 Stratified sampling
78.3 Cluster sampling
78.4 Multistage sampling
78.5 Survey weights correcting known design probabilities
78.6 Design effects on estimator variance
78.7 Why naive IID analysis fails under complex sampling


Part XVIII — Prediction under time, transport, and changing regimes

79. How time ordering changes the statistical problem

79.1 Past information versus future outcomes
79.2 Serial dependence between nearby observations
79.3 Trend representing persistent directional movement
79.4 Seasonality representing recurring temporal structure
79.5 Forecast horizons and expanding uncertainty
79.6 Why random train-test splitting fails for many forecasting problems

80. How state-space models represent hidden systems evolving through time

80.1 Latent states generating observable measurements
80.2 Transition models describing state evolution
80.3 Observation models describing noisy measurement
80.4 Filtering current hidden states from past observations
80.5 Smoothing past hidden states using later information
80.6 Forecasting future states from estimated dynamics

81. How transport asks whether a relationship survives outside the observed domain

81.1 Transport across populations
81.2 Transport across institutions
81.3 Transport across geographic locations
81.4 Transport across historical periods
81.5 Transport across operating conditions
81.6 Why internal validation cannot establish external validity

82. How distribution shift breaks relationships learned from historical data

82.1 Covariate shift in observed predictor distributions
82.2 Label shift in outcome prevalence
82.3 Conditional shift in relationships between predictors and outcomes
82.4 Concept drift in the predictive mapping itself
82.5 Structural breaks caused by new mechanisms
82.6 Regime change invalidating historical comparability

83. How deployed predictions alter the system that generated the training data

83.1 Predictions changing human decisions
83.2 Decisions changing which outcomes become observable
83.3 Selection effects created by model deployment
83.4 Strategic adaptation by people subject to prediction
83.5 Self-fulfilling predictions
83.6 Self-defeating predictions
83.7 Feedback loops that generate a new data distribution


Part XIX — Distribution-free and weak-assumption predictive guarantees

84. How conformal prediction constructs finite-sample prediction sets from exchangeability

84.1 Nonconformity scores measuring unusual candidate outcomes
84.2 Calibration observations used to construct prediction thresholds
84.3 Marginal coverage under exchangeability
84.4 Split conformal prediction
84.5 Full conformal prediction
84.6 Regression prediction intervals
84.7 Classification prediction sets
84.8 Why marginal coverage does not imply conditional validity

85. How robust prediction weakens dependence on exact parametric assumptions

85.1 Robust losses reducing sensitivity to extreme observations
85.2 Heavy-tailed outcome distributions
85.3 Contamination models
85.4 Distributionally robust optimization
85.5 Worst-case performance over uncertainty sets
85.6 The tradeoff between robustness and efficiency


Part XX — Structure construction without a supplied prediction target

86. What unsupervised procedures actually optimize when no response variable is supplied

86.1 Similarity defined by a chosen representation
86.2 Distance defined by a chosen geometry
86.3 Density defined relative to a chosen coordinate system
86.4 Compression defined by a reconstruction criterion
86.5 Clusters defined by an optimization objective
86.6 Why the algorithm always supplies structure even when the source does not

87. How clustering partitions observations according to imposed similarity criteria

87.1 K-means clustering around representative centers
87.2 Hierarchical clustering through recursive merging or splitting
87.3 Density-based clustering around high-density regions
87.4 Model-based clustering through mixture distributions
87.5 Sensitivity of clusters to scaling and distance
87.6 Stability of clusters across samples and perturbations

88. How discovered structure can be tested against external evidence

88.1 Stability under repeated samples
88.2 Reproducibility under alternative representations
88.3 Predictive consequences on variables not used for construction
88.4 Agreement with independently measured structure
88.5 Experimental discrimination where intervention is possible
88.6 Why stable, interpretable clusters still need not represent real kinds


Part XXI — Detecting when a statistical model has stopped deserving trust

89. How residual behavior can reveal model inadequacy

89.1 Systematic residual patterns
89.2 Heterogeneous residual variance
89.3 Temporal residual dependence
89.4 Spatial residual dependence
89.5 Extreme residuals and unrepresented mechanisms
89.6 Why residual conformity cannot prove the model is correct

90. How calibration decay can expose changing predictive relationships

90.1 Drift in predicted probability reliability
90.2 Subgroup-specific calibration failure
90.3 Calibration failure concentrated in new regions of state space
90.4 Distinguishing random fluctuation from persistent deterioration
90.5 When recalibration is sufficient
90.6 When recalibration merely hides deeper structural failure

91. How novelty detection can expose observations outside historical experience

91.1 Distance from previously observed regions
91.2 Density-based novelty scores
91.3 Representation-dependent anomaly detection
91.4 New combinations of familiar variables
91.5 New variables or mechanisms absent from the original ontology
91.6 Why novelty scores cannot detect what the representation cannot express

92. How to distinguish parameter drift from representation failure

92.1 Same structure with changing parameter values
92.2 Same variables with changing conditional relationships
92.3 New hidden mechanism producing familiar observations
92.4 Missing variable becoming operationally important
92.5 Observation-system failure mimicking distribution shift
92.6 Criteria for updating, retraining, replacing, or abandoning the model


Part XXII — The jurisdictional boundary of statistical estimation and prediction

93. Why prediction does not by itself explain the mechanism producing an outcome

93.1 Predictive association without causal mechanism
93.2 Mechanistic explanation requiring additional structural claims
93.3 Equivalent predictive models with incompatible explanations
93.4 Why explanatory adequacy requires evidence beyond test-set accuracy

94. Why observing a predictor does not establish the effect of changing it

94.1 Conditioning on X versus intervening on X
94.2 Confounding by common causes
94.3 Selection effects created by treatment assignment
94.4 Counterfactual outcomes that are never jointly observable
94.5 Why predictive importance can be useless for intervention design

95. How experiments create information that passive observation may not contain

95.1 Deliberate perturbation of candidate causal variables
95.2 Randomization breaking systematic treatment selection
95.3 Control groups providing counterfactual reference distributions
95.4 Factorial experiments discriminating interactions
95.5 Sequential experimentation adapting to accumulated evidence
95.6 Why experiment design belongs upstream of estimation even when statistics analyzes the result

96. Why prediction does not specify which action should be taken

96.1 Predictions describing consequences without expressing preferences
96.2 Utility and cost functions ranking possible outcomes
96.3 Constraints limiting feasible actions
96.4 Decision rules mapping predictions to actions
96.5 Why greater predictive accuracy can still produce worse decisions
96.6 How decision theory begins where pure prediction ends

97. Why prediction differs from controlling a dynamic system

97.1 Forecasting future state under existing inputs
97.2 Selecting inputs to alter future state
97.3 Feedback from current outcomes into future actions
97.4 Controllability as distinct from predictability
97.5 Stable control despite imperfect prediction
97.6 Excellent prediction despite little ability to control

98. Why statistical analysis differs from engineering the successor system

98.1 Estimating properties of systems that already exist
98.2 Generating candidate systems that do not yet exist
98.3 Choosing design variables rather than merely observing predictors
98.4 Using experiments to discriminate among candidate designs
98.5 Optimization under physical and operational constraints
98.6 Statistics as evidence inside engineering rather than a substitute for engineering


Part XXIII — Failure modes arranged by where they enter the inferential chain

99. Source failures that occur before observation begins

99.1 Relevant mechanisms absent from the conceptual ontology
99.2 Relevant populations excluded from the imagined domain
99.3 New mechanisms outside historical experience
99.4 Rare regimes never previously realized
99.5 Structural changes that make old distinctions obsolete

100. Observation failures that corrupt what becomes available as evidence

100.1 Sensor failure and measurement drift
100.2 Coding and transcription errors
100.3 Reporting distortion
100.4 Selection into observation
100.5 Informative missingness
100.6 Administrative incentives changing measured behavior

101. Representation failures that corrupt the statistical world built from observations

101.1 Wrong variables representing the wrong distinctions
101.2 Wrong categories collapsing substantively different states
101.3 Wrong aggregation hiding local structure
101.4 Proxy substitution replacing the target with convenience
101.5 Hidden-state collapse
101.6 Representations incapable of expressing the actual mechanism

102. Comparability failures that make observed cases poor evidence for one another

102.1 Hidden population heterogeneity
102.2 Changing institutional environment
102.3 Temporal regime change
102.4 Geographic heterogeneity
102.5 Network interference between supposedly separate units
102.6 Deployment-induced changes in behavior

103. Estimation failures occurring after the problem has been correctly represented

103.1 Weak or absent identification
103.2 Excessive sampling variability
103.3 Systematic estimator bias
103.4 Numerical instability
103.5 Poor optimization
103.6 Selection-induced optimism

104. Prediction failures that appear only when the model encounters new cases

104.1 Overfitting to sample-specific variation
104.2 Leakage from unavailable future information
104.3 Calibration failure
104.4 Tail failure
104.5 Subgroup-specific failure
104.6 Distribution shift
104.7 New regimes outside the training experience

105. Evaluation failures that make weak models appear successful

105.1 Wrong performance metric
105.2 Wrong test population
105.3 Random splitting when temporal splitting is required
105.4 Repeated benchmark reuse
105.5 Hidden test contamination
105.6 Aggregate performance concealing catastrophic local failure
105.7 Statistical success concealing operational failure


Part XXIV — The irreducible limits on statistical claims

106. What statistical estimation can legitimately establish under explicit conditions

106.1 Claims conditional on the observation mechanism
106.2 Claims conditional on the represented variables
106.3 Claims conditional on inferential comparability
106.4 Claims conditional on identification assumptions
106.5 Claims conditional on the probability model or resampling structure
106.6 Claims conditional on stability across the intended domain

107. What statistical estimation cannot manufacture from absent information

107.1 Variables that were never observed
107.2 Mechanisms excluded from the representation
107.3 Populations excluded from the observation system
107.4 Counterfactual outcomes unsupported by causal structure
107.5 Future regimes absent from historical experience
107.6 Correct ontology merely from greater computational power

108. The non-collapses that define the field's jurisdiction

108.1 Source system != observed trace
108.2 Observation != recorded representation
108.3 Observed sample != target population
108.4 Estimand != estimator != realized estimate
108.5 Good fit != correct source description
108.6 Association != mechanism
108.7 Predictive importance != causal importance
108.8 Prediction != intervention
108.9 Prediction != decision
108.10 Prediction != control
108.11 Prediction != engineering design
108.12 Training performance != deployment performance
108.13 Model-conditional uncertainty != total epistemic uncertainty
108.14 Larger dataset != better observability
108.15 More flexible model != better representation
108.16 Better prediction != better target
108.17 Historical validity != future validity
108.18 Statistical precision != structural correctness


Part XXV — The complete admissibility test for an estimation or prediction claim

109. Questions that must be answered about the source before analyzing the dataset

109.1 What system generated the phenomena being studied?
109.2 Which possible states of that system are relevant to the question?
109.3 Which relevant states could the observation system actually detect?
109.4 Which mechanisms could remain invisible despite extensive data collection?
109.5 What evidence exists that the source remained sufficiently stable during observation?

110. Questions that must be answered about how observations became data

110.1 What observation mechanism produced each recorded variable?
110.2 Which populations, locations, and periods were covered?
110.3 Which cases could never enter the dataset?
110.4 What measurement errors or transformations occurred before recording?
110.5 What incentives affected measurement, reporting, or classification?
110.6 Which missingness mechanisms remain plausible?

111. Questions that must be answered before treating observations as comparable evidence

111.1 Why should one observed case inform another case?
111.2 Which aspects of the generating conditions are assumed stable?
111.3 Which forms of heterogeneity are explicitly modeled?
111.4 Which dependencies violate simple independent-sampling assumptions?
111.5 Across which populations, times, locations, or regimes is comparison intended?
111.6 What changes would make that comparison invalid?

112. Questions that must be answered before choosing an estimator or predictor

112.1 What quantity is actually being estimated or predicted?
112.2 Is the recorded label the true target or merely a proxy?
112.3 Is the target identifiable from the available observations?
112.4 Which assumptions create that identification?
112.5 What information does the proposed representation discard?
112.6 What loss or consequence determines what counts as predictive success?

113. Questions that must be answered before accepting reported uncertainty

113.1 Which uncertainty sources are included in the calculation?
113.2 Which uncertainty sources are conditioned away by the model?
113.3 How sensitive are results to alternative plausible representations?
113.4 How sensitive are results to alternative plausible models?
113.5 How sensitive are results to observation or missingness assumptions?
113.6 Does the reported interval describe parameter, predictive, model, or total uncertainty?

114. Questions that must be answered before claiming generalization

114.1 Why should the fitted relationship persist beyond the observed sample?
114.2 What evidence supports comparability between training and deployment domains?
114.3 Has evaluation reproduced the temporal, spatial, hierarchical, or institutional structure of deployment?
114.4 Has the model been tested on genuinely external observations?
114.5 Which subgroups or state-space regions remain poorly supported?
114.6 Which future changes would invalidate the generalization argument?

115. Questions that must be answered before continuing to use a deployed model

115.1 What observations would contradict the model's expected behavior?
115.2 Which residual patterns indicate parameter drift rather than random variation?
115.3 Which observations suggest a missing variable or new mechanism?
115.4 When is recalibration sufficient?
115.5 When is retraining required?
115.6 When must the representation itself be replaced?
115.7 When must the observation system be redesigned?
115.8 Under what conditions should the model be abandoned rather than repaired?

116. Questions that determine whether the problem has left the jurisdiction of prediction

116.1 Is the user asking what is likely to happen?
116.2 Is the user asking why it happens?
116.3 Is the user asking what would happen under intervention?
116.4 Is the user asking which experiment would discriminate competing explanations?
116.5 Is the user asking which action should be chosen?
116.6 Is the user asking how to control a dynamic system?
116.7 Is the user asking how to design a new system?
116.8 Is a statistical result being promoted into a claim that requires a different discipline?


Comments

Popular posts from this blog

Semiotics Rebooted

ORSI: The Telic Geometry of Meaning

THE COLLAPSE ENGINE: AI, Capital, and the Terminal Logic of 2025