Statistical Estimation and Prediction
Statistical Estimation and Prediction
Table of Contents
Part I — The pre-statistical situation: what must exist before statistics is possible
1. Why statistical reasoning begins with multiple possible states of the world
1.1 How uncertainty requires more than one possible state
1.2 How a realized state differs from the alternatives that could have occurred
1.3 Why the set of imagined possibilities can itself be incomplete
1.4 How impossible, possible, plausible, and observed states differ
1.5 Why probability cannot be introduced before possibilities have been distinguished
1.6 What happens when the assumed possibility space excludes the actual state
2. How one realized world becomes only partially accessible to an observer
2.1 The difference between what exists and what becomes observable
2.2 How observation exposes some properties while hiding others
2.3 Why every observation system performs selection and transformation
2.4 How measurement resolution determines which distinctions survive observation
2.5 Why repeated measurement does not necessarily recover hidden structure
2.6 How information can be permanently destroyed before statistical analysis begins
3. How distinctions turn an observed world into representable differences
3.1 What it means to distinguish one observed state from another
3.2 How equality, ordering, magnitude, and distance create different representational structures
3.3 Why some distinctions exist in the source but not in the representation
3.4 How discretization creates categories that do not exist naturally
3.5 How measurement resolution collapses distinct source states into identical records
3.6 Why indistinguishability in data does not imply identity in the source
4. How multiplicity creates the possibility of statistical comparison
4.1 Why statistics requires more than an isolated realization
4.2 How variation can arise across units, occasions, locations, and conditions
4.3 Why repeated measurements and repeated systems are different forms of multiplicity
4.4 How temporal, spatial, hierarchical, and network multiplicity differ
4.5 Why multiple observations do not automatically constitute repeated instances of one problem
4.6 How apparent repetition can conceal changes in the generating system
5. Why comparability is the foundational inferential commitment
5.1 What it means to let one observation inform another
5.2 Which properties must remain sufficiently stable for comparison to be meaningful
5.3 How comparability can hold globally, locally, conditionally, or only partially
5.4 Why heterogeneous observations may still be inferentially related
5.5 How comparison fails when relevant generating conditions change
5.6 Why comparability is asserted before it is formalized probabilistically
5.7 How every statistical inference depends on an implicit relation among cases
Part II — Defining the question before constructing the mathematics
6. How an investigator converts incomplete observation into an inferential problem
6.1 The difference between what is observed and what someone wants to know
6.2 How scientific, operational, and commercial purposes create different targets
6.3 Why the same dataset can support many incompatible inferential questions
6.4 How an inferential problem is created by pairing observations with a target
6.5 Why the target is imposed by inquiry rather than supplied automatically by data
7. How to distinguish the true target from the recorded quantity used in its place
7.1 What makes a target directly observable, indirectly observable, or latent
7.2 How labels become proxies for quantities that actually matter
7.3 Why an available outcome variable may be easier to measure but conceptually wrong
7.4 How target definitions determine what counts as success
7.5 How target drift occurs when institutional objectives change
7.6 Why a perfectly predicted proxy can still answer the wrong question
8. How consequences determine what kinds of prediction errors matter
8.1 Why numerical error and practical consequence are different objects
8.2 How asymmetric consequences change the preferred prediction
8.3 Why ranking, classification, forecasting, and probability estimation require different objectives
8.4 How different users can rationally prefer different predictions from the same information
8.5 Why a prediction problem cannot be fully specified without stating its intended use
8.6 Where statistical prediction ends and decision-making begins
Part III — Constitutional distinctions that prevent statistical categories from collapsing
9. Why the source system must never be identified with its observations
9.1 The source as the system that generates observable consequences
9.2 The observation as a partial trace of the source
9.3 Why measurement is not recovery of the underlying state
9.4 How unobserved mechanisms can dominate observed outcomes
9.5 Why more records cannot compensate for systematically missing dimensions
10. Why an observation must not be confused with its recorded representation
10.1 How observations become encoded values
10.2 How coding schemes introduce distinctions and collapse distinctions
10.3 How aggregation changes the object being analyzed
10.4 How transformations preserve some information and destroy other information
10.5 Why the stored dataset is already several transformations removed from the source
11. Why the observed sample must not be identified with the intended population
11.1 How observed cases differ from the domain of intended inference
11.2 Why sampling frames determine who can possibly enter the data
11.3 How coverage error differs from sampling variability
11.4 Why convenience data can be statistically precise and inferentially irrelevant
11.5 How transport beyond the observed sample requires additional assumptions
12. Why estimand, estimator, and estimate are fundamentally different objects
12.1 The estimand as the quantity the analysis seeks
12.2 The estimator as the rule that maps observations to a claim
12.3 The estimate as the realized output of that rule
12.4 Why many estimators can target the same estimand
12.5 Why the same estimator can behave differently under different source structures
12.6 How estimator performance depends on assumptions about the observation process
13. Why goodness of fit must not be mistaken for truth
13.1 How multiple incompatible structures can fit the same observations
13.2 Why perfect interpolation does not establish correct representation
13.3 How underdetermination arises from finite observations
13.4 Why model compatibility is weaker than source correspondence
13.5 How predictive adequacy and explanatory adequacy can diverge
14. Why dependence, association, mechanism, and intervention must remain distinct
14.1 How variables can move together without one generating the other
14.2 How conditional dependence differs from causal dependence
14.3 How common causes create predictive relationships without direct mechanisms
14.4 How reverse direction produces misleading interpretations
14.5 Why observing a variable differs from deliberately changing it
14.6 Why predictive importance does not imply intervention value
15. Why model-conditional uncertainty is not total uncertainty
15.1 Uncertainty about parameters inside a fixed model
15.2 Uncertainty about future observations inside a fixed distribution
15.3 Uncertainty about which model should be used
15.4 Uncertainty created by representation choice
15.5 Uncertainty created by missing mechanisms
15.6 Uncertainty created by future regime changes
15.7 Why conventional confidence statements usually cover only part of the epistemic problem
Part IV — How observation systems manufacture the data available to statistics
16. How different observation systems expose different fragments of reality
16.1 Sensors that transform physical states into recorded signals
16.2 Surveys that transform responses into coded variables
16.3 Administrative systems that record activity for institutional purposes
16.4 Transaction systems that observe only events passing through a platform
16.5 Human classification systems that embed judgment into labels
16.6 Automated logging systems that record what engineers chose to instrument
16.7 Experimental systems that deliberately create informative variation
17. How measurement quality constrains everything downstream
17.1 Defining measurement rules before analyzing their outputs
17.2 Distinguishing resolution, precision, accuracy, and reproducibility
17.3 How calibration links an instrument to an external reference
17.4 How instrument drift changes measurements over time
17.5 How systematic measurement error differs from random variation
17.6 Why measurement error can attenuate, distort, or manufacture relationships
18. How coverage determines which parts of the source can ever become evidence
18.1 Geographic regions that never enter the observation system
18.2 Time periods that are absent or poorly observed
18.3 Populations excluded by the sampling frame
18.4 Rare events absent because the observation period is too short
18.5 Extreme states absent because instruments saturate or systems fail
18.6 Why missing coverage cannot be repaired merely by increasing sample size
19. How selection determines which realizations become visible
19.1 Random selection and its inferential advantages
19.2 Convenience selection and hidden population distortions
19.3 Self-selection by observed units
19.4 Selection created by institutional procedures
19.5 Survivorship as conditioning on continued existence
19.6 Selection created by previous predictions or decisions
19.7 Why selection can reverse apparently stable relationships
20. How missingness can occur at several structurally different levels
20.1 Missing values within otherwise observed cases
20.2 Missing variables that were never measured
20.3 Missing units that never entered the observation system
20.4 Missing regions of the state space
20.5 Informative missingness caused by the hidden quantity itself
20.6 Censoring in which the target is only partially revealed
20.7 Truncation in which some cases cannot appear at all
20.8 When no statistical adjustment can recover the missing information
21. How incentives alter what gets reported, measured, and recorded
21.1 Reporting incentives that distort measured outcomes
21.2 Administrative targets that change classification behavior
21.3 Compensation systems that alter recorded performance
21.4 Strategic actors who respond to measurement rules
21.5 Goodhart effects when a measure becomes an objective
21.6 Adversarial behavior against scoring and screening systems
21.7 Why the observation mechanism may change after deployment
Part V — Constructing the represented statistical world
22. How observed distinctions become statistical variables
22.1 Numerical variables representing magnitude
22.2 Ordered variables representing rank without known distance
22.3 Categorical variables representing discrete distinctions
22.4 Binary variables representing deliberately collapsed alternatives
22.5 Derived variables created from other measurements
22.6 Proxy variables standing in for inaccessible targets
22.7 Latent variables representing unobserved structure
23. How units of analysis determine what counts as one observation
23.1 Individuals, organizations, events, and transactions as possible units
23.2 Repeated measurements nested within the same unit
23.3 Units nested within groups, institutions, or regions
23.4 Temporal units whose dependence changes with separation
23.5 Spatial units whose relationships depend on proximity
23.6 Network units whose dependence follows connections rather than distance
23.7 Why choosing the wrong unit can manufacture pseudo-replication
24. How representation choices determine what relationships can be expressed
24.1 Encoding categorical states into analyzable quantities
24.2 Scaling variables without changing their substantive meaning
24.3 Transforming skewed or multiplicative quantities
24.4 Constructing interactions between represented properties
24.5 Creating basis expansions for nonlinear relationships
24.6 Aggregating observations across time or groups
24.7 How representation can manufacture apparent simplicity
24.8 How representation can make real structure impossible to recover
25. How domains define where a statistical claim is intended to hold
25.1 The source domain that generates possible realizations
25.2 The observed domain that actually enters the data
25.3 The training domain used for constructing a model
25.4 The validation domain used for selecting among alternatives
25.5 The test domain used for final evaluation
25.6 The deployment domain where predictions will actually be used
25.7 The target population about which conclusions are intended
25.8 How domain mismatch becomes a generalization problem
Part VI — Probability as a formal calculus over represented uncertainty
26. Why probability is introduced only after representation and comparability exist
26.1 The need to represent uncertainty across possible outcomes
26.2 The need to combine information coherently across related observations
26.3 The difference between uncertainty in the source and uncertainty in the representation
26.4 Why probabilities operate on represented events rather than directly on reality
26.5 What probability adds that descriptive counting alone cannot provide
27. How events define the propositions to which probability can be assigned
27.1 Events as subsets of represented possibilities
27.2 Simple events and compound events
27.3 Mutually exclusive and overlapping events
27.4 Exhaustive collections of events
27.5 Event spaces generated by observable distinctions
27.6 Why changing representation changes the event space
28. How probability assignments obey a coherent mathematical structure
28.1 Nonnegative weighting of possible events
28.2 Normalization across the represented possibility space
28.3 Additivity across mutually exclusive events
28.4 Conditional probability after information becomes available
28.5 Competing interpretations of probability
28.6 Why probabilistic coherence does not guarantee empirical adequacy
29. How random variables connect represented quantities to probability
29.1 Random variables as mappings from possible outcomes to values
29.2 Discrete random variables and countable outcomes
29.3 Continuous random variables and numerical ranges
29.4 Vector-valued random variables representing multiple measurements
29.5 Indicator variables representing events numerically
29.6 Why the same source can support many different random-variable representations
30. How joint, marginal, and conditional distributions encode relationships
30.1 Joint distributions over several represented quantities
30.2 Marginalization by ignoring parts of the joint structure
30.3 Conditional distributions after observing additional information
30.4 Dependence as failure of probabilistic factorization
30.5 Independence as a strong structural simplification
30.6 Conditional independence as dependence mediated through other variables
31. How expectation summarizes probabilistic quantities
31.1 Expected value as probability-weighted averaging
31.2 Conditional expectation after information becomes available
31.3 Expectation of transformed random quantities
31.4 Linearity of expectation and why it matters
31.5 Variance as expected squared deviation
31.6 Covariance as expected joint deviation
31.7 Why expectation is the bridge from probability to many estimators and predictors
Part VII — Formalizing inferential relationships among observations
32. How exchangeability formalizes one important kind of comparability
32.1 Why order may sometimes carry no inferential information
32.2 Finite exchangeability across observations
32.3 Infinite exchangeability and mixture representations
32.4 Conditional exchangeability after accounting for known structure
32.5 Why exchangeability is weaker than assuming identical mechanisms
32.6 Where exchangeability fails in real systems
33. How independence simplifies inference and where it becomes unrealistic
33.1 Independent observations under repeated sampling
33.2 Conditional independence after controlling relevant information
33.3 Dependence induced by shared environments
33.4 Dependence created by repeated measurements
33.5 Dependence created by networks and spatial proximity
33.6 Consequences of pretending dependent observations are independent
34. How stationarity formalizes stability across time
34.1 Stable marginal behavior through time
34.2 Stable joint behavior under time translation
34.3 Local stationarity over restricted periods
34.4 Seasonal structure that violates simple stationarity
34.5 Structural breaks that destroy historical comparability
34.6 Why future prediction requires some form of temporal persistence
35. How hierarchical structure creates partial rather than complete comparability
35.1 Observations nested within shared groups
35.2 Group-specific variation around common structure
35.3 Partial pooling across related groups
35.4 Random effects as probabilistic representations of group heterogeneity
35.5 Why complete pooling and complete separation are limiting cases
35.6 How multilevel models encode graded inferential relationships
Part VIII — From repeated evidence to stable statistical quantities
36. Why empirical quantities can stabilize as comparable observations accumulate
36.1 Sample averages as functions of repeated observations
36.2 The law of large numbers as stabilization under assumptions
36.3 Weak and strong forms of convergence
36.4 Why convergence does not repair systematic observation bias
36.5 Why larger samples improve precision without necessarily improving validity
37. How concentration describes the probability of large deviations from typical behavior
37.1 Deviations of empirical averages from expectations
37.2 The role of boundedness and variance assumptions
37.3 Concentration inequalities as finite-sample guarantees
37.4 Why high-dimensional settings make concentration especially important
37.5 How dependence weakens familiar concentration results
38. How central-limit behavior creates approximate distributions for estimation error
38.1 Standardized sums of many contributions
38.2 Conditions supporting normal approximation
38.3 Why convergence in distribution differs from convergence in probability
38.4 How central-limit approximations produce standard errors
38.5 Where heavy tails, dependence, or small samples break the approximation
Part IX — Defining the estimation problem precisely
39. How population properties become explicit inferential targets
39.1 Means, variances, quantiles, and probabilities as distributional functionals
39.2 Regression functions as conditional population quantities
39.3 Risk measures as summaries of tail behavior
39.4 Latent quantities that require additional assumptions
39.5 Structural quantities that cannot be identified from association alone
40. How estimands specify exactly what is being asked of the data
40.1 Specifying a quantity before choosing an estimator
40.2 Distinguishing descriptive estimands from predictive estimands
40.3 Population-wide estimands versus subgroup-specific estimands
40.4 Conditional estimands indexed by observed information
40.5 Why changing the estimand changes the scientific question
41. Why identification must be solved before estimation
41.1 What it means for observations to distinguish competing parameter values
41.2 Global identification across the entire parameter space
41.3 Local identification near the true parameter value
41.4 Partial identification when data determine only a set of possibilities
41.5 Non-identification when observationally equivalent explanations remain
41.6 How additional assumptions can create identification
41.7 Why stronger assumptions can manufacture apparent certainty
42. How estimators map finite observations into claims about unknown quantities
42.1 Statistics as functions of the observed data
42.2 Estimators as rules defined before observing the realized sample
42.3 Estimates as sample-specific numerical outputs
42.4 Deterministic and randomized estimators
42.5 Plug-in estimation using estimated distributions
42.6 Why estimator choice encodes assumptions and priorities
Part X — Constructing and comparing estimators
43. How empirical averages generate the simplest estimators
43.1 Estimating means from sample averages
43.2 Estimating probabilities from empirical frequencies
43.3 Estimating moments from empirical moments
43.4 Estimating conditional quantities through stratification
43.5 Why empirical estimation becomes difficult in high dimensions
44. How estimating equations generalize moment-based estimation
44.1 Matching theoretical and empirical moments
44.2 Solving estimating equations for unknown parameters
44.3 Overidentified systems with more equations than parameters
44.4 Generalized method-of-moments reasoning
44.5 Sensitivity to poorly chosen moment conditions
45. How likelihood converts a probabilistic model into an estimation criterion
45.1 Likelihood as compatibility of parameters with observed data
45.2 Log-likelihood as an additive optimization objective
45.3 Maximum likelihood estimation
45.4 Score functions and information
45.5 Likelihood under model misspecification
45.6 Why maximum likelihood optimizes within the assumed family rather than proving it correct
46. How Bayesian estimation combines a prior representation with observed evidence
46.1 Prior distributions over uncertain quantities
46.2 Likelihood contributions from observed data
46.3 Posterior distributions after conditioning
46.4 Posterior expectations, medians, and modes as estimators
46.5 Posterior predictive distributions
46.6 Sensitivity to prior and likelihood assumptions
46.7 Why Bayesian coherence does not eliminate structural uncertainty
47. How estimator quality is judged under repeated hypothetical samples
47.1 Bias as systematic displacement from the estimand
47.2 Variance as sensitivity to the realized sample
47.3 Mean squared error combining bias and variance
47.4 Consistency as convergence toward the estimand
47.5 Efficiency as relative use of available information
47.6 Robustness against contamination and model deviation
47.7 Stability under small perturbations of the observed data
Part XI — Quantifying uncertainty about estimates
48. How sampling distributions describe estimator variability
48.1 Repeated-sample thought experiments
48.2 Exact sampling distributions when derivable
48.3 Approximate sampling distributions when exact forms are unavailable
48.4 Standard errors as scales of estimator variability
48.5 Why sampling uncertainty is conditional on the assumed data-generating structure
49. How confidence intervals convert estimator variability into coverage statements
49.1 Constructing intervals from pivotal quantities
49.2 Interpreting repeated-sample coverage correctly
49.3 Why confidence level is not posterior probability
49.4 Approximate intervals based on asymptotic normality
49.5 Bootstrap confidence intervals
49.6 How model misspecification destroys nominal coverage
50. How Bayesian credible intervals answer a different uncertainty question
50.1 Posterior probability assigned to parameter regions
50.2 Equal-tail credible intervals
50.3 Highest-posterior-density regions
50.4 Prior sensitivity of credible regions
50.5 Why credible intervals and confidence intervals can coincide numerically but differ conceptually
51. How hypothesis tests turn claims into procedures for detecting disagreement with data
51.1 Null hypotheses as restricted statistical worlds
51.2 Alternative hypotheses representing departures of interest
51.3 Test statistics designed to expose those departures
51.4 Reference distributions under the null
51.5 p-values as tail probabilities under an assumed reference model
51.6 Type I and Type II errors
51.7 Power as the probability of detecting specified departures
51.8 Why statistical significance does not establish substantive importance
52. How repeated searching manufactures apparently strong evidence
52.1 Multiple testing across many hypotheses
52.2 Data-dependent hypothesis generation
52.3 Selective reporting of favorable results
52.4 Model search followed by naive inference
52.5 Winner's curse among selected effects
52.6 False discovery control
52.7 Selective inference after data-dependent selection
Part XII — Defining the prediction problem
53. How prediction differs from estimating a population property
53.1 Predicting an unobserved value for a particular case
53.2 Predicting a probability distribution rather than a single number
53.3 Predicting class membership
53.4 Predicting relative rank
53.5 Predicting event timing
53.6 Predicting sequences and trajectories
53.7 Why estimation and prediction can prefer different models
54. How information available at prediction time constrains legitimate predictors
54.1 Predictors known before the target occurs
54.2 Predictors observed only after the target occurs
54.3 Leakage through future or post-outcome information
54.4 Prediction from partial histories
54.5 Real-time prediction under delayed measurements
54.6 Why deployment-time information must match evaluation-time information
55. How loss functions translate prediction errors into consequences
55.1 Squared loss emphasizing large numerical errors
55.2 Absolute loss emphasizing typical absolute deviation
55.3 Classification loss penalizing incorrect categories
55.4 Logarithmic loss evaluating full probability assignments
55.5 Asymmetric loss for unequal consequences
55.6 Cost-sensitive prediction under heterogeneous stakes
55.7 Why the optimal predictor changes when the loss changes
56. How optimal prediction follows from the conditional distribution under stated assumptions
56.1 Conditional expectation under squared-error loss
56.2 Conditional median under absolute-error loss
56.3 Conditional mode under simple classification loss
56.4 Probability thresholds under asymmetric classification costs
56.5 Bayes risk as minimum expected loss within the assumed representation
56.6 Irreducible uncertainty remaining even under the optimal predictor
56.7 Why Bayes-optimal does not mean optimal outside the assumed probabilistic world
Part XIII — Constructing predictive functions from finite observations
57. How candidate function classes restrict which predictive relationships can be learned
57.1 Linear function classes
57.2 Additive function classes
57.3 Local function classes
57.4 Partition-based function classes
57.5 Kernel-based function classes
57.6 Compositional neural function classes
57.7 Why every function class encodes representational exclusions
58. How empirical risk minimization converts prediction into an optimization problem
58.1 Empirical loss calculated over observed training cases
58.2 Selecting a function that minimizes observed loss
58.3 Population risk that remains unobserved
58.4 Optimization error versus statistical error
58.5 Local minima, saddle points, and numerical approximation
58.6 Why successful optimization does not imply successful generalization
59. How approximation, estimation, and optimization errors arise from different sources
59.1 Approximation error from an inadequate function class
59.2 Estimation error from finite observations
59.3 Optimization error from incomplete numerical search
59.4 Representation error from missing predictive information
59.5 Observation error inherited from the data-production system
59.6 Why reducing one error source can expose another
60. How regularization deliberately restricts fitting to improve future performance
60.1 Complexity penalties added to empirical objectives
60.2 Ridge penalties shrinking parameter magnitudes
60.3 Lasso penalties encouraging sparse solutions
60.4 Elastic-net penalties combining shrinkage and sparsity
60.5 Early stopping as optimization-dependent regularization
60.6 Architectural restrictions as implicit regularization
60.7 Why regularization trades flexibility for stability rather than creating truth
Part XIV — Evaluating whether predictive performance generalizes
61. Why training performance cannot estimate future performance without correction
61.1 Optimism created by evaluating on fitted observations
61.2 Training error as a biased estimate of future error
61.3 Validation data for model and hyperparameter selection
61.4 Test data reserved for final evaluation
61.5 How repeated test-set use converts test data into training information
62. How resampling estimates performance when observations are limited
62.1 Simple holdout evaluation
62.2 K-fold cross-validation
62.3 Leave-one-out cross-validation
62.4 Repeated cross-validation
62.5 Bootstrap estimation of prediction error
62.6 Nested cross-validation for model-selection pipelines
62.7 Why resampling still depends on comparability between folds and deployment
63. How evaluation design must reproduce the structure of deployment
63.1 Random splitting for exchangeable observations
63.2 Grouped splitting for clustered observations
63.3 Temporal splitting for future prediction
63.4 Spatial splitting for geographic generalization
63.5 Leave-domain-out validation for transport problems
63.6 External validation on independently collected data
63.7 Why the wrong split can manufacture convincing but useless performance
64. How predictive metrics answer different operational questions
64.1 Mean squared error for numerical prediction
64.2 Mean absolute error for robust numerical prediction
64.3 Accuracy for symmetric categorical decisions
64.4 Precision and recall under imbalanced outcomes
64.5 Sensitivity and specificity under diagnostic framing
64.6 ROC curves and threshold-independent ranking behavior
64.7 Precision-recall curves for rare outcomes
64.8 Ranking metrics when ordering matters more than classification
65. How calibration evaluates whether predicted probabilities mean what they claim
65.1 Marginal calibration across all predictions
65.2 Calibration curves across probability ranges
65.3 Conditional calibration within relevant subgroups
65.4 Calibration under changing prevalence
65.5 Calibration drift after deployment
65.6 Why discrimination can remain strong while probability estimates become wrong
66. How aggregate evaluation can conceal catastrophic local failure
66.1 Average performance hiding subgroup errors
66.2 Tail performance hidden by mean metrics
66.3 Rare catastrophic events overwhelmed by common easy cases
66.4 Unequal consequences across populations
66.5 Performance degradation concentrated near regime boundaries
66.6 Why operational evaluation often needs a failure distribution rather than one score
Part XV — Representation, dimension, and latent structure
67. How high dimensionality changes the geometry of statistical estimation
67.1 Growth of possible configurations with additional variables
67.2 Sparsity of observations in high-dimensional spaces
67.3 Distance concentration and failure of intuitive neighborhood concepts
67.4 Parameter proliferation relative to available observations
67.5 Why structural assumptions become more important as dimension increases
68. How dimension reduction trades retained information for simpler representation
68.1 Principal components as directions of maximal variation
68.2 Low-rank approximations of high-dimensional observations
68.3 Factor models representing shared variation
68.4 Sparse representations preserving selected dimensions
68.5 Nonlinear embeddings
68.6 Why useful compression is not evidence that compressed dimensions are ontologically real
69. How latent-variable models posit hidden structure to explain observed relationships
69.1 Hidden common factors generating observable dependence
69.2 Mixture components representing unobserved subpopulations
69.3 Hidden states evolving through time
69.4 Latent classes explaining categorical heterogeneity
69.5 Identification problems in latent-variable models
69.6 Why multiple latent structures can explain the same observations
Part XVI — Major predictive model families as implementations of earlier principles
70. How linear and generalized linear models impose global parametric structure
70.1 Ordinary linear regression for conditional means
70.2 Multiple regression with several represented predictors
70.3 Interaction terms representing conditional effects
70.4 Polynomial and basis expansions representing nonlinear trends
70.5 Logistic regression for conditional class probabilities
70.6 Generalized linear models for non-Gaussian responses
70.7 Diagnostics for systematic residual structure
71. How local prediction methods infer from nearby observations
71.1 Nearest-neighbor prediction
71.2 Kernel-weighted local averaging
71.3 Local polynomial regression
71.4 Bandwidth choice controlling neighborhood size
71.5 Failure of locality in high-dimensional spaces
71.6 Why local similarity depends entirely on representation
72. How partitioning methods divide represented space into predictive regions
72.1 Recursive binary partitioning
72.2 Split criteria based on predictive improvement
72.3 Tree depth as model complexity
72.4 Pruning to reduce unstable partitions
72.5 Interpretability of tree-based decision rules
72.6 Instability of individual trees under small data changes
73. How ensemble methods combine unstable predictors into more reliable systems
73.1 Bagging through repeated resampling and averaging
73.2 Random forests through feature-randomized trees
73.3 Boosting through sequential correction of predictive residuals
73.4 Gradient boosting as functional optimization
73.5 Diversity among component predictors
73.6 Why ensemble accuracy can reduce interpretability
74. How margin and kernel methods construct separating boundaries
74.1 Linear separating hyperplanes
74.2 Maximum-margin classification
74.3 Soft margins allowing classification violations
74.4 Kernel functions representing implicit feature transformations
74.5 Support vectors determining the separating boundary
74.6 Sensitivity to kernel and regularization choices
75. How neural predictors construct highly flexible compositional functions
75.1 Layered function composition
75.2 Nonlinear activation functions
75.3 Gradient-based parameter fitting
75.4 Representation learning inside predictive optimization
75.5 Overparameterization beyond classical parameter-count intuition
75.6 Explicit and implicit regularization
75.7 Scaling behavior with data, parameters, and computation
75.8 Why high predictive capacity does not eliminate observation or target failure
Part XVII — Estimation under incomplete and structured observation
76. How missing-data methods depend on assumptions about why values are missing
76.1 Missing completely at random
76.2 Missing at random conditional on observed information
76.3 Missing not at random
76.4 Complete-case analysis
76.5 Weighting adjustments
76.6 Multiple imputation
76.7 Sensitivity analysis for unverifiable missingness assumptions
77. How censoring and truncation alter what can be learned about outcomes
77.1 Right censoring of incomplete event times
77.2 Left censoring below detection thresholds
77.3 Interval censoring between observation occasions
77.4 Truncation excluding entire cases from observation
77.5 Survival functions and hazard functions
77.6 Kaplan-Meier estimation under censoring assumptions
77.7 Regression with censored outcomes
78. How complex sampling designs alter estimation and uncertainty
78.1 Unequal-probability sampling
78.2 Stratified sampling
78.3 Cluster sampling
78.4 Multistage sampling
78.5 Survey weights correcting known design probabilities
78.6 Design effects on estimator variance
78.7 Why naive IID analysis fails under complex sampling
Part XVIII — Prediction under time, transport, and changing regimes
79. How time ordering changes the statistical problem
79.1 Past information versus future outcomes
79.2 Serial dependence between nearby observations
79.3 Trend representing persistent directional movement
79.4 Seasonality representing recurring temporal structure
79.5 Forecast horizons and expanding uncertainty
79.6 Why random train-test splitting fails for many forecasting problems
80. How state-space models represent hidden systems evolving through time
80.1 Latent states generating observable measurements
80.2 Transition models describing state evolution
80.3 Observation models describing noisy measurement
80.4 Filtering current hidden states from past observations
80.5 Smoothing past hidden states using later information
80.6 Forecasting future states from estimated dynamics
81. How transport asks whether a relationship survives outside the observed domain
81.1 Transport across populations
81.2 Transport across institutions
81.3 Transport across geographic locations
81.4 Transport across historical periods
81.5 Transport across operating conditions
81.6 Why internal validation cannot establish external validity
82. How distribution shift breaks relationships learned from historical data
82.1 Covariate shift in observed predictor distributions
82.2 Label shift in outcome prevalence
82.3 Conditional shift in relationships between predictors and outcomes
82.4 Concept drift in the predictive mapping itself
82.5 Structural breaks caused by new mechanisms
82.6 Regime change invalidating historical comparability
83. How deployed predictions alter the system that generated the training data
83.1 Predictions changing human decisions
83.2 Decisions changing which outcomes become observable
83.3 Selection effects created by model deployment
83.4 Strategic adaptation by people subject to prediction
83.5 Self-fulfilling predictions
83.6 Self-defeating predictions
83.7 Feedback loops that generate a new data distribution
Part XIX — Distribution-free and weak-assumption predictive guarantees
84. How conformal prediction constructs finite-sample prediction sets from exchangeability
84.1 Nonconformity scores measuring unusual candidate outcomes
84.2 Calibration observations used to construct prediction thresholds
84.3 Marginal coverage under exchangeability
84.4 Split conformal prediction
84.5 Full conformal prediction
84.6 Regression prediction intervals
84.7 Classification prediction sets
84.8 Why marginal coverage does not imply conditional validity
85. How robust prediction weakens dependence on exact parametric assumptions
85.1 Robust losses reducing sensitivity to extreme observations
85.2 Heavy-tailed outcome distributions
85.3 Contamination models
85.4 Distributionally robust optimization
85.5 Worst-case performance over uncertainty sets
85.6 The tradeoff between robustness and efficiency
Part XX — Structure construction without a supplied prediction target
86. What unsupervised procedures actually optimize when no response variable is supplied
86.1 Similarity defined by a chosen representation
86.2 Distance defined by a chosen geometry
86.3 Density defined relative to a chosen coordinate system
86.4 Compression defined by a reconstruction criterion
86.5 Clusters defined by an optimization objective
86.6 Why the algorithm always supplies structure even when the source does not
87. How clustering partitions observations according to imposed similarity criteria
87.1 K-means clustering around representative centers
87.2 Hierarchical clustering through recursive merging or splitting
87.3 Density-based clustering around high-density regions
87.4 Model-based clustering through mixture distributions
87.5 Sensitivity of clusters to scaling and distance
87.6 Stability of clusters across samples and perturbations
88. How discovered structure can be tested against external evidence
88.1 Stability under repeated samples
88.2 Reproducibility under alternative representations
88.3 Predictive consequences on variables not used for construction
88.4 Agreement with independently measured structure
88.5 Experimental discrimination where intervention is possible
88.6 Why stable, interpretable clusters still need not represent real kinds
Part XXI — Detecting when a statistical model has stopped deserving trust
89. How residual behavior can reveal model inadequacy
89.1 Systematic residual patterns
89.2 Heterogeneous residual variance
89.3 Temporal residual dependence
89.4 Spatial residual dependence
89.5 Extreme residuals and unrepresented mechanisms
89.6 Why residual conformity cannot prove the model is correct
90. How calibration decay can expose changing predictive relationships
90.1 Drift in predicted probability reliability
90.2 Subgroup-specific calibration failure
90.3 Calibration failure concentrated in new regions of state space
90.4 Distinguishing random fluctuation from persistent deterioration
90.5 When recalibration is sufficient
90.6 When recalibration merely hides deeper structural failure
91. How novelty detection can expose observations outside historical experience
91.1 Distance from previously observed regions
91.2 Density-based novelty scores
91.3 Representation-dependent anomaly detection
91.4 New combinations of familiar variables
91.5 New variables or mechanisms absent from the original ontology
91.6 Why novelty scores cannot detect what the representation cannot express
92. How to distinguish parameter drift from representation failure
92.1 Same structure with changing parameter values
92.2 Same variables with changing conditional relationships
92.3 New hidden mechanism producing familiar observations
92.4 Missing variable becoming operationally important
92.5 Observation-system failure mimicking distribution shift
92.6 Criteria for updating, retraining, replacing, or abandoning the model
Part XXII — The jurisdictional boundary of statistical estimation and prediction
93. Why prediction does not by itself explain the mechanism producing an outcome
93.1 Predictive association without causal mechanism
93.2 Mechanistic explanation requiring additional structural claims
93.3 Equivalent predictive models with incompatible explanations
93.4 Why explanatory adequacy requires evidence beyond test-set accuracy
94. Why observing a predictor does not establish the effect of changing it
94.1 Conditioning on X versus intervening on X
94.2 Confounding by common causes
94.3 Selection effects created by treatment assignment
94.4 Counterfactual outcomes that are never jointly observable
94.5 Why predictive importance can be useless for intervention design
95. How experiments create information that passive observation may not contain
95.1 Deliberate perturbation of candidate causal variables
95.2 Randomization breaking systematic treatment selection
95.3 Control groups providing counterfactual reference distributions
95.4 Factorial experiments discriminating interactions
95.5 Sequential experimentation adapting to accumulated evidence
95.6 Why experiment design belongs upstream of estimation even when statistics analyzes the result
96. Why prediction does not specify which action should be taken
96.1 Predictions describing consequences without expressing preferences
96.2 Utility and cost functions ranking possible outcomes
96.3 Constraints limiting feasible actions
96.4 Decision rules mapping predictions to actions
96.5 Why greater predictive accuracy can still produce worse decisions
96.6 How decision theory begins where pure prediction ends
97. Why prediction differs from controlling a dynamic system
97.1 Forecasting future state under existing inputs
97.2 Selecting inputs to alter future state
97.3 Feedback from current outcomes into future actions
97.4 Controllability as distinct from predictability
97.5 Stable control despite imperfect prediction
97.6 Excellent prediction despite little ability to control
98. Why statistical analysis differs from engineering the successor system
98.1 Estimating properties of systems that already exist
98.2 Generating candidate systems that do not yet exist
98.3 Choosing design variables rather than merely observing predictors
98.4 Using experiments to discriminate among candidate designs
98.5 Optimization under physical and operational constraints
98.6 Statistics as evidence inside engineering rather than a substitute for engineering
Part XXIII — Failure modes arranged by where they enter the inferential chain
99. Source failures that occur before observation begins
99.1 Relevant mechanisms absent from the conceptual ontology
99.2 Relevant populations excluded from the imagined domain
99.3 New mechanisms outside historical experience
99.4 Rare regimes never previously realized
99.5 Structural changes that make old distinctions obsolete
100. Observation failures that corrupt what becomes available as evidence
100.1 Sensor failure and measurement drift
100.2 Coding and transcription errors
100.3 Reporting distortion
100.4 Selection into observation
100.5 Informative missingness
100.6 Administrative incentives changing measured behavior
101. Representation failures that corrupt the statistical world built from observations
101.1 Wrong variables representing the wrong distinctions
101.2 Wrong categories collapsing substantively different states
101.3 Wrong aggregation hiding local structure
101.4 Proxy substitution replacing the target with convenience
101.5 Hidden-state collapse
101.6 Representations incapable of expressing the actual mechanism
102. Comparability failures that make observed cases poor evidence for one another
102.1 Hidden population heterogeneity
102.2 Changing institutional environment
102.3 Temporal regime change
102.4 Geographic heterogeneity
102.5 Network interference between supposedly separate units
102.6 Deployment-induced changes in behavior
103. Estimation failures occurring after the problem has been correctly represented
103.1 Weak or absent identification
103.2 Excessive sampling variability
103.3 Systematic estimator bias
103.4 Numerical instability
103.5 Poor optimization
103.6 Selection-induced optimism
104. Prediction failures that appear only when the model encounters new cases
104.1 Overfitting to sample-specific variation
104.2 Leakage from unavailable future information
104.3 Calibration failure
104.4 Tail failure
104.5 Subgroup-specific failure
104.6 Distribution shift
104.7 New regimes outside the training experience
105. Evaluation failures that make weak models appear successful
105.1 Wrong performance metric
105.2 Wrong test population
105.3 Random splitting when temporal splitting is required
105.4 Repeated benchmark reuse
105.5 Hidden test contamination
105.6 Aggregate performance concealing catastrophic local failure
105.7 Statistical success concealing operational failure
Part XXIV — The irreducible limits on statistical claims
106. What statistical estimation can legitimately establish under explicit conditions
106.1 Claims conditional on the observation mechanism
106.2 Claims conditional on the represented variables
106.3 Claims conditional on inferential comparability
106.4 Claims conditional on identification assumptions
106.5 Claims conditional on the probability model or resampling structure
106.6 Claims conditional on stability across the intended domain
107. What statistical estimation cannot manufacture from absent information
107.1 Variables that were never observed
107.2 Mechanisms excluded from the representation
107.3 Populations excluded from the observation system
107.4 Counterfactual outcomes unsupported by causal structure
107.5 Future regimes absent from historical experience
107.6 Correct ontology merely from greater computational power
108. The non-collapses that define the field's jurisdiction
108.1 Source system != observed trace
108.2 Observation != recorded representation
108.3 Observed sample != target population
108.4 Estimand != estimator != realized estimate
108.5 Good fit != correct source description
108.6 Association != mechanism
108.7 Predictive importance != causal importance
108.8 Prediction != intervention
108.9 Prediction != decision
108.10 Prediction != control
108.11 Prediction != engineering design
108.12 Training performance != deployment performance
108.13 Model-conditional uncertainty != total epistemic uncertainty
108.14 Larger dataset != better observability
108.15 More flexible model != better representation
108.16 Better prediction != better target
108.17 Historical validity != future validity
108.18 Statistical precision != structural correctness
Part XXV — The complete admissibility test for an estimation or prediction claim
109. Questions that must be answered about the source before analyzing the dataset
109.1 What system generated the phenomena being studied?
109.2 Which possible states of that system are relevant to the question?
109.3 Which relevant states could the observation system actually detect?
109.4 Which mechanisms could remain invisible despite extensive data collection?
109.5 What evidence exists that the source remained sufficiently stable during observation?
110. Questions that must be answered about how observations became data
110.1 What observation mechanism produced each recorded variable?
110.2 Which populations, locations, and periods were covered?
110.3 Which cases could never enter the dataset?
110.4 What measurement errors or transformations occurred before recording?
110.5 What incentives affected measurement, reporting, or classification?
110.6 Which missingness mechanisms remain plausible?
111. Questions that must be answered before treating observations as comparable evidence
111.1 Why should one observed case inform another case?
111.2 Which aspects of the generating conditions are assumed stable?
111.3 Which forms of heterogeneity are explicitly modeled?
111.4 Which dependencies violate simple independent-sampling assumptions?
111.5 Across which populations, times, locations, or regimes is comparison intended?
111.6 What changes would make that comparison invalid?
112. Questions that must be answered before choosing an estimator or predictor
112.1 What quantity is actually being estimated or predicted?
112.2 Is the recorded label the true target or merely a proxy?
112.3 Is the target identifiable from the available observations?
112.4 Which assumptions create that identification?
112.5 What information does the proposed representation discard?
112.6 What loss or consequence determines what counts as predictive success?
113. Questions that must be answered before accepting reported uncertainty
113.1 Which uncertainty sources are included in the calculation?
113.2 Which uncertainty sources are conditioned away by the model?
113.3 How sensitive are results to alternative plausible representations?
113.4 How sensitive are results to alternative plausible models?
113.5 How sensitive are results to observation or missingness assumptions?
113.6 Does the reported interval describe parameter, predictive, model, or total uncertainty?
114. Questions that must be answered before claiming generalization
114.1 Why should the fitted relationship persist beyond the observed sample?
114.2 What evidence supports comparability between training and deployment domains?
114.3 Has evaluation reproduced the temporal, spatial, hierarchical, or institutional structure of deployment?
114.4 Has the model been tested on genuinely external observations?
114.5 Which subgroups or state-space regions remain poorly supported?
114.6 Which future changes would invalidate the generalization argument?
115. Questions that must be answered before continuing to use a deployed model
115.1 What observations would contradict the model's expected behavior?
115.2 Which residual patterns indicate parameter drift rather than random variation?
115.3 Which observations suggest a missing variable or new mechanism?
115.4 When is recalibration sufficient?
115.5 When is retraining required?
115.6 When must the representation itself be replaced?
115.7 When must the observation system be redesigned?
115.8 Under what conditions should the model be abandoned rather than repaired?
116. Questions that determine whether the problem has left the jurisdiction of prediction
116.1 Is the user asking what is likely to happen?
116.2 Is the user asking why it happens?
116.3 Is the user asking what would happen under intervention?
116.4 Is the user asking which experiment would discriminate competing explanations?
116.5 Is the user asking which action should be chosen?
116.6 Is the user asking how to control a dynamic system?
116.7 Is the user asking how to design a new system?
116.8 Is a statistical result being promoted into a claim that requires a different discipline?
Comments
Post a Comment