Validation ledgers#
Validation is a release requirement, not a retrospective exercise. These records connect each public implementation to its hypotheses, primary-paper formula, calibration, literal or trusted oracle, numerical edge cases, invariants, and known differences from SHT 0.1.9.
Classical formulas, high-dimensional trace and diagonal statistics, Behrens–Fisher approximations, randomization, Bayes factors, multi-group procedures, and the validation-blocked sparse maximum-test audit.
Alternative-tail coverage, robust scale handling, degenerate samples, and independent SciPy comparisons.
Null-covariance whitening, literal U-statistics, two-sided projection tests, multi-group formulas, and conditional-regression Bayes factors.
Joint normal-population likelihood ratios, corrected rejection tails, component combinations, and stable exact quadrature.
Nonidentity-null whitening, high-dimensional centering, classical LRT checks, and unequal-sample trace estimators.
Exact enumeration, corrected Monte Carlo inference, floating-point ties, scale stability, and a literal distance-statistic oracle.
Shapiro approximations, moment formulas, finite-sample null-size audits, and Monte Carlo defaults.
Interpoint moments, quantile transforms, support boundaries, and asymptotic versus Monte Carlo calibration.
Dirichlet likelihoods, symmetric and general optimization, strict boundaries, and Wilks-regime checks.
Multivariate-mean method ledgers#
Cai–Liu–Xia — validation-blocked; not public
What the evidence labels mean#
Evidence |
Question answered |
|---|---|
Formula ledger |
Does the code evaluate the intended finite-sample quantity? |
Hand fixture |
Can a reader reproduce at least one result independently? |
Trusted oracle |
Does a separate established implementation agree where contracts overlap? |
Literal reference |
Does an intentionally simple implementation reproduce the optimized path? |
Invariance check |
Does the result respect transformations implied by the mathematics? |
Numerical stress test |
Does finite precision preserve inferential ordering and physical units? |
Null simulation |
Is an approximation calibrated in the regime where it is advertised? |
No single row is sufficient by itself. The appropriate combination depends on the method and its calibration. A public asymptotic option remains labeled as an approximation when finite-sample simulation does not support a stronger claim.
The release-simulation runner records the maintained scenario and random-stream contracts for the null and targeted alternative audits, and explains why historical Fisher rows are excluded from the 0.1.0 release evidence.
Legacy corrections#
The legacy audit records confirmed defects and how pySHT addresses them. Important corrections include null-covariance whitening, literal scale-equivariant covariance U-statistics, genuine two-sided Wu–Li tests, corrected likelihood-ratio tails, fixed CPH group indexing, defined AJB/RJB defaults with Monte Carlo uncertainty, fixed auxiliary randomness during permutation, log-domain exact and Bayesian calculations, and removal of the invalid distribution-equality asymptotic branch.