{"publication":{"abstract":"Discrete information-theoretic measures summarize uncertainty and dependence without requiring a prespecified functional form, but their plugin estimators have null distributions that depend on sample size and contingency-table structure. This complicates inference across heterogeneous analyses and dataset and can make repeated permutation costly.\n\nWe present a likelihood-ratio framework for the discrete multinomial plugin setting. When the reported information quantity is the Kullback–Leibler divergence (KL) between the empirical distribution and the maximum-likelihood fit under a prespecified nested null, the exact identity $G^2=2N\\ln(2)D_{\\mathrm{KL},2}$ converts information in bits to the likelihood-ratio deviance. Wilks' theorem then supplies an asymptotic chi-squared reference distribution, with degrees of freedom determined by the model comparison and prespecified support.\n\nThe framework applies directly to mutual information, conditional mutual information, fixed-reference KL goodness of fit, entropy deficit, and conditional entropy deficit. It yields distinct outputs for distinct purposes: a first-order-corrected information estimate in bits for effect magnitude, and a likelihood-ratio $p$-value. It also gives the leading plugin correction a dual interpretation as estimator-bias correction and asymptotic null-mean centering.\n\nWe use simulations to evaluate mutual information (MI) and conditional mutual information (CMI) null moments, distributional calibration, sparse-table breakdown, and agreement with permutation across changes in sample size and conditioning cardinality. Sparse expected cell counts, rather than low degrees of freedom, were the principal finite-sample failure mode: The normal-equivalent representation handles small degrees of freedom, but it cannot repair a chi-squared approximation limited by sparsity. The particular count at which this transition appeared is specific to the simulation grids and is not proposed as a universal rule. The component identities and asymptotic results are classical; the contribution is their synthesis and operationalization for large, heterogeneous collections of discrete information-theoretic tests.","body":"# Introduction\n\nInformation-theoretic measures are attractive because they quantify uncertainty and statistical dependence without requiring a prespecified relationship between those variables (e.g., linear, additive, or monotone). For a discrete variable $Y$, entropy $\\mathrm{H}(Y)$ measures uncertainty about its state. Mutual information $\\mathrm{I}(X;Y)$ measures the reduction in uncertainty about one variable supplied by observing the other, and conditional mutual information $\\mathrm{I}(X;Y\\mid Z)$ measures the reduction in uncertainty provided by the set. Each of these can be formulated as a Kullback–Leibler ($D_{\\mathrm{KL}}$) divergence, which quantifies the difference between two distributions. These are measures of information (conventionally bits when logarithms are base two) and are comparable [](https://doi.org/10.1002/j.1538-7305.1948.tb01338.x) [](https://doi.org/10.1214/aoms/1177729694) [](https://doi.org/10.1002/047174882X).\n\nOne benefit of these values is the quantification of statistical dependence when the form of that dependence is unknown. In human genetics, for example, genotype–phenotype relationships may be additive, dominant, recessive, multiallelic, nonmonotone, or heterogeneous across sex, ancestry, or environment. A conventional single-variant genome-wide association test is highly efficient when its additive model is correct, but arbitrary categorical dependence requires additional model terms (often prohibitive at typical study scales) or an omnibus test [](https://doi.org/10.1038/nrg1916) [](https://doi.org/10.1038/nrg2579) [](https://doi.org/10.1038/s43586-021-00056-9) [](https://doi.org/10.1086/519795). Mutual information (MI) and conditional mutual information (CMI) provide such omnibus summaries. They can therefore complement regression-based association analyses by testing for dependence without requiring the analyst to specify its form in advance.\n\nA second application is repeated conditional-independence testing. Constraint-based causal-discovery, Bayesian-network learning, and ordered-trajectory methods repeatedly evaluate hypotheses of the form\n\n$$\n\nX \\perp Y \\mid Z,\n\n$$\n\nwhile the size and state space of $Z$ change across steps [](https://doi.org/10.7551/mitpress/1754.001.0001) [](https://jmlr.org/papers/v7/decampos06a.html) [](https://jmlr.org/papers/v22/19-600.html) [](https://proceedings.mlr.press/v124/runge20a.html). CMI is a standard statistic for categorical conditional-independence testing [](https://jmlr.org/papers/v7/decampos06a.html) [](https://jmlr.org/papers/v22/19-600.html). Local permutation-based testing (computationally expensive) or a fixed significance threshold across the changing state space (providing differing false positive rates across strata) are widely used [](https://proceedings.mlr.press/v84/runge18a.html).\n\nIn both cases, the practical obstacle to widespread use of information-theoretic measures isn’t their computation. It’s calculating well-calibrated (and thus comparable) significance for these estimates. In a genome-wide scan, loci can differ in sample size, number of genetic states at a given locus, missingness, and conditioning structure [](https://doi.org/10.1038/nrg1916) [](https://doi.org/10.1038/s43586-021-00056-9) [](https://doi.org/10.1086/519795). Each of these differences causes different significance thresholds for a locus–disease association. In causal discovery, the number of conditioning configurations can grow from one test to the next, again resulting in different significance thresholds as the trajectories progress. The paradigm presented here provides calibrated test and ranking statistics for any information-theoretic value that can be formulated as a Kullback–Leibler divergence (KL), including MI, CMI, entropy, and conditional entropy.\n\n## The problem\n\nFor discrete plugin estimators, the null mean and dispersion depend on both sample size $N$ and the dimensionality ($C$) of the relevant contingency-table model. The plugin MI of independent variables is upwardly biased; plugin entropy is negatively biased; and the null mean and variance of CMI grow as conditioning introduces additional strata. These effects, including mean corrections, are established in the finite-sample information-estimation literature [](https://api.semanticscholar.org/CorpusID:125662170) [](https://doi.org/10.1137/1104033) [](https://doi.org/10.1162/neco.1995.7.2.399) [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272). These biases have impacts on the utility of information-theoretic values. A fixed threshold on a raw information value can correspond to different false-positive probabilities in tables with different dimensions, and mean correction alone does not fix this problem.\n\nPermutation, exact testing, and other resampling procedures provide general empirical calibration strategies, and they remain important when the estimator or sampling regime lies outside ordinary likelihood-ratio asymptotics [](https://doi.org/10.3390/e16052839) [](https://proceedings.mlr.press/v84/runge18a.html) [](https://proceedings.mlr.press/v89/marx19a.html). Monte Carlo permutation $p$-values also require careful finite-resolution calculation [](https://doi.org/10.2202/1544-6115.1585). These approaches are costly and can be prohibitive for large-scale data, such as current biobank datasets (often 500k samples and 1M parameters). Moreover, different information-theoretic measures are often presented with separate bias corrections, normalizations, or significance procedures, obscuring their shared statistical structure. The resulting methodological landscape can make information-theoretic inference appear more fragmented than it is.\n\n## Our approach: A synthesis of classical results\n\nWe therefore sought to establish an easily calculated, calibrated test statistic and evidence score that could be applied to MI, CMI, and other information theoretic values to enable genome-wide analyses with well-calibrated ranking statistics and significance thresholds.\n\nThe paper distinguishes between two inferential objects ([Table 1](#inferential-objects)).\n\n::::::figure{#inferential-objects type=\"table\" label=\"Table 1\"}\n| **Quantity**         | **Scientific role**                                       | **Typical output**                                                   |\n| -------------------- | --------------------------------------------------------- | -------------------------------------------------------------------- |\n| Information estimate | Magnitude of departure from a reference model             | MI, CMI, entropy deficit, or KL divergence in bits                   |\n| Calibrated evidence  | Compatibility of the observed departure with a null model | Likelihood-ratio $p$-value or normal-equivalent evidence ($z$) score |\n\n:::::figcaption\n**Table 1.** **Information theoretic effect size vs. calibrated evidence.**\n:::::\n\n::::::\n\nA transformed evidence score is not an information-theoretic effect size. Conversely, an information estimate in bits does not by itself provide a significance threshold.\n\nThe mathematical ingredients needed to provide calibrated evidence for a statistical dependency in this case are classical. Shannon established entropy as a measure of uncertainty [](https://doi.org/10.1002/j.1538-7305.1948.tb01338.x). Kullback and Leibler introduced relative information as a measure of statistical discrimination [](https://doi.org/10.1214/aoms/1177729694), and Kullback developed its role in hypothesis testing and contingency-table analysis [](https://books.google.com/books?id=XeRQAAAAMAAJ). In the multinomial setting relevant here, the log-likelihood ratio between the empirical distribution and the maximum-likelihood distribution under a null model is exactly a sample-size-scaled KL divergence; this relationship is also standard in categorical-data treatments of likelihood-ratio testing [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1007/978-1-4612-4578-0) [](https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html). Independently, Wilks showed that likelihood-ratio statistics converge to chi-squared distributions under regularity conditions [](https://doi.org/10.1214/aoms/1177732360).\n\nConnecting these two results provides the central inferential step:\n\n$$\n\nG^2 = 2N\\ln(2)\\thinspace D_{\\mathrm{KL},2}(\\widehat p\\Vert \\widehat q_0) \\xrightarrow{d} \\chi^2_{\\nu},\n\n$$\n\nwhere $\\widehat p$ is the empirical distribution, $\\widehat q_0$ is the maximum-likelihood distribution under the scientific null, and $\\nu$ is the difference in model dimension. The equality is algebraic; the chi-squared distribution is asymptotic.\n\nThe contribution of the present work is the synthesis and its consequences. If an information statistic can be formulated as a KL divergence, the saturated-versus-null model pair and prespecified support determine the form and degrees of freedom of the test, the observed table determines $G^2$, and Wilks calibration supplies the corresponding null moments, parametric $p$-value, and common evidence (z-score) representation. MI, CMI, fixed-reference KL, maximum-entropy testing, and conditional maximum-entropy testing then become instances of one workflow rather than unrelated procedures. The section \"[Framework criterion](#framework-criterion)\" states the applicability criterion and clarifies why neighboring quantities (e.g., interaction information) require different inferential constructions.\n\nThe same synthesis clarifies classical first-order plugin corrections. Miller and Basharin developed leading entropy corrections, while later work treated transmitted information, conditional information, and the uncertainty of measured information quantities [](https://api.semanticscholar.org/CorpusID:125662170) [](https://doi.org/10.1137/1104033) [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272). Under the chi-squared reference,\n\n$$\n\n\\mathbb{E}_0[G^2] \\approx \\nu,\n\n$$\n\nthe corresponding fixed-support information statistic has a leading null expectation $\\nu/(2N\\ln 2)$. Subtracting this expectation, the standard small sample bias correction mentioned above, is algebraically identical to centering $G^2$ on the mean of its asymptotic null. This doesn't replace the conventional interpretation as a first-order estimator-bias correction; it provides a second, likelihood-ratio interpretation of the same term.\n\n## Relation to prior work and scope\n\nThe likelihood-ratio interpretation of information statistics has a long history. McGill showed that sample-transmitted information could be used to measure and test association in multidimensional contingency tables [](https://doi.org/10.1007/BF02289159). Kullback developed the representation of multinomial likelihood-ratio statistics as sample-size-scaled discrimination information [](https://books.google.com/books?id=XeRQAAAAMAAJ), and the same connection underlies later information-theoretic treatments of contingency-table testing, partial association, and power [](https://doi.org/10.1111/j.2517-6161.1969.tb00808.x) [](https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html).\n\nFinite-sample estimation of entropy and mutual information is, similarly, an established subject. Miller and Basharin derived leading entropy-bias terms [](https://api.semanticscholar.org/CorpusID:125662170) [](https://doi.org/10.1137/1104033); later work developed corrections and uncertainty calculations for transmitted information, conditional information, entropy, and MI [](https://doi.org/10.1162/neco.1995.7.2.399) [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272). Roulston is especially close to the present standardization argument: He derived a transformation intended to map the MI of unrelated variables to a zero-mean, unit-variance normal test statistic as an analytic alternative to surrogate-data testing [](https://doi.org/10.1016/S0167-2789(97)00117-6). Chance correction and standardization have also been developed for clustering-comparison null models [](https://jmlr.org/papers/v11/vinh10a.html) [](https://proceedings.mlr.press/v32/romano14.html).\n\nFor conditional independence in categorical data, the conditional or stratified likelihood-ratio test is also known. It appears in log-linear analyses of partial association [](https://doi.org/10.1111/j.2517-6161.1969.tb00808.x), in information-theoretic Bayesian-network learning [](https://jmlr.org/papers/v7/decampos06a.html), and in modern work that explicitly treats the usual asymptotic CMI test as a standard baseline whose power deteriorates with increasing conditioning dimension [](https://jmlr.org/papers/v22/19-600.html). Exact and resampling-based alternatives remain important in finite, sparse, dependent, continuous, or mixed-data settings [](https://doi.org/10.3390/e16052839) [](https://proceedings.mlr.press/v84/runge18a.html) [](https://proceedings.mlr.press/v89/marx19a.html). Recently, Goehle [](https://doi.org/10.1007/s41884-025-00170-7) obtained generalized chi-squared approximations for measured KL divergences between random exponential-family distributions using information geometry and a Bayesian central-limit argument; that setting is adjacent to, but distinct from, the saturated-versus-null multinomial likelihood ratios considered here.\n\nAccordingly, the contribution of this paper is a synthesis and operational framework rather than a new likelihood-ratio identity or limiting theorem. The paper places the following elements in one explicit workflow: the null-model KL representation; Wilks calibration; null moments in information units; the dual interpretation of leading bias correction as asymptotic null-mean centering; support-aware degrees of freedom for conditional tables; and normal-equivalent representation of the resulting likelihood-ratio evidence. It applies that workflow consistently to MI, CMI, fixed-reference KL goodness of fit, entropy deficit, and conditional entropy deficit, and evaluates its finite-sample behavior in the large heterogeneous discrete-table regime motivated by the intended applications.\n\nThe framework applies only when the reported information quantity can be formulated as a KL divergence. It does not imply that every divergence or every information-theoretic functional is itself a likelihood-ratio statistic. For example, interaction information is a signed difference of MI and CMI rather than a KL divergence to one null-constrained model; inference requires a joint sampling distribution or an explicitly nested model comparison and is outside the present scope. Continuous estimators, data-adaptive discretization, boundary cases, and growing-dimensional tables likewise require estimator- or regime-specific treatment.\n\n## Contributions and organization\n\nThe paper makes five contributions:\n\n1. It applies the framework consistently to MI, CMI, entropy deficit, conditional entropy deficit, and fixed-reference KL goodness of fit.\n2. It shows that classical first-order bias corrections for information-theoretic values are algebraically the information-scaled versions of asymptotic chi-squared null-mean centering.\n3. It separates effect magnitude in bits from statistical evidence represented by $G^2$, a $p$-value, or a normal-equivalent score.\n4. It states an operational applicability criterion (see \"[Framework criterion](#framework-criterion)\"): An information quantity admits the direct construction when it is the KL divergence between the empirical distribution and the maximum-likelihood distribution under a prespecified nested multinomial null fitted to the same observations. The criterion delineates direct applicability and clarifies why signed composite statistics, the simple asymmetric KL divergence between separately estimated empirical distributions, nonmultinomial estimators, and nonregular or data-adaptive regimes require different inferential constructions.\n5. It evaluates finite-sample calibration across changes in sample size and conditioning cardinality and localizes the failure mode: Sparsity of expected cell counts, rather than low degrees of freedom, drives the observed breakdown. This distinction follows the structure of the construction, since the probability-integral transform addresses chi-squared skewness at any degrees of freedom but inherits whatever error is already present in the chi-squared approximation to $G^2$. The analytic and permutation comparisons agree at the resolution available. The particular count at which the transition appeared is specific to these grids and is reported as a warning region, not a decision boundary.\n\nThe next section states the framework in its general multinomial form. Subsequent sections treat MI, CMI, entropy, conditional entropy, and fixed-reference KL as worked examples before evaluating the framework numerically.\n\n# The Kullback–Wilks framework\n\nThe framework requires only two steps. Kullback’s information-theoretic representation converts the relevant plugin statistic into a multinomial likelihood-ratio deviance, and Wilks' theorem supplies its asymptotic null distribution [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://doi.org/10.1214/aoms/1177732360). The remaining quantities — null moments, first-order correction, $p$-value, and optional standardized evidence coordinates — follow directly.\n\n## Information divergence as a multinomial likelihood ratio\n\nConsider $N$ independent observations allocated among $K$ cells. Let\n\n$$\n\n\\widehat p_i=\\frac{n_i}{N}\n\n$$\n\nbe the saturated multinomial maximum-likelihood estimate, and let $\\widehat q_{0i}$ be the maximum-likelihood cell probability under a null model $\\mathcal M_0$. The null model may be fixed, as in a goodness-of-fit test against a prespecified distribution, or it may be fitted subject to constraints, as in independence or conditional independence. This saturated-versus-null likelihood-ratio construction is standard in multinomial and categorical-data analysis [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1007/978-1-4612-4578-0) [](https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html).\n\n**Proposition 1 (Kullback representation).** For the saturated multinomial alternative and a null-constrained maximum-likelihood distribution $\\widehat q_0$,\n\n$$\n\nG^2 =2\\lbrace \\ell(\\widehat p)-\\ell(\\widehat q_0)\\rbrace =2N\\thinspace D_{\\mathrm{KL}}(\\widehat p\\Vert \\widehat q_0),\n\n$$\n\nwhen the divergence uses natural logarithms. In bits,\n\n$$\n\nG^2 =2N\\ln(2)\\thinspace D_{\\mathrm{KL},2}(\\widehat p\\Vert \\widehat q_0).\n\n$$\n\n_Proof._ For the multinomial log likelihood, terms independent of the model probabilities cancel:\n\n$$\n\n\\begin{aligned} 2\\lbrace \\ell(\\widehat p)-\\ell(\\widehat q_0)\\rbrace &=2\\sum_i n_i\\log\\left(\\frac{\\widehat p_i}{\\widehat q_{0i}}\\right)\\cr &=2N\\sum_i \\widehat p_i\\log\\left(\\frac{\\widehat p_i}{\\widehat q_{0i}}\\right)\\cr &=2N\\thinspace D_{\\mathrm{KL}}(\\widehat p\\Vert \\widehat q_0). \\end{aligned}\n\n$$\n\nChanging from natural to base-two logarithms introduces the factor $\\ln(2)$.\n\nThe equation above is exact for every observed table. It should not be read as saying that an arbitrary KL divergence between two estimated distributions is automatically a likelihood-ratio statistic. The reference distribution must be the distribution implied by the null likelihood for the same observations, and the alternative considered here is the empirical multinomial model [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html).\n\n## Wilks calibration\n\n**Proposition 2 (Wilks calibration).** Suppose the null model is correctly specified and nested within the saturated multinomial model, with fixed dimension, identifiable parameters, and true probabilities in the interior of the parameter space. Then under $H_0$,\n\n$$\n\nG^2 \\xrightarrow{d} \\chi^2_{\\nu}, \\qquad \\nu=\\dim(\\mathcal M_1)-\\dim(\\mathcal M_0),\n\n$$\n\nwhere $\\mathcal M_1$ is the saturated alternative [](https://doi.org/10.1214/aoms/1177732360) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635).\n\nThe conditions are worth separating because they delimit the ordinary Wilks regime. The numerical study in the section \"[The CDF representation is normal when the chi-squared input is adequate](#the-cdf-representation-is-normal-when-the-chi-squared-input-is-adequate)\" examines primarily the expected-cell-count condition rather than attempting to characterize every possible failure mode:\n\n1. **Fixed model dimension.** The number of cells, and hence $\\dim(\\mathcal M_1)$ and $\\dim(\\mathcal M_0)$, are held fixed as $N\\to\\infty$. State spaces that grow with the sample fall outside fixed-dimensional theory.pro\n2. **Identifiable nested models.** $\\mathcal M_0$ is nested within the saturated model, and both are identifiable.\n3. **Interior parameters.** The true probabilities lie in the interior of the parameter space. Structural zeros and boundary cases violate this and require non-standard limiting theory [](https://doi.org/10.1080/01621459.1987.10478472).\n4. **Adequate expected cell counts.** The chi-squared approximation requires cells to be sufficiently populated for the multinomial normal limit to hold. This is the condition we varried most directly in the numerical study; in the grids we examined, sparse expected counts produced the first visible deterioration in calibration [](https://doi.org/10.1080/01621459.1978.10481567) [](https://doi.org/10.1080/01621459.1980.10477473) [](https://doi.org/10.1080/01621459.1986.10478294) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635).\n5. **Prespecified partitions.** Categories, bins, and conditioning strata are fixed in advance. Partitions selected from the same data introduce effective parameters that $\\nu$ doesn't count, making the nominal test anticonservative.\n\nAll five conditions delimit the ordinary chi-squared reference distribution rather than the algebraic identity in Proposition 1. That identity remains exact for an observed table whenever the saturated and null likelihood fits are well defined.\n\nConditions 1-3 and 5 are required for Wilks’ theorem to apply, but condition 4 limits the accuracy of the application of this synthesis. In other contexts, there are corrections for this limitation. Here, we demonstrate that this boundary is empirically near 18–20 expected observations per cell.\n\nThe primary inferential output is the upper-tail probability\n\n$$\n\np_{\\chi^2} =\\operatorname{Pr}\\negthinspace \\left(\\chi^2_{\\nu}\\geq G^2_{\\mathrm{obs}}\\right).\n\n$$\n\nEach analysis uses its own $N$, null model, and $\\nu$. Once this calibration is valid, the resulting $p$-values have the usual common-null interpretation even when the underlying tables differ.\n\n## Null moments in information units\n\nLet\n\n$$\n\n\\widehat {\\mathrm{D}}=D_{\\mathrm{KL},2}(\\widehat p\\Vert \\widehat q_0)\n\n$$\n\nbe the information statistic in bits. Because a chi-squared random variable has mean $\\nu$ and variance $2\\nu$, Propositions 1 and 2 imply the asymptotic null moments\n\n$$\n\n\\begin{aligned} \\mathbb{E}_0[\\widehat {\\mathrm{D}}] &\\approx \\frac{\\nu}{2N\\ln 2}, \\cr \\operatorname{Var}_0(\\widehat {\\mathrm{D}}) &\\approx \\frac{2\\nu}{(2N\\ln 2)^2}. \\end{aligned}\n\n$$\n\nThese expressions expose the dependence of raw plugin information on both sample size and model dimension. Increasing $N$ contracts the null in information units; increasing the number of free departures from the null shifts and broadens it. Related bias and uncertainty calculations for entropy, transmitted information, conditional information, and MI appear throughout the finite-sample information-estimation literature [](https://api.semanticscholar.org/CorpusID:125662170) [](https://doi.org/10.1137/1104033) [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272).\n\n## First-order correction as asymptotic null centering\n\n**Corollary 1 (dual interpretation of the leading correction).** For the fixed-support information functionals treated in the following sections, define the first-order plugin correction\n\n$$\n\n\\widehat {\\mathrm{D}}_{\\mathrm{BC}} =\\widehat {\\mathrm{D}}-\\frac{\\nu}{2N\\ln 2},\n\n$$\n\nthen\n\n$$\n\n\\widehat {\\mathrm{D}}_{\\mathrm{BC}} =\\frac{G^2-\\nu}{2N\\ln 2}.\n\n$$\n\nThus, the leading information-scale correction is algebraically identical to centering the likelihood-ratio deviance on the mean of its asymptotic chi-squared null.\n\nThe identity in the equation above is exact after the correction term is defined. The statement that $\\nu$ equals the finite-sample expectation of $G^2$ isn't exact; it follows from the asymptotic reference distribution. Under the regular fixed-support expansions used for entropy, MI, and conditional information, the correction has two compatible interpretations [](https://api.semanticscholar.org/CorpusID:125662170) [](https://doi.org/10.1137/1104033) [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1162/089976603321780272):\n\n1. It removes the leading plugin-estimator bias\n2. Under the null, it subtracts the leading chi-squared mean in information units.\n\nThis framing does not recast the historical corrections as something other than bias corrections. It shows that the same term also has a likelihood-ratio interpretation.\n\n## Three representations: Information, deviance, and standardized evidence\n\nThe framework yields related quantities with distinct roles.\n\n**Information effect.** $\\widehat {\\mathrm{D}}$ or its first-order-corrected form $\\widehat {\\mathrm{D}}_{\\mathrm{BC}}$ remains in bits. It quantifies the magnitude of departure from the reference model. The correction can be negative near a true boundary of zero; a negative corrected estimate does not imply negative population KL divergence.\n\n**Likelihood-ratio deviance.** $G^2=2N\\ln(2)\\widehat {\\mathrm{D}}$ is the canonical test statistic. It incorporates both effect magnitude and sample size and is the natural scale for nested-model comparisons and deviance decompositions.\n\n**Moment-standardized statistic.** The affine transformation\n\n$$\n\nZ_{\\mathrm{DN}} =\\frac{G^2-\\nu}{\\sqrt{2\\nu}}\n\n$$\n\nhas mean zero and variance one under the chi-squared reference. We retain the name _dimensionality normalization_ (DN) for continuity with the numerical development of this work. DN removes the first two null-moment dependencies but retains chi-squared skewness\n\n$$\n\n\\operatorname{skew}(Z_{\\mathrm{DN}})=\\sqrt{\\frac{8}{\\nu}}.\n\n$$\n\nIt should therefore not be treated as standard normal when $\\nu$ is small.\n\n**Normal-equivalent evidence score.** Let $F_{\\chi^2_{\\nu}}$ and $\\Phi$ denote the chi-squared and standard-normal cumulative distribution functions (CDFs). Define\n\n$$\n\nZ_{\\mathrm{CDF}} =\\Phi^{-1}\\negthinspace \\left(F_{\\chi^2_{\\nu}}(G^2)\\right) =\\Phi^{-1}(1-p_{\\chi^2}).\n\n$$\n\nRoulston previously proposed an analytic normal transformation of MI for significance testing [](https://doi.org/10.1016/S0167-2789(97)00117-6). The present use of the chi-squared CDF differs principally in making the null-model and degrees-of-freedom dependence explicit and in applying the same evidence coordinate across the broader family of likelihood-ratio statistics treated here. If $G^2$ followed the continuous chi-squared reference exactly, the probability-integral transform would make $Z_{\\mathrm{CDF}}$ exactly standard normal for every $\\nu>0$. For finite contingency tables, the reference is asymptotic and the observed statistic is discrete; consequently, $Z_{\\mathrm{CDF}}$ is standard normal only to the extent that the chi-squared approximation is accurate. It's a monotone representation of the same evidence as $p_{\\chi^2}$ and adds no inferential content.\n\n## Framework criterion\n\nAn information-theoretic quantity belongs to this framework when it is the KL divergence between the empirical distribution and the maximum-likelihood distribution under a multinomial null. Different quantities correspond to different null models ([Table 2](#quantities-hypotheses)).\n\n::::::figure{#quantities-hypotheses type=\"table\" label=\"Table 2\"}\n| **Quantity**                   | **Null model**                           | **Scientific question**                                  |\n| ------------------------------ | ---------------------------------------- | -------------------------------------------------------- |\n| Mutual information             | $X\\perp Y$                               | Is there marginal dependence?                            |\n| Conditional mutual information | $X\\perp Y\\mid Z$                         | Is there dependence within conditioning strata?          |\n| Fixed-reference KL             | $p=q_0$                                 | Does the distribution differ from a specified reference? |\n| Entropy deficit                | $p$ uniform on fixed support             | Is uncertainty below its support-constrained maximum?    |\n| Conditional entropy deficit    | $p(x\\mid z)$ uniform within each stratum | Is uncertainty below its conditional maximum?            |\n\n:::::figcaption\n**Table 2.** **Null models and relevant hypotheses for Information quantities**\n:::::\n\n::::::\n\nThe criterion also separates the quantities above from neighboring ones. Four cases fall outside the direct construction, each for a distinct reason.\n\n1. **Signed or composite statistics.** Interaction information is $\\mathrm{I}(X;Y)-\\mathrm{I}(X;Y\\mid Z)$, a difference of two statistics computed on the same observations. It isn't a single nonnegative KL divergence to one null-constrained model, and because the two estimates are correlated, its null variance isn't the sum of theirs. Calibrated inference requires their joint sampling distribution or an explicitly nested model comparison [](https://doi.org/10.1007/BF02289159) [](https://doi.org/10.1111/j.2517-6161.1969.tb00808.x).\n2. **Simple divergences between separately estimated empirical distributions.** The asymmetric quantity $D_{\\mathrm{KL}}(\\widehat p_1\\Vert \\widehat p_2)$ is not itself the likelihood-ratio deviance for a two-sample homogeneity test. That test instead uses each sample’s divergence from a pooled null-constrained estimate.\n3. **Nonmultinomial estimators.** Continuous nearest-neighbor and related estimators don't have the exact multinomial likelihood-ratio representation used here and require estimator-specific null theory or resampling [](https://proceedings.mlr.press/v84/runge18a.html).\n4. **Nonregular or data-adaptive regimes.** Boundary parameters, rapidly growing state spaces, and categories or partitions selected using the same data may invalidate the ordinary Wilks reference even when a likelihood-ratio statistic can still be defined [](https://doi.org/10.1080/01621459.1987.10478472).\n\n# Mutual information as a test of independence\n\nThis section instantiates the framework for the null $X\\perp Y$. The null-model specification and prespecified support determine the form and degrees of freedom of the test; the observed table determines its likelihood-ratio statistic. Let $X$ and $Y$ have $k_x$ and $k_y$ states, respectively, and let $n_{ij}$ denote the observed counts in a $k_x\\times k_y$ contingency table. Write $n_{i\\cdot}$ and $n_{\\cdot j}$ for the row and column totals, and let $N=\\sum_{ij}n_{ij}$.\n\nThe plugin estimator in bits is\n\n$$\n\n\\widehat {\\mathrm{I}}(X;Y) =\\sum_{i,j}\\widehat p_{ij} \\log_2\\left(\\frac{\\widehat p_{ij}} {\\widehat p_{i\\cdot}\\widehat p_{\\cdot j}}\\right).\n\n$$\n\nIts population target is the KL divergence between the joint distribution and the product of its marginals. The corresponding null is\n\n$$\n\nH_0:X\\perp Y.\n\n$$\n\nUnder this null, the maximum-likelihood expected count is\n\n$$\n\nE_{ij}=\\frac{n_{i\\cdot}n_{\\cdot j}}{N}.\n\n$$\n\nThe information-based contingency-table test is classical and appears in early work on transmitted information, Kullback’s statistical formulation, later information identities for contingency tables, and standard categorical-data treatments [](https://doi.org/10.1007/BF02289159) [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635).\n\n**Proposition 3 (mutual information and the $G$ test).** For every observed contingency table,\n\n$$\n\nG^2_{\\mathrm{MI}} =2\\sum_{i,j}n_{ij}\\log\\left(\\frac{n_{ij}}{E_{ij}}\\right) =2N\\ln(2)\\thinspace \\widehat {\\mathrm{I}}(X;Y).\n\n$$\n\nThe equality is an immediate instance of Proposition 1. It's exact at every sample size and for every true distribution. Correctness of the null reference distribution is asymptotic [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635).\n\nFor a full-support table, the saturated joint model has $k_xk_y-1$ free parameters, whereas the independence model has $(k_x-1)+(k_y-1)$ free parameters. Therefore\n\n$$\n\n\\nu_{\\mathrm{MI}}=(k_x-1)(k_y-1),\n\n$$\n\nand under the regularity conditions of Wilks' theorem,\n\n$$\n\nG^2_{\\mathrm{MI}}\\xrightarrow{d}\\chi^2_{\\nu_{\\mathrm{MI}}}.\n\n$$\n\nThe resulting inferential quantities are\n\n$$\n\n\\begin{aligned} p_{\\mathrm{MI}} &=\\operatorname{Pr}\\negthinspace \\left(\\chi^2_{\\nu_{\\mathrm{MI}}} \\geq G^2_{\\mathrm{MI}}\\right), \\cr Z_{\\mathrm{CDF,MI}} &=\\Phi^{-1}(1-p_{\\mathrm{MI}}). \\end{aligned}\n\n$$\n\nThe $p$-value is the primary hypothesis-testing output. The normal-equivalent score is useful for ranking, visualization, or integration with analyses expressed on standard-normal evidence scales; it's a representation of the same calibrated tail probability [](https://doi.org/10.1016/S0167-2789(97)00117-6).\n\n## Null moments and the first-order plugin correction\n\nThe equations above give\n\n$$\n\n\\begin{aligned} \\mathbb{E}_0[\\widehat {\\mathrm{I}}] &\\approx \\frac{\\nu_{\\mathrm{MI}}}{2N\\ln2}, \\cr \\operatorname{Var}_0(\\widehat {\\mathrm{I}}) &\\approx \\frac{2\\nu_{\\mathrm{MI}}}{(2N\\ln2)^2}. \\end{aligned}\n\n$$\n\nThe leading correction for MI is therefore\n\n$$\n\n\\widehat {\\mathrm{I}}_{\\mathrm{BC}} =\\widehat {\\mathrm{I}}-\\frac{\\nu_{\\mathrm{MI}}}{2N\\ln2} =\\frac{G^2_{\\mathrm{MI}}-\\nu_{\\mathrm{MI}}}{2N\\ln2}.\n\n$$\n\nThis is the MI instance of Corollary 1. Leading MI and conditional-information corrections of this form are derived directly in the later information-estimation literature [](https://doi.org/10.1162/neco.1995.7.2.399) [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272). Under fixed support and positive cell probabilities, it removes the leading $O(N^{-1})$ plugin bias; it isn't generally an exactly unbiased finite-sample estimator.\n\nThe corresponding DN statistic is\n\n$$\n\nZ_{\\mathrm{DN,MI}} =\\frac{\\widehat {\\mathrm{I}}_{\\mathrm{BC}}} {\\sqrt{\\operatorname{Var}_0(\\widehat {\\mathrm{I}})}} =\\frac{G^2_{\\mathrm{MI}}-\\nu_{\\mathrm{MI}}} {\\sqrt{2\\nu_{\\mathrm{MI}}}}.\n\n$$\n\nThis standardizes the first two reference moments, but not the full distribution. In particular, a $2\\times2$ independence test has one degree of freedom and therefore retains substantial chi-squared skewness after affine standardization.\n\n## Interpretation in large-scale association analysis\n\nFor a locus–phenotype table, $\\widehat {\\mathrm{I}}_{\\mathrm{BC}}$ is a first-order-corrected effect estimate in bits: It estimates how much knowing the locus reduces uncertainty about the phenotype, without assigning a direction or imposing an additive genotype model. The pair $(p_{\\mathrm{MI}},Z_{\\mathrm{CDF,MI}})$ measures evidence against marginal independence. This separation is analogous in role to reporting a regression coefficient together with its test statistic, while recognizing that the information-theoretic effect and a regression coefficient answer different scientific questions [](https://doi.org/10.1038/nrg1916) [](https://doi.org/10.1038/nrg2579) [](https://doi.org/10.1038/s43586-021-00056-9).\n\nLocus-specific changes in $N$, missingness, or observed genotype support alter the information-scale null moments. They don't prevent common evidence thresholds, provided each table is calibrated with its own likelihood-ratio statistic and valid degrees of freedom. When categories are removed because they're unobserved, however, the analyst must distinguish prespecified structural absence from random sampling zeros; the latter can signal a boundary or sparse-table regime where Wilks' approximation is unreliable [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1080/01621459.1987.10478472).\n\n# Conditional mutual information as a stratified likelihood-ratio test\n\nThis section instantiates the framework for the null $X\\perp Y\\mid Z$. The conditional-independence model and prespecified stratum-specific supports determine the form and degrees of freedom of the test; the observed stratified table determines its likelihood-ratio statistic. Let $X$, $Y$, and $Z$ be discrete variables. The plugin conditional mutual information in bits is\n\n$$\n\n\\widehat {\\mathrm{I}}(X;Y\\mid Z) =\\sum_z \\widehat p_z\\thinspace \\widehat {\\mathrm{I}}_z(X;Y),\n\n$$\n\nwhere $\\widehat {\\mathrm{I}}_z(X;Y)$ is the plugin MI calculated within stratum $Z=z$ and $\\widehat p_z=N_z/N$. Equivalently,\n\n$$\n\n\\widehat {\\mathrm{I}}(X;Y\\mid Z) =\\sum_{x,y,z}\\widehat p_{xyz} \\log_2\\left[ \\frac{\\widehat p_{xy\\mid z}} {\\widehat p_{x\\mid z}\\widehat p_{y\\mid z}} \\right].\n\n$$\n\nThe scientific null is conditional independence,\n\n$$\n\nH_0:X\\perp Y\\mid Z.\n\n$$\n\nCMI and its asymptotic reference distribution are standard tools for categorical conditional-independence testing and Bayesian-network learning [](https://doi.org/10.1002/047174882X) [](https://jmlr.org/papers/v7/decampos06a.html) [](https://jmlr.org/papers/v22/19-600.html) [](https://doi.org/10.7551/mitpress/1754.001.0001).\n\n**Proposition 4 (CMI as a stratified $G$ statistic).** Let $G^2_z$ be the likelihood-ratio statistic for independence of $X$ and $Y$ within stratum $z$. Then\n\n$$\n\nG^2_{\\mathrm{CMI}} =\\sum_z G^2_z =2N\\ln(2)\\thinspace \\widehat {\\mathrm{I}}(X;Y\\mid Z).\n\n$$\n\n_Proof._ Applying Proposition 3 within each stratum gives\n\n$$\n\nG^2_z=2N_z\\ln(2)\\thinspace \\widehat {\\mathrm{I}}_z(X;Y).\n\n$$\n\nSumming across strata and using $N_z=N\\widehat p_z$ yields\n\n$$\n\n\\begin{aligned} \\sum_zG^2_z &=2\\ln(2)\\sum_zN_z\\widehat {\\mathrm{I}}_z(X;Y)\\cr &=2N\\ln(2)\\sum_z\\widehat p_z\\widehat {\\mathrm{I}}_z(X;Y)\\cr &=2N\\ln(2)\\thinspace \\widehat {\\mathrm{I}}(X;Y\\mid Z). \\end{aligned}\n\n$$\n\nThis representation is the classical conditional or stratified $G$ test, closely related to likelihood-ratio tests of partial association in log-linear models [](https://doi.org/10.1111/j.2517-6161.1969.tb00808.x) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://jmlr.org/papers/v7/decampos06a.html) [](https://jmlr.org/papers/v22/19-600.html). The role of the present framework is to connect its likelihood-ratio calibration explicitly to the plugin CMI estimator, its null moments, and its information-scale correction.\n\n## Degrees of freedom\n\nSuppose first that every one of the $k_z$ strata contains the same $k_x$ and $k_y$ admissible states. The saturated joint model has\n\n$$\n\nk_xk_yk_z-1\n\n$$\n\nfree parameters. Under conditional independence,\n\n$$\n\np(x,y,z)=p(z)p(x\\mid z)p(y\\mid z),\n\n$$\n\nand the model dimension is\n\n$$\n\n(k_z-1)+k_z(k_x-1)+k_z(k_y-1).\n\n$$\n\nThe difference is\n\n$$\n\n\\nu_{\\mathrm{CMI}} =k_z(k_x-1)(k_y-1) =k_z\\nu_{\\mathrm{MI}}.\n\n$$\n\nUnder $H_0$ and Wilks' conditions,\n\n$$\n\nG^2_{\\mathrm{CMI}} \\xrightarrow{d} \\chi^2_{\\nu_{\\mathrm{CMI}}}.\n\n$$\n\nThese parameter counts are standard consequences of the conditional-independence/log-linear model [](https://doi.org/10.1111/j.2517-6161.1969.tb00808.x) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1007/978-1-4612-4578-0).\n\nWhen the admissible supports differ across prespecified strata, the appropriate count is\n\n$$\n\n\\nu_{\\mathrm{CMI}} =\\sum_{z\\in\\mathcal Z_+} (k_{x,z}-1)(k_{y,z}-1),\n\n$$\n\nwhere $\\mathcal Z_+$ denotes the strata included in the model. This formula shouldn't be applied mechanically after deleting random zero-count categories: Such deletions can make the model selection data-adaptive and conceal boundary behavior. In sparse genomic tables, the nominal degrees of freedom and the adequacy of Wilks' approximation must therefore be assessed together [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1080/01621459.1987.10478472).\n\n## Null moments and calibrated evidence\n\nThe asymptotic information-scale moments are\n\n$$\n\n\\begin{aligned} \\mathbb{E}_0[\\widehat {\\mathrm{I}}(X;Y\\mid Z)] &\\approx \\frac{\\nu_{\\mathrm{CMI}}}{2N\\ln2}, \\cr \\operatorname{Var}_0[\\widehat {\\mathrm{I}}(X;Y\\mid Z)] &\\approx \\frac{2\\nu_{\\mathrm{CMI}}}{(2N\\ln2)^2}. \\end{aligned}\n\n$$\n\nIn the full-support balanced case, the null mean is $k_z$ times the MI null mean for the same $k_x$, $k_y$, and $N$, and the null standard deviation is $\\sqrt{k_z}$ times larger. This is the table-dimensional inflation that makes raw MI and CMI unsuitable as common evidence measures.\n\nThe first-order-corrected CMI is\n\n$$\n\n\\widehat {\\mathrm{I}}_{\\mathrm{BC}}(X;Y\\mid Z) =\\widehat {\\mathrm{I}}(X;Y\\mid Z) -\\frac{\\nu_{\\mathrm{CMI}}}{2N\\ln2},\n\n$$\n\nwith the leading conditional-information correction supported directly by finite-sample information-estimation work [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://jmlr.org/papers/v22/19-600.html) [](https://doi.org/10.1162/089976603321780272). The inferential outputs are\n\n$$\n\n\\begin{aligned} p_{\\mathrm{CMI}} &=\\operatorname{Pr}\\negthinspace \\left( \\chi^2_{\\nu_{\\mathrm{CMI}}} \\geq G^2_{\\mathrm{CMI}} \\right), \\cr Z_{\\mathrm{DN,CMI}} &=\\frac{G^2_{\\mathrm{CMI}}-\\nu_{\\mathrm{CMI}}} {\\sqrt{2\\nu_{\\mathrm{CMI}}}}, \\cr Z_{\\mathrm{CDF,CMI}} &=\\Phi^{-1}(1-p_{\\mathrm{CMI}}). \\end{aligned}\n\n$$\n\nCalibrated MI and CMI $p$-values, or their normal-equivalent coordinates, can be compared as evidence against their respective null hypotheses even when $N$ and table dimensions differ. They shouldn't be interpreted as measurements of the same population quantity. MI tests marginal independence, whereas CMI tests conditional independence. In particular,\n\n$$\n\nZ_{\\mathrm{CDF,CMI}}-Z_{\\mathrm{CDF,MI}}\n\n$$\n\nis not a calibrated test of effect modification, confounding, or interaction.\n\n## Relation to conditional analyses in genetics and causal discovery\n\nIn genetics, $I(G;D\\mid C)$ asks whether genotype $G$ contains disease information within levels of a conditioning variable $C$, such as sex or a prespecified population label. It is an omnibus conditional-dependence test. It can be sensitive to additive, dominant, recessive, multiallelic, nonmonotone, or context-dependent patterns, but it pays the degrees-of-freedom cost associated with that flexibility [](https://doi.org/10.1038/nrg2579) [](https://doi.org/10.1038/nrg1916). A fully categorical regression containing all genotype main effects and genotype-by-condition interactions tests a closely related alternative with the same parameter count under the corresponding saturated categorical parameterization [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635). Thus, the advantage isn't a cost-free interaction test; it's a functional-form-agnostic representation of arbitrary discrete conditional dependence.\n\nIn constraint-based causal discovery, the same calibration permits a fixed significance level to be applied while the conditioning table changes. Procedures that already compute test-specific $p$-values from valid null distributions are already calibrated in this sense. The framework is most relevant when raw CMI thresholds are used, when analytic calibration hasn't been made explicit, or when repeated permutation is computationally prohibitive [](https://jmlr.org/papers/v7/decampos06a.html) [](https://jmlr.org/papers/v22/19-600.html) [](https://proceedings.mlr.press/v84/runge18a.html) [](https://proceedings.mlr.press/v124/runge20a.html) [](https://proceedings.mlr.press/v89/marx19a.html). Continuous nearest-neighbor CMI estimators remain outside the present theory [](https://proceedings.mlr.press/v84/runge18a.html).\n\n# Fixed-reference KL divergence and entropy-based tests\n\nThe previous sections use null-constrained distributions estimated from the observed marginals. A simpler instance of the framework compares the empirical distribution with a fixed reference model. Entropy and conditional entropy enter as special cases when that reference is uniform.\n\n## Fixed-reference KL goodness of fit\n\nLet $X$ have $k$ prespecified states, empirical probabilities $\\widehat p_i$, and a fixed reference distribution $q_0$ with $q_{0i}>0$ on the admissible support. The plugin divergence in bits is\n\n$$\n\n\\widehat {\\mathrm{D}}_q =D_{\\mathrm{KL},2}(\\widehat p\\Vert q_0) =\\sum_i\\widehat p_i\\log_2\\left(\\frac{\\widehat p_i}{q_{0i}}\\right).\n\n$$\n\nThe null is\n\n$$\n\nH_0:p=q_0.\n\n$$\n\nBy Proposition 1,\n\n$$\n\nG^2_q=2N\\ln(2)\\thinspace \\widehat {\\mathrm{D}}_q.\n\n$$\n\nIf $q_0$ is completely specified, the null has no fitted multinomial parameters and\n\n$$\n\n\\nu_q=k-1.\n\n$$\n\nIf $q_0$ belongs to a regular $r$-parameter family whose parameters are fitted from the same observations, then the usual goodness-of-fit count is\n\n$$\n\n\\nu_q=k-1-r,\n\n$$\n\nprovided the fitted model is identifiable and regular. These are standard multinomial goodness-of-fit results [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1007/978-1-4612-4578-0) [](https://doi.org/10.1111/j.2517-6161.1984.tb01318.x).\n\nThe framework does not identify an arbitrary divergence between two independently estimated empirical distributions with a likelihood-ratio statistic. A two-sample homogeneity test, for example, uses each sample’s divergence from a pooled null-constrained estimate rather than simply $D_{\\mathrm{KL}}(\\widehat p_1\\Vert \\widehat p_2)$. Such tests can be handled by likelihood-ratio theory, but the asymmetric pairwise divergence is not itself the relevant deviance [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1007/978-1-4612-4578-0).\n\n## Entropy deficit as a maximum-entropy test\n\nLet $u_i=1/k$ be uniform on the fixed support. Shannon entropy in bits is\n\n$$\n\n\\widehat {\\mathrm{H}}(X)=-\\sum_i\\widehat p_i\\log_2\\widehat p_i.\n\n$$\n\nThe divergence from uniformity is\n\n$$\n\n\\begin{aligned} D_{\\mathrm{KL},2}(\\widehat p\\Vert u) &=\\sum_i\\widehat p_i\\log_2(k\\widehat p_i)\\cr &=\\log_2 k-\\widehat {\\mathrm{H}}(X). \\end{aligned}\n\n$$\n\nDefine the entropy deficit\n\n$$\n\n\\widehat\\Delta_H =\\log_2 k-\\widehat {\\mathrm{H}}(X).\n\n$$\n\nThen\n\n$$\n\nG^2_H =2N\\ln(2)\\thinspace \\widehat\\Delta_H \\xrightarrow{d}\\chi^2_{k-1}\n\n$$\n\nunder\n\n$$\n\nH_0:p=u.\n\n$$\n\nThis is the ordinary multinomial likelihood-ratio goodness-of-fit test for uniformity expressed as a parametric test of maximum entropy on a fixed discrete support [](https://doi.org/10.1002/j.1538-7305.1948.tb01338.x) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1007/978-1-4612-4578-0). It isn't a generic test of the composite hypothesis $H(p)=h_0$ for an arbitrary nonmaximal value $h_0$. Parametric inference comparing Shannon diversities and entropies has other established formulations [](https://doi.org/10.1016/0022-5193(70)90124-4); the specific result here is the maximum-entropy/uniformity case.\n\nThe null mean of the entropy deficit is\n\n$$\n\n\\mathbb{E}_0[\\widehat\\Delta_H] \\approx\\frac{k-1}{2N\\ln2}.\n\n$$\n\nSubtracting this mean from the deficit is equivalent to adding it to the plugin entropy:\n\n$$\n\n\\begin{aligned} \\widehat\\Delta_{H,\\mathrm{BC}} &=\\widehat\\Delta_H-\\frac{k-1}{2N\\ln2},\\cr \\widehat {\\mathrm{H}}_{\\mathrm{MM}} &=\\widehat {\\mathrm{H}}+\\frac{k-1}{2N\\ln2}. \\end{aligned}\n\n$$\n\nThe equation above is the Miller–Madow first-order entropy correction; closely related entropy-bias expansions were derived by Basharin and treated more generally in later work [](https://api.semanticscholar.org/CorpusID:125662170) [](https://doi.org/10.1137/1104033) [](https://doi.org/10.1162/089976603321780272). The sign is opposite to that for MI because the likelihood-ratio statistic is the _deficit_ $\\log_2k-\\widehat {\\mathrm{H}}$, not entropy itself. This is the entropy analog of Corollary 1: The leading entropy correction and asymptotic null centering of the deficit are the same operation written on opposite sides of the identity in the equation above.\n\n## Conditional entropy deficit\n\nLet $Z$ define prespecified strata and let $X$ have $k_{x,z}$ admissible states in stratum $z$. The plugin conditional entropy is\n\n$$\n\n\\widehat {\\mathrm{H}}(X\\mid Z) =\\sum_z\\widehat p_z\\widehat {\\mathrm{H}}(X\\mid Z=z).\n\n$$\n\nThe maximum conditional entropy allowed by the prespecified stratum-specific supports is\n\n$$\n\nH_{\\max}(X\\mid Z) =\\sum_z\\widehat p_z\\log_2 k_{x,z}.\n\n$$\n\nThe conditional entropy deficit is\n\n$$\n\n\\begin{aligned} \\widehat\\Delta_{H\\mid Z} &=H_{\\max}(X\\mid Z)-\\widehat {\\mathrm{H}}(X\\mid Z)\\cr &=\\sum_z\\widehat p_z D_{\\mathrm{KL},2}\\negthinspace \\left(\\widehat p_{X\\mid z}\\Vert u_z\\right), \\end{aligned}\n\n$$\n\nwhere $u_z$ is uniform on the admissible support in stratum $z$. Therefore,\n\n$$\n\nG^2_{H\\mid Z} =2N\\ln(2)\\thinspace \\widehat\\Delta_{H\\mid Z} =\\sum_z 2N_z\\ln(2) D_{\\mathrm{KL},2}\\negthinspace \\left(\\widehat p_{X\\mid z}\\Vert u_z\\right).\n\n$$\n\nUnder the null that $X$ is uniform within every stratum,\n\n$$\n\nG^2_{H\\mid Z} \\xrightarrow{d} \\chi^2_{\\nu_{H\\mid Z}}, \\qquad \\nu_{H\\mid Z}=\\sum_z(k_{x,z}-1).\n\n$$\n\nWhen every stratum has the same $k_x$ states, this reduces to $k_z(k_x-1)$. The construction is a direct stratified multinomial goodness-of-fit test [](https://doi.org/10.1002/047174882X) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1111/j.2517-6161.1969.tb00808.x).\n\nThis test asks whether uncertainty is maximal within every conditioning stratum. It's distinct from asking whether conditioning changes entropy. The latter question is\n\n$$\n\nH(X)-H(X\\mid Z)=I(X;Z),\n\n$$\n\nand is tested by the MI likelihood-ratio statistic described in the section \"[Mutual information as a test of independence](#mutual-information-as-a-test-of-independence).\"\n\nBelow we provide a summary table of the principal uses of this synthesis and articulate the estimand and the appropriate representation, null hypothesis, and degrees of freedom ([Table 3](#framework-summary)).\n\n## Summary of principal instances\n\n::::::figure{#framework-summary type=\"table\" label=\"Table 3\"}\n| **Quantity in bits**               | **KL representation**                                                                                               | **Null hypothesis**          | **Degrees of freedom**               |\n| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------- | ---------------------------- | ------------------------------------ |\n| $\\widehat {\\mathrm{I}}(X;Y)$       | $D_{\\mathrm{KL},2}(\\widehat p_{XY}\\Vert \\widehat p_X\\widehat p_Y)$                                              | $X\\perp Y$                   | $(k_x-1)(k_y-1)$                   |\n| $\\widehat {\\mathrm{I}}(X;Y\\mid Z)$ | $\\sum_z\\widehat p_z$ $D_{\\mathrm{KL},2}(\\widehat p_{XY\\mid z}\\Vert \\widehat p_{X\\mid z}\\widehat p_{Y\\mid z})$ | $X\\perp Y\\mid Z$             | $\\sum_z$ $(k_{x,z}-1)(k_{y,z}-1)$ |\n| $\\widehat {\\mathrm{D}}_q$         | $D_{\\mathrm{KL},2}(\\widehat p\\Vert q_0)$                                                                          | $p=q_0$                     | $k-1-r$                              |\n| $\\widehat\\Delta_H$                | $D_{\\mathrm{KL},2}(\\widehat p\\Vert u)$                                                                             | uniformity / maximum entropy | $k-1$                                |\n| $\\widehat\\Delta_{H\\mid Z}$        | $\\sum_z\\widehat p_zD_{\\mathrm{KL},2}(\\widehat p_{X\\mid z}\\Vert u_z)$                                           | conditional uniformity       | $\\sum_z(k_{x,z}-1)$                |\n\n:::::figcaption\n**Table 3.** **Principal information-theoretic instances of the likelihood-ratio framework.**\n\nDegrees of freedom assume fixed, fully represented supports unless a general sum is shown.\n:::::\n\n::::::\n\nIn every row,\n\n$$\n\nG^2=2N\\ln(2)\\times\\text{information statistic},\n\n$$\n\nand the asymptotic $p$-value is read from the chi-squared distribution with the row-specific degrees of freedom.\n\n# Numerical validation\n\nThe numerical study evaluates the consequences of the Kullback–Wilks synthesis in the setting that motivated the framework: discrete MI and CMI calculated from contingency tables with varying sample sizes and conditioning dimensions. The simulations are not intended to re-establish the classical likelihood-ratio results. They ask whether the asymptotic reference is accurate over the examined finite-sample regimes, whether the predicted null moments match empirical behavior, and where sparse cells cause the calibration to fail.\n\nThe existing experiments focus on MI and CMI. Fixed-reference KL and entropy-deficit tests are direct multinomial goodness-of-fit cases, and we don't simulate them separately in the present draft. Their inclusion in the mathematical framework doesn't depend on the MI/CMI simulations.\n\n## Software and reproducibility\n\n::::::div{.info-box}\n**Code**, including a Python library and a script for figure generation, is available in this [GitHub repository](https://github.com/Arcadia-Science/arcadia-info-standardization/tree/v1.0.1) (DOI: [10.5281/zenodo.22214680](https://doi.org/10.5281/zenodo.22214680))\n::::::\n\nWe wrote all simulations in Python using NumPy (v1.26.4), SciPy (v1.11.4), and Matplotlib (v3.10.9) [](https://doi.org/10.1038/s41586-020-2649-2) [](https://doi.org/10.1038/s41592-019-0686-2) [](https://doi.org/10.1109/MCSE.2007.55). Random draws used NumPy’s `default_rng` with seed 42. We reused the same estimator and transformation functions across experiments; parameter sweeps and null generators varied by figure.\n\n## Estimators and statistics\n\nFor a $k_x\\times k_y$ table, plugin MI was calculated as\n\n$$\n\n\\widehat {\\mathrm{I}}=\\sum_{i,j}\\widehat p_{ij} \\log_2\\left( \\frac{\\widehat p_{ij}} {\\widehat p_{i\\cdot}\\widehat p_{\\cdot j}} \\right).\n\n$$\n\nFor a $k_x\\times k_y\\times k_z$ table, plugin CMI was calculated as the stratum-weighted sum\n\n$$\n\n\\widehat {\\mathrm{I}}(X;Y\\mid Z) =\\sum_z\\widehat p_z\\widehat {\\mathrm{I}}_z(X;Y).\n\n$$\n\nFor each estimate, the simulations formed\n\n$$\n\n\\begin{aligned} G^2&=2N\\ln(2)\\thinspace \\widehat {\\mathrm{I}},\\cr \\widehat {\\mathrm{I}}_{\\mathrm{BC}}&=\\widehat {\\mathrm{I}}-\\frac{\\nu}{2N\\ln2},\\cr Z_{\\mathrm{DN}}&=\\frac{G^2-\\nu}{\\sqrt{2\\nu}},\\cr Z_{\\mathrm{CDF}}&=\\Phi^{-1}\\negthinspace \\left(F_{\\chi^2_{\\nu}}(G^2)\\right). \\end{aligned}\n\n$$\n\nFor numerical stability, the CDF values were clipped to $[10^{-15},1-10^{-15}]$ before applying $\\Phi^{-1}$. In the full-support CMI simulations,\n\n$$\n\n\\nu=k_z(k_x-1)(k_y-1).\n\n$$\n\nThe historical DN parameterization describes the information-scale variance as proportional to $C/N^2$, where $C=k_xk_y$. Equating that representation with the equation above gives\n\n$$\n\n\\sigma_0= \\sqrt{ \\frac{2(k_x-1)(k_y-1)} {k_xk_y(2\\ln2)^2} }.\n\n$$\n\nFor the $3\\times2$ table used throughout, $\\sigma_0\\approx0.589$.\n\n## Null data generation\n\nFor most experiments, we used a genotype–disease null. We generated a genotype $G\\in\\lbrace 0,1,2\\rbrace$ as the sum of two Bernoulli allele draws, giving $k_x=3$. We generated disease $D\\in\\lbrace 0,1\\rbrace$ as Bernoulli$(0.5)$, giving $k_y=2$. We assigned a population label to each observation, one of $k_z$ strata. The baseline allele frequency was 0.30. A linear frequency gradient and optional Gaussian perturbation could be introduced across strata, but we set both to zero for the null figures, so genotype and disease were independent both marginally and within strata.\n\nFor matched comparisons across conditioning cardinalities, we generated a partition at the largest $k_z$ and coarsened to obtain the smaller values. Thus, the same observations underlay the paired values at different $k_z$ in a simulation replicate.\n\nTwo experiments used alternative null generators. The degrees-of-freedom experiment sampled cell counts from a uniform multinomial over the $k_xk_yk_z$ cells. The validity-regime experiment generated each variable as an independent, uniform categorical variable. For power comparisons on $2\\times2$ tables, we fixed the marginals, numerically computed the joint cell probability corresponding to a target nonzero MI using `scipy.optimize.brentq`, and sampled counts from the resulting multinomial.\n\n## Monte Carlo evaluation\n\nAt each parameter setting, we generated independent replicate tables (default 500), computed the relevant statistics, and compared their empirical distributions with the predicted references. Evaluation included:\n\n* quantile–quantile plots against the scaled chi-squared null for raw MI/CMI and against $\\mathcal N(0,1)$ for $Z_{\\mathrm{CDF}}$;\n* Kolmogorov–Smirnov comparisons with the stated reference distribution;\n* empirical mean, variance, skewness, and kurtosis;\n* three estimates of chi-squared degrees of freedom: $\\overline{G^2}$, $\\operatorname{Var}(G^2)/2$, and maximum-likelihood fitting of a chi-squared distribution; and\n* comparison with a label-permutation test in selected $2\\times2$ alternatives.\n\nNon-rejection by a KS test is treated as evidence of agreement at the resolution of the simulation, not proof that the finite-sample distribution equals its asymptotic reference. In particular, the available replicate counts can't directly validate genome-wide tail probabilities such as $5\\times10^{-8}$.\n\nMonte Carlo permutation $p$-values should use a finite-sample correction such as $(b+1)/(m+1)$, where $b$ is the number of permuted statistics at least as large as the observed statistic and $m$ is the number of random permutations [](https://doi.org/10.2202/1544-6115.1585).\n\nExperiments\n\n**Correction progression across conditioning sizes ([Figure 1](#correction-progression)).** At $N=100{,}000$, $k_z\\in\\lbrace 10,30,50,70,100\\rbrace$, and 500 replicates, we compared the statistic at a baseline $k_z=10$ with the statistic at each larger $k_z$ under four representations: raw information, first-order plugin centering (Basharin), DN, and the chi-squared-CDF transform.\n\n**Scaled chi-squared null across $k_z$ and $N$ ([Figure 2](#chi-k) and [Figure 3](#chi-n)).** Using the genotype–disease null, we compared $G^2=2N\\ln(2)\\widehat {\\mathrm{I}}$ with $\\chi^2_{\\nu}$. [Figure 2](#chi-k) varies $k_z\\in\\lbrace 6,10,30,100\\rbrace$ at $N=100{,}000$. [Figure 3](#chi-n) varies $N\\in\\lbrace 1{,}000,5{,}000,20{,}000,100{,}000\\rbrace$ at $k_z=6$. Each setting uses 500 replicates.\n\n**Variance constant and degrees of freedom ([Figure 4](#variance-df)).** We compared the predicted $\\sigma_0$ with empirical estimates for a $3\\times2$ table across $k_z\\in\\lbrace 6,10,20,50,100,200\\rbrace$ at $N=100{,}000$. We split three hundred tables into ten groups for repeated variance estimates. In a separate uniform-multinomial experiment, we evaluated $\\nu_{\\mathrm{CMI}}=k_z\\nu_{\\mathrm{MI}}$ across $k_z\\in\\lbrace 2,3,5,6,10,15,20,30,50,100\\rbrace$ at $N=10{,}000$ using 2,000 replicates.\n\n**Normal-equivalent calibration ([Figure 5](#zcdf-normal)).** We compared $Z_{\\mathrm{CDF}}$ with $\\mathcal N(0,1)$ across $N\\in\\lbrace 5{,}000,20{,}000,100{,}000\\rbrace$ and $k_z\\in\\lbrace 6,10,30,100\\rbrace$ using 1,000 replicates per setting.\n\n**Sparse-cell validity regime ([Figure 6](#validity)).** We generated independent-uniform draws across $k_z\\in\\lbrace 6,10,30,50,70\\rbrace$ and $N\\in\\lbrace 100,500,1{,}000,5{,}000,20{,}000,100{,}000\\rbrace$ using 500 replicates. We recorded the KS statistic and average expected count per cell, $N/(k_xk_yk_z)$.\n\n**Tail behavior ([Figure 7](#tails)).** At $N=100{,}000$, $k_z\\in\\lbrace 6,10,30,100\\rbrace$, and 10,000 replicates, we compared raw information, DN, and CDF-transformed values in the upper empirical tail.\n\n**Analytic vs. permutation inference ([Figure 8](#permutation)).** For $2\\times2$ tables with target MI in $\\lbrace 0.001,0.003,0.005,0.01,0.02\\rbrace$ bits and $N\\in\\lbrace 1{,}000,2{,}000,5{,}000,20{,}000\\rbrace$, we used 100 replicate tables and 200 permutations per table to compare detection power and per-table $p$-values.\n\n# Numerical results\n\n## Progressive removal of null dependence on table dimension\n\nRaw plugin information changed systematically with conditioning cardinality in the matched-table comparison ([Figure 1](#correction-progression), A). Across the $k_z$ values we examined, the normalized mean RMSE was $0.3274\\pm0.2428$, and the normalized slope deviation was 1.0914. Subtracting the first-order plugin term removed most of the mean shift but left the changing dispersion intact ([Figure 1](#correction-progression), B; normalized mean RMSE $0.0632\\pm0.0407$, normalized slope deviation 1.0914). Dividing the centered statistic by the chi-squared null standard deviation aligned the first two moments ([Figure 1](#correction-progression), C; normalized mean RMSE $0.01262\\pm0.0083$, normalized slope deviation 0.0266), while the residual chi-squared skewness remained degree-of-freedom dependent. Mapping the full chi-squared reference through its CDF produced the closest agreement across conditioning sizes ([Figure 1](#correction-progression), D; normalized mean RMSE $0.0116\\pm0.0060$, normalized slope deviation 0.0393).\n\nThese panels illustrate the different roles of the operations. First-order plugin centering addresses the leading null location, DN aligns location and scale, and the CDF transform expresses each table’s chi-squared tail coordinate on the same standard-normal reference. Only the final step addresses the full reference shape. These operations have direct antecedents in the finite-sample information and analytic MI-significance literature [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1016/S0167-2789(97)00117-6) [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272); the simulations here validate their implementation and operating regime in the present CMI setting rather than the classical identities themselves.\n\n::::::figure{#correction-progression align=\"center\" type=\"image\" label=\"Figure 1\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_fa6fdb65_ebb0334966a5\" width=\"100%\" alt=\"Quantile–quantile plots comparing the distributions of raw and various versions of corrected conditional mutual information demonstrating that the corrections described make CMI of differing degrees of freedom comparable.\"}\n:::::\n\n:::::figcaption\n**Figure 1.** **Progression from raw information to a common null-evidence coordinates across conditioning dimensions.**\n\nThe panels compare CMI-null realizations as the number of conditioning strata changes.\n\n(A) Raw plugin information.\n\n(B) First-order plugin correction (labeled “Basharin”).\n\n(C) First and second order correction (dimensionality normalization, $Z_{\\mathrm{DN}}$).\n\n(D) Chi-squared-CDF normal-equivalent score, $Z_{\\mathrm{CDF}}$. Simulations used $N=100{,}000$ and $3\\times2\\times k_z$ tables with $k_z\\in\\lbrace 10,30,50,70,100\\rbrace$. The final panel should be interpreted as alignment of evidence under the respective nulls, not equality of CMI effect magnitudes across differing $k_z$ dimensions.\n:::::\n\n::::::\n\n## The scaled plugin statistic follows the predicted chi-squared null\n\nAcross changes in conditioning cardinality, the empirical plugin statistic agreed with the scaled chi-squared reference and with the directly calculated likelihood-ratio statistic ([Figure 2](#chi-k)). All reported KS comparisons in this dense-table experiment had $p>0.1$. The same pattern held across changes in sample size at fixed $k_z=6$ ([Figure 3](#chi-n)). These results reproduce the expected consequence of the classical Kullback representation and Wilks calibration: The information estimate and $G^2$ are algebraically the same statistic in different units, while the quality of the chi-squared approximation depends on the sampling regime [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://doi.org/10.1214/aoms/1177732360) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635).\n\n::::::figure{#chi-k align=\"center\" type=\"image\" label=\"Figure 2\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_b5e71833_93807469602b\" width=\"100%\" alt=\"Quartile–quartile plots and real-vs-theoretical (chi-square) distribution plots demonstrating real and theoretical (chi-squared) distributions of MI match across differing degrees of freedom.\"}\n:::::\n\n:::::figcaption\n**Figure 2.** **The plugin CMI statistic follows its scaled chi-squared reference across conditioning-table sizes.**\n\nAt left, each row shows Empirical vs. theoretical quantiles. At right, each row shows Empirical densities and the corresponding scaled chi-squared densities. $N=100{,}000$; table size $3\\times2\\times k_z$; $k_z\\in\\lbrace 6,10,30,100\\rbrace$.\n\n(A) $k_z$ = 6\n\n(B) $k_z$ = 10\n\n(C)$k_z$ = 30\n\n(D) $k_z$ = 100\n:::::\n\n::::::\n\n::::::figure{#chi-n align=\"center\" type=\"image\" label=\"Figure 3\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_f4ca7263_ebb0334966a5\" width=\"100%\" alt=\"Quantile–quantile plots and real-vs-theoretical (chi-square) distribution plots demonstrating real and theoretical (chi-square) distributions of MI match across differing sample sizes.\"}\n:::::\n\n:::::figcaption\n**Figure 3.** **The plugin CMI statistic follows its scaled chi-squared reference across sample sizes.**\n\nAt left, each row shows empirical vs. theoretical quantiles. At right, each row shows empirical densities and the corresponding scaled chi-squared densities.Table size $3\\times2\\times6$; $N\\in\\lbrace 1{,}000,5{,}000,20{,}000,100{,}000\\rbrace$.\n\n(A) $n=$ 1,000\n\n(B) $n=$ 5,000\n\n(C) $n=$ 20,000\n\n(D) $n=$ 100,000\n:::::\n\n::::::\n\n## Empirical variance and degrees of freedom match the framework\n\nThe empirical information-scale standard-deviation constant for the $3\\times2$ table was 0.588, compared with the theoretical value 0.589 from the equation above ([Figure 4](#variance-df), A). Across conditioning cardinalities, degrees of freedom estimated from the mean, variance, and maximum-likelihood chi-squared fit closely tracked the theoretical value $k_z(k_x-1)(k_y-1)$ ([Figure 4](#variance-df), B). The regression of empirical on theoretical degrees of freedom had slope 1.017 and correlation $r=1.00$; the reported regressions had $p<0.001$.\n\nThis experiment is best viewed as a numerical check on the null-moment consequences of the framework rather than an independent derivation of the variance correction. Once the chi-squared reference is accepted, both the variance and the degree-of-freedom scaling follow analytically; related MI variance and error calculations have been derived by other routes [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272).\n\n::::::figure{#variance-df align=\"center\" type=\"image\" label=\"Figure 4\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_8b78de00_ebb0334966a5\" width=\"100%\" alt=\"Violin plots demonstrating the empirical and theoretical (chi-squared) standard deviation are comparable across differing degrees of freedom. Correlation plot demonstrating that the empirical and theoretical (chi-squared) degrees of freedom are comparable across differing contingency table structures for CMI.\"}\n:::::\n\n:::::figcaption\n**Figure 4.** **Empirical variance and degrees of freedom agree with the theoretical chi-squared moments.**\n\n(A) Repeated estimates of the information-scale standard-deviation constant across $k_z\\in\\lbrace 6,10,20,50,100,200\\rbrace$ for $3\\times2\\times k_z$ tables at $N=100{,}000$. The empirical mean (0.588, solid black line) and theoretical value (0.589, dashed salmon line) nearly coincide. Distributional medians indicated by white bars.\n\n(B) Empirical degrees of freedom estimated from the mean, variance, and maximum-likelihood chi-squared fit vs. $k_z(k_x-1)(k_y-1)$. MLE: Maximum likelihood estimation.\n:::::\n\n::::::\n\n## The CDF representation is normal when the chi-squared input is adequate\n\nThe normal-equivalent score was close to $\\mathcal N(0,1)$ across most of the sample-size and table-dimension combinations examined ([Figure 5](#zcdf-normal)). Sparse combinations departed from the reference, most clearly at $k_z=100$ and $N=5{,}000$. This behavior is consistent with the framework: The CDF mapping handles chi-squared skewness at small or moderate degrees of freedom, but it can't correct a poor chi-squared approximation caused by sparse cells.\n\n::::::figure{#zcdf-normal align=\"center\" type=\"image\" label=\"Figure 5\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_7f686c12_93807469602b\" width=\"100%\" alt=\"Quartile–quartile plots demonstrating that the PIT corrected MI values are normally distributed across a broad range of sample sizes and degrees of freedom up to a limit.\"}\n:::::\n\n:::::figcaption\n**Figure 5.** **Normal-equivalent evidence scores across sample size and conditioning cardinality.**\n\nQQ plots compare empirical $Z_{\\mathrm{CDF}}$ values with $\\mathcal N(0,1)$. Columns vary $N\\in\\lbrace 5{,}000,20{,}000,100{,}000\\rbrace$; rows vary $k_z\\in\\lbrace 6,10,30,100\\rbrace$. Deviations concentrate in the sparsest combinations, where the chi-squared approximation to $G^2$ is least accurate.\n:::::\n\n::::::\n\nIn the balanced validity grid we examined here, departures from normal-equivalent calibration became common as average expected counts fell below approximately 18–20 per cell ([Figure 6](#validity)). This provides an empirical warning region for these table geometries, not a general operating boundary or a stand-alone rule for replacing permutation. The transition is specific to the null generators, table geometries, replicate counts, and KS diagnostic used here. Calibration can also depend on the minimum and distribution of expected counts, marginal imbalance, structural zeros, the target tail probability, and other departures from regular multinomial conditions, as established in the broader sparse-multinomial literature [](https://doi.org/10.1080/01621459.1978.10481567) [](https://doi.org/10.1080/01621459.1980.10477473) [](https://doi.org/10.1080/01621459.1986.10478294) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635).\n\n::::::figure{#validity align=\"center\" type=\"image\" label=\"Figure 6\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_4e516315_ebb0334966a5\" width=\"100%\" alt=\"Grid plots showing the relationship between cell sparsity (sample size vs degrees of freedom) and the accuracy of the chi-squared synthesis presented where corrections begin to be inaccurate at ~18 samples per cell.\"}\n:::::\n\n:::::figcaption\n**Figure 6.** **Sparse-cell boundaries for applicability of this paradigm.**\n\n(A) KS diagnostic for $Z_{\\mathrm{CDF}}$ across $N$ and $k_z$. Blue denotes $p>0.05$ and orange $p<0.05$.\n\n(B) Average expected count per table cell, $N/(k_xk_yk_z)$. The transition near 20 counts per cell is specific to this grid and marks deterioration of the chi-squared input approximation, not failure of the probability-integral transform itself.\n:::::\n\n::::::\n\nThe upper-tail comparison likewise showed progressive improvement from raw information to DN and then to the CDF representation ([Figure 7](#tails)). DN retained degree-of-freedom-dependent chi-squared tails, whereas the CDF representation followed the normal reference closely in the dense regimes examined.\n\n::::::figure{#tails align=\"center\" type=\"image\" label=\"Figure 7\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_e34cf878_ebb0334966a5\" width=\"100%\" alt=\"Quartile–quartile plots demonstrating the conformity of the tails of distributions to the theoretical expectations under differing degrees of freedom with differing corrections.\"}\n:::::\n\n:::::figcaption\n**Figure 7.** **Upper-tail behavior of raw, moment-standardized, and cumulative distribution function (CDF)-transformed statistics.**\n\nColumns show raw information, DN, and $Z_{\\mathrm{CDF}}$; rows vary conditioning cardinality. The CDF representation accounts for the full chi-squared reference shape, while DN retains residual skewness at lower degrees of freedom.\n:::::\n\n::::::\n\n## Analytic and permutation inference agree at the available resolution\n\nAcross the $2\\times2$ alternatives examined, analytic and permutation estimates of detection power were close ([Figure 8](#permutation), A). Per-table $p$-values agreed where the 200-permutation procedure had sufficient resolution ([Figure 8](#permutation), B). The horizontal floor reflects the finite resolution of the chosen number of permutations; below that floor, the comparison cannot validate the analytic tail probability. Exact and permutation-based MI tests provide important alternatives in regimes where asymptotic calibration is unsuitable [](https://doi.org/10.3390/e16052839) [](https://doi.org/10.2202/1544-6115.1585).\n\nThe experiment supports analytic likelihood-ratio calibration as a computational alternative in well-populated discrete tables, but it does not establish equivalence in genome-wide extreme tails. More permutations, importance sampling, or substantially larger null simulations would be required for direct empirical assessment at such thresholds.\n\n::::::figure{#permutation align=\"center\" type=\"image\" label=\"Figure 8\"}\n\n:::::image{src=\"https://thestacks-01.s3.amazonaws.com/publications/method-arcadia-info-standardization/media_5f4b6247_a5d6c390d180\" width=\"94%\" alt=\"Paired plots showing the relationship between the parametric statistics provided by our paradigm and permutation-based statistics with regard to the power of detection and the p-values.\"}\n:::::\n\n:::::figcaption\n**Figure 8.** **Analytic chi-squared and permutation inference in $2\\times2$ alternatives.**\n\n(A) Detection power from the permutation test vs. the analytic likelihood-ratio test across sample sizes and target MI values.\n\n(B) Per-table permutation and analytic $p$-values. The dotted horizontal line marks the permutation-resolution floor. Only panels A and B of the original figure are shown because they compare like inferential quantities.\n:::::\n\n::::::\n\n## Summary of the numerical evidence\n\nTogether, the simulations support four conclusions within the regimes we examined. First, the plugin MI and CMI statistics are numerically identical to the corresponding likelihood-ratio deviances after unit conversion. Second, their null means, variances, and degrees of freedom follow the chi-squared predictions in dense tables. Third, the normal-equivalent CDF representation aligns heterogeneous reference distributions when the underlying chi-squared approximation is adequate. Fourth, sparse cells — rather than small degrees of freedom alone — are the principal finite-sample failure mode observed here.\n\n# LLM usage\n\nWe used Claude (Opus 4.8) to help write, clean up, comment, and review our code. We also used Claude to write text, suggest wording ideas, and help clarify and streamline text. Claude also suggested papers on relevant science, we did further reading, and we cited some of this literature. We also used ChatGPT (GPT-5.6 Sol) to review our code. We reviewed all AI-assisted content and take responsibility for its accuracy and integrity.\n\n# Discussion\n\n## A framework built from established results\n\nThis work combines two classical facts. Kullback’s formulation identifies the relevant plugin information statistic as a multinomial likelihood-ratio deviance in information units, and Wilks' theorem supplies its asymptotic chi-squared null [](https://books.google.com/books?id=XeRQAAAAMAAJ) [](https://doi.org/10.1214/aoms/1177732360). The synthesis is simple:\n\n$$\n\n\\text{null model} \\longrightarrow D_{\\mathrm{KL},2}(\\widehat p\\Vert \\widehat q_0) \\longrightarrow G^2 \\longrightarrow \\chi^2_{\\nu} \\longrightarrow p.\n\n$$\n\nThe mathematical ingredients aren't new. Information-based contingency-table tests, analytic MI significance transformations, KL-based hypothesis identities, and categorical CMI tests all predate this work [](https://doi.org/10.1007/BF02289159) [](https://doi.org/10.1016/S0167-2789(97)00117-6) [](https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html) [](https://jmlr.org/papers/v7/decampos06a.html) [](https://jmlr.org/papers/v22/19-600.html). The contribution is to make their combined implications explicit across a broader family of information-theoretic measures, to distinguish effect magnitude from calibrated evidence, and to validate the resulting workflow in the heterogeneous MI/CMI regimes that motivate large-scale application. The adjacent information-geometric treatment of measured KL divergence by Goehle [](https://doi.org/10.1007/s41884-025-00170-7) provides additional evidence that chi-squared structure recurs broadly, although its probabilistic construction differs from the saturated-versus-null multinomial setting used here.\n\nThe framework is driven by the scientific null rather than by the name of the information measure. Independence yields MI; conditional independence yields CMI; a fixed reference yields ordinary KL goodness of fit; uniformity yields an entropy-deficit test; and stratum-wise uniformity yields a conditional entropy-deficit test. The same inferential engine applies because each statistic is the KL term in a saturated-versus-null multinomial likelihood ratio.\n\n## Bias correction and likelihood-ratio centering\n\nThe connection between first-order plugin correction and likelihood-ratio centering is a conceptual bridge between information estimation and hypothesis testing. The term $\\nu/(2N\\ln2)$ is familiar from entropy, transmitted-information, and conditional-information bias expansions [](https://api.semanticscholar.org/CorpusID:125662170) [](https://doi.org/10.1137/1104033) [](https://doi.org/10.1080/0954898X.1996.11978656) [](https://doi.org/10.1016/S0167-2789(98)00269-3) [](https://doi.org/10.1162/089976603321780272). The Kullback–Wilks synthesis shows that, under the null, the same term is the asymptotic chi-squared mean expressed in bits. Thus,\n\n$$\n\n\\widehat {\\mathrm{D}}_{\\mathrm{BC}} =\\widehat {\\mathrm{D}}-\\frac{\\nu}{2N\\ln2} =\\frac{G^2-\\nu}{2N\\ln2}.\n\n$$\n\nThe distinction between estimator correction and significance calibration remains important. Subtracting the null mean doesn't produce a $p$-value, because the variance, skewness, and higher moments still depend on $\\nu$. DN removes the first two moment dependencies, but the full chi-squared survival function is required for calibrated likelihood-ratio inference. The CDF-derived normal score is merely a coordinate transformation of that $p$-value.\n\n## Effect magnitude, evidence, and heterogeneous analyses\n\nThe framework gives a familiar two-output structure. Corrected information in bits describes the effect magnitude, whereas the likelihood-ratio $p$-value describes evidence. In a genetic analysis, this is analogous in role — though not in interpretation — to reporting both a regression coefficient and its inferential statistic. A locus can have a large estimated information effect, but weak evidence because it's rare or poorly sampled, or a very small information effect with overwhelming evidence in a large cohort.\n\nProperly calibrated $p$-values already provide a common null interpretation across tests with different $N$ and degrees of freedom. The value of $Z_{\\mathrm{CDF}}$ is representational: It expresses the same tail probability on a standard-normal coordinate useful for Manhattan-style plots, evidence ranking, or downstream methods that operate on $Z$ scores. Significance thresholds should still be defined in $p$-value space, including Bonferroni, false-discovery-rate, or other multiplicity procedures, and then transformed to $Z$ only for display if desired [](https://doi.org/10.1111/j.2517-6161.1995.tb02031.x).\n\nMI and CMI evidence scores can be placed on this common coordinate, but they test different null hypotheses. A stronger conditional score than marginal score means that the conditional-independence test provides stronger standardized evidence against its own null; it is not, by itself, a test that the conditioning variable modifies the locus effect. Formal inference for\n\n$$\n\nI(X;Y)-I(X;Y\\mid Z)\n\n$$\n\nrequires the joint sampling distribution and covariance of the two estimates. Interaction information originated as a signed multivariate information quantity rather than a single KL divergence to one null model [](https://doi.org/10.1007/BF02289159); its calibration is deliberately deferred to separate work.\n\n## Relationship to regression-based association testing\n\nInformation-theoretic association tests should be viewed as complementary to regression, not as a generic replacement. A typical one-degree-of-freedom additive regression test concentrates power on a specific alternative and will generally be more efficient when that alternative is correct. MI and CMI distribute power across a larger omnibus alternative that includes additive, dominant, recessive, multiallelic, nonmonotone, and context-dependent patterns. Their advantage is sensitivity to unspecified discrete dependence, not uniformly greater power [](https://doi.org/10.1038/nrg1916) [](https://doi.org/10.1038/nrg2579) [](https://doi.org/10.1038/s43586-021-00056-9).\n\nThe inferential contribution in genetics is therefore infrastructural. The framework gives information-theoretic analyses the same broad workflow used in regression-based genome-wide studies:\n\n$$\n\n\\begin{aligned} \\text{effect estimate} &\\longrightarrow \\text{test statistic} \\longrightarrow \\text{reference distribution}\\cr &\\longrightarrow p \\longrightarrow \\text{multiplicity control and visualization}. \\end{aligned}\n\n$$\n\nThis permits locus-specific changes in sample size, observed support, and conditioning cardinality to be handled through the appropriate test-specific null rather than through repeated locus-specific permutation when Wilks' approximation is adequate.\n\n## Conditional-independence testing\n\nFor discrete data, CMI is exactly the stratified likelihood-ratio statistic in information units. The formula\n\n$$\n\n\\nu_{\\mathrm{CMI}} =\\sum_z(k_{x,z}-1)(k_{y,z}-1)\n\n$$\n\nmakes the effect of changing conditioning state spaces explicit. A fixed raw-CMI threshold is therefore not a fixed significance threshold as the conditioning set changes. A fixed significance level based on the correct $\\chi^2$ reference is calibrated in the ordinary asymptotic sense. Both the classical log-linear literature and modern machine-learning work treat this categorical conditional-independence test as established [](https://doi.org/10.1111/j.2517-6161.1969.tb00808.x) [](https://jmlr.org/papers/v7/decampos06a.html) [](https://jmlr.org/papers/v22/19-600.html) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635).\n\nThis observation doesn't imply that all existing causal-discovery implementations are miscalibrated: Procedures that calculate a valid test-specific $p$-value or use an appropriate resampling null already address the changing reference distribution. The present result supplies a transparent and inexpensive analytic route for the discrete plugin estimator. Nearest-neighbor and other continuous CMI estimators require estimator-specific null theory or resampling and are outside the framework developed here [](https://proceedings.mlr.press/v84/runge18a.html) [](https://proceedings.mlr.press/v124/runge20a.html) [](https://proceedings.mlr.press/v89/marx19a.html).\n\n## Entropy and conditional entropy\n\nThe entropy result is best interpreted as a test of maximum entropy on a fixed support. Because\n\n$$\n\n\\log_2 k-\\widehat {\\mathrm{H}}(X)=D_{\\mathrm{KL},2}(\\widehat p\\Vert u),\n\n$$\n\nthe ordinary multinomial likelihood-ratio test of uniformity is also a parametric test that entropy is below its support-constrained maximum. The conditional extension tests uniformity within every stratum. Neither is a general test that entropy equals an arbitrary value. Entropy and diversity inference predates this formulation [](https://doi.org/10.1016/0022-5193(70)90124-4) [](https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635) [](https://doi.org/10.1007/978-1-4612-4578-0); the contribution here is to place the maximum-entropy and conditional maximum-entropy cases explicitly inside the same likelihood-ratio framework as MI and CMI.\n\nThis formulation can nevertheless be useful when maximum uncertainty is scientifically meaningful, and it clarifies the Miller–Madow sign. The likelihood-ratio statistic is built from entropy deficit; subtracting its asymptotic null mean corresponds to adding the familiar correction to entropy itself.\n\n## Limitations\n\nThe central limitation is asymptotic calibration. The algebraic conversion from information to deviance is exact, but the chi-squared null is not. In the grids examined, sparse cells were the principal observed failure mode, and departures became common near 18–20 average observations per cell. That transition is a simulation-specific warning region, not a universal threshold. However, it is likely that this represents the boundary where Wilks’ condition 4 is violated and there are, in the context of two-dimensional contingency tables, known corrections for the small sample bias that results from sparsity. Here, where we work with both two- and three-dimensional tables, application of these other methods to three-dimensional tables would require extension and is, thus, beyond the scope of this work. Thus, calibration depends on the full expected-count distribution, marginal imbalance, the presence of structural or sampling zeros, and the tail probability of interest [](https://doi.org/10.1080/01621459.1978.10481567) [](https://doi.org/10.1080/01621459.1980.10477473) [](https://doi.org/10.1080/01621459.1986.10478294).\n\nThe ordinary Wilks regime also assumes fixed model dimension and prespecified categories. If bins, genotype groupings, or conditioning strata are selected because they maximize the observed information, the nominal degrees of freedom don't account for selection, and the test is anticonservative. Parameters on boundaries and other nonstandard likelihood settings require different asymptotic treatment [](https://doi.org/10.1080/01621459.1987.10478472). Likewise, state spaces that grow rapidly with $N$ fall outside fixed-dimensional theory.\n\nThe framework is restricted to KL terms that are exact multinomial likelihood-ratio deviances. Other smooth divergence statistics can have chi-squared asymptotics after appropriate scaling, but they don't share the exact Kullback identity used here and aren't automatically covered; the Cressie–Read power-divergence family is the standard example [](https://doi.org/10.1111/j.2517-6161.1984.tb01318.x) [](https://doi.org/10.1007/978-1-4612-4578-0). Arbitrary KL divergence between two fitted empirical distributions is also not sufficient: The reference distribution must be the null-constrained likelihood fit for the same test.\n\nFinally, the simulations demonstrate agreement over accessible probability ranges, not direct calibration at extremely small genome-wide thresholds. Analytic extrapolation to those tails rests on the adequacy of the likelihood-ratio approximation. Representative large-scale null simulations, higher-order corrections, or specialized tail methods would strengthen application-specific validation.\n\n## Extensions\n\nSeveral familiar quantities satisfy the same inclusion criterion. Total correlation is the KL divergence between a multivariate joint distribution and the product of its marginals and therefore yields an omnibus test of mutual independence [](https://doi.org/10.1147/rd.41.0066). Discrete transfer entropy is a form of conditional mutual information and inherits the corresponding stratified test when the temporal states and conditioning structure are fixed in advance [](https://doi.org/10.1103/PhysRevLett.85.461).\n\nOther quantities fall outside the direct framework. Interaction information is signed and cannot be a single KL divergence [](https://doi.org/10.1007/BF02289159). Directed information measures that sum multiple dependent CMI terms may require joint covariance or a larger explicitly nested likelihood comparison [](https://scispace.com/pdf/causality-feedback-and-directed-information-3emhzj969g.pdf). Developing parametric inference for these composite statistics is a natural next step.\n\n## Conclusion\n\nKullback provides the information-theoretic representation; Wilks provides the asymptotic calibration. Connecting them shows that a family of discrete information measures can be analyzed through one likelihood-ratio workflow. The resulting framework provides a concise explanation for plugin null moments, a likelihood-ratio interpretation of classical first-order corrections, a common treatment of entropy-based nulls, and a practical route from information estimates to calibrated evidence in large heterogeneous analyses.\n\n::::::bibtex\n@article{deCampos2006,\\\ntitle = {A Scoring Function for Learning Bayesian Networks Based on Mutual Information and Conditional Independence Tests},\\\nauthor = {de Campos, Luis M.},\\\njournal = {Journal of Machine Learning Research},\\\nvolume = {7},\\\npages = {2149--2187},\\\nyear = {2006},\\\nurl = {https://jmlr.org/papers/v7/decampos06a.html}\\\n}\n::::::\n\n::::::bibtex\n@article{kubkowski2021,\\\ntitle = {How to Gain on Power: Novel Conditional Independence Tests Based on Short Expansion of Conditional Mutual Information},\\\nauthor = {Kubkowski, Mariusz and Mielniczuk, Jan and Teisseyre, Pawe{\\l}},\\\njournal = {Journal of Machine Learning Research},\\\nvolume = {22},\\\nnumber = {62},\\\npages = {1--57},\\\nyear = {2021},\\\nurl = {https://jmlr.org/papers/v22/19-600.html}\\\n}\n::::::\n\n::::::bibtex\n@inproceedings{runge2020,\\\ntitle = {Discovering Contemporaneous and Lagged Causal Relations in Autocorrelated Nonlinear Time Series Datasets},\\\nauthor = {Runge, Jakob},\\\nbooktitle = {Proceedings of the 36th Conference on Uncertainty in Artificial Intelligence},\\\nseries = {Proceedings of Machine Learning Research},\\\nvolume = {124},\\\npages = {1388--1397},\\\nyear = {2020},\\\nurl = {https://proceedings.mlr.press/v124/runge20a.html}\\\n}\n::::::\n\n::::::bibtex\n@inproceedings{runge2018,\\\ntitle = {Conditional Independence Testing Based on a Nearest-Neighbor Estimator of Conditional Mutual Information},\\\nauthor = {Runge, Jakob},\\\nbooktitle = {Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics},\\\nseries = {Proceedings of Machine Learning Research},\\\nvolume = {84},\\\npages = {938--947},\\\nyear = {2018},\\\nurl = {https://proceedings.mlr.press/v84/runge18a.html}\\\n}\n::::::\n\n::::::bibtex\n@incollection{miller1955,\\\ntitle = {Note on the Bias of Information Estimates},\\\nauthor = {Miller, George A.},\\\neditor = {Quastler, Henry},\\\nbooktitle = {Information Theory in Psychology: Problems and Methods},\\\npublisher = {Free Press},\\\naddress = {Glencoe, Illinois},\\\npages = {95--100},\\\nyear = {1955},\n\nurl = {https://api.semanticscholar.org/CorpusID:125662170}\\\n}\n::::::\n\n::::::bibtex\n@inproceedings{marxVreeken2019,\\\ntitle = {Testing Conditional Independence on Discrete Data Using Stochastic Complexity},\\\nauthor = {Marx, Alexander and Vreeken, Jilles},\\\nbooktitle = {Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics},\\\nseries = {Proceedings of Machine Learning Research},\\\nvolume = {89},\\\npages = {496--505},\\\nyear = {2019},\\\nurl = {https://proceedings.mlr.press/v89/marx19a.html}\\\n}\n::::::\n\n::::::bibtex\n@misc{kullback1959,\\\ntitle = {Information Theory and Statistics},\\\nauthor = {Kullback, Solomon},\\\nyear = {1959},\n\nurl = {https://books.google.com/books?id=XeRQAAAAMAAJ}\\\n}\n::::::\n\n::::::bibtex\n@book{agresti2013,\\\ntitle = {Categorical Data Analysis},\\\nauthor = {Agresti, Alan},\\\npublisher = {John Wiley & Sons},\\\naddress = {Hoboken, New Jersey},\\\nedition = {3},\\\nyear = {2013},\\\nurl = {https://www.wiley.com/en-us/Categorical+Data+Analysis%2C+3rd+Edition-p-9780470463635}\\\n}\n::::::\n\n::::::bibtex\n@article{cheng2008,\\\ntitle = {Information Identities and Testing Hypotheses: Power Analysis for Contingency Tables},\\\nauthor = {Cheng, Philip E. and Liou, Michelle and Aston, John A. D. and Tsai, Arthur C.},\\\njournal = {Statistica Sinica},\\\nvolume = {18},\\\nnumber = {2},\\\npages = {535--558},\\\nyear = {2008},\\\nurl = {https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html}\\\n}\n::::::\n\n::::::bibtex\n@article{vinh2010,\\\ntitle = {Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance},\\\nauthor = {Vinh, Nguyen Xuan and Epps, Julien and Bailey, James},\\\njournal = {Journal of Machine Learning Research},\\\nvolume = {11},\\\npages = {2837--2854},\\\nyear = {2010},\\\nurl = {https://jmlr.org/papers/v11/vinh10a.html}\\\n}\n::::::\n\n::::::bibtex\n@inproceedings{romano2014,\\\ntitle = {Standardized Mutual Information for Clustering Comparisons: One Step Further in Adjustment for Chance},\\\nauthor = {Romano, Simone and Bailey, James and Nguyen, Vinh and Verspoor, Karin},\\\nbooktitle = {Proceedings of the 31st International Conference on Machine Learning},\\\nseries = {Proceedings of Machine Learning Research},\\\nvolume = {32},\\\nnumber = {2},\\\npages = {1143--1151},\\\nyear = {2014},\\\nurl = {https://proceedings.mlr.press/v32/romano14.html}\\\n}\n::::::\n\n::::::bibtex\n@inproceedings{massey1990,\\\ntitle = {Causality, Feedback and Directed Information},\\\nauthor = {Massey, James L.},\\\nbooktitle = {Proceedings of the International Symposium on Information Theory and Its Applications},\\\npages = {303--305},\\\nyear = {1990},\n\nurl = {https://scispace.com/pdf/causality-feedback-and-directed-information-3emhzj969g.pdf}\\\n}\n::::::","contributors":[{"user_id":6,"role":"Supervision","first_name":"Prachee","last_name":"Avasthi","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Prachee-Arcadia-headshot.png","in_byline":1,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":2407,"role":"Visualization","first_name":"Audrey","last_name":"Bell","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Audrey-Arcadia-headshot.png","in_byline":0,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4225,"role":"Critical feedback","first_name":"Erin","last_name":"McGeever","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Erin-Arcadia-headshot.png","in_byline":0,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4216,"role":"Conceptualization","first_name":"David G.","last_name":"Mets","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Dave-Mets-Arcadia-headshot.png","in_byline":1,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4216,"role":"Experimental design","first_name":"David G.","last_name":"Mets","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Dave-Mets-Arcadia-headshot.png","in_byline":1,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4216,"role":"Formal analysis","first_name":"David G.","last_name":"Mets","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Dave-Mets-Arcadia-headshot.png","in_byline":1,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4216,"role":"Methodology","first_name":"David G.","last_name":"Mets","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Dave-Mets-Arcadia-headshot.png","in_byline":1,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4216,"role":"Software","first_name":"David G.","last_name":"Mets","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Dave-Mets-Arcadia-headshot.png","in_byline":1,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4216,"role":"Visualization","first_name":"David G.","last_name":"Mets","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Dave-Mets-Arcadia-headshot.png","in_byline":0,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4216,"role":"Writing","first_name":"David G.","last_name":"Mets","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Dave-Mets-Arcadia-headshot.png","in_byline":1,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]},{"user_id":4213,"role":"Validation","first_name":"Austin H.","last_name":"Patton","avatar":"https://thestacks-01.s3.amazonaws.com/users/avatars/Austin-Arcadia-headshot.png","in_byline":0,"priority":null,"affiliations":[{"org_id":2,"name":"Arcadia Science","slug":"arcadia-science","ror_id":"052zd4v68","avatar":"avatars/orgs/org_2_1766013565.jpg"}]}],"created_at":"2026-08-31T20:23:49.000Z","doi":"10.57844/arcadia-y5aa-6u5u","version_doi":null,"feedback_form_embed_src":"https://publishing-tools.arcadiascience.com/typeform/AacF8YKM?pub_name=rectUK1fyZTU6RNl2","id":177,"license":"CC BY","linked_assets":[{"type":"Code","url":"https://github.com/Arcadia-Science/arcadia-info-standardization/tree/v1.0.1","name":""}],"orgs":[{"org_id_int":2,"priority":null,"slug":"arcadia-science","name":"Arcadia Science","avatar":"avatars/orgs/org_2_1766013565.jpg"}],"pdf_url":"https://thestacks-01.s3.us-west-2.amazonaws.com/publications/method-arcadia-info-standardization/method-arcadia-info-standardization_v2.pdf","references":[{"url":"https://doi.org/10.1080/0954898x.1996.11978656","text":"Panzeri S, Treves A. (1996). Analytical estimates of limited sampling biases in different information measures."},{"url":"https://doi.org/10.1002/j.1538-7305.1948.tb01338.x","text":"Shannon CE. (1948). A Mathematical Theory of Communication."},{"url":"https://doi.org/10.1016/s0167-2789(98)00269-3","text":"Roulston MS. (1999). Estimating the errors on measured entropy and mutual information."},{"url":"https://doi.org/10.1214/aoms/1177729694","text":"Kullback S, Leibler RA. (1951). On Information and Sufficiency."},{"url":"https://doi.org/10.1162/089976603321780272","text":"Paninski L. (2003). Estimation of Entropy and Mutual Information."},{"url":"https://doi.org/10.1002/047174882x","text":"Cover TM, Thomas JA. (2005). Elements of Information Theory."},{"url":"https://doi.org/10.3390/e16052839","text":"Pethel S, Hahs D. (2014). Exact Test of Independence Using Mutual Information."},{"url":"https://doi.org/10.1038/nrg1916","text":"Balding DJ. (2006). A tutorial on statistical methods for population association studies."},{"url":"https://proceedings.mlr.press/v89/marx19a.html","text":"Marx A., Vreeken J. (2019). Testing Conditional Independence on Discrete Data Using Stochastic Complexity."},{"url":"https://doi.org/10.1038/nrg2579","text":"Cordell HJ. (2009). Detecting gene–gene interactions that underlie human diseases."},{"url":"https://doi.org/10.2202/1544-6115.1585","text":"Phipson B, Smyth GK. (2010). Permutation P-values Should Never Be Zero: Calculating Exact P-values When Permutations Are Randomly Drawn."},{"url":"https://doi.org/10.1038/s43586-021-00056-9","text":"Uffelmann E, Huang QQ, Munung NS, de Vries J, Okada Y, Martin AR, Martin HC, Lappalainen T, Posthuma D. (2021). Genome-wide association studies."},{"url":"https://books.google.com/books?id=xerqaaaamaaj","text":"Kullback S. (1959). Information Theory and Statistics."},{"url":"https://doi.org/10.1086/519795","text":"Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, Maller J, Sklar P, de Bakker PI, Daly MJ, Sham PC. (2007). PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses."},{"url":"https://www.wiley.com/en-us/categorical+data+analysis%2c+3rd+edition-p-9780470463635","text":"Agresti A. (2013). Categorical Data Analysis."},{"url":"https://doi.org/10.7551/mitpress/1754.001.0001","text":"Spirtes P, Glymour C, Scheines R. (2001). Causation, Prediction, and Search."},{"url":"https://doi.org/10.1007/978-1-4612-4578-0","text":"Read TRC, Cressie NAC. (1988). Goodness-of-Fit Statistics for Discrete Multivariate Data."},{"url":"https://jmlr.org/papers/v7/decampos06a.html","text":"de Campos L. M. (2006). A Scoring Function for Learning Bayesian Networks Based on Mutual Information and Conditional Independence Tests."},{"url":"https://www3.stat.sinica.edu.tw/statistica/j18n2/j18n27/j18n27.html","text":"Cheng P. E., Liou M., Aston J. A. D., Tsai A. C. (2008). Information Identities and Testing Hypotheses: Power Analysis for Contingency Tables."},{"url":"https://jmlr.org/papers/v22/19-600.html","text":"Kubkowski M., Mielniczuk J., Teisseyre P. (2021). How to Gain on Power: Novel Conditional Independence Tests Based on Short Expansion of Conditional Mutual Information."},{"url":"https://doi.org/10.1214/aoms/1177732360","text":"Wilks SS. (1938). The Large-Sample Distribution of the Likelihood Ratio for Testing Composite Hypotheses."},{"url":"https://proceedings.mlr.press/v124/runge20a.html","text":"Runge J. (2020). Discovering Contemporaneous and Lagged Causal Relations in Autocorrelated Nonlinear Time Series Datasets."},{"url":"https://doi.org/10.1007/bf02289159","text":"McGill WJ. (1954). Multivariate Information Transmission."},{"url":"https://proceedings.mlr.press/v84/runge18a.html","text":"Runge J. (2018). Conditional Independence Testing Based on a Nearest-Neighbor Estimator of Conditional Mutual Information."},{"url":"https://doi.org/10.1111/j.2517-6161.1969.tb00808.x","text":"Goodman LA. (1969). On Partitioning χ2 and Detecting Partial Association in Three-Way Contingency Tables."},{"url":"https://api.semanticscholar.org/corpusid:125662170","text":"Miller G. A. (1955). Note on the Bias of Information Estimates."},{"url":"https://doi.org/10.1016/s0167-2789(97)00117-6","text":"Roulston MS. (1997). Significance testing of information theoretic functionals."},{"url":"https://doi.org/10.1137/1104033","text":"Basharin GP. (1959). On a Statistical Estimate for the Entropy of a Sequence of Independent Random Variables."},{"url":"https://jmlr.org/papers/v11/vinh10a.html","text":"Vinh N. X., Epps J., Bailey J. (2010). Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance."},{"url":"https://doi.org/10.1162/neco.1995.7.2.399","text":"Treves A, Panzeri S. (1995). The Upward Bias in Measures of Information Derived from Limited Data Samples."},{"url":"https://proceedings.mlr.press/v32/romano14.html","text":"Romano S., Bailey J., Nguyen V., Verspoor K. (2014). Standardized Mutual Information for Clustering Comparisons: One Step Further in Adjustment for Chance."},{"url":"https://doi.org/10.1007/s41884-025-00170-7","text":"Goehle G. (2025). Approximation of the measured Kullback–Leibler divergence as a generalized chi-squared random variable."},{"url":"https://doi.org/10.1080/01621459.1987.10478472","text":"Self SG, Liang K. (1987). Asymptotic Properties of Maximum Likelihood Estimators and Likelihood Ratio Tests under Nonstandard Conditions."},{"url":"https://doi.org/10.1080/01621459.1978.10481567","text":"Larntz K. (1978). Small-Sample Comparisons of Exact Levels for Chi-Squared Goodness-of-Fit Statistics."},{"url":"https://doi.org/10.1080/01621459.1980.10477473","text":"Koehler KJ, Larntz K. (1980). An Empirical Investigation of Goodness-of-Fit Statistics for Sparse Multinomials."},{"url":"https://doi.org/10.1080/01621459.1986.10478294","text":"Koehler KJ. (1986). Goodness-of-Fit Tests for Log-Linear Models in Sparse Contingency Tables."},{"url":"https://doi.org/10.1111/j.2517-6161.1984.tb01318.x","text":"Cressie N, Read TR. (1984). Multinomial Goodness-Of-Fit Tests."},{"url":"https://doi.org/10.1016/0022-5193(70)90124-4","text":"Hutcheson K. (1970). A test for comparing diversities based on the shannon formula."},{"url":"https://doi.org/10.1038/s41586-020-2649-2","text":"Harris CR, Millman KJ, van der Walt SJ, Gommers R, Virtanen P, Cournapeau D, Wieser E, Taylor J, Berg S, Smith NJ, Kern R, Picus M, Hoyer S, van Kerkwijk MH, Brett M, Haldane A, del Río JF, Wiebe M, Peterson P, Gérard-Marchant P, Sheppard K, Reddy T, Weckesser W, Abbasi H, Gohlke C, Oliphant TE. (2020). Array programming with NumPy."},{"url":"https://doi.org/10.1038/s41592-019-0686-2","text":"Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, Burovski E, Peterson P, Weckesser W, Bright J, van der Walt SJ, Brett M, Wilson J, Millman KJ, Mayorov N, Nelson ARJ, Jones E, Kern R, Larson E, Carey CJ, Polat İ, Feng Y, Moore EW, VanderPlas J, Laxalde D, Perktold J, Cimrman R, Henriksen I, Quintero EA, Harris CR, Archibald AM, Ribeiro AH, Pedregosa F, van Mulbregt P, Contributors S1, Vijaykumar A, Bardelli AP, Rothberg A, Hilboll A, Kloeckner A, Scopatz A, Lee A, Rokem A, Woods CN, Fulton C, Masson C, Häggström C, Fitzgerald C, Nicholson DA, Hagen DR, Pasechnik DV, Olivetti E, Martin E, Wieser E, Silva F, Lenders F, Wilhelm F, Young G, Price GA, Ingold G-L, Allen GE, Lee GR, Audren H, Probst I, Dietrich JP, Silterra J, Webber JT, Slavič J, Nothman J, Buchner J, Kulick J, Schönberger JL, de Miranda Cardoso JV, Reimer J, Harrington J, Rodríguez JLC, Nunez-Iglesias J, Kuczynski J, Tritz K, Thoma M, Newville M, Kümmerer M, Bolingbroke M, Tartre M, Pak M, Smith NJ, Nowaczyk N, Shebanov N, Pavlyk O, Brodtkorb PA, Lee P, McGibbon RT, Feldbauer R, Lewis S, Tygier S, Sievert S, Vigna S, Peterson S, More S, Pudlik T, Oshima T, Pingel TJ, Robitaille TP, Spura T, Jones TR, Cera T, Leslie T, Zito T, Krauss T, Upadhyay U, Halchenko YO, Vázquez-Baeza Y. (2020). SciPy 1.0: fundamental algorithms for scientific computing in Python."},{"url":"https://doi.org/10.1109/mcse.2007.55","text":"Hunter JD. (2007). Matplotlib: A 2D Graphics Environment."},{"url":"https://doi.org/10.1111/j.2517-6161.1995.tb02031.x","text":"Benjamini Y, Hochberg Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing."},{"url":"https://doi.org/10.1147/rd.41.0066","text":"Watanabe S. (1960). Information Theoretical Analysis of Multivariate Correlation."},{"url":"https://doi.org/10.1103/physrevlett.85.461","text":"Schreiber T. (2000). Measuring Information Transfer."},{"url":"https://scispace.com/pdf/causality-feedback-and-directed-information-3emhzj969g.pdf","text":"Massey J. L. (1990). Causality, Feedback and Directed Information."}],"slug":"method-arcadia-info-standardization","social_posts_count":null,"social_posts_embed_src":"https://publishing-tools.arcadiascience.com/twitter/104432725709568164","state":"PUBLISHED","subtitle":"We've developed a framework bridging information theory and frequentist statistics. This synthesis of classical results maps any discrete estimator that can be represented as a Kullback–Leibler divergence to a G-test, providing parametric, scalable evidence tests.","tags":["method","feedback requested"],"title":"A likelihood-ratio framework for inference with discrete information-theoretic measures","version_desc":"Fixed numbering error in propositions: proposition 4 to proposition 3, proposition 5 to proposition 4","version_number":2,"versions":[{"id":493,"version_number":2,"version_desc":"Fixed numbering error in propositions: proposition 4 to proposition 3, proposition 5 to proposition 4","doi":null,"created_at":"2026-09-02T21:49:26.000Z"},{"id":486,"version_number":1,"version_desc":null,"doi":null,"created_at":"2026-08-31T20:23:49.000Z"}]}}