anderson#
- scipy.stats.anderson(x, dist='norm', *, method='interpolate')[source]#
Anderson-Darling test for data coming from a particular distribution.
The Anderson-Darling test tests the null hypothesis that a sample is drawn from a population that follows a particular distribution. For the Anderson-Darling test, the critical values depend on which distribution is being tested against. This function works for normal, exponential, logistic, weibull_min, or Gumbel (Extreme Value Type I) distributions.
- Parameters:
- xarray_like
Array of sample data.
- dist{‘norm’, ‘expon’, ‘logistic’, ‘gumbel’, ‘gumbel_l’, ‘gumbel_r’, ‘extreme1’, ‘weibull_min’}, optional
The type of distribution to test against. The default is ‘norm’. The names ‘extreme1’, ‘gumbel_l’ and ‘gumbel’ are synonyms for the same distribution.
- methodstr or instance of
MonteCarloMethod Defines the method used to compute the p-value. If method is
"interpolated", the p-value is interpolated from pre-calculated tables (without extrapolating). If method is an instance ofMonteCarloMethod, the p-value is computed usingscipy.stats.monte_carlo_testwith the provided configuration options and other appropriate settings.
- Returns:
- resultAndersonResult
If method is provided, this is an object with the following attributes:
- statisticfloat
The Anderson-Darling test statistic.
- pvalue: float
The p-value corresponding with the test statistic, calculated according to the specified method.
See also
kstestThe Kolmogorov-Smirnov test for goodness-of-fit.
Notes
For
weibull_min, maximum likelihood estimation is known to be challenging. If the test returns successfully, then the first order conditions for a maximum likelihood estimate have been verified and the critical values correspond relatively well to the significance levels, provided that the sample is sufficiently large (>10 observations [7]). However, for some data - especially data with no left tail -andersonis likely to result in an error message. In this case, consider performing a custom goodness of fit test usingscipy.stats.monte_carlo_test.References
[2]Stephens, M. A. (1974). EDF Statistics for Goodness of Fit and Some Comparisons, Journal of the American Statistical Association, Vol. 69, pp. 730-737.
[3]Stephens, M. A. (1976). Asymptotic Results for Goodness-of-Fit Statistics with Unknown Parameters, Annals of Statistics, Vol. 4, pp. 357-369.
[4]Stephens, M. A. (1977). Goodness of Fit for the Extreme Value Distribution, Biometrika, Vol. 64, pp. 583-588.
[5]Stephens, M. A. (1977). Goodness of Fit with Special Reference to Tests for Exponentiality , Technical Report No. 262, Department of Statistics, Stanford University, Stanford, CA.
[6]Stephens, M. A. (1979). Tests of Fit for the Logistic Distribution Based on the Empirical Distribution Function, Biometrika, Vol. 66, pp. 591-595.
[7]Richard A. Lockhart and Michael A. Stephens “Estimation and Tests of Fit for the Three-Parameter Weibull Distribution” Journal of the Royal Statistical Society.Series B(Methodological) Vol. 56, No. 3 (1994), pp. 491-500, Table 0.
[8]D’Agostino, Ralph B. (1986). “Tests for the Normal Distribution”. In: Goodness-of-Fit Techniques. Ed. by Ralph B. D’Agostino and Michael A. Stephens. New York: Marcel Dekker, pp. 122-141. ISBN: 0-8247-7487-6.
Examples
Test the null hypothesis that a random sample was drawn from a normal distribution (with unspecified mean and standard deviation).
>>> import numpy as np >>> from scipy.stats import anderson >>> rng = np.random.default_rng() >>> data = rng.random(size=35) >>> res = anderson(data, dist='norm', method='interpolate') >>> res.statistic np.float64(0.9887620209957291) >>> res.pvalue np.float64(0.012111200538380142)
The p-value is approximately 0.012,, so the null hypothesis may be rejected at a significance level of 2.5%, but not at a significance level of 1%.