PLAYGROUND

DISTRIBUTIONS / DISCRETE / LOGARITHMIC

Logarithmic distribution

What it represents

The Logarithmic distribution is a positive count law with probability decreasing like p^x/x. Despite its name, it comes from a logarithmic series and models group or species sizes.

Historical clue

Fisher, Corbet, and Williams introduced it in 1943 to describe species-abundance counts in ecological samples.

Relationships that clarify its use

It can arise by conditioning a compound Poisson model on a positive count under particular mixtures. It is related to abundance and overdispersed count families.

Data examples

  • number of individuals per species in ecology
  • group sizes, incidences, and positive counts with many small groups

Modelling warning

Do not confuse it with a continuous lognormal distribution or the logarithm of a count; they are different objects.

From a mathematical series to observed species

The normalizing constant comes from the series for -ln [1-p], hence the names Logarithmic and log-series. Fisher, Corbet, and Williams brought it into ecology in 1943 to describe abundance: many species appear once while a few contain numerous individuals. Its parameter controls how slowly that sequence of group sizes declines.

There is also a structural connection that a histogram cannot reveal. If the number of groups is Poisson and each positive group size is Logarithmic, the total may have a Negative Binomial distribution. This compound-sum representation explains its appearance in clustered populations. It also warns that a dominant count of one does not identify the model, since several count laws can look that way.

Decision guide

A good candidate when: positive counts represent group or species sizes with many ones and a decreasing tail, with zero impossible.

Compare it with: Geometric and Negative Binomial. The Logarithmic law arises naturally as cluster size in compound constructions and is not a generic model for counts containing zeros.

References

  • SciPy reference: scipy.stats.logser — definition and parameterization
  • Johnson, N. L., Kemp, A. W. & Kotz, S. (2005). Univariate Discrete Distributions, 3rd ed. Wiley.
  • Fisher, R. A., Corbet, A. S. & Williams, C. B. (1943). The relation between the number of species and the number of individuals in a random sample of an animal population. Journal of Animal Ecology, 12(1), 42–58.

Logarithmic Distribution: equations and calculator

Distribution defintion

X∼Logarithmic(p)X\sim\mathrm{Logarithmic}\left(p\right)

Distribution domain

x∈N⩾1≡{1,2,… }x\in\mathbb{N}_{\geqslant 1}\equiv \left\{1,2,\dots\right\}

Parameters domain and parameters constraints

p∈(0,1)⊆Rp\in\left(0,1\right)\subseteq\mathbb{R}

Cumulative distribution function

FX(x)=∑i=0x1−ln⁡(1−p)piiF_{X}\left(x\right)=\sum_{i=0}^{x}\frac{1}{-\ln(1 - p)} \frac{p^i}{i}

Probability mass function

fX(x)=1−ln⁡(1−p)pxxf_{X}\left(x\right)=\frac{1}{-\ln(1 - p)} \frac{p^x}{x}

Percent point function/Sample

FX−1(u)=arg⁡min⁡x∣FX(x)−u∣F^{-1}_{X}\left(u\right)=\arg\min_{x}\left| F_{X}\left(x\right)-u \right|

Parametric centered moments

E[Xk]=μk′=∑x=0∞xkfX(x)=(k−1)!−ln⁡(1−p)(p1−p)kE[X^k]=\mu'_{k}=\sum_{x=0}^{\infty}x^{k}f_{X}\left(x\right)=\frac{(k - 1)!}{-\ln(1 - p)} \left(\frac{p}{1 - p}\right)^k

Parametric mean

Mean(X)=μ1′=1−ln⁡(1−p)p1−p\mathrm{Mean}(X)=\mu'_{1}=\frac{1}{-\ln(1 - p)} \frac{p}{1 - p}

Parametric variance

Variance(X)=(μ2′−μ1′2)=−p2+pln⁡(1−p)(1−p)2(ln⁡(1−p))2\mathrm{Variance}(X)=(\mu'_{2}-\mu'^{2}_{1})=-\frac{p^2 + p\ln(1-p)}{(1-p)^2(\ln(1-p))^2}

Parametric skewness

Skewness(X)=μ3′−3μ2′μ1′+2μ1′3(μ2′−μ1′2)1.5\mathrm{Skewness}(X)=\frac{\mu'_{3}-3\mu'_{2}\mu'_{1}+2\mu'^{3}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{1.5}}

Parametric kurtosis

Kurtosis(X)=μ4′−4μ1′μ3′+6μ1′2μ2′−3μ1′4(μ2′−μ1′2)2\mathrm{Kurtosis}(X)=\frac{\mu'_{4}-4\mu'_{1}\mu'_{3}+6\mu'^{2}_{1}\mu'_{2}-3\mu'^{4}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{2}}

Parametric median

Median(X)=FX−1(0.5)\mathrm{Median}(X)=F^{-1}_{X}\left(0.5\right)

Parametric mode

Mode(X)=1\mathrm{Mode}(X)=1

Additional information and definitions

Computing an analytic expression for the inverse of the cumulative distribution function is not feasible. However, it is possible to calculate the Percentile Point Function by approximating it to the nearest integer.
u:Uniform[0,1] random varibleu:\text{Uniform[0,1] random varible}