PLAYGROUND

DISTRIBUTIONS / CONTINUOUS / LOGNORMAL

Lognormal distribution

What it represents

The Lognormal distribution describes a positive variable whose logarithm is normal. It is the classic model for quantities formed by many multiplicative factors rather than additive effects.

Historical clue

Gibrat systematized the model in proportional-growth studies, and Aitchison and Brown gave it a comprehensive treatment.

Relationships that clarify its use

It is a limit of Generalized Gamma and appears as a limit of some multiplicative families. It resembles Gamma and Weibull centrally, but its tail is heavier than Weibull and lighter than comparable-index Pareto.

Data examples

  • income, wealth, firm, and city sizes
  • particle concentration, repair times, and positive biological variables

Modelling warning

Do not log-transform away zeros, censoring, or retransformation bias: returning from log scale to the original scale requires care.

Arithmetic changes when effects multiply

Adding small effects often leads to a Normal law; multiplying positive factors leads, through a parallel argument, to Lognormal. On a logarithmic scale, products become sums. The family consequently appears in particle sizes, concentrations, and quantities exposed to proportional growth.

The median is e^mu in the unshifted form, while the mean lies above it and also depends on dispersion. The geometric mean better estimates the multiplicative centre. A long right tail does not prove lognormality: Pareto and other heavy-tailed laws can look similar in moderate samples while extrapolating very differently. Comparing logarithms, high quantiles, and generating mechanisms prevents the histogram from deciding alone.

Decision guide

A good candidate when: positive values arise from multiplicative factors and their logarithms are approximately symmetric.

Compare it with: Gamma and Weibull. Inspect log-scale residuals and high quantiles; positivity and skewness alone do not establish a Lognormal model.

References

Lognormal Distribution: equations and calculator

Distribution defintion

X∼LogNormal(μ,σ)X\sim\mathrm{LogNormal}\left(\mu,\sigma\right)

Distribution domain

x∈(−∞,∞)x\in\left(-\infty,\infty\right)

Parameters domain and parameters constraints

μ∈R,σ∈R+\mu\in\mathbb{R},\sigma\in\mathbb{R}^{+}

Cumulative distribution function

FX(x)=12[1+erf⁡(ln⁡(x)−μσ2)]F_{X}\left(x\right)=\frac{1}{2}\left[1+\operatorname{erf}\left(\frac{\ln (x)-\mu}{\sigma\sqrt{2}}\right)\right]

Probability density function

fX(x)=1xσ2π exp⁡(−(ln⁡(x)−μ)22σ2)f_{X}\left(x\right)=\frac 1 {x\sigma\sqrt{2\pi}}\ \exp\left(-\frac{\left(\ln\left(x\right) -\mu\right)^2}{2\sigma^2}\right)

Percent point function/Sample

FX−1(u)=exp⁡(μ+2σ2erf⁡−1(2u−1))F^{-1}_{X}\left(u\right)=\exp(\mu+\sqrt{2\sigma^2}\operatorname{erf}^{-1}(2u-1))

Non-central parametric moments

μk′=E[Xk]=∫−∞∞xkfX(x)dx=ekμ+k2σ2/2\mu'_{k}=E[X^k]=\int_{-\infty}^{\infty}x^{k}f_{X}\left(x\right)dx=e^{k\mu+k^2\sigma^2/2}

Parametric mean

Mean(X)=μ1′=eμ+σ22\mathrm{Mean}(X)=\mu'_{1}=e^{\mu + \frac{\sigma^2}{2}}

Parametric variance

Variance(X)=μ2′−μ1′2=e2μ+σ2(eσ2−1)\mathrm{Variance}(X)=\mu'_{2}-\mu'^{2}_{1}=e^{2\mu+\sigma^2}(e^{\sigma^2}-1)

Parametric skewness

Skewness(X)=μ3′−3μ2′μ1′+2μ1′3(μ2′−μ1′2)1.5=(eσ2+2)eσ2−1\mathrm{Skewness}(X)=\frac{\mu'_{3}-3\mu'_{2}\mu'_{1}+2\mu'^{3}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{1.5}}=(e^{\sigma^2}+2)\sqrt{e^{\sigma^2}-1}

Parametric kurtosis

Kurtosis(X)=μ4′−4μ1′μ3′+6μ1′2μ2′−3μ1′4(μ2′−μ1′2)2=e4σ2+2e3σ2+3e2σ2−3\mathrm{Kurtosis}(X)=\frac{\mu'_{4}-4\mu'_{1}\mu'_{3}+6\mu'^{2}_{1}\mu'_{2}-3\mu'^{4}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{2}}=e^{4\sigma^2}+2e^{3\sigma^2}+3e^{2\sigma^2}-3

Parametric median

Median(X)=exp⁡(μ)\mathrm{Median}(X)=\exp(\mu)

Parametric mode

Mode(X)=exp⁡(μ−σ2)\mathrm{Mode}(X)=\exp(\mu-\sigma^2)

Additional information and definitions

μ:Location parameter\mu:\text{Location parameter}
σ:Scale parameter\sigma:\text{Scale parameter}
u:Uniform[0,1] random varibleu:\text{Uniform[0,1] random varible}