PLAYGROUND

DISTRIBUTIONS / CONTINUOUS / PARETO FIRST KIND

Pareto First Kind distribution

From phenomenon to model

In 1896 Vilfredo Pareto observed a concentration regularity in wealth: the number of people above a level approximately followed a power law. The distribution was later separated from the empirical law.

What to look at

The support starts at loc + xm in Phitter’s form. xm sets the minimum and alpha the tail index; the mean exists only when alpha > 1.

Useful relationships

It is the direct case of regularly varying tails and a special or limiting case of several Burr and Beta-prime families. Generalized Pareto shares tail language but models excesses above a threshold.

Names you may also encounter

It is also called Pareto Type I, classical Pareto, or Pareto law.

When it makes sense

Pareto is suited to an upper tail rather than necessarily the full range. In catastrophe and insurance work, alpha determines how much weight very large events receive.

What not to assume

A straight log-log plot does not validate Pareto; select a threshold and quantify tail-index uncertainty.

The straight line Pareto found among high incomes

While studying European tax tables, Vilfredo Pareto observed that the proportion of people above a high income declined approximately as a power. On logarithmic axes, the tail becomes a line whose slope is tied to alpha. His central source was the 1896 and 1897 Cours d’économie politique.

The first-kind Pareto distribution formalizes that tail above a strict minimum. It does not claim that the entire income distribution, including its centre, follows the same power. Nor is it synonymous with the 80/20 rule, which corresponds only to particular choices and was popularized later. Since the existence of a mean and variance depends on alpha, extrapolation without parameter uncertainty can turn a few maxima into enormous economic conclusions.

Decision guide

A good candidate when: interest lies in values above a real minimum and the tail approximately follows a power law.

Compare it with: Generalized Pareto for exceedances and Burr or Lognormal for the full body. Check threshold stability: Pareto rarely describes the entire sample.

References

  • SciPy reference: scipy.stats.pareto — definition and parameterization

  • Pareto, V. (1896–1897). Cours d’économie politique. F. Rouge.

  • Johnson, N. L., Kotz, S. & Balakrishnan, N. (1994). Continuous Univariate Distributions, 2nd ed., Vol. 1. Wiley.

  • Clauset, A., Shalizi, C. R. & Newman, M. E. J. (2009). Power-law distributions in empirical data. SIAM Review, 51(4), 661–703.

Pareto First Kind Distribution: equations and calculator

Distribution defintion

X∼ParetoFirstKind(xm,α,Loc)X\sim\mathrm{ParetoFirstKind}\left(x_\mathrm{m},\alpha,\text{Loc}\right)

Distribution domain

x∈[Loc+xm,∞)x\in [\text{Loc}+x_\mathrm{m},\infty)

Parameters domain and parameters constraints

xm∈R+,α∈R+,Loc∈Rx_\mathrm{m}\in\mathbb{R}^{+},\alpha\in\mathbb{R}^{+},\text{Loc}\in\mathbb{R}

Cumulative distribution function

FX(x)=1−(xmx−Loc)αF_{X}\left(x\right)=1-\left(\frac{x_\mathrm{m}}{x-\text{Loc}}\right)^\alpha

Probability density function

fX(x)=αxmα(x−Loc)α+1f_{X}\left(x\right)=\frac{\alpha x_\mathrm{m}^\alpha}{(x-\text{Loc})^{\alpha+1}}

Percent point function/Sample

FX−1(u)=Loc+xm(1−u)−1αF^{-1}_{X}\left(u\right)=\text{Loc}+x_\mathrm{m} {(1-u)}^{-\frac{1}{\alpha}}

Non-central parametric moments

μ~k′=E[X~k]=∫xm∞xkfX~(x)dx={∞if α≤kαxmkα−kif α>k\tilde{\mu}'_{k}=E[\tilde{X}^k]=\int_{x_\mathrm{m}}^{\infty}x^{k}f_{\tilde{X}}\left(x\right)dx=\left\{\begin{array}{cl}\infty & \text{if } \alpha\le k \\ \frac{\alpha x_\mathrm{m}^k}{\alpha-k} & \text{if } \alpha>k \end{array} \right.

Parametric mean

Mean(X)=Loc+μ~1′=Loc+αxmα−1if α>1\mathrm{Mean}(X)=\text{Loc}+\tilde{\mu}'_{1}=\text{Loc}+\dfrac{\alpha x_\mathrm{m}}{\alpha-1} \quad \text{if }\alpha>1

Parametric variance

Variance(X)=(μ~2′−μ~1′2)=xm2α(α−1)2(α−2)if α>2\mathrm{Variance}(X)=(\tilde{\mu}'_{2}-\tilde{\mu}'^{2}_{1})=\dfrac{x_\mathrm{m}^2\alpha}{(\alpha- 1)^2(\alpha-2)} \quad \text{if }\alpha>2

Parametric skewness

Skewness(X)=μ~3′−3μ~2′μ~1′+2μ~1′3(μ~2′−μ~1′2)1.5=2(1+α)α−3α−2αif α>3\mathrm{Skewness}(X)=\frac{\tilde{\mu}'_{3}-3\tilde{\mu}'_{2}\tilde{\mu}'_{1}+2\tilde{\mu}'^{3}_{1}}{(\tilde{\mu}'_{2}-\tilde{\mu}'^{2}_{1})^{1.5}}=\frac{2(1+\alpha)}{\alpha-3}\sqrt{\frac{\alpha-2}{\alpha}}\quad \text{if } \alpha>3

Parametric kurtosis

Kurtosis(X)=μ~4′−4μ~1′μ~3′+6μ~1′2μ~2′−3μ~1′4(μ~2′−μ~1′2)2=6(α3+α2−6α−2)α(α−3)(α−4)if α>4\mathrm{Kurtosis}(X)=\frac{\tilde{\mu}'_{4}-4\tilde{\mu}'_{1}\tilde{\mu}'_{3}+6\tilde{\mu}'^{2}_{1}\tilde{\mu}'_{2}-3\tilde{\mu}'^{4}_{1}}{(\tilde{\mu}'_{2}-\tilde{\mu}'^{2}_{1})^{2}}=\frac{6(\alpha^3+\alpha^2-6\alpha-2)}{\alpha(\alpha-3)(\alpha-4)}\quad \text{if }\alpha>4

Parametric median

Median(X)=Loc+xm2α\mathrm{Median}(X)=\text{Loc}+x_\mathrm{m} \sqrt[\alpha]{2}

Parametric mode

Mode(X)=Loc+xm\mathrm{Mode}(X)=\text{Loc}+x_\mathrm{m}

Additional information and definitions

X~∼ParetoFirstKind(xm,α,0)\tilde{X}\sim\mathrm{ParetoFirstKind}\left(x_\mathrm{m},\alpha,0\right)
Loc:Location parameter\text{Loc}:\text{Location parameter}
xm:Scale parameterx_\mathrm{m}:\text{Scale parameter}
u:Uniform[0,1] random varibleu:\text{Uniform[0,1] random varible}