PLAYGROUND

DISTRIBUTIONS / DISCRETE / NEGATIVE BINOMIAL

Negative Binomial distribution

Quick view

The Negative Binomial counts failures before observing r successes. In Phitter’s parameterization, r is the target number of successes and p success probability per trial.

If you are coming from another distribution

At r=1 it connects to Geometric under the failure-count convention. As a Gamma–Poisson mixture it explains variance larger than the mean.

History and terminology

The family grew from Pascal trial and counting problems. Modern use expanded in epidemiology, ecology, and insurance as an overdispersed alternative to Poisson.

A familiar situation

It suits data where event rates vary between units or unobserved heterogeneity exists. Dispersion may then represent mixture rather than temporal dependence.

Fitting with care

State whether the variable counts failures or total trials and whether p means success or failure; both conventions circulate.

Two stories for the same shape

In Pascal’s story, an experiment stops after r successes and the variable counts accumulated failures. In the modern regression story, a Poisson count has a rate that varies across units according to a Gamma distribution. Integrating out that unknown rate also yields a Negative Binomial law. The first story is about waiting; the second is about heterogeneity and overdispersion.

Those interpretations do not make parameterizations interchangeable. Some texts count successes, others failures, and the second parameter may be stated as a probability, a mean, or a dispersion. Always verify support and convention before comparing results. For longitudinal data, variance above the mean can arise from serial dependence rather than Gamma mixing alone.

Decision guide

A good candidate when: counts have variance greater than the mean, or failures are counted until a fixed number of successes occurs.

Compare it with: Poisson, zero-inflated models, and Beta-binomial according to mechanism. Check parameter conventions: libraries differ on successes, failures, and probability.

References

  • SciPy reference: scipy.stats.nbinom — definition and parameterization
  • Johnson, N. L., Kemp, A. W. & Kotz, S. (2005). Univariate Discrete Distributions, 3rd ed. Wiley.
  • Feller, W. (1968). An Introduction to Probability Theory and Its Applications, Vol. 1, 3rd ed. Wiley.
  • Hilbe, J. M. (2011). Negative Binomial Regression, 2nd ed. Cambridge University Press.

Negative Binomial Distribution: equations and calculator

Distribution defintion

X∼NegativeBinomial(r,p)X\sim\mathrm{NegativeBinomial}\left(r,p\right)

Distribution domain

x∈N≡{0,1,2,… }x\in\mathbb{N}\equiv \left\{0,1,2,\dots\right\}

Parameters domain and parameters constraints

r∈N⩾1,p∈(0,1)⊆Rr\in\mathbb{N}_{\geqslant 1},p\in\left(0,1\right)\subseteq\mathbb{R}

Cumulative distribution function

FX(x)=I(p,r,x+1)F_{X}\left(x\right)=I(p,r,x+1)

Probability mass function

fX(x)=(r+x−1x)pr(1−p)xf_{X}\left(x\right)=\binom{r+x-1}{x}p^r(1-p)^x

Percent point function/Sample

FX−1(u)=arg⁡min⁡x∣FX(x)−u∣F^{-1}_{X}\left(u\right)=\arg\min_{x}\left| F_{X}\left(x\right)-u \right|

Parametric centered moments

E[Xk]=μk′=∑x=0∞xkfX(x)E[X^k]=\mu'_{k}=\sum_{x=0}^{\infty}x^{k}f_{X}\left(x\right)

Parametric mean

Mean(X)=μ1′=r(1−p)p\mathrm{Mean}(X)=\mu'_{1}=\frac{r(1-p)}{p}

Parametric variance

Variance(X)=(μ2′−μ1′2)=r(1−p)p2\mathrm{Variance}(X)=(\mu'_{2}-\mu'^{2}_{1})=\frac{r(1-p)}{p^2}

Parametric skewness

Skewness(X)=μ3′−3μ2′μ1′+2μ1′3(μ2′−μ1′2)1.5=2−pr (1−p)\mathrm{Skewness}(X)=\frac{\mu'_{3}-3\mu'_{2}\mu'_{1}+2\mu'^{3}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{1.5}}=\frac{2-p}{\sqrt{r\,(1-p)}}

Parametric kurtosis

Kurtosis(X)=μ4′−4μ1′μ3′+6μ1′2μ2′−3μ1′4(μ2′−μ1′2)2=3+6r+p2r (1−p)\mathrm{Kurtosis}(X)=\frac{\mu'_{4}-4\mu'_{1}\mu'_{3}+6\mu'^{2}_{1}\mu'_{2}-3\mu'^{4}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{2}}=3+\frac{6}{r} + \frac{p^2}{r\,(1-p)}

Parametric median

Median(X)=FX−1(0.5)\mathrm{Median}(X)=F^{-1}_{X}\left(0.5\right)

Parametric mode

Mode(X)=⌊(r−1) (1−p)/p⌋\mathrm{Mode}(X)=\lfloor(r-1)\,(1-p)/p\rfloor

Additional information and definitions

Computing an analytic expression for the inverse of the cumulative distribution function is not feasible. However, it is possible to calculate the Percentile Point Function by approximating it to the nearest integer.
u:Uniform[0,1] random varibleu:\text{Uniform[0,1] random varible}
I(x,a,b):Regularized incomplete beta functionI\left(x,a,b\right):\text{Regularized incomplete beta function}
⌊x⌋:Floor function\lfloor{x}\rfloor: \text{Floor function}