PLAYGROUND

DISTRIBUTIONS / CONTINUOUS / T STUDENT

T Student distribution

What it describes

Student’s t arises when a sample mean is standardized while population variance is estimated from a small normal sample. Its heavier tails represent additional scale-estimation uncertainty.

It spans the real line. df controls tails: few degrees of freedom allow more extremes; as df grows, t approaches Normal.

History and names

William Sealy Gosset published in 1908 under the pseudonym “Student” to protect Guinness confidentiality rules. His problem was inference for means with small samples.

How it connects to other distributions

Its square is F with one numerator degree of freedom. Noncentral t handles alternatives with an effect; Normal is its large-df limit.

Where it appears

  • t intervals and tests for means
  • linear regression, correlation, and unknown-variance estimation

Fitting cautions

Do not equate df with sample size without accounting for estimated parameters; t does not automatically fix dependence or heteroskedasticity.

The brewery that made the bell uncertain

William Sealy Gosset worked at the Guinness brewery and published under the pen name Student. His problem was practical: drawing conclusions from small samples when variance also had to be estimated. That extra uncertainty thickens the distribution’s tails.

Degrees of freedom provide a bridge. With few of them, extreme values receive much more probability and some moments do not even exist. As they increase, t approaches the Normal distribution. A t law can also provide robust regression errors, but there the degrees of freedom govern tail weight and need not represent a sample size. Confusing those two uses creates a false inferential story even when the fitted density is identical.

Decision guide

A good candidate when: errors are symmetric but contain more extremes than a Normal model, and degrees of freedom can represent that tail intensity.

Compare it with: Normal as the limiting case and Cauchy at ν=1. With skewness, a symmetric t may inflate both tails to explain extremes on only one side.

References

  • SciPy reference: scipy.stats.t — definition and parameterization
  • Johnson, N. L., Kotz, S. & Balakrishnan, N. (1994). Continuous Univariate Distributions, 2nd ed., Vol. 1. Wiley.
  • Student (W. S. Gosset) (1908). The probable error of a mean. Biometrika, 6(1), 1–25.
  • Fisher, R. A. (1928). Statistical Methods for Research Workers. Oliver and Boyd.

T Student Distribution: equations and calculator

Distribution defintion

X∼TStudent(df)X\sim\mathrm{TStudent}\left(\text{df}\right)\\

Distribution domain

x∈(−∞,∞)x\in\left(-\infty,\infty\right)

Parameters domain and parameters constraints

df∈R+\text{df}\in\mathbb{R}^{+}\\

Cumulative distribution function

FX(x)=I(x+x2+df2x2+df,df2,df2)F_{X}\left(x\right)=I\left(\frac{x+\sqrt{x^{2}+\text{df}}}{2\sqrt{x^{2}+\text{df}}},\frac{\text{df}}{2},\frac{\text{df}}{2}\right)\\

Probability density function

fX(x)=(1+x2/df)−(1+df)/2df×Beta(12,df2)f_{X}\left(x\right)=\frac{\left(1+x^{2}/\text{df}\right)^{-(1+\text{df})/2}}{\sqrt{\text{df}}\times \text{Beta}\left(\frac{1}{2},\frac{\text{df}}{2}\right)}\\

Percent point function/Sample

FX−1(u)={df(1−I−1(u,df/2,df/2))I−1(u,df/2,df/2)if  u≥12−df(1−I−1(u,df/2,df/2))I−1(u,df/2,df/2)if  u<12F^{-1}_{X}\left(u\right)=\left\{\begin{array}{cl} \sqrt{\frac{\text{df}(1-I^{-1}\left(u,\text{df}/2,\text{df}/2\right))}{I^{-1}\left(u,\text{df}/2,\text{df}/2\right)}} & \text{if } \ u \geq \frac{1}{2} \\ -\sqrt{\frac{\text{df}(1-I^{-1}\left(u,\text{df}/2,\text{df}/2\right))}{I^{-1}\left(u,\text{df}/2,\text{df}/2\right)}} & \text{if } \ u < \frac{1}{2} \end{array} \right.\\

Non-central parametric moments

μk′=E[Xk]=∫−∞∞xkfX(x)dx={0if  k odd ∧ 0<k<dfdfk2 ∏i=1k/22i−1df−2iif  k even ∧ 0<k<df\mu'_{k}=E[X^k]=\int_{-\infty }^{\infty }x^{k}f_{X}\left(x\right)dx=\left\{\begin{array}{cl} 0 & \text{if } \ k\text{ odd} \ \wedge \ 0 < k < \text{df} \\ \text{df}^{\frac{k}{2}} \,\prod_{i=1}^{k/2}\frac{2i-1}{\text{df}-2i} & \text{if } \ k\text{ even} \ \wedge \ 0 < k < \text{df} \end{array} \right.\\

Parametric mean

Mean(X)=μ1′=0\mathrm{Mean}(X)=\mu'_{1}=0\\

Parametric variance

Variance(X)=μ2′−μ1′2={df/(df+2)if  df>2undefinedif  df≤2\mathrm{Variance}(X)=\mu'_{2}-\mu'^{2}_{1}=\left\{\begin{array}{cl} \text{df}/(\text{df}+2) & \text{if } \ \text{df} > 2 \\ \text{undefined} & \text{if } \ \text{df} \leq 2 \end{array} \right.\\

Parametric skewness

Skewness(X)=μ3′−3μ2′μ1′+2μ1′3(μ2′−μ1′2)1.5={0if  df>3undefinedif  df≤3\mathrm{Skewness}(X)=\frac{\mu'_{3}-3\mu'_{2}\mu'_{1}+2\mu'^{3}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{1.5}}=\left\{\begin{array}{cl} 0 & \text{if } \ \text{df} > 3 \\ \text{undefined} & \text{if } \ \text{df} \leq 3 \end{array} \right.\\

Parametric kurtosis

Kurtosis(X)=μ4′−4μ1′μ3′+6μ1′2μ2′−3μ1′4(μ2′−μ1′2)2={3+6/(df−4)if  df>4undefinedif  df≤4\mathrm{Kurtosis}(X)=\frac{\mu'_{4}-4\mu'_{1}\mu'_{3}+6\mu'^{2}_{1}\mu'_{2}-3\mu'^{4}_{1}}{(\mu'_{2}-\mu'^{2}_{1})^{2}}=\left\{\begin{array}{cl} 3+6/(\text{df}-4) & \text{if } \ \text{df} > 4 \\ \text{undefined} & \text{if } \ \text{df} \leq 4 \end{array} \right.\\

Parametric median

Median(X)=0\mathrm{Median}(X)=0\\

Parametric mode

Mode(X)=0\mathrm{Mode}(X)=0

Additional information and definitions

u:Uniform[0,1] random varibleu:\text{Uniform[0,1] random varible}
I(x,a,b):Regularized incomplete beta functionI\left(x,a,b\right):\text{Regularized incomplete beta function}
I−1(x,a,b):Inverse of regularized incomplete beta functionI^{-1}\left(x,a,b\right):\text{Inverse of regularized incomplete beta function}
Beta(x,y):Beta function\text{Beta}\left(x,y\right):\text{Beta function}