Online probability distribution fitting tool

Phitter fits your sample against 87 theoretical probability distributions, estimates the parameters of each one, ranks the candidates, and shows you the evidence behind the ranking. It is free, needs no account, and runs entirely inside your browser.

How the fit works

You bring a sample of observations. Phitter estimates the parameters of every candidate distribution from that sample, computes the theoretical density or mass function implied by those parameters, and compares it against the empirical distribution of your data.

The ranking first favours candidates with more goodness-of-fit tests that do not reject them. For continuous fits, ties are ordered by lower BIC, then AIC, then sum of squared errors (SSE). For discrete fits, ties are ordered by lower SSE; the discrete results do not report AIC or BIC. These criteria help compare candidates, but they do not identify a true model automatically.

The whole calculation is the Phitter Python kernel, executed locally through Pyodide and WebAssembly. Your observations are never uploaded to a fitting server, and the same kernel is available as a Python package if you would rather run the fit in a notebook or a pipeline.

The data you can bring

Paste values directly, or import a TXT file with one observation per line, a CSV, an XLSX spreadsheet, or a file served from a public URL. Phitter handles continuous samples and discrete counts, and you choose which family to fit before running the analysis.

The web application accepts samples of at least 20 numeric observations. That is an input requirement, not a guarantee of a reliable fit. The amount of data needed depends on the differences between models, the precision you need, and whether your decision concerns typical values or rare events. Small samples can leave several candidates difficult to distinguish.

The three goodness-of-fit tests

The Kolmogorov–Smirnov statistic measures the largest vertical distance between the empirical cumulative distribution of your sample and the fitted one. It is generally more sensitive in the centre than in the tails; it can still detect tail discrepancies.

The Anderson–Darling statistic, reported for continuous fits, weights cumulative discrepancies more heavily in the tails. This can be useful for examining waiting times, losses or failures, but no test is most powerful against every kind of departure.

The chi-square test bins the data and compares observed against expected counts in each bin. It is the natural test for discrete data, and for continuous data its result depends on how the bins were chosen.

A test can reject a model or fail to reject it; neither outcome proves which distribution generated the data. P-values and critical values also depend on their assumptions. Estimating parameters from the tested sample, discrete observations, and small expected bin counts can require adjusted calibration. Read the results with the plots and the test assumptions in mind.

87 distributions, continuous and discrete

Phitter covers 77 continuous and 10 discrete distributions, from the ones every course starts with to families that rarely appear in general-purpose tools: Johnson SB and SU, Burr, Dagum, generalised extreme value, non-central chi-square, Kumaraswamy and its exponentiated form.

Every distribution also has its own page with the density or mass function, the cumulative function, the quantile function, the parametric moments, and an interactive calculator you can use without fitting anything.

A ranking is a starting point, not a verdict

The distribution at the top of the table leads under the ranking rules described above. That does not establish that it generated your data. Compare a group of plausible candidates before selecting a model.

Before you commit to a model, compare the histogram with the fitted density, inspect the Q–Q plot, and ask whether the family makes sense for the process. If a quantity cannot be negative, examine how much probability a candidate assigns below zero and whether that approximation is acceptable for your use.

Frequently asked questions

Is Phitter free?
Yes. The web application is free to use with no account, no row limit, and no paid tier. The Python kernel behind it is open source under the MIT licence.
Does my data leave my browser?
No. The fit runs locally through Pyodide and WebAssembly, so your observations are processed on your own machine and are never sent to a Phitter fitting server.
How many distributions does Phitter fit?
87 in total: 77 continuous distributions and 10 discrete ones, each with its own parameter estimation, goodness-of-fit results and interactive charts.
Which goodness-of-fit tests does Phitter run?
Continuous fits report Kolmogorov–Smirnov, Anderson–Darling and chi-square results, with AIC and BIC in the ranking table. Discrete fits report Kolmogorov–Smirnov and chi-square results, with SSE. Test calibration depends on the data and on how the parameters were estimated; a non-rejection does not establish that the model is correct.
What file formats can I import?
TXT with one observation per line, CSV, XLSX, and files served from a public URL. You can also paste values straight into the input.
Can I run the same fit from Python?
Yes. Phitter is also available as a Python package for notebooks and pipelines. To compare a Python fit with the web application, preserve the input data and match the candidate distributions, fitting settings, and kernel and dependency versions. Different versions or settings can change results.
How much data do I need?
The web application requires at least 20 numeric observations. The sample size needed for a useful statistical conclusion depends on the models, the differences you need to detect, and the quantities you want to estimate. Meeting the input minimum does not guarantee enough information, especially about the tails.

Phitter has a peer-reviewed software paper in the Journal of Open Source Software, and its kernel, formulas and implementation are public.