DATES / Normality Analysis User Guide

Normality & Distribution Analysis User Guide

The Normality module evaluates whether continuous quantitative measurements follow a Gaussian normal distribution using 10 statistical significance tests, skewness and kurtosis diagnostics, interactive diagnostic plots, and integrated automatic data transformation.

1. INTRODUCTION

The Normality & Distribution Analysis module provides comprehensive statistical validation of distributional assumptions across scientific disciplines (biology, chemistry, physics, medicine, psychology, engineering, and environmental science).

Most standard parametric statistical procedures—such as Analysis of Variance (ANOVA), Student's t-tests, Pearson correlation, and linear regression—require the underlying population error terms to be normally distributed. Running parametric tests on severely non-normal data can result in inflated Type I error rates or reduced statistical power. The Normality module tests continuous variables across experimental groups, provides skewness and kurtosis metrics, generates diagnostic plots (Q-Q plots, histograms, density curves, box plots, violin plots), and offers automated data transformation to achieve normality.

When to use this module:

2. AVAILABLE OPTIONS & SETTINGS

The sidebar control panel and top action bar provide settings for test selection, alpha significance thresholds, decimal precision, and automated transformations:

Control / Option What it does Why it is used When to use / select
Test Selection Dropdown Selects a specific statistical normality test (e.g., Shapiro-Wilk, Anderson-Darling, Kolmogorov-Smirnov) or runs All Tests simultaneously. Chooses the hypothesis testing algorithm used to calculate test statistics and p-values. Select All Tests for comprehensive multi-test confirmation, or pick a specific test suited to your sample size.
Alpha Level Selector Toggles the significance threshold between 5% (0.05) and 1% (0.01). Sets the critical alpha cutoff for rejecting the null hypothesis of normality. Select 5% for standard scientific studies; select 1% for stringent clinical or industrial quality bounds.
Decimal Rounding Sets output numerical precision selector (0, 1, 2, 3, or 4 decimal places). Formats calculated test statistics, p-values, skewness, and kurtosis in result tables. Located in the top header bar; adjust to match target publication standards.
Transformation Toggle Toggles automated transformation mode On or Off and opens the transformation config modal. Automatically applies mathematical transformations if data fails normality at the chosen alpha level. Turn On when preparing non-normal datasets for downstream parametric models.
Variables (Traits) Selection Pill selectors choosing quantitative continuous columns to analyze for normality. Specifies target continuous measurement variables for distribution testing. Select one or more continuous numeric variables from your uploaded dataset.

3. INPUT DATA FORMAT

The module accepts tabular spreadsheets (Excel .xlsx, .xls, or CSV .csv) with the following structure:

Sample Input Table (Continuous Measurements across Experimental Groups)

Below is a representative dataset containing generalized grouping factors (Group, Condition) and quantitative measurement variables (Response_Value, Concentration):

Normality_Input_Data.xlsx Sheet: Trial_Measurements
Group Condition Replicate Response_Value Concentration
Group-01ControlR145.2012.40
Group-01ControlR246.8013.10
Group-01ControlR344.9012.80
Group-01TreatedR158.4025.60
Group-01TreatedR261.2028.10
Group-01TreatedR359.1026.40
Group-02ControlR141.5010.20
Group-02ControlR243.1011.00
Group-02ControlR342.0010.80

4. METHODS / MODES

The Normality module provides 10 standard statistical test algorithms to evaluate distributional symmetry and tail behavior:

Statistical Normality Test Optimal Sample Size Range Null Hypothesis (H0) & Decision Rule Recommended Science Application
Shapiro-Wilk Small to Medium (n = 3 to 50) H0: Data is normally distributed. Reject H0 if p-value < Alpha. Gold standard test for biological, agricultural, and small-sample experimental data.
Anderson-Darling All sample sizes (n ≥ 5) H0: Data follows normal distribution. Gives higher weight to distribution tails. Ideal when tail behavior and extreme values are critical (environmental, engineering trials).
Kolmogorov-Smirnov Large samples (n > 50) H0: Sample distribution matches continuous normal baseline. General goodness-of-fit comparison against theoretical cumulative distribution.
Lilliefors Medium to Large samples Modifies K-S test when population mean and variance are unknown. Useful for continuous physical and chemical measurements with estimated parameters.
Jarque-Bera Large samples (n > 100) H0: Sample skewness and excess kurtosis match a normal distribution (skewness = 0, kurtosis = 3). Commonly applied in financial modeling, econometrics, and large sensor streams.
D'Agostino-Pearson Medium to Large (n ≥ 20) Combines skewness and kurtosis tests into an omnibus chi-square statistic. Versatile test for clinical, physiological, and industrial quality control data.
Cramér-von Mises Medium to Large samples Evaluates distance between empirical and theoretical cumulative distributions. High statistical power for symmetric, continuous physical measurements.
Pearson Chi-Square Large samples binned into classes H0: Frequency counts across binned intervals match expected normal counts. Applied to binned discrete or continuous industrial measurements.
Shapiro-Francia Medium to Large samples (n = 5 to 5000) Modification of Shapiro-Wilk optimized for platykurtic and leptokurtic data. Effective for heavy-tailed biological and ecological measurement series.
Ryan-Joiner Small to Medium samples Calculates correlation between observed values and normal scores (similar to Shapiro-Wilk). Widely used alternative in engineering quality control and quality assurance studies.

5. RESULTS

Upon clicking RUN ANALYSIS, the module displays a multi-tab analytical suite comprising main statistical summary tables, interactive visualization plots, and automated text interpretation.

Sample Result 1: Main Normality Test Summary Table

The main summary table reports sample size, mean, median, standard deviation, skewness, kurtosis, test statistics, p-values, and automated conclusion badges:

Normality_Test_Results.xlsx Test: Shapiro-Wilk (Alpha = 0.05)
Variable Group Sample Size (n) Mean Std Dev Skewness Kurtosis Test Statistic p-Value Normality Status
Response_ValueGroup-011245.8502.1400.124-0.3420.9680.7420Pass (Normal)
Response_ValueGroup-021242.1001.9800.085-0.5100.9740.8250Pass (Normal)
ConcentrationGroup-011218.4006.8501.4522.1800.8120.0084Fail (Non-Normal)

Sample Result 2: Diagnostic Plot Suite

The Plots tab generates 5 interactive visual diagnostic plots for each variable:

6. QUICK WORKFLOW

  1. Upload Dataset: Upload your .xlsx or .csv spreadsheet using the sidebar upload control.
  2. Select Test & Alpha: Choose your desired normality test (e.g., Shapiro-Wilk or All Tests) and set the significance level (5% or 1%) in the top bar.
  3. Select Variables: Click the pill selectors in the sidebar to choose quantitative continuous variables for analysis.
  4. Run Analysis: Click RUN ANALYSIS to compute statistics, p-values, and plot visualizations.
  5. Review Results & Plots: Inspect the summary table, toggle to the Plots tab to inspect Q-Q plots, and view automated text interpretations.
  6. Export Outputs: Export summary tables as Excel (.xlsx) or DOCX, or download plot figures as publication-ready images or PowerPoint presentations.

7. IMPORTANT NOTES

Interpretation Rule for Normality Hypotheses

Unlike standard hypothesis testing where p-value < Alpha indicates a positive finding, in normality testing the null hypothesis (H0) states that the data is normal. Therefore, a p-value > Alpha (e.g., > 0.05) means you fail to reject normality (Data is Normal), whereas a p-value < Alpha indicates significant deviation from normality.

Cite DATES in Research Papers

If you use the DATES Normality & Distribution Analysis module for statistical validation in published scientific work, please cite it as follows:

@software{dates_app_2026, author = {DATES Development Team}, title = {DATES: Data Analysis and Trial Evaluation System}, year = {2026}, url = {https://dates-app.org}, note = {Data Preparation & Normality Analysis Modules} }