General dataset — training data

Any CSV with a SampleID column and numeric variable columns. Class is optional — include it for classification, leave it out for pure exploratory PCA.

Add samples (bulk)

CSV: SampleID (or first column), optional Class, and 2+ numeric columns.

Training dataset

General dataset — models

Explore with PCA any time (no Class needed). Train a classifier only once your dataset has Class labels with 2+ groups.

Data preprocessing

i
None
No row adjustment.
By sum
Each sample divided by its own total — corrects overall concentration/dilution differences between samples.
By median
Each sample divided by its own median — more resistant to a few very large peaks than "by sum".
Quantile
Forces every sample to the same value distribution. Best for very wide datasets (1000+ variables).
i
None
No element-wise change.
Log10 / Log2
Compresses large values, pulls in right-skew — the standard choice for concentration-like data spanning orders of magnitude.
Square root / Cube root
Milder than log — stabilizes variance without compressing as aggressively.
Variance stabilizing
Behaves like log for large values and linear near zero — handles data that includes zeros better than a plain log.
Arcsin
The standard transform for proportion/fraction data (e.g. relative abundances) — stabilizes variance near 0 and 1.
Box-Cox
Finds the power transform that makes a column closest to normally distributed. Needs positive values (shifted automatically if needed).
Winsorize
Caps (doesn't remove) the most extreme low/high values per column at a chosen percentile — tempers outliers without deleting the sample.
Johnson
Fits a flexible Johnson-family curve per column to push it toward normality — more aggressive than Box-Cox, works on data Box-Cox can't (including negative values).
i
None
No column adjustment.
Center
Subtracts each column's mean — centers on zero, no change to spread.
Standardize (n\u20131)
Center, then divide by the sample standard deviation — the common default ("auto-scaling"). Every variable ends up equally weighted regardless of its original units.
Standardize (n)
Same, using the population standard deviation (divide by n instead of n\u20131) \u2014 nearly identical for larger sample counts.
/ Std. dev. (n\u20131 or n)
Divides by the standard deviation only \u2014 no centering, so a variable's original scale relative to zero is kept.
Pareto scaling
Divide by the square root of the standard deviation \u2014 a middle ground between no scaling and full standardizing; keeps large-signal variables somewhat more influential.
Range scaling
Center, then divide by (max \u2212 min).
Rescale 0\u20131 / 0\u2013100
Plain min-max rescaling into a fixed range \u2014 easiest to interpret, but sensitive to outliers (an extreme value stretches the whole scale).
Binarize (0/1)
Every nonzero value becomes 1, zero stays 0 \u2014 reduces to simple presence/absence.
Sign (\u22121/0/1)
Keeps only the direction of each value, not its size.
iThe p-value cutoff used for ANOVA-based biomarker validation and (if chosen) the Johnson normality check below. 0.05 is the common food-industry default; halal-specific work sometimes uses a stricter 0.01.

Applied in this order — normalize samples (rows) → transform values → scale variables (columns) — to every PCA/biplot/PLS-DA/statistical-report/univariate-plot computation below. Model training keeps its own separate scaling per cross-validation fold, unaffected by this.

Preview preprocessing effect

Before vs after the pipeline above — overall density curve + grouped box-and-whisker. Works with or without Class labels.

General dataset — predict sample

Uses the active model + diagnostic ratio for this dataset (only available once you've trained one with Class labels).

Checking active model…

Predict samples (bulk)

CSV with SampleID and the same numeric columns used for training.

Prediction history