Halal Authentication Platform

Loading engine… Local app — data stored in this browser
1

Training data iAny CSV with a SampleID column and numeric variable columns. Class is optional — include it for classification, leave it out for pure exploratory PCA.

Upload and tag the samples your models will learn from

Add samples (bulk)

CSV: SampleID (or first column), optional Class, and 2+ numeric columns.

Training dataset

2

Models iExplore with PCA any time (no Class needed). Train a classifier only once your dataset has Class labels with 2+ groups.

Explore the data, pick biomarkers, and train classifiers

Suggest preprocessing iLooks at your actual uploaded data (skew, scale differences between variables, how much sample totals vary) and suggests a starting pipeline — not a mandate. You can always change any setting below yourself.

Data preprocessing

i
None
No row adjustment.
By sum
Each sample divided by its own total — corrects overall concentration/dilution differences between samples.
By median
Each sample divided by its own median — more resistant to a few very large peaks than "by sum".
Quantile
Forces every sample to the same value distribution. Best for very wide datasets (1000+ variables).
i
None
No element-wise change.
Log10 / Log2
Compresses large values, pulls in right-skew — the standard choice for concentration-like data spanning orders of magnitude.
Square root / Cube root
Milder than log — stabilizes variance without compressing as aggressively.
Variance stabilizing
Behaves like log for large values and linear near zero — handles data that includes zeros better than a plain log.
Arcsin
The standard transform for proportion/fraction data (e.g. relative abundances) — stabilizes variance near 0 and 1.
Box-Cox
Finds the power transform that makes a column closest to normally distributed. Needs positive values (shifted automatically if needed).
Winsorize
Caps (doesn't remove) the most extreme low/high values per column at a chosen percentile — tempers outliers without deleting the sample.
Johnson
Fits a flexible Johnson-family curve per column to push it toward normality — more aggressive than Box-Cox, works on data Box-Cox can't (including negative values).
i
None
No column adjustment.
Center
Subtracts each column's mean — centers on zero, no change to spread.
Standardize (n–1)
Center, then divide by the sample standard deviation — the common default ("auto-scaling"). Every variable ends up equally weighted regardless of its original units.
Standardize (n)
Same, using the population standard deviation (divide by n instead of n–1) — nearly identical for larger sample counts.
/ Std. dev. (n–1 or n)
Divides by the standard deviation only — no centering, so a variable's original scale relative to zero is kept.
Pareto scaling
Divide by the square root of the standard deviation — a middle ground between no scaling and full standardizing; keeps large-signal variables somewhat more influential.
Range scaling
Center, then divide by (max − min).
Rescale 0–1 / 0–100
Plain min-max rescaling into a fixed range — easiest to interpret, but sensitive to outliers (an extreme value stretches the whole scale).
Binarize (0/1)
Every nonzero value becomes 1, zero stays 0 — reduces to simple presence/absence.
Sign (−1/0/1)
Keeps only the direction of each value, not its size.
iThe p-value cutoff used for ANOVA-based biomarker validation and (if chosen) the Johnson normality check below. 0.05 is the common food-industry default; halal-specific work sometimes uses a stricter 0.01.
i
ON (default)
The normalization, transformation and scaling above are used to pick the biomarkers and to train every model. Inside cross-validation they are re-learned from each fold's training samples only (nothing leaks from the held-out samples). Each saved model keeps what it learned, so unknown samples get exactly the same treatment. Normalization is computed over ALL variables first, then the biomarkers are picked, so the prediction file must contain every training variable.
OFF
The old behaviour: models and biomarker selection work on the raw areas, with each model only standardizing its own columns. Use it to compare.
What actually changes results
SVM, kNN, ANN and PLS-DA standardize their inputs internally anyway, so the Scaling choice barely moves them. Normalization and transformation (for example log10) are what really change the models. Scaling matters mostly for Random Forest and Naive Bayes.

Switch it off to reproduce the old behaviour. Scaling has little effect on SVM / kNN / ANN / PLS-DA (they standardize internally); normalization and transformation are what change them.

Applied in this order — normalize samples (rows) → transform values → scale variables (columns) — to every PCA/biplot/PLS-DA/statistical-report/univariate-plot computation below. With the switch above ON, the same settings are also used for biomarker selection (Step 2) and model training (Step 3).

3

Predict sample iUses the active model(s) for this dataset — only available once you've trained and activated at least one with Class labels.

Classify new, unknown samples with your active model(s)

Halal / Non-Halal determination iThe model predicts a class name (e.g. "Lard", "Beef Tallow") — this turns that into a Halal/Non-Halal verdict: if the predicted class name contains any of these keywords, the sample is marked Non-Halal. Comma-separated, case-insensitive. Change this any time without retraining — it's a labeling rule applied AFTER prediction, not part of the model itself.

Predict samples (bulk)

CSV with SampleID and the same numeric columns used for training.

Prediction history