Signal Lab.
Make data.
Test the model.
A milling-machine experiment you can take apart. Generate telemetry, train a failure classifier and see what survives the test.
Same test.
Different training data.
A source-data model sets the reference. A second model learns only from the records you generate. Both are evaluated on the same 2,000 held-out examples.
Model utility
Average precision (AP) measures how well failures rank above normal examples. Higher is better.
Inspect the rows.
Check the trade-offs.
Changing class balance or noise can help one goal and hurt another. Inspect the marginal distributions and structural checks alongside model utility.
Distribution comparison
Histograms are normalized within each dataset. KS distance is the largest difference between cumulative distributions. Smaller is closer.
Data checks
TRAIN → SYNTHETICPositive temperatures and speed, non-negative torque and wear, process temperature above air temperature, and a valid product type.
Generated record preview
First 12 records from the current generated dataset.
| Record | Type | Air · K | Process · K | Speed · rpm | Torque · Nm | Wear · min | Label |
|---|---|---|---|---|---|---|---|
| Training the first experiment… | |||||||
Give the model
a different machine state.
Change six inputs and run both trained models. These are classification scores, not calibrated failure probabilities or a maintenance recommendation.
Machine state
What-if inputs are not validated operating conditions. Values outside a model’s training range are flagged.
Synthetic model
LIVE INFERENCELargest feature contributions to this score. Correlated features and interactions make these unsuitable as causal explanations.
The experiment
comes with its workings.
One reproducible pipeline. A stated source, a fixed split, fitted generation, trained models and downloadable results.
Split the source
6,000 train / 2,000 validation / 2,000 test. Stratified by failure label. Split seed 6841.
Fit & generate
Six label/type groups. Empirical numeric marginals and correlated Gaussian latent samples. Noise is normalized.
Train two models
The same class-balanced logistic regression, quadratic interactions, Adam optimizer and 160 training epochs.
Evaluate & inspect
Shared source test set. Average precision and ROC AUC. Validation selects each initial F1 threshold.
What enters the model?
Product type, air temperature, temperature gap, rotational speed, torque and tool wear. The numeric inputs are expanded into linear, squared and pairwise interaction features. Each model fits its own standardization on its training data.
Record IDs, product IDs and failure-subtype target columns are excluded. The composite failure label is the prediction target.
What do the metrics mean?
Average precision summarizes the precision-recall curve. The no-skill reference is the test set’s failure rate. ROC AUC measures ranking across both classes. Neither describes performance on real machinery.
Alert precision is the fraction of flagged records that have failure labels. Recall is the fraction of failure-labelled records that are flagged. The threshold changes these two quantities.
What are the limits?
The UCI source is synthetic, not a customer dataset. The random split measures benchmark classification, not prediction of future failures or performance at an unseen plant. Repeated test-set exploration is not independent final validation.
The generator does not implement differential privacy or reproduce every original failure rule. Label-conditioned sampling can produce ambiguous records. Structural checks do not certify semantic correctness, confidentiality or safe equipment operation.
Class-balanced training scores are not calibrated probabilities. This demonstrator is a portfolio experiment and cannot determine a maintenance action.
Where does the data come from?
AI4I 2020 Predictive Maintenance Dataset, Stephan Matzka (2020), distributed by the UCI Machine Learning Repository. DOI: 10.24432/C5HS5C ↗.
Licensed under Creative Commons Attribution 4.0. Changes: compact encoding, removal of identifier and subtype predictors, a seeded stratified split, and new samples fitted to training records. The data card records the source CSV’s SHA-256.
Export this experiment.
The recipe, measured results and model parameters all correspond to the completed run shown above.
Source data cardJSON · provenance & limitations↓Generated data is a derived benchmark artifact. Keep the source attribution with your copy.
Have a dataset
that needs to do more?
We can scope the generator, the model and the evaluation around your actual problem.
The architecture.
The measured trade-offs.
Read the fixed benchmark results, alert workload, sensitivity analysis and proposed company pilot.
Read the case study ↗