in silico
Phase 1: In Silico Validation
- Objective
- Determine whether physically-valid spectral augmentations (Gaussian broadening, baseline drift, scattering slope, dilution) preserve chemical signal (SNR > 1) and whether SSL pretext tasks can theoretically outperform supervised baselines at n_label ≤ 50, using synthetic and open-access UV-Vis spectra, before any wet-lab investment.
- Estimated cost
- €500-2000 (compute credits if cloud, otherwise €0 with local GPU)
- Estimated duration
- 4-8 weeks
- Success criteria
- SNR of NO3- band after augmentation at sigma=8 nm, drift=0.05 AU · SNR > 1 (chemical signal not destroyed) · (Ratio of integrated NO3- band absorbance to nuisance variance in 200-240 nm)
- RMSE reduction SSL vs best supervised at n_label=25 · ≥ 15% relative reduction, Wilcoxon p < 0.05 across 20 seeds · (Paired Wilcoxon signed-rank on 20 seed-paired RMSE values)
- Optimal sigma* location · sigma* in [2,5] nm with RMSE at sigma* ≥ 15% lower than at sigma=0 and sigma=8 · (Quadratic fit to RMSE vs sigma, ANOVA + Tukey HSD)
- Effective rank increase SSL vs supervised at n_label=25 · ≥ 2 dimensions increase out of 128 · (Exponential of entropy of singular value distribution, Wilcoxon test)
- Band reconstruction error after pretraining · Nitrate < 0.01 AU, CDOM < 0.02 AU, ≥ 30% reduction vs random init · (Linear decoder RMSE on held-out spectra, paired t-test)
- Go if
- At least 3 of 5 success criteria met, including RMSE reduction ≥ 15% at n_label=25 AND SNR > 1 at sigma=8 nm
- No-go if
- SNR < 1 at sigma ≤ 4 nm (augmentation destroys chemical signal) OR RMSE reduction < 5% at n_label=25 across all pretext tasks
- Pivot if
- RMSE reduction 5-15% OR optimal sigma* outside [2,5] nm: pivot to weaker augmentations (sigma ≤ 4 nm) or perturbation-prediction only, and re-run Phase 1 with adjusted boundaries
- Risks
- Synthetic generator does not match real UV-Vis spectra statistics, leading to over-optimistic SSL resultsProbability: highMitigation: Calibrate generator against ≥ 3 open real datasets; compute maximum mean discrepancy (MMD) between synthetic and real spectra; if MMD > threshold, add real unlabeled spectra to pretraining corpus
- Contrastive collapse despite VICReg regularization at batch 512Probability: mediumMitigation: Monitor effective rank during pretraining; if rank < 10, increase VICReg variance weight or switch to MoCo momentum encoder
- Supervised baselines (PLS, SVR) are already near-optimal at n_label=25, leaving no room for SSL improvementProbability: mediumMitigation: Tune PLS n_components and SVR C/gamma via nested cross-validation; if baselines are near-optimal, pivot to harder regime (n_label=5-10) or noisier labels
- Compute budget insufficient for 20 seeds × 7 methods × 6 n_label valuesProbability: lowMitigation: Use Hydra for parallel sweeps, reduce seeds to 10 for exploratory runs, use mixed precision and gradient accumulation