Raman analysis results depend heavily on preprocessing. On the same set of spectra, changing the baseline correction or the intensity normalisation is enough to shift reported accuracy by several percentage points. When every group uses its own pipeline without recording parameters, the numbers in their papers are not comparable.
The RamanBench 2026 benchmark standardised 74 public datasets covering 325,668 spectra and reported one crucial finding: none of the architectures tested generalised across datasets. In other words, high accuracy on one dataset says nothing about real-world usability on a different instrument, different samples or a different laboratory.
The main cause lies not in the models but in the data: acquisition conditions recorded incompletely and preprocessing applied inconsistently. That is why BIRALAB puts this standard ahead of publishing a dataset at all — standardising how measurements are recorded and processed is a precondition for cross-instrument generalisation to be solvable.
Omitting any of the groups in the table below renders the data unusable for cross-instrument generalisation — the group's principal research direction.