BIRALAB

BIRALAB preprocessing standard

The mandatory metadata fields for every measurement and the recommended Raman preprocessing pipeline — published openly so other groups can compare results on a common basis.

CC BY 4.0

Why a standard is needed

Raman analysis results depend heavily on preprocessing. On the same set of spectra, changing the baseline correction or the intensity normalisation is enough to shift reported accuracy by several percentage points. When every group uses its own pipeline without recording parameters, the numbers in their papers are not comparable.

The RamanBench 2026 benchmark standardised 74 public datasets covering 325,668 spectra and reported one crucial finding: none of the architectures tested generalised across datasets. In other words, high accuracy on one dataset says nothing about real-world usability on a different instrument, different samples or a different laboratory.

The main cause lies not in the models but in the data: acquisition conditions recorded incompletely and preprocessing applied inconsistently. That is why BIRALAB puts this standard ahead of publishing a dataset at all — standardising how measurements are recorded and processed is a precondition for cross-instrument generalisation to be solvable.

Omitting any of the groups in the table below renders the data unusable for cross-instrument generalisation — the group's principal research direction.

Eight mandatory field groups for every measurement

Every spectrum in a BIRALAB dataset must carry at least the following fields. This is what determines scientific quality: a group strong in AI but careless in recording acquisition conditions produces irreproducible data — the most common and hardest-to-detect failure in spectroscopy.

Table of the eight field groups that must be recorded for every Raman measurement
Field groupMandatory fields
InstrumentManufacturer, model, unit serial number, firmware version
Laser sourceWavelength (nm), power at sample (mW), linewidth
OpticsObjective, numerical aperture, spot size, grating (lines/mm)
AcquisitionIntegration time, number of accumulations, wavenumber range, spectral resolution
CalibrationReference material used for wavenumber calibration, most recent calibration date
SampleSample type, preparation, temperature, humidity, time since collection
PreprocessingBaseline subtraction, normalisation, smoothing and cosmic ray removal methods — parameters stated explicitly
LabelLabel source (histopathology, HPLC, reference chromatography), annotator, confidence

Recommended preprocessing pipeline

The order of the steps matters: reordering them changes the result. For each step, record the method and every parameter in the dataset metadata.

  1. Step 1

    Cosmic ray removal

    Detect anomalous sharp spikes by comparing repeated acquisitions or with a median filter along the wavenumber axis. Record the detection threshold and the number of points replaced. Do this first, because a surviving cosmic spike distorts every later step.

  2. Step 2

    Instrument background subtraction

    Subtract dark current and the background of the holder, cuvette or substrate, measured at the same integration time. Record which background spectrum was used.

  3. Step 3

    Wavenumber-axis calibration

    Calibrate against a reference material with known bands, for example silicon at 520.7 cm⁻¹, polystyrene or paracetamol. Record the material, the peaks used for fitting and the calibration date. This step decides whether two instruments can be compared at all.

  4. Step 4

    Wavenumber range cropping

    Keep the fingerprint region appropriate to the problem, typically 400–1800 cm⁻¹ for biological samples. Record the retained range; never crop by intuition after inspecting classification results.

  5. Step 5

    Baseline correction

    Use a method with publicly stated parameters, for example asymmetric least squares, SNIP or polynomial fitting. Record the method, polynomial order or smoothing coefficient, and iteration count.

  6. Step 6

    Smoothing

    Apply Savitzky–Golay smoothing with the window length and polynomial order stated. Too wide a window erases narrow discriminative peaks, so re-check peak widths after smoothing.

  7. Step 7

    Intensity normalisation

    Use SNV, unit-vector normalisation or normalisation to an internal reference peak. Record the chosen method; this step has the strongest effect on whether a model transfers to another instrument.

  8. Step 8

    Resampling onto a common wavenumber grid

    Interpolate onto a common wavenumber grid at 1 cm⁻¹ spacing and record the interpolation method. Mandatory if the data will be pooled or compared across instruments.

Applying and citing the standard

Apply the standard as two separable parts. Recording: add the eight field groups above to the metadata file accompanying every measurement, one row per spectrum. Processing: run the eight steps in order and store every parameter along with the random seed, so that someone else re-running the pipeline obtains the same result.

The standard is released under the open CC BY 4.0 licence. You are free to apply, modify and redistribute it, provided you credit the source. If you modify the pipeline, state clearly how it differs from the original so results remain comparable.

Citing the standard

In the methods section of your paper, state the version of the standard you used and cite it as follows:

BIRALAB (2026). BIRALAB Raman spectral preprocessing standard, version 1.0. Biomedical Raman & AI Analytics Laboratory, International School, Vietnam National University, Hanoi. https://biralab.org/en/data-hub/standards

Comment on the standard

This standard exists for the community to share, so it needs input from groups measuring real spectra on real instruments. If you find a missing field, a step that does not suit your sample type, or you want to jointly standardise data across two laboratories, get in touch.