BIRALAB

Data Hub

BIRALAB open Raman spectral data, preprocessing standard and open-source tools — released with complete acquisition metadata so third parties can reproduce the results.

What the Data Hub is for

The Data Hub is where BIRALAB publishes Raman spectral data, its preprocessing standard and open-source tools. The goal is not storage but becoming the standardisation layer of Vietnam's Raman ecosystem: when other groups use our data and tools as a reference, a shared basis for comparison emerges and domestic research results become reproducible.

Everything here is released under an open licence, with full acquisition conditions so users can judge fitness for purpose before downloading. We publish no data containing patient-identifying information; clinical data is shared only after full de-identification and ethics committee approval.

When other groups, at home and abroad, use your data and tools as their reference standard, a leading position is established objectively, independently of publication counts.
BIRALAB open-data strategy

Three ways into the Data Hub

  • Open datasets

    Raman spectral sets with acquisition metadata, a data dictionary, licence terms and a citable DOI.

    Browse datasets
  • Open-source tools

    Preprocessing libraries, classification models and edge-deployment code, each with usage documentation.

    Browse tools
  • Preprocessing standard

    The mandatory metadata fields for every measurement and the recommended spectral preprocessing pipeline, step by step.

    Read the standard

Datasets

Each dataset states its spectrum count, instrument, version and usage licence.

Tools

Open source that runs as published, frozen at the version tag cited in the corresponding paper.

  • Language: PythonLicence: MIT

    raman-skincancer-augmentation

    Open-source implementation of physically constrained data augmentation and interpretable machine learning for class-imbalanced Raman diagnosis of skin cancer. Full preprocessing chain, augmentation pipeline, PCA-30 + logistic regression classifier, and leakage-safe cross-validation. Regenerates every figure and metric of the paper.