DG

COBRA dataset

6,000+ curated whole-slide images of basal cell carcinoma, benign skin and other skin cancers, public and browsable in a web viewer.

Dataset author, viewer builder · 2024 – present · visit →

Most public pathology datasets are a download link and a CSV. Before you see a single slide you have to download hundreds of gigabytes, pick a viewer and write a loader. I wanted COBRA to be as accessible a dataset as possible. As part of that I built this viewer, so anyone can quickly browse through the dataset before downloading anything.

The COBRA slide viewer with a basal cell carcinoma case open: case list and filters on the left, the pathology report above the slide, and six serial tissue sections in the viewer.cobra.daangeijs.nl

Browse the cases and open any slide in your browser.

Open the viewer →

What COBRA is

COBRA stands for Classification Of Basal cell carcinoma, Risky skin tumors and Abnormalities. It is the dataset I collected during my PhD at Radboudumc: several thousand whole-slide images of skin biopsies and excisions, released publicly for research.

Most slides are basal cell carcinoma and non-malignant skin. The rest is everything a BCC model should not confuse with a BCC: squamous cell carcinoma, melanoma, lymphoma, Merkel cell carcinoma, cutaneous metastases and a long list of rare adnexal tumors. I added those on purpose. A BCC model is only useful in practice if it also knows when it is looking at something else, and you need that something else in your data to test it.

Every slide has a diagnosis label. Most come with the original pathology report. A subset of resections has pixel-level annotations of tissue, tumor and epidermis.

The viewer

The viewer is built on OpenLayers. For my Simple WSI viewer I used OpenSeadragon, but I have come to prefer OpenLayers. It streams the tiles straight from storage, so the whole thing is a relatively simple frontend without a backend.

Using the data

The viewer is for looking. If you want to download the dataset in bulk or work with it in your own code, use the COBRA data toolkit. That is also the record to cite when you use the dataset.

The dataset is released under a CC BY 4.0 license, so you can use it for academic and commercial work as long as you give credit. The toolkit code is Apache 2.0.

Citing the dataset

Cite
@dataset{geijs2025cobra,
  author    = {Geijs, Daan},
  title     = {COBRA data toolkit: A whole-slide image dataset of basal cell carcinoma and diverse skin malignancies for computational dermatopathology},
  year      = {2025},
  publisher = {Zenodo},
  version   = {1.0.0-alpha},
  doi       = {10.5281/zenodo.18106134},
  url       = {https://doi.org/10.5281/zenodo.18106134},
}