Skip to content

Add BRAX_Dataset for the BRAX Brazilian chest X-ray dataset - #194

Open
AmlanMishra2004 wants to merge 3 commits into
mlmed:mainfrom
AmlanMishra2004:add-brax-dataset
Open

AmlanMishra2004 wants to merge 3 commits into
mlmed:mainfrom
AmlanMishra2004:add-brax-dataset

Conversation

@AmlanMishra2004

@AmlanMishra2004 AmlanMishra2004 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds BRAX_Dataset, a loader for BRAX v1.1.0 (Reis et al., Scientific Data 2022, https://doi.org/10.1038/s41597-022-01608-8): 40,967 chest radiographs from Hospital Israelita Albert Einstein, São Paulo, labeled from Brazilian Portuguese reports with a Portuguese adaptation of the CheXpert labeler. It is one of the few non-US/European CXR datasets and a natural external test set for CheXpert/MIMIC-trained models, but xrv has no loader for it.

Because the labels use CheXpert's encoding, BRAX_Dataset follows CheX_Dataset/MIMIC_Dataset closely:

  • Same 13 pathologies in the same order (sorted, then "Pleural Effusion" renamed to "Effusion"), so it merges with CheXpert/MIMIC without relabelling.
  • -1 (uncertain) becomes NaN; "No Finding" zeroes every other label except Support Devices.
  • imgpath is the release's images/ PNG folder, csvpath is master_spreadsheet_update.csv. Images are 8-bit grayscale PNGs (normalize(maxval=255)).
  • views filters on ViewPosition; the ~27% of rows with no ViewPosition get "UNKNOWN" via limit_to_selected_views.
  • csv gets the usual patientid, age_years (BRAX gives 5-year age groups, with the oldest as the string "85 or more", mapped to 85), sex_male, sex_female.

License / access: BRAX is under the PhysioNet Credentialed Health Data License 1.5.0 (credentialed account, CITI training, signed DUA). This PR ships no BRAX data or labels; csvpath is required, like MIMIC_Dataset. The docstring states the license and access requirements.

Counts: the paper and PhysioNet page say 19,351 patients and 24,959 studies, but the official master_spreadsheet_update.csv has 18,442 distinct PatientIDs and 23,276 distinct AccessionNumbers (the image count, 40,967, matches). The docstring therefore only states the image count.

Two BRAX-specific choices worth flagging:

  • max_aspect_ratio=10: 12 labeled images in the release are thin strips (12-118 px on one side, 19:1 to 213:1 aspect ratio) rather than chest radiographs; the bytes match the release's own SHA256SUMS.txt, so this is in the source data. Every other image is below 9:1. Two of them also exceed Pillow's decompression-bomb pixel limit. They are dropped by default using the CSV's Rows/Columns; max_aspect_ratio=None keeps them. This is similar in spirit to the known-bad image list in PC_Dataset, but doesn't hardcode filenames.
  • unique_patients uses drop_duplicates("PatientID") rather than groupby("PatientID").first(). groupby().first() takes the first non-null value per column, so for a patient with several images the resulting row can combine labels from different images (e.g. image 1 is "No Finding", image 2 has Cardiomegaly=1, and the merged row gets both). drop_duplicates keeps one real image. Happy to switch to groupby().first() for consistency with the other loaders if you prefer.

Testing

  • 2 new tests in tests/test_dataloaders.py using a small synthetic master_spreadsheet_update.csv: default view/unique-patient filtering, the aspect-ratio filter (and turning it off), UNKNOWN views, -1 to NaN, No Finding zeroing, that a second image's labels don't leak into a patient's row, "85 or more" ages, and image loading from the images/ path.
  • Full tests/test_dataloaders.py suite passes (23/23).
  • Verified against the real BRAX v1.1.0 label CSV: views=["*"], unique_patients=False loads 40,955 images (40,967 without the aspect filter), views=["PA"] loads 13,958 images / 12,640 unique patients, and all views give 18,442 unique patients (every patient ID in the CSV). Labels matched an independent loader image-for-image across ["*"], ["PA"] and ["PA", "AP"]. A real release PNG loads as (1, 1649, 1024) in [-1024, 1024].
  • pep8.sh selection (autopep8 --select=E1,E2,E3,W1,W2) reports nothing on the new code.

Made with Cursor

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant