Repository navigation
Add BRAX_Dataset for the BRAX Brazilian chest X-ray dataset - #194
Open
AmlanMishra2004 wants to merge 3 commits into
Open
AmlanMishra2004 wants to merge 3 commits into
AmlanMishra2004 wants to merge 3 commits into
Conversation
AmlanMishra2004
force-pushed
the
add-brax-dataset
branch
from
October 1, 2026 17:06
fa322e4 to
5d5f501
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
BRAX_Dataset, a loader for BRAX v1.1.0 (Reis et al., Scientific Data 2022, https://doi.org/10.1038/s41597-022-01608-8): 40,967 chest radiographs from Hospital Israelita Albert Einstein, São Paulo, labeled from Brazilian Portuguese reports with a Portuguese adaptation of the CheXpert labeler. It is one of the few non-US/European CXR datasets and a natural external test set for CheXpert/MIMIC-trained models, but xrv has no loader for it.Because the labels use CheXpert's encoding,
BRAX_DatasetfollowsCheX_Dataset/MIMIC_Datasetclosely:-1(uncertain) becomesNaN; "No Finding" zeroes every other label except Support Devices.imgpathis the release'simages/PNG folder,csvpathismaster_spreadsheet_update.csv. Images are 8-bit grayscale PNGs (normalize(maxval=255)).viewsfilters onViewPosition; the ~27% of rows with noViewPositionget"UNKNOWN"vialimit_to_selected_views.csvgets the usualpatientid,age_years(BRAX gives 5-year age groups, with the oldest as the string"85 or more", mapped to 85),sex_male,sex_female.License / access: BRAX is under the PhysioNet Credentialed Health Data License 1.5.0 (credentialed account, CITI training, signed DUA). This PR ships no BRAX data or labels;
csvpathis required, likeMIMIC_Dataset. The docstring states the license and access requirements.Counts: the paper and PhysioNet page say 19,351 patients and 24,959 studies, but the official
master_spreadsheet_update.csvhas 18,442 distinctPatientIDs and 23,276 distinctAccessionNumbers (the image count, 40,967, matches). The docstring therefore only states the image count.Two BRAX-specific choices worth flagging:
max_aspect_ratio=10: 12 labeled images in the release are thin strips (12-118 px on one side, 19:1 to 213:1 aspect ratio) rather than chest radiographs; the bytes match the release's ownSHA256SUMS.txt, so this is in the source data. Every other image is below 9:1. Two of them also exceed Pillow's decompression-bomb pixel limit. They are dropped by default using the CSV'sRows/Columns;max_aspect_ratio=Nonekeeps them. This is similar in spirit to the known-bad image list inPC_Dataset, but doesn't hardcode filenames.unique_patientsusesdrop_duplicates("PatientID")rather thangroupby("PatientID").first().groupby().first()takes the first non-null value per column, so for a patient with several images the resulting row can combine labels from different images (e.g. image 1 is "No Finding", image 2 has Cardiomegaly=1, and the merged row gets both).drop_duplicateskeeps one real image. Happy to switch togroupby().first()for consistency with the other loaders if you prefer.Testing
tests/test_dataloaders.pyusing a small syntheticmaster_spreadsheet_update.csv: default view/unique-patient filtering, the aspect-ratio filter (and turning it off),UNKNOWNviews,-1toNaN, No Finding zeroing, that a second image's labels don't leak into a patient's row,"85 or more"ages, and image loading from theimages/path.tests/test_dataloaders.pysuite passes (23/23).views=["*"], unique_patients=Falseloads 40,955 images (40,967 without the aspect filter),views=["PA"]loads 13,958 images / 12,640 unique patients, and all views give 18,442 unique patients (every patient ID in the CSV). Labels matched an independent loader image-for-image across["*"],["PA"]and["PA", "AP"]. A real release PNG loads as(1, 1649, 1024)in[-1024, 1024].pep8.shselection (autopep8 --select=E1,E2,E3,W1,W2) reports nothing on the new code.Made with Cursor