To generate masking configurations for a dataset, use:
make generate_masking_configurations ORIGINAL_DATA="path/to/your/dataset.csv" CLASS_LABEL="target_column_name" TOTAL_CONFIGURATIONS=10 CONFIG_DIR="configs"ORIGINAL_DATA- Path to the original CSV datasetCLASS_LABEL- Name of the target variable (excluded from masking)TOTAL_CONFIGURATIONS- Number of masking configurations to generateCONFIG_DIR- Directory where configurations will be stored (default: configs)
The command will generate YAML configuration files in the configs directory. Each configuration includes:
- Dataset path information
- Masking functions to be applied to each attribute
Example configuration:
dataset:
original_path: path/to/your/dataset.csv
target_variable: target_column_name
masking:
attributes:
age:
- function: generalize
params:
M: 3
income:
- function: suppress
params: {}To apply a masking configuration and generate a masked dataset, use:
make mask_data MASKING_CONFIG="configs/config_1.yaml" MASKED_DATA="output/masked_dataset.csv"MASKING_CONFIG- Path to a YAML masking configuration fileMASKED_DATA- Path to the output masked dataset (including directory)
The command will apply the specified masking configuration to the original dataset (defined in the configuration file) and save the masked data to the specified output location.
Before performing data reconstruction, you need to compute distributions. Use:
make compute_distribution DATASET="path/to/dataset.csv" MARGINALS="path/to/marginals.pkl" JOINT_DISTRIBUTION="path/to/joint_distribution.pkl" CLASS_LABEL="target_column_name"DATASET- Path to the input dataset CSV fileMARGINALS- Path to save marginals (frequency distributions for each attribute)JOINT_DISTRIBUTION- Path to save joint distributions with class labelCLASS_LABEL- Name of the target/class variable column
The command computes:
- Marginals: frequency counts for each value in each attribute
- Joint distributions: joint distributions between each attribute and the class label
These distributions are saved as pickle files and are required for the reconstruction algorithms.
To evaluate the PUD between masked and reconstructed distributions, use:
make evaluate_pud METRIC="chi2" MASKED_JOINT_DIST="path/to/masked.pkl" RECONSTRUCTED_JOINT_DIST="path/to/reconstructed.pkl"METRIC- Metric to use for evaluation (options: chi2, mi, tvd, g3)MASKED_JOINT_DIST- Path to masked joint distribution pickle fileRECONSTRUCTED_JOINT_DIST- Path to reconstructed joint distribution pickle file
The command calculates and displays PUD across all attributes
Available metrics:
chi2- Chi-square statisticmi- Mutual Informationtvd- Total Variation Distanceg3- G3 metric
To run a machine learning model on a dataset and evaluate its performance, use:
make model DATA="path/to/dataset.csv" MODEL="path/to/model_file.py" TARGET="target_column_name"DATA- Path to the dataset CSV file to evaluateMODEL- Path to the model Python file to use for evaluationTARGET- Name of the target variable column
The command will train the specified model on the dataset and output its accuracy.