This project provides a modular pipeline for preparing parallel sentence-level corpora for machine translation research.
It standardises preprocessing across datasets, making it easier to train and evaluate translation models on consistent, high-quality data.
The pipeline is:
- Modular – enable or disable steps as needed.
- Configurable – control all behaviour via a YAML config file.
- Reproducible – consistent outputs for large-scale experiments.
- Scalable – simple to adapt for parallelisation in an HPC environment.
- Input handling: Read corpora in plain text (
.txt), tab-separated (.tsv) or translation memory (.tmx) formats. - Embeddings: Compute multilingual sentence embeddings (e.g. LaBSE (default), SONAR (can optionally be added)).
- Language ID: Calculate probability that segments are in the desired language (using GlotLID).
- Filtering: Filter by user-defined embedding scores and language probability thresholds.
- Deduplication: Remove duplicate sentence pairs and fuzzy matches across corpora.
- Bifixer: Enables optional Bifixer-based processing if Bifixer is installed. Requires separate installation. By default, deduplication and segmentation are skipped.
- Normalisation: Standardise punctuation, spacing, and casing. Includes easy-to extend language specific normalisation.
- Python ≥3.8
- Sufficient disk space (embeddings can be large).
- (Optional) HPC cluster or multi-core machine for parallel processing.
Clone the repository and set up a virtual environment:
git clone https://github.com/langtech-bsc/ParaCLEAN
cd ParaCLEAN
python -m venv venv
source venv/bin/activate # or venv\Scripts\activate on Windows
pip install -r requirements.txtRun the pipeline with a configuration file:
python pipeline.py --config config_multi.yamlProgress and outputs are logged to the console.
Each step writes intermediate .tsv files to the specified output directory.
The pipeline supports both single-corpus and multi-corpus runs.
- Single-corpus runs: one corpus passes sequentially through all selected steps.
- Multi-corpus runs: several corpora are processed individually up to the scoring steps, then merged for filtering, deduplication, and normalisation.
input Reads and normalises the raw input format.
embeddings Computes sentence embeddings for filtering.
langid Runs language identification on both sides.
filter Applies thresholds for similarity and language probability.
dedup Removes exact and near-duplicate sentence pairs.
bifixer Runs optional Bifixer cleaning (requires Bifixer installed).
normalise Applies final punctuation and spacing normalisation.
# ===========================
# Pipeline configuration file
# Example: single corpus run
# ===========================
# Output directory (will be created if missing)
output: "data/Europarl"
# Languages
l1: "es"
l2: "de"
# Embedding model (choices: labse, sonar if downloaded)
model: "labse"
model_path: null # optional, if you want to point to a local model path.
# Filtering thresholds
alignment_score: 0.75
langid_l1_prob: 0.5
langid_l2_prob: 0.5
# Pipeline steps to run (in order)
steps:
- input
- embeddings
- langid
- filter
- dedup
- bifixer
- normalise
# Optional Bifixer flags. More info can be found at https://github.com/bitextor/bifixer
bifixer_flags: ["--ignore_segmentation", "--ignore_duplicates"]
# Input corpus (single)
input: ["data-storage/Europarl.es-de.es", "data-storage/Europarl.es-de.de"]
format: "plain_text" # or "tsv" or "tmx"# ===========================
# Pipeline configuration file
# Example: multi-corpus run
# ===========================
output: "testing/multi"
l1: "Catalan"
l2: "cmn_Hani"
model: "labse"
model_path: null
alignment_score: 0.75
langid_l1_prob: 0.5
langid_l2_prob: 0.5
steps:
- filter
- dedup
- bifixer
- normalise
bifixer_flags: ["--ignore_segmentation", "--ignore_duplicates"]
inputs:
- name: "TED2020"
type: "plain_text"
start_from: "testing/multi/TED2020.embeddings.tsv"
steps: ["langid"]
- name: "QED"
type: "plain_text"
start_from: "testing/multi/QED.embeddings.tsv"
steps: ["langid"]Each corpus runs its own per-corpus steps (input, embeddings, langid) before merging.
The merged dataset then passes through filter, dedup, bifixer, and normalise.
Each step writes a .tsv file in the specified output directory. Intermediate files are named after their processing step, e.g.:
Europarl.embeddings.tsv
Europarl.langid.tsv
Europarl.filtered.tsvThe final output (after normalise) contains four columns:
- Source language (original)
- Target language (original)
- Source language (normalised)
- Target language (normalised)
Language codes can be specified flexibly and are resolved internally. The following are equivalent for Catalan:
- Catalan
- ca
- ca_Latn
- cat
SONAR embeddings: inputs longer than 514 tokens are truncated.
Disk space: embeddings can be large; ensure sufficient space.
Extensibility: filtering and normalisation rules can be customised by editing the corresponding modules in steps/ or adding language-specific rules in normalisation
To add a new processing step:
- Create a module in steps/ e.g.
steps/my_step.py). - Define a function with a consistent interface (
input_path,output_path, etc.). - Register it in
pipeline.pywithin thestep_fnsdictionary.
Example:
"my_step": lambda p: my_step.run(current, p + ".my_step.tsv", l1, l2)This project is released under the Apache 2.0 License. For details, please see the license tab here.
Note: The ACL Anthology entry is forthcoming. In the meantime, please cite via the LREC proceedings.
Mash, A., Bohman, E. P., & Melero, M. (2026). ParaCLEAN: Improving Translation Quality
through Systematic Parallel Data Cleaning. In Proceedings of the Fifteenth Language
Resources and Evaluation Conference (LREC 2026), pp. 6630–6640, Palma, Mallorca, Spain.
European Language Resources Association (ELRA).
https://doi.org/10.63317/36e3vfurjna4
BibTeX
@inproceedings{mash-etal-2026-paraclean,
title = "{P}ara{CLEAN}: Improving Translation Quality through Systematic Parallel Data Cleaning",
author = "Mash, Audrey and Bohman, Ella Paulina and Melero, Maite",
booktitle = "Proceedings of the Fifteenth Language Resources and Evaluation Conference",
pages = "6630--6640",
month = may,
year = "2026",
address = "Palma, Mallorca, Spain",
publisher = "European Language Resources Association",
url = "https://lrec.elra.info/lrec2026-main-527",
doi = "10.63317/36e3vfurjna4",
issn = "2522-2686",
isbn = "978-2-493814-49-4",
}This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA.
This project builds upon components and concepts from: