Quantization pipelines under one library.
Under development
- Overview
- Features
- Supported Methods
- Installation
- Usage
- Project Structure
- System Architecture
- Workflow Process
- Contributing
- License
quantool provides a unified interface for various model quantization methods, allowing users to compress and accelerate machine learning models with ease.
- Unified CLI for multiple quantization algorithms
- Plugin-based architecture for extendability
- Built-in support for logging experiments with MLflow and Weights & Biases
- Modular core pipeline for preprocessing, quantization and evaluation
- GGUF
- AWQ
- GPTQ
- SmoothQuant
- GPTQv2
- EXL2
- AQLM
- HIGGS
- MLX
git clone https://github.com/langtech-bsc/quantool.git
cd quantool
pip install -e .To install llama.cpp only:
git clone https://github.com/langtech-bsc/quantool.git
cd quantool
pip install -e '.[llama-cpp]'Basic CLI example:
quantool config.yamlRun quantool --help for a full list of options.
quantool uses dataclass-based argument groups. You can pass a YAML config containing values for these groups or override individual fields on the CLI. The most common argument groups are:
-
ModelArguments
model_id(str) — Path or HF repo ID of the pretrained model.tokenizer_name(str, optional) — Tokenizer path/name if different from model.cache_dir(str, optional) — Cache directory for model/tokenizer.
-
QuantizationArguments
method(str) — Quantization method (e.g.gptq,awq,gguf).quant_level(str, optional) — Specific quant set label.quantization_config(dict) — Method-specific configuration (passed as a JSON string on the CLI).
-
CalibrationArguments
dataset_id/dataset_path— HF dataset id or local dataset path used for calibration.dataset_config(str) — Dataset config name for HF datasets.split(str) — Split to use (e.g.train).dataset_cache_dir(str, optional) — Cache dir for datasets.sample_size(int or null) — Number of calibration examples (null = full split).load_in_pipeline(bool) — If true, pipeline loads and preprocesses the dataset and passes it to the quantizer.preprocess_fn(str, optional) — Optional preprocessing function (module.func) to run on examples before calibration.
-
ExportArguments
output_path(str) — Where to save the exported quantized model.push_to_hub(bool) — Whether to push the result to the HF Hub.
-
EvaluationArguments
enable_evaluation(bool) — Run evaluation after quantization.eval_dataset(str) — Dataset to use for evaluation.
This README lists the most used fields; consult the dataclass definitions in src/quantool/args/quantization_args.py for the full set.
- Run with a YAML configuration file (recommended for repeatability):
quantool test_gptq_config.yaml
# Or, when running via python module entrypoint:
python -m quantool.entrypoints.cli test_gptq_config.yaml- Run by passing individual CLI arguments (overrides YAML or can be used standalone).
Notes:
- Field names are taken from dataclass attribute names. For dictionary fields like
quantization_configsupply a JSON string on the CLI. - If multiple dataclasses expose the same field name a conflict may occur; prefer YAML in that case.
- Field names are taken from dataclass attribute names. For dictionary fields like
Example (override some options on the CLI):
quantool \
--model_id facebook/llama-7b \
--method gguf \
--quantization_config '{"llama_cpp_path": "/path/to/llama.cpp"}' \
--dataset_id wikitext \
--sample_size 512Or using the python -m entrypoint:
python -m quantool.entrypoints.cli \
--model_id facebook/llama-7b \
--method gguf \
--quantization_config '{"llama_cpp_path": "/path/to/llama.cpp"}' \
--dataset_id wikitext \
--sample_size 512If you need complex nested method configuration, prefer a YAML file and pass it as the single positional argument.
Use a YAML config file to run quantization with llama_cpp to do gguf quantization:
# config.yaml
model_id: "facebook/llama-7b"
method: "gguf"
quantization_config:
llama_cpp_path: "/path/to/llama.cpp/executable".
├── assets/ # Diagrams and images
├── src/
│ └── quantool/ # Library source
│ ├── args/
│ ├── core/
│ ├── entrypoints/
│ ├── loggers/
│ ├── methods/
│ └── utils/
├── tests/ # Unit and integration tests
│ ├── entrypoints/
│ ├── methods/
│ └── utils/
├── pyproject.toml
└── test_gguf_config.yaml # Example configuration
graph LR
CLI[CLI Entrypoint] --> ArgsParser[Args Parser]
ArgsParser --> Registry[Method Registry]
Registry --> Method[Quantization Method Plugin]
Method --> Core[Core Pipeline]
Core --> Output[Quantized Model Output]
CLI --> Loggers{Logging Integrations}
Loggers --> MLflow[MLflow]
Loggers --> WAndB[Weights & Biases]
flowchart TD
A[Start: User invokes CLI] --> B[Parse Arguments]
B --> C[Resolve Method from Registry]
C --> D[Load Model and Data]
D --> E[Execute Quantization Pipeline]
E --> F[Log Metrics & Artifacts]
F --> G[Save Quantized Model]
G --> H[End]
Contributions are welcome! Please read CONTRIBUTING.md for guidelines.
Distributed under the Apache-2.0 License. See LICENSE for more information.
