Sub-Millisecond, Microjoule Edge Inference for Indoor Environment Identification via Layer-Wise Mixed-Precision Quantization
Hamza A. Abushahla, Muhammed Noshin, Dr. Mohamed I. AlHajri, and Dr. Nazar T. Ali
This repository contains code and resources for the paper: Sub-Millisecond, Microjoule Edge Inference for Indoor Environment Identification via Layer-Wise Mixed-Precision Quantization.
Figure 1: System overview of the proposed indoor environment identification framework. The pipeline illustrates the complete workflow, including data preprocessing, model training and quantization, and deployment on the MAX78002.
This work presents the first hardware-validated framework that integrates Quantization-Aware Training (QAT) with layer-wise Mixed-Precision Quantization (MPQ) to enable sub-millisecond, microjoule indoor environment identification on the MAX78002 microcontroller. The main contributions are:
-
We evaluate CNN-based indoor environment identification on resource-constrained IoT hardware by deploying QAT-enabled, layer-wise MPQ models across multiple bandwidth settings, showing that MPQ consistently outperforms uniform quantization.
-
We provide hardware-grounded, on-device deployment characterization by separating weight loading from inference, evaluating clocking modes, and quantifying their impact on energy–latency trade-offs.
-
At 99.21% accuracy, the best MPQ configuration achieves 127.7 µs latency and 27 µJ per inference, while reducing model size by 77% and delivering 10% faster inference and 22% lower inference energy than uniform INT8; it also reduces weight-loading time and energy by 85.2% and 55.3%, respectively.
-
Under a relaxed 98% accuracy requirement, compact MPQ configurations achieve 75.5 µs latency and 15.6 µJ per inference, with reductions of 87% in model size, 46.7% in inference time, 55% in inference energy, 91.5% in weight-loading time, and 74.8% in weight-loading energy, enabling flexible accuracy–efficiency trade-offs.
This repository is organized as overlays on top of the official Analog Devices ai8x toolchain:
-
envs/max_linux.yml/max_mac.yml: recommended conda environments for reproducibility
-
requirements.txt(repo root): fallback pip requirements (Python 3.11) -
training/Overlay files forai8x-training:- dataset (
data/) - dataloader (
datasets/) - model definitions (
models/) - training policies (LR schedule + QAT/MPQ policies) (
policies/) - scripts and sweep drivers (
scripts/, plus sweep scripts at thetraining/root)
- dataset (
-
synthesis/Overlay files forai8x-synthesis:- izer configs / network YAMLs
- generation scripts for hardware project generation
-
inference/Example exported projects and ready-to-run MAX78002 deployments:- generated izer project folders (e.g.,
indoor_env_1d_51_q8824/) - example quantized checkpoints (when applicable)
- generated izer project folders (e.g.,
-
results/Raw experiment outputs and processed summaries used in the paper:- CSV logs for simulation sweeps (QAT / PTQ)
- aggregated summary tables (e.g., mean/std across seeds)
-
figs/Figures used in this repo/README -
hawqv2/(git submodule → HAWQ-V2) Standalone HAWQ-V2 initializer, tools, and bundled example runs. After cloning this repo, rungit submodule update --init --recursive(or clone with--recursive).
Clone this repository with submodules so hawqv2/ is populated (HAWQ-V2 lives in a separate repo):
git clone --recursive <URL-of-this-repository>
# or, if you already cloned without --recursive:
git submodule update --init --recursiveFrom the repo root:
conda env create -f envs/max_linux.yml
# or
conda env create -f envs/max_mac.yml
conda activate maxCreate a fresh conda env called max (Python 3.11), then install the pip requirements inside it:
conda create -n max python=3.11 -y
conda activate max
python -m pip install --upgrade pip
pip install -r requirements.txtWe recommend a single workspace that contains both toolchains:
mkdir max_workdir && cd max_workdir
git clone --recursive https://github.com/analogdevicesinc/ai8x-training.git
git clone --recursive https://github.com/analogdevicesinc/ai8x-synthesis.gitExpected layout:
max_workdir/
├── ai8x-training/
└── ai8x-synthesis/
Copy the contents of this repo’s training/ into your local ai8x-training/ (merge folders; do not nest).
Example mapping (will expand as repo evolves):
Target folder (inside ai8x-training/) |
Copy from (this repo) | Purpose |
|---|---|---|
data/ |
training/data/indoor_environment/ |
Dataset (.mat files) |
datasets/ |
training/datasets/ |
Dataloader(s) |
models/ |
training/models/ |
ai8x model(s) |
policies/ |
training/policies/ |
LR schedule + QAT/MPQ policies |
scripts/ / sweeps/ |
training/scripts/, training/sweeps/ |
training/eval + sweep drivers |
Similarly, merge this repo’s synthesis/ into your local ai8x-synthesis/.
All commands below assume you are inside:
cd max_workdir/ai8x-trainingQAT-enabled run:
python train.py --epochs 10 --batch-size 256 \
--optimizer Adam --lr 0.001 --weight-decay 0.0005 \
--use-bias --deterministic \
--model ai85indoorenvnetv2 --dataset IndoorEnvironment_1D --data data/indoor_environment \
--compress policies/schedule-indoor-env.yaml \
--qat-policy policies/qat_policy_indoor_v2.yaml \
--input-1d-length 101 \
--device MAX78002 --name indoor_runPTQ-only (no QAT):
Set --qat-policy to None (or remove it, depending on your local script conventions).
Float checkpoint for HAWQ: If you want a clean FP32 checkpoint for HAWQ-V2 selection, use:
./scripts/train_float_for_hawq.sh 101 42This runs standard training with --qat-policy None and prints the resulting best.pth.tar path to use as HAWQ input.
Outputs (logs directory):
checkpoint.pth.tar,best.pth.tar: float checkpointsqat_checkpoint.pth.tar,qat_best.pth.tar: checkpoints after QAT starts Note: these are not yet “izer-ready” until quantization/export steps are run by the provided scripts.
MPQ configuration:
Layer-wise bitwidths are controlled in the relevant QAT policy YAML (e.g., qat_policy_indoor_v2.yaml).
Script: train_indoor_1D_mixed_sweep.py
What it does:
- Enumerates all 891 MPQ configs (3^4 over {8,4,2} bits for conv1/conv2/fc1/fc2)
- Sweeps multiple input lengths (α) and multiple seeds
- Runs full QAT, then quantizes and evaluates each run
- Writes detailed and aggregated CSV summaries
By default the sweep uses 5 seeds starting at 42 (seeds 42–46 inclusive), matching the script’s positional defaults. Override with python train_indoor_1D_mixed_sweep.py <num_seeds> <start_seed>.
Run:
python train_indoor_1D_mixed_sweep.pyDefault output folder (example):
ai8x_seed_runs_out/
├── logs_mixed/
├── checkpoints_mixed/
├── policies/sweep/
├── mixed_precision_sweep_results.csv
└── mixed_precision_sweep_summary.csv
Script (in this repo): training/train_indoor_1D_hawq_qat_sweep.py
What it does:
- Consumes HAWQ-retained candidates from
export_hawq_candidates.py. - Runs QAT training,
quantize.py, and eval from inside the ai8x-training tree (same layout as upstreamtrain.py). - Writes
hawq_qat_sweep_results.csvandhawq_qat_sweep_summary.csvunder--out-dir(default relative paths resolve under ai8x-training, not under this repository).
Layout: copy the sweep driver into the root of ai8x-training (the directory that contains train.py). The script assumes ai8x-synthesis is a sibling folder of ai8x-training (e.g. …/testMax/ai8x-training and …/testMax/ai8x-synthesis) so it can find quantize.py.
One-time (or after editing the script here):
export HAWQ_REPO_ROOT=/path/to/Indoor-Environment-Quantization # e.g. …/Indoor-Environment-Quantization
export AI8X_TRAINING_ROOT=/path/to/ai8x-training # e.g. …/testMax/ai8x-training
cp "${HAWQ_REPO_ROOT}/training/train_indoor_1D_hawq_qat_sweep.py" "${AI8X_TRAINING_ROOT}/"1) Export frontier candidates (from this repo, max env):
eval "$(conda shell.bash hook)"
conda activate max
export HAWQ_REPO_ROOT=/path/to/Indoor-Environment-Quantization
cd "${HAWQ_REPO_ROOT}" || exit 1
ALPHAS=(101 91 81 71 61 51 41 31 21 11 5)
HAWQ_DIR=hawqv2/hawqv2_runs/indoor_alpha_sweep_pareto_seed42
ITEMS=()
for A in "${ALPHAS[@]}"; do
ITEMS+=(--item "${A}:${HAWQ_DIR}/results/L${A}.json")
done
python hawqv2/tools/export_hawq_candidates.py \
"${ITEMS[@]}" \
--candidate-set frontier \
--output "${HAWQ_DIR}/hawq_frontier_candidates.csv"2) Dry-run then run the sweep (from ai8x-training; uses the CSV path inside HAWQ_REPO_ROOT):
export HAWQ_REPO_ROOT=/path/to/Indoor-Environment-Quantization
export AI8X_TRAINING_ROOT=/path/to/ai8x-training
ALPHAS=(101 91 81 71 61 51 41 31 21 11 5)
cd "${AI8X_TRAINING_ROOT}" || exit 1
python -u train_indoor_1D_hawq_qat_sweep.py \
--configs-csv "${HAWQ_REPO_ROOT}/hawqv2/hawqv2_runs/indoor_alpha_sweep_pareto_seed42/hawq_frontier_candidates.csv" \
--input-lengths "${ALPHAS[@]}" \
--num-seeds 5 \
--start-seed 42 \
--workers 4 \
--epochs 10 \
--z-score 2.0 \
--out-dir hawq_qat_frontier_out \
--dry-run
# Full run: remove --dry-run when ready
python -u train_indoor_1D_hawq_qat_sweep.py \
--configs-csv "${HAWQ_REPO_ROOT}/hawqv2/hawqv2_runs/indoor_alpha_sweep_pareto_seed42/hawq_frontier_candidates.csv" \
--input-lengths "${ALPHAS[@]}" \
--num-seeds 5 \
--start-seed 42 \
--workers 4 \
--epochs 10 \
--z-score 2.0 \
--out-dir hawq_qat_frontier_outQuick sanity check (single length, one seed, still from AI8X_TRAINING_ROOT):
cd "${AI8X_TRAINING_ROOT}" || exit 1
python -u train_indoor_1D_hawq_qat_sweep.py \
--configs-csv "${HAWQ_REPO_ROOT}/hawqv2/hawqv2_runs/indoor_alpha_sweep_pareto_seed42/hawq_frontier_candidates.csv" \
--input-lengths 101 \
--num-seeds 1 \
--start-seed 42 \
--workers 0 \
--epochs 10 \
--z-score 2.0 \
--out-dir hawq_qat_frontier_out \
--dry-run3) ACE on the reduced QAT summary (back in this repo; rebuild ITEMS as in step 1 if you opened a new shell):
cd "${HAWQ_REPO_ROOT}" || exit 1
python hawqv2/tools/select_ace_from_hawq.py \
"${ITEMS[@]}" \
--eval-csv "${AI8X_TRAINING_ROOT}/hawq_qat_frontier_out/hawq_qat_sweep_summary.csv" \
--eval-source-label QAT \
--candidate-set frontier \
--target-acc 99.2 \
--beta1 1.0 \
--beta2 0.0 \
--limit 20This route evaluates only the HAWQ size--(\Omega) frontier with QAT, while the full 891-point QAT sweep remains the reference/oracle.
Script: train_indoor_1D_mixed_sweep_ptq.py
What it does:
- Trains or loads an FP32 checkpoint for each input length.
- Generates a per-config ai8x quantization policy with layer-wise bit overrides for
conv1,conv2,fc1, andfc2. - Runs PTQ calibration with
calibrate_ptq_simple.py. - Runs ai8x
quantize.pyto produce the quantized checkpoint. - Evaluates the quantized checkpoint with
train.py --evaluate -8using the same policy.
Run (example):
python -u train_indoor_1D_mixed_sweep_ptq.py \
--num-seeds 5 \
--start-seed 42 \
--input-lengths 101 \
--epochs 10 \
--z-score 2.0 \
--calib-split trainOutputs (example):
ai8x_ptq_sweep_out/
├── ptq_sweep_results.csv
├── ptq_sweep_summary.csv
├── logs_ptq/
└── checkpoints_ptq/
Quantized checkpoints follow:
- Uniform:
_q8.pth.tar,_q4.pth.tar,_q2.pth.tar - Mixed precision:
_qmixed.pth.tar
Example:
indoor_mixed_seed_46__L101__8_8_8_8_*_qat_best_q8.pth.tar
Make sure you are using the same Python environment set up for the ai8x toolchain.
Go to ai8x-synthesis:
cd ../ai8x-synthesisEdit the generation script (example):
scripts/gen_indoor_1d.sh
Set inside the script:
LENGTH=101CONFIG="8-8-8-8"(or MPQ config such as8-8-2-4)CHECKPOINT=<path-to-quantized-checkpoint>
Make the script executable (only once):
chmod +x scripts/gen_indoor_1d.shRun the script from the ai8x-synthesis root:
./scripts/gen_indoor_1d.shThe script internally calls ai8xize.py to generate a deployable C/C++ project, e.g.:
python ai8xize.py \
--test-dir "$TARGET" \
--prefix "$PREFIX" \
--checkpoint-file "$CHECKPOINT" \
--config-file networks/indoorenvnet-v2-chw-${LENGTH}.yaml \
--sample-input tests/sample_indoorenvironment_1d_${LENGTH}.npy \
--overwrite --softmax --compact-data --mexpress --max-speed --energyA C/C++ project folder will be created (example):
HW_Evaluation/indoor_env_1d_101_q8888/
This includes:
main.c(entry point)cnn.c,cnn.h(generated network)- Makefile / Eclipse launch files (depending on izer output)
Import the generated project into Eclipse after setting up the Analog Devices MSDK:
For paper-style evaluation, we modify main.c to report:
- inference latency
- energy/power measurements (as used in our evaluation scripts)
This repo includes example exported projects and checkpoints under:
inference/
Example:
inference/indoor_env_1d_51_q8824/inference/indoor_env_1d_91_q8824/inference/indoor_mixed_seed_46__L51__8_8_2_4_indoor_mixed_seed_46__L51__8_8_2_4_qat_best_qmixed.pth.tarinference/indoor_mixed_seed_46__L91__8_8_2_4_indoor_mixed_seed_46__L91__8_8_2_4_qat_best_qmixed.pth.tar
If you use our work for your own research, please cite us with the below:
@article{abushahla2026submillisecond,
title={Sub-Millisecond, Microjoule Edge Inference for Indoor Environment Identification via Layer-Wise Mixed-Precision Quantization},
author={Abushahla, Hamza A. and Noshin, Muhammed and AlHajri, Mohamed I. and Ali, Nazar T.},
journal={IEEE Internet of Things Journal},
year={2026},
publisher={IEEE},
doi={10.1109/JIOT.2026.3710680}
}You can also reach out through email to:
- Hamza Abushahla - b00090279@alumni.aus.edu
- Dr. Mohamed AlHajri - mialhajri@aus.edu
