This guide explains how to run a Grail validator. Validators verify miners’ rollouts with GRAIL proofs, compute miner scores over rolling windows, and set weights on-chain. For this release the active environments are coding (MBPP, HumanEval) and math (GSM8K, MATH); the trainer publishes the active environment per checkpoint.
- Introduction
- Prerequisites
- Quick Start
- Configuration
- Running the Validator
- Operations
- Troubleshooting
- Reference
Grail validators:
- Track Bittensor blocks and operate on complete windows.
- Fetch miners’ window files from object storage using per-miner read credentials committed on-chain.
- Verify each rollout with the Verifier (GRAIL proof + environment evaluation + signature checks).
- Score miners by unique successful solutions and estimated valid rollouts.
- Normalize and set weights on-chain each window.
- Linux with NVIDIA GPU drivers installed
- Docker and Docker Compose installed
- NVIDIA Container Toolkit installed (required for GPU passthrough to Docker containers)
- 1x NVIDIA GPU (A100 or H100 recommended, 40GB+ VRAM) for model inference and proof verification
- The coding/math environments run entirely on this GPU; no separate kernel evaluation GPU is required for this release.
- A second GPU is only needed if you opt in to the
triton_kernelenvironment (not active in this release).
- At least 40GB RAM recommended for optimal performance
- Bittensor wallet (cold/hot) registered on the target subnet
- Cloudflare R2 (or S3-compatible) bucket and credentials
- Create a Bucket: Name it the same as your account ID and set the region to ENAM.
- Optional: WandB account for monitoring
For detailed hardware specifications, see compute.min.yaml.
# Clone and enter
git clone https://github.com/one-covenant/grail
cd grail
# Configure environment
cp .env.example .env
# Edit .env with your wallet names, network, and R2 credentials
# If you have a large storage mount (e.g., /ephemeral, /mnt/data), set it in .env:
# GRAIL_HOST_STORAGE_PATH=/ephemeral
# GRAIL_CACHE_DIR=/ephemeral/grail_cache
# Otherwise the default ~/grail-storage is used.
# Run validator with Docker Compose
docker compose --env-file .env -f docker/docker-compose.validator.yml up -d# Clone and enter
git clone https://github.com/one-covenant/grail
cd grail
# Create venv and install
uv venv && source .venv/bin/activate
uv sync
# Configure environment
cp .env.example .env
# Edit .env with your wallet names, network, and R2 credentials
# Run validator
grail validate
# Run validator in debug mode
grail -vv validate
# Run in test mode (validates only your own files)
grail -vv validate --test-modeSet these in .env (see .env.example). This file will be used with Docker Compose:
- Network & subnet
BT_NETWORK(finney|test|custom)BT_CHAIN_ENDPOINT(whenBT_NETWORK=custom)NETUID(target subnet id)
- Wallets
BT_WALLET_COLD(coldkey name)BT_WALLET_HOT(hotkey name)
- Model (dynamically loaded from R2 checkpoints)
- The model and the active environment are loaded automatically from each window's R2 checkpoint.
- The trainer publishes the per-checkpoint sampling policy (
max_tokens,temperature,top_p,top_k,repetition_penalty) inCheckpointMetadata.generation_params. The validator readsmax_tokensfrom there and caps the expected completion length atmin(metadata.max_tokens, MAX_NEW_TOKENS_PROTOCOL_CAP); rollouts longer than that failtermination_valid. MAX_NEW_TOKENS_PROTOCOL_CAP = 8192(ingrail/protocol/constants.py) is the protocol cap; the trainer's per-checkpointmax_tokensis the actual limit.- Rollouts per problem is fixed at 16 (
ROLLOUTS_PER_PROBLEM). - No manual model or sampling configuration required.
- Object storage (R2/S3)
R2_BUCKET_ID,R2_ACCOUNT_ID- Dual credentials (recommended):
- Read-only:
R2_READ_ACCESS_KEY_ID,R2_READ_SECRET_ACCESS_KEY - Write:
R2_WRITE_ACCESS_KEY_ID,R2_WRITE_SECRET_ACCESS_KEY
- Read-only:
- Monitoring
GRAIL_MONITORING_BACKEND(wandb|null)WANDB_API_KEY,WANDB_PROJECT,WANDB_ENTITY,WANDB_MODE
- Kernel evaluation (only needed if the
triton_kernelenvironment is active — not the case in this release)GRAIL_GPU_EVAL(true|false, default: true) — enable GPU-based kernel correctness evaluationKERNEL_EVAL_GPU_IDS(default: 1) — physical GPU index for kernel eval (GPU 0 runs the model). Only read whentriton_kernelis the active environment; safe to leave unset for the coding/math envs.KERNEL_EVAL_BACKEND(persistent|subprocess|basilica, default: persistent) — evaluation backend.persistentreuses the CUDA context (~40x faster),subprocessisolates each eval,basilicauses Basilica cloud GPU workers (not yet implemented)KERNEL_EVAL_TIMEOUT(default: 60) — per-kernel evaluation timeout in seconds
- Optional
HF_TOKEN,HF_USERNAME(for Hugging Face dataset publishing)
Use a registered wallet. Set BT_WALLET_COLD/BT_WALLET_HOT to the names you created with btcli.
Bucket requirement: Name it the same as your account ID; set the region to ENAM.
Validators load local write credentials and use miners’ read credentials fetched from chain to download their files.
Set GRAIL_MONITORING_BACKEND=wandb to enable metrics; otherwise use null.
For centralized logging, use Promtail to ship logs to Loki. This approach is more robust than in-process log shipping and prevents stalling under network pressure.
Promtail is disabled by default. To enable it, use Docker Compose profiles.
Setup:
-
Set environment variables in
.env:PROMTAIL_ENABLE=true PROMTAIL_LOKI_URL=http://your-loki-server:3100/loki/api/v1/push GRAIL_ENV=prod GRAIL_LOG_FILE=/var/log/grail/grail.log # Optional: log rotation settings GRAIL_LOG_MAX_SIZE=100MB GRAIL_LOG_BACKUP_COUNT=5 -
Deploy with Promtail enabled using the
--profile promtailflag:docker compose --env-file .env --profile promtail -f docker/docker-compose.validator.yml up -d
Or set the profile in your environment:
export COMPOSE_PROFILES=promtail docker compose --env-file .env -f docker/docker-compose.validator.yml up -d -
To run without Promtail (default behavior):
docker compose --env-file .env -f docker/docker-compose.validator.yml up -d
-
Verify logs appear in Grafana with the configured labels.
Architecture:
- App writes logs to console (Rich) and file (
GRAIL_LOG_FILE) - Promtail tails the log file and ships to Loki with buffering, backoff, and timeouts
- No network I/O in the app's logging path prevents stalling
Troubleshooting:
- Check Promtail logs:
docker logs grail-promtail - Verify log file exists:
docker exec grail-validator ls -la /var/log/grail/ - Test Loki connectivity:
docker exec grail-promtail wget -O- ${PROMTAIL_LOKI_URL}
Public dashboards:
- WandB Dashboard: set
WANDB_ENTITY=tplrandWANDB_PROJECT=grailto publish validator logs to the public W&B project for detailed metrics and historical data. View at https://wandb.ai/tplr/grail. - Grafana Dashboard: Real-time system logs, validator performance metrics, and network statistics are available at https://grail-grafana.tplr.ai/.
# Start validator with Docker Compose (includes Watchtower for automatic updates)
docker compose --env-file .env -f docker/docker-compose.validator.yml up -d
# View validator logs
docker logs -f grail-validator
# View Watchtower logs (automatic updater)
docker logs -f watchtowerMost behavior is configured via .env.
The Docker Compose configuration includes Watchtower, which automatically:
- Checks for new validator images every 30 seconds
- Pulls the latest
ghcr.io/one-covenant/grail:latestimage - Gracefully restarts your validator with the new version
- Cleans up old images to save disk space
This ensures your validator always runs the latest stable version without manual intervention.
# Stop validator and Watchtower
docker compose -f docker/docker-compose.validator.yml down
# Update manually (if Watchtower is stopped)
docker pull ghcr.io/one-covenant/grail:latest
docker compose --env-file .env -f docker/docker-compose.validator.yml up -d
# Monitor resources
docker stats grail-validator
nvidia-smi # GPU usageThe implementation lives in grail/cli/validate.py.
- Determine current block; compute windows of length
WINDOW_LENGTH. - Process the previous complete window:
target_window = current - WINDOW_LENGTH. - Load checkpoint: Download the appropriate model checkpoint for the target window from R2.
- For each hotkey (test mode: just self; otherwise all active hotkeys in metagraph):
- Build expected file:
grail/windows/{hotkey}-window-{target_window}.parquet. - Retrieve miner’s read credentials from chain (
GrailChainManager). - Check existence and download with read creds; fallback to local creds if needed.
- Build expected file:
For each downloaded inference:
- Required fields:
window_start,nonce,block_hash,commit,proof,challenge,hotkey,signature, plus environment-specific fields. - Window and block-hash must match the validator's
target_windowand hash. - Nonce must be unique within a miner's window file.
- Signature check: hotkey verifies
challenge = seed + block_hash + nonce. - Seed is reconstructed as
{wallet_addr}-{target_window_hash}-{rollout_group}. GRPO group size is fixed at 16 (ROLLOUTS_PER_PROBLEM). - Challenge randomness for GRAIL proof: mix drand randomness (current round) with
target_window_hash; fallback to hash-only on failures. - Verifier (
grail/grail.py) checks token validity, sketch proof, and model identity. - For Triton Kernel environment: kernel correctness is verified by re-executing the generated kernel on-GPU against reference implementations.
Sampling and batching:
- If total rollouts ≤ MAX_SAMPLES_PER_MINER → verify all.
- Else sample complete GRPO groups (~10%) with early stopping if failures exceed threshold.
Per miner, compute estimated_unique rollouts over recent windows (extrapolated to account for sampling). If UNIQUE_ROLLOUTS_CAP is enabled, unique rollouts are capped at this value.
Apply superlinear curve (SUPERLINEAR_EXPONENT = 4.0):
score = estimated_unique ** SUPERLINEAR_EXPONENT
Normalize to weights across miners; set on-chain with set_weights. An emission burn mechanism (GRAIL_BURN_PERCENTAGE = 90%) redirects a portion of emissions to the burn UID.
- Upload all valid rollouts to R2/S3 for training (
upload_valid_rollouts).
- No files found: ensure miners committed read creds on-chain; verify bucket name (should be the same as your account ID) and permissions.
- Frequent verification failures: check drand connectivity; fall back with
--no-drand. - Weight setting fails: wallet funding/permissions and network connectivity.
- Test mode confusion:
--test-modevalidates only your own files; disable for production.
Validator Not Starting:
- Check logs:
docker logs grail-validator - Verify wallet path: Ensure
~/.bittensoris accessible - Check hardware support: Ensure your platform's floating point precision is within tolerance thresholds
Validator exits with torch.cuda.is_available() is False / GPU not visible inside the container:
The validator refuses to load the model on CPU by design: an 8B model on CPU in FP32 saturates every core and exhausts host RAM (we have seen this freeze a 180 GB host). If you see a RuntimeError: get_model() called with device=None but torch.cuda.is_available() is False in docker logs grail-validator, the container is running without GPU passthrough. Diagnose in this order:
-
Confirm the host sees the GPUs:
nvidia-smi
If this fails on the host, reinstall the NVIDIA drivers before touching Docker.
-
Confirm Docker can pass GPUs into a container (the preflight check):
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
If this fails but
nvidia-smiworks on the host,nvidia-container-toolkitis either not installed or not wired into the Docker daemon. Install or reconfigure it, then restart Docker:sudo nvidia-ctk runtime configure --runtime=docker sudo systemctl restart docker
Re-run the preflight until it prints your GPUs, then
docker compose -f docker/docker-compose.validator.yml up -d. -
Confirm the validator container itself sees the GPUs:
docker exec grail-validator nvidia-smiIf the host preflight passes but this fails, check that
CUDA_VISIBLE_DEVICESis not being set to an empty string anywhere in.env, and thatNVIDIA_VISIBLE_DEVICES=allis still present in the validator service env block ofdocker/docker-compose.validator.yml.
Watchtower Not Updating:
- Check registry access: `