A follow-along, host-runnable tutorial. Every step is a command you can paste. Baseline:
CrewAI 1.15.1, Python 3.10–3.12, local qwen3.6 via Ollama at
http://<your-ollama-host>:11434. Substitute your own Ollama host/model where noted.
CrewAI moves fast. This tutorial is pinned to 1.15.1. If you bump the version, re-verify the decorator import path (
crewai.project) and the LLM/Ollama env names.
Set these once in your terminal session. Every kubectl apply, envsubst, and
curl command in this tutorial uses them — no hardcoded IPs or tokens anywhere.
# ── Ollama ────────────────────────────────────────────────────────────────────
# IP or hostname of your Ollama host (no protocol, no port)
export OLLAMA_HOST=<your-ollama-host> # e.g. 192.168.1.100 or ollama.local
export OLLAMA_MODEL=qwen3.6 # model tag you have pulled on that host
# ── Dynatrace ─────────────────────────────────────────────────────────────────
export DT_TENANT_URL=https://<your-tenant>.live.dynatrace.com
export DT_API_TOKEN=<operator-token> # scopes: see README § Dynatrace tokens
export DT_INGEST_TOKEN=<data-ingest-token> # scopes: metrics.ingest + logs.ingest
# + openTelemetryTrace.ingest
# ── GitHub (optional — dev agent GitHub MCP) ──────────────────────────────────
# Leave blank to run without GitHub integration (dev agent writes to ./output/).
export GITHUB_PAT=<your-fine-grained-PAT> # fine-grained PAT: Contents + PRs read/write
export GITHUB_REPO=owner/repo-name # target repo slug (e.g. acme/my-app)
# ── OTel Collector (in-cluster) ───────────────────────────────────────────────
# If you deploy the in-cluster Collector (Step 5), use its cluster-internal DNS.
# Override if your Collector lives on a different host/port.
export OTEL_COLLECTOR_ENDPOINT=http://otel-gateway-collector.observability.svc.cluster.local:4318Tip: Save these exports to a file (e.g.
vars.env) andsource vars.envat the start of each session. Addvars.envto your.gitignore— never commit real tokens.
python3 --version # 3.10 – 3.12
pip install uv # or: curl -LsSf https://astral.sh/uv/install.sh | sh
# Confirm Ollama is reachable and the model is pulled:
curl http://${OLLAMA_HOST}:11434/api/tags | grep ${OLLAMA_MODEL}
# If the model is missing, pull it on the Ollama host: ollama pull qwen3.6git clone https://github.com/isItObservable/CrewAI.git
cd CrewAI/bmad-crew
uv venv && source .venv/bin/activate
uv pip install -e . # installs crewai==1.15.1, crewai-tools, fastapi, openlit
crewai version # expect 1.15.1Prefer the scaffold flow to see how CrewAI generates a project?
crewai create crew demoproduces the sameconfig/agents.yaml+config/tasks.yaml+crew.pyshape this repo uses. We ship a ready-made BMAD crew so you can go straight to running it.
cp .env.example .env
# Edit .env and set your Ollama host (you set $OLLAMA_HOST in Step 0):
sed -i.bak "s|http://<your-ollama-host>:11434|http://${OLLAMA_HOST}:11434|" .env
cat .env | grep OLLAMA_BASE_URL # verify the substitutionThe crew builds one shared LLM from these env vars (crew.py: build_llm()), so every
agent talks to local qwen. No OpenAI/Anthropic key is needed — that's the differentiator.
Eight BMAD personas, each a CrewAI Agent (config/agents.yaml), and eight
hand-offs, each a Task (config/tasks.yaml), run as a sequential process:
analyst → pm → ux → po → architect → sm → dev → qa
Each task's context names the upstream task(s) whose output feeds it, so the hand-offs
are explicit in both the code and (later) the trace. The final QA task writes
output/qa_report.md.
Open config/agents.yaml to see the personas, and crew.py to see how @agent /
@task / @crew wire them with Process.sequential.
One crew, one process — read this before you scale it.
kickoff()runs all eight agents in a single Python process: taskcontexthand-offs are in-memory, so the natural unit of deployment is ONE container/pod (that's whatk8s/ships, behind one Service). Per-agent containers are not a native CrewAI pattern. If you outgrow it, the escape hatch is to split the crew into per-agent services — each a single-agent crew behind its own Deployment — and hand work off over HTTP. You gain independent scaling and failure domains; you pay with network hops, more moving parts, and losing the free in-process context. Measure first (the dashboards in the observability chapter) — for a crew this size, one pod is the honest architecture.
Zero-argument run uses a built-in sample brief (a Slack stand-up bot):
uv run bmad-crewOr drive it with your own brief:
uv run bmad-crew \
--project "Observability status page" \
--brief "We need a public status page that reflects our SLOs in real time..."Watch the console (verbose=True): each agent reasons, then hands its output to the
next. When it finishes, read the full delivery:
cat output/qa_report.md# HITL mode (default) — agents pause for human validation at 3 checkpoints:
# 1. After analyst: validate requirements (BLK-N blockers surfaced one at a time)
# 2. After architect: validate tech stack
# 3. After QA: final SHIP / NO-SHIP sign-off
# Decisions written to output/decisions.memlog.md
uv run bmad-crew
# Fully automated (no pauses — for CI / scripted runs):
uv run bmad-crew --no-hitl
# With GitHub MCP (dev agent creates branch, commits files, opens PR):
uv run bmad-crew --github-repo owner/myrepo --no-hitl
# Requires GITHUB_PERSONAL_ACCESS_TOKEN and GITHUB_MCP_MODE in env.
# See docs/github-mcp-demo.md for the full setup.Optional — dynamic delegation (the CrewAI "wow" beat). Add
--hierarchicalto run a BMad-Orchestrator manager that delegates dynamically (Process.hierarchical). It's heavier on the model's tool-calling and non-deterministic — great for a contrast segment, but sequential is the backbone.
CrewAI is trace-rich but token-blind by default. We use
OpenLIT to emit gen_ai.* spans and GenAI metrics. It's already wired in
observability.py and initialised at startup — you just point it at a collector:
# Run a local OTel Collector (or use your cluster's). Then:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
uv run bmad-crew
# console prints: [bmad-crew] OpenTelemetry instrumentation: ONLeave OTEL_EXPORTER_OTLP_ENDPOINT unset to run without telemetry (no crash).
OTEL_SDK_DISABLED=true is the hard off-switch. The Collector → Dynatrace pipeline and
the dashboards live in observability/README.md.
Three drop-in options emit to the same Collector → Dynatrace pipeline. Switch by setting
CREWAI_INSTRUMENTATION (no rebuild — the image ships all three):
| Value | Library | Notes |
|---|---|---|
openlit (default) |
openlit | gen_ai.* spans + GenAI metrics out of the box |
openllmetry |
traceloop-sdk | Traceloop ecosystem; token attr names differ slightly |
native |
opentelemetry-sdk only | Full control; no third-party auto-instrumentation |
export CREWAI_INSTRUMENTATION=native # switch path; restart crew
uv run bmad-crew
# console: [bmad-crew] Instrumentation path: Native event-bus (crewai_otel)The images are already built and publicly available on GHCR. The deployment manifests already reference the correct tag — you do not need to build or push anything.
| Image | Registry |
|---|---|
| Backend (FastAPI crew server) | ghcr.io/isitobservable/bmad-crew:sha-ba7cb21 |
| Frontend (CopilotKit UI) | ghcr.io/isitobservable/bmad-crew-ui:sha-ba7cb21 |
You can verify the images are accessible before deploying:
docker pull ghcr.io/isitobservable/bmad-crew:sha-ba7cb21
docker pull ghcr.io/isitobservable/bmad-crew-ui:sha-ba7cb21Forking and modifying the code? Push your changes to
mainon your fork..github/workflows/build-image.ymlbuilds and publishes both images automatically using the repo's built-inGITHUB_TOKEN— no personal access token required. After the first successful run, set each package to Public (Package → Settings → Change visibility) so the cluster can pull without animagePullSecret. Then update theimage:tag ink8s/deployment.yamlandui/k8s/deployment.yamlto match the newsha-<7-char-git-sha>shown in the Actions run.
Steps 8–9 need a cluster you can kubectl apply to. If you already have one
and kubectl get nodes returns healthy nodes, skip ahead.
If you need a cluster, the simplest option is a local one with
kind (Kubernetes-in-Docker — no cloud account or
homelab required). Full steps including MetalLB for LoadBalancer IPs are in
docs/cluster-setup.md.
# Quick-start (see docs/cluster-setup.md for MetalLB setup):
kind create cluster --name bmad-crew
kubectl get nodesOllama reachability. Pods must reach your Ollama host over the network. If Ollama runs on the same machine as Docker, use
host.docker.internal(macOS/Windows) or172.17.0.1(Linux) instead oflocalhostforOLLAMA_HOST.
Deploy in this order: observability infrastructure first, then the application.
The ConfigMap files use ${...} placeholders — envsubst fills them in at apply time
so no secrets or IPs ever touch the repo.
Prerequisite:
envsubstships with thegettextpackage. macOS:brew install gettext. Ubuntu/Debian:apt-get install gettext-base.
# Add the Dynatrace Helm chart repo (once)
helm repo add dynatrace https://raw.githubusercontent.com/Dynatrace/dynatrace-operator/main/config/helm/repos/stable
helm repo update
# Install the operator (CRDs + controller only — no DynaKube yet)
helm install dynatrace-operator dynatrace/dynatrace-operator \
-n dynatrace --create-namespace
# Wait for the operator to be ready
kubectl -n dynatrace rollout status deploy/dynatrace-operator
# Create the secret the DynaKube CR references under `tokens:`
kubectl create secret generic observable-crewai \
-n dynatrace \
--from-literal=apiToken="${DT_API_TOKEN}" \
--from-literal=dataIngestToken="${DT_INGEST_TOKEN}" \
--dry-run=client -o yaml | kubectl apply -f -
# Apply the DynaKube CR (substitutes $DT_TENANT_URL)
envsubst < k8s/dynakube.yaml | kubectl apply -f -
# Wait for the ActiveGate pod to come up
kubectl -n dynatrace wait pod -l app.kubernetes.io/name=dynatrace-activegate \
--for=condition=Ready --timeout=180sThe ActiveGate handles cluster-level telemetry (K8s metrics, workload inventory). All application traces and logs flow through the OTel Collector in the next step.
The Collector is managed by the OpenTelemetry Operator. Install the operator once, then apply the gateway CRD — the operator handles the Deployment, Service, and RBAC.
# 1. cert-manager (operator dependency)
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.16.0/cert-manager.yaml
kubectl -n cert-manager rollout status deploy/cert-manager --timeout=120s
# 2. OTel Operator
helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm repo update
helm install opentelemetry-operator open-telemetry/opentelemetry-operator \
-n opentelemetry-operator-system --create-namespace \
--set manager.collectorImage.repository=otel/opentelemetry-collector-k8s
kubectl -n opentelemetry-operator-system rollout status deploy/opentelemetry-operator --timeout=120s
# 3. Dynatrace OTLP auth secret
kubectl create namespace observability --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret generic dynatrace-otlp \
-n observability \
--from-literal=endpoint="${DT_TENANT_URL}" \
--from-literal=token="Api-Token ${DT_INGEST_TOKEN}" \
--dry-run=client -o yaml | kubectl apply -f -
# 4. Deploy the gateway collector CRD
kubectl apply -f observability/collector/k8s/otel-gateway.yaml
kubectl -n observability rollout status deploy/otel-gateway-collector --timeout=120sThe operator creates a Service named otel-gateway-collector in the observability
namespace — this is the endpoint the crew pods send telemetry to. The Collector
enriches spans with Kubernetes metadata, normalises gen_ai.* attributes across all
three instrumentation paths, and forwards everything to Dynatrace.
A second Collector runs as a DaemonSet — one pod per node — to tail container
log files directly from the host filesystem (/var/log/pods). It parses the
CRI-O / containerd / Docker log format, extracts pod metadata from the file path,
enriches with k8sattributes, and exports logs to Dynatrace alongside the OTLP
signals from the crew pods.
kubectl apply -f observability/collector/k8s/otel-logs-collector.yaml
# The operator creates a DaemonSet named otel-logs-collector-collector
kubectl -n observability rollout status daemonset/otel-logs-collector-collector --timeout=120sWhy two collectors? The gateway (
otel-gateway, Deployment) receives OTLP pushed by the application pods — traces, metrics, and structured logs. The log collector (DaemonSet) pulls raw container stdout/stderr from the node filesystem, which catches logs from any pod regardless of whether it speaks OTLP. Both export to the samedynatrace-otlpsecret and the same Dynatrace tenant.
# 1. Namespace
kubectl apply -f k8s/namespace.yaml
# 2. ConfigMap — substitute $OLLAMA_HOST at apply time
envsubst < k8s/configmap.yaml | kubectl apply -f -
# 3. Secret — inject your GitHub PAT directly (never commit the real value)
kubectl create secret generic bmad-crew-secrets \
-n bmad-crew \
--from-literal=GITHUB_PERSONAL_ACCESS_TOKEN="${GITHUB_PAT}" \
--from-literal=OPENAI_API_KEY="" \
--dry-run=client -o yaml | kubectl apply -f -
# 4. Storage, Deployment, Service, HPA
kubectl apply -f k8s/pvc.yaml
kubectl apply -f k8s/deployment.yaml
kubectl apply -f k8s/service.yaml
kubectl apply -f k8s/hpa.yaml
# 5. Wait for rollout
kubectl -n bmad-crew rollout status deploy/bmad-crewTwo things that bite in k8s:
- Egress to Ollama. Pods must reach
http://${OLLAMA_HOST}:11434. Confirm routing from the cluster before you apply — if the model is unreachable,/kickoffhangs on the first LLM call. - Memory persistence. CrewAI memory (LanceDB) is on local disk and ephemeral in a
pod — the
PVCmounts it at/app/.crewai. Only relevant if you enablememory=True.
The browser-based frontend. ui/k8s/configmap.yaml also contains ${OLLAMA_HOST} —
use envsubst the same way:
# ConfigMap (substitutes $OLLAMA_HOST for the CopilotKit Ollama adapter)
envsubst < ui/k8s/configmap.yaml | kubectl apply -f -
# Deployment and Service
kubectl apply -f ui/k8s/deployment.yaml
kubectl apply -f ui/k8s/service.yaml
kubectl -n bmad-crew rollout status deploy/bmad-crew-uiTraffic flow: Browser → /api/crew/* → Next.js rewrite → FastAPI pod (ClusterIP,
internal). The FastAPI backend is never directly exposed to the internet — all external
traffic enters through the UI's LoadBalancer service.
Get the LoadBalancer IP assigned to the UI service and open it in your browser:
# Wait for EXTERNAL-IP to appear (cloud cluster / MetalLB):
kubectl get svc bmad-crew-ui -n bmad-crew
# NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
# bmad-crew-ui LoadBalancer 10.96.x.x <EXTERNAL-IP> 80:31234/TCP 2m
# Grab the IP and open the browser:
BMAD_UI_IP=$(kubectl get svc bmad-crew-ui -n bmad-crew \
-o jsonpath='{.status.loadBalancer.ingress[0].ip}')
echo "CopilotKit UI → http://${BMAD_UI_IP}"
open "http://${BMAD_UI_IP}" # macOS
# xdg-open "http://${BMAD_UI_IP}" # LinuxIP still
<pending>? The cluster's LoadBalancer controller hasn't assigned an address yet. On a bare-metal cluster (no cloud LB) you need MetalLB. As a quick alternative, use port-forward:kubectl port-forward svc/bmad-crew-ui -n bmad-crew 3000:80 # Open http://localhost:3000
Type a project brief in the chat sidebar, watch the Agent Timeline panel light up agent by agent (click any completed agent card to expand its full output), and read the QA report in the output panel once the pipeline finishes.
kubectl -n bmad-crew port-forward svc/bmad-crew 8080:80 &
curl -s localhost:8080/kickoff/async \
-H 'content-type: application/json' \
-d '{"project":"Status page","brief":"A public SLO status page..."}' | jq .
# returns {"run_id": "..."} — poll /run/{run_id}/result or stream /stream/{run_id}/kickoff (sync) still works for quick tests; /kickoff/async + /stream/{run_id} is
what the UI uses for live streaming.
After triggering at least one crew run, confirm all three signal types landed in Dynatrace:
Distributed Traces
Open the Dynatrace UI → Distributed Traces and filter by service.name = bmad-crew.
Each crew run produces one trace with this structure:
invoke_workflow BmadCrew ← root span (full crew duration)
create_agent <role> ×8 ← agent initialisation
invoke_agent <role> ×8 ← one per BMAD agent
chat qwen3.6 ×N ← one per LLM call (carries token counts)
execute_tool <name> ×M ← GitHub tool calls (create_branch, etc.)
mcp tools/call ×P ← MCP protocol calls
Token metrics (gen_ai.usage.input_tokens / gen_ai.usage.output_tokens) are
attributes on every chat qwen3.6 span. A typical full pipeline consumes
≈ 130k input / 240k output tokens across 8 LLM calls.
Logs
Filter logs by service.name = bmad-crew. You should see agent lifecycle events
(agent started, task completed) and any tool errors.
Dashboards
Import the pre-built dashboards from observability/dashboards/ in the repo:
- CrewAI Efficiency — token consumption per agent, LLM latency
- CrewAI Health — error rate, span counts, tool call breakdown
- CrewAI + CopilotKit — end-to-end view combining UI and backend telemetry
If spans are missing: check that OTEL_EXPORTER_OTLP_ENDPOINT in the ConfigMap
points at the Collector (http://otel-gateway-collector.observability.svc.cluster.local:4318),
that the Collector pod is running, and that the Dynatrace export token has
metrics.ingest + traces.ingest scopes. Walk the full pipeline in
observability/README.md.
| Symptom | Cause / fix |
|---|---|
crewai version ≠ 1.15.1 |
uv pip install 'crewai==1.15.1' — the tutorial is pinned |
| Hangs on first agent | Ollama unreachable — check OLLAMA_BASE_URL and curl .../api/tags |
OPENAI_API_KEY demanded |
You enabled memory=True without a local embedder — set the Ollama embedder (research §3.4) |
| Tool-calling / delegation flaky | qwen tool-calling can wobble; fall back to sequential (drop --hierarchical) |
| No spans in Dynatrace | OTEL_EXPORTER_OTLP_ENDPOINT unset/wrong, or Collector not forwarding |
import crewai.project fails |
Version drift — verify the decorator import path for your CrewAI version |
| UI pod CrashLoopBackoff | BMAD_CREW_URL wrong or Ollama unreachable from ui pod — check configmap and Ollama egress |
| Chat sidebar shows no response | COPILOTKIT_ADAPTER=ollama but OLLAMA_BASE_URL unreachable from ui pod |
| Double crew run on kickoff | kickoff_crew must be client-side only (useCopilotAction in page.tsx) — route.ts must have empty actions array |
| HITL never pauses | BMAD_HUMAN_IN_LOOP=false in ConfigMap — set to true for interactive mode |
- Enable memory with a local Ollama embedder for cross-run recall (research §3.4).
- Add custom tools (
@tool/BaseTool) so the dev agent can read the repo or run tests. - Wrap the crew in a Flow (
@start/@listen/@router) for retries and branching. - Contrast sequential vs hierarchical delegation on camera.