Skip to content

Latest commit

 

History

History
515 lines (387 loc) · 21.3 KB

File metadata and controls

515 lines (387 loc) · 21.3 KB

TUTORIAL — Rebuild the BMAD crew in CrewAI, then observe it

A follow-along, host-runnable tutorial. Every step is a command you can paste. Baseline: CrewAI 1.15.1, Python 3.10–3.12, local qwen3.6 via Ollama at http://<your-ollama-host>:11434. Substitute your own Ollama host/model where noted.

CrewAI moves fast. This tutorial is pinned to 1.15.1. If you bump the version, re-verify the decorator import path (crewai.project) and the LLM/Ollama env names.


0. Prerequisites

0.1 Set your environment variables

Set these once in your terminal session. Every kubectl apply, envsubst, and curl command in this tutorial uses them — no hardcoded IPs or tokens anywhere.

# ── Ollama ────────────────────────────────────────────────────────────────────
# IP or hostname of your Ollama host (no protocol, no port)
export OLLAMA_HOST=<your-ollama-host>       # e.g. 192.168.1.100 or ollama.local
export OLLAMA_MODEL=qwen3.6                 # model tag you have pulled on that host

# ── Dynatrace ─────────────────────────────────────────────────────────────────
export DT_TENANT_URL=https://<your-tenant>.live.dynatrace.com
export DT_API_TOKEN=<operator-token>        # scopes: see README § Dynatrace tokens
export DT_INGEST_TOKEN=<data-ingest-token>  # scopes: metrics.ingest + logs.ingest
                                            #         + openTelemetryTrace.ingest

# ── GitHub (optional — dev agent GitHub MCP) ──────────────────────────────────
# Leave blank to run without GitHub integration (dev agent writes to ./output/).
export GITHUB_PAT=<your-fine-grained-PAT>  # fine-grained PAT: Contents + PRs read/write
export GITHUB_REPO=owner/repo-name         # target repo slug (e.g. acme/my-app)

# ── OTel Collector (in-cluster) ───────────────────────────────────────────────
# If you deploy the in-cluster Collector (Step 5), use its cluster-internal DNS.
# Override if your Collector lives on a different host/port.
export OTEL_COLLECTOR_ENDPOINT=http://otel-gateway-collector.observability.svc.cluster.local:4318

Tip: Save these exports to a file (e.g. vars.env) and source vars.env at the start of each session. Add vars.env to your .gitignore — never commit real tokens.

0.2 Verify tooling and Ollama

python3 --version           # 3.10 – 3.12
pip install uv              # or: curl -LsSf https://astral.sh/uv/install.sh | sh

# Confirm Ollama is reachable and the model is pulled:
curl http://${OLLAMA_HOST}:11434/api/tags | grep ${OLLAMA_MODEL}
# If the model is missing, pull it on the Ollama host:  ollama pull qwen3.6

1. Install CrewAI

git clone https://github.com/isItObservable/CrewAI.git
cd CrewAI/bmad-crew
uv venv && source .venv/bin/activate
uv pip install -e .          # installs crewai==1.15.1, crewai-tools, fastapi, openlit
crewai version               # expect 1.15.1

Prefer the scaffold flow to see how CrewAI generates a project? crewai create crew demo produces the same config/agents.yaml + config/tasks.yaml + crew.py shape this repo uses. We ship a ready-made BMAD crew so you can go straight to running it.


2. Configure the model (local qwen, no API key)

cp .env.example .env
# Edit .env and set your Ollama host (you set $OLLAMA_HOST in Step 0):
sed -i.bak "s|http://<your-ollama-host>:11434|http://${OLLAMA_HOST}:11434|" .env
cat .env | grep OLLAMA_BASE_URL   # verify the substitution

The crew builds one shared LLM from these env vars (crew.py: build_llm()), so every agent talks to local qwen. No OpenAI/Anthropic key is needed — that's the differentiator.


3. Understand the crew (what you're running)

Eight BMAD personas, each a CrewAI Agent (config/agents.yaml), and eight hand-offs, each a Task (config/tasks.yaml), run as a sequential process:

analyst → pm → ux → po → architect → sm → dev → qa

Each task's context names the upstream task(s) whose output feeds it, so the hand-offs are explicit in both the code and (later) the trace. The final QA task writes output/qa_report.md.

Open config/agents.yaml to see the personas, and crew.py to see how @agent / @task / @crew wire them with Process.sequential.

One crew, one process — read this before you scale it. kickoff() runs all eight agents in a single Python process: task context hand-offs are in-memory, so the natural unit of deployment is ONE container/pod (that's what k8s/ ships, behind one Service). Per-agent containers are not a native CrewAI pattern. If you outgrow it, the escape hatch is to split the crew into per-agent services — each a single-agent crew behind its own Deployment — and hand work off over HTTP. You gain independent scaling and failure domains; you pay with network hops, more moving parts, and losing the free in-process context. Measure first (the dashboards in the observability chapter) — for a crew this size, one pod is the honest architecture.


4. Run the crew locally

Zero-argument run uses a built-in sample brief (a Slack stand-up bot):

uv run bmad-crew

Or drive it with your own brief:

uv run bmad-crew \
  --project "Observability status page" \
  --brief "We need a public status page that reflects our SLOs in real time..."

Watch the console (verbose=True): each agent reasons, then hands its output to the next. When it finishes, read the full delivery:

cat output/qa_report.md
# HITL mode (default) — agents pause for human validation at 3 checkpoints:
#   1. After analyst: validate requirements (BLK-N blockers surfaced one at a time)
#   2. After architect: validate tech stack
#   3. After QA: final SHIP / NO-SHIP sign-off
# Decisions written to output/decisions.memlog.md
uv run bmad-crew

# Fully automated (no pauses — for CI / scripted runs):
uv run bmad-crew --no-hitl

# With GitHub MCP (dev agent creates branch, commits files, opens PR):
uv run bmad-crew --github-repo owner/myrepo --no-hitl
# Requires GITHUB_PERSONAL_ACCESS_TOKEN and GITHUB_MCP_MODE in env.
# See docs/github-mcp-demo.md for the full setup.

Optional — dynamic delegation (the CrewAI "wow" beat). Add --hierarchical to run a BMad-Orchestrator manager that delegates dynamically (Process.hierarchical). It's heavier on the model's tool-calling and non-deterministic — great for a contrast segment, but sequential is the backbone.


5. Add observability (OpenTelemetry → Dynatrace)

CrewAI is trace-rich but token-blind by default. We use OpenLIT to emit gen_ai.* spans and GenAI metrics. It's already wired in observability.py and initialised at startup — you just point it at a collector:

# Run a local OTel Collector (or use your cluster's). Then:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
uv run bmad-crew
# console prints: [bmad-crew] OpenTelemetry instrumentation: ON

Leave OTEL_EXPORTER_OTLP_ENDPOINT unset to run without telemetry (no crash). OTEL_SDK_DISABLED=true is the hard off-switch. The Collector → Dynatrace pipeline and the dashboards live in observability/README.md.

5.1 Choose your instrumentation path

Three drop-in options emit to the same Collector → Dynatrace pipeline. Switch by setting CREWAI_INSTRUMENTATION (no rebuild — the image ships all three):

Value Library Notes
openlit (default) openlit gen_ai.* spans + GenAI metrics out of the box
openllmetry traceloop-sdk Traceloop ecosystem; token attr names differ slightly
native opentelemetry-sdk only Full control; no third-party auto-instrumentation
export CREWAI_INSTRUMENTATION=native   # switch path; restart crew
uv run bmad-crew
# console: [bmad-crew] Instrumentation path: Native event-bus (crewai_otel)

6. Container images (pre-built — no action needed)

The images are already built and publicly available on GHCR. The deployment manifests already reference the correct tag — you do not need to build or push anything.

Image Registry
Backend (FastAPI crew server) ghcr.io/isitobservable/bmad-crew:sha-ba7cb21
Frontend (CopilotKit UI) ghcr.io/isitobservable/bmad-crew-ui:sha-ba7cb21

You can verify the images are accessible before deploying:

docker pull ghcr.io/isitobservable/bmad-crew:sha-ba7cb21
docker pull ghcr.io/isitobservable/bmad-crew-ui:sha-ba7cb21

Forking and modifying the code? Push your changes to main on your fork. .github/workflows/build-image.yml builds and publishes both images automatically using the repo's built-in GITHUB_TOKEN — no personal access token required. After the first successful run, set each package to Public (Package → Settings → Change visibility) so the cluster can pull without an imagePullSecret. Then update the image: tag in k8s/deployment.yaml and ui/k8s/deployment.yaml to match the new sha-<7-char-git-sha> shown in the Actions run.


7. Provision a Kubernetes cluster

Steps 8–9 need a cluster you can kubectl apply to. If you already have one and kubectl get nodes returns healthy nodes, skip ahead.

If you need a cluster, the simplest option is a local one with kind (Kubernetes-in-Docker — no cloud account or homelab required). Full steps including MetalLB for LoadBalancer IPs are in docs/cluster-setup.md.

# Quick-start (see docs/cluster-setup.md for MetalLB setup):
kind create cluster --name bmad-crew
kubectl get nodes

Ollama reachability. Pods must reach your Ollama host over the network. If Ollama runs on the same machine as Docker, use host.docker.internal (macOS/Windows) or 172.17.0.1 (Linux) instead of localhost for OLLAMA_HOST.


8. Deploy to Kubernetes

Deploy in this order: observability infrastructure first, then the application. The ConfigMap files use ${...} placeholders — envsubst fills them in at apply time so no secrets or IPs ever touch the repo.

Prerequisite: envsubst ships with the gettext package. macOS: brew install gettext. Ubuntu/Debian: apt-get install gettext-base.

8.1 Deploy the Dynatrace Operator and cluster monitor

# Add the Dynatrace Helm chart repo (once)
helm repo add dynatrace https://raw.githubusercontent.com/Dynatrace/dynatrace-operator/main/config/helm/repos/stable
helm repo update

# Install the operator (CRDs + controller only — no DynaKube yet)
helm install dynatrace-operator dynatrace/dynatrace-operator \
  -n dynatrace --create-namespace

# Wait for the operator to be ready
kubectl -n dynatrace rollout status deploy/dynatrace-operator

# Create the secret the DynaKube CR references under `tokens:`
kubectl create secret generic observable-crewai \
  -n dynatrace \
  --from-literal=apiToken="${DT_API_TOKEN}" \
  --from-literal=dataIngestToken="${DT_INGEST_TOKEN}" \
  --dry-run=client -o yaml | kubectl apply -f -

# Apply the DynaKube CR (substitutes $DT_TENANT_URL)
envsubst < k8s/dynakube.yaml | kubectl apply -f -

# Wait for the ActiveGate pod to come up
kubectl -n dynatrace wait pod -l app.kubernetes.io/name=dynatrace-activegate \
  --for=condition=Ready --timeout=180s

The ActiveGate handles cluster-level telemetry (K8s metrics, workload inventory). All application traces and logs flow through the OTel Collector in the next step.

8.2 Deploy the OpenTelemetry Collector (OTel Operator)

The Collector is managed by the OpenTelemetry Operator. Install the operator once, then apply the gateway CRD — the operator handles the Deployment, Service, and RBAC.

# 1. cert-manager (operator dependency)
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.16.0/cert-manager.yaml
kubectl -n cert-manager rollout status deploy/cert-manager --timeout=120s

# 2. OTel Operator
helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm repo update
helm install opentelemetry-operator open-telemetry/opentelemetry-operator \
  -n opentelemetry-operator-system --create-namespace \
  --set manager.collectorImage.repository=otel/opentelemetry-collector-k8s
kubectl -n opentelemetry-operator-system rollout status deploy/opentelemetry-operator --timeout=120s

# 3. Dynatrace OTLP auth secret
kubectl create namespace observability --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret generic dynatrace-otlp \
  -n observability \
  --from-literal=endpoint="${DT_TENANT_URL}" \
  --from-literal=token="Api-Token ${DT_INGEST_TOKEN}" \
  --dry-run=client -o yaml | kubectl apply -f -

# 4. Deploy the gateway collector CRD
kubectl apply -f observability/collector/k8s/otel-gateway.yaml
kubectl -n observability rollout status deploy/otel-gateway-collector --timeout=120s

The operator creates a Service named otel-gateway-collector in the observability namespace — this is the endpoint the crew pods send telemetry to. The Collector enriches spans with Kubernetes metadata, normalises gen_ai.* attributes across all three instrumentation paths, and forwards everything to Dynatrace.

8.3 Deploy the node log collector (DaemonSet)

A second Collector runs as a DaemonSet — one pod per node — to tail container log files directly from the host filesystem (/var/log/pods). It parses the CRI-O / containerd / Docker log format, extracts pod metadata from the file path, enriches with k8sattributes, and exports logs to Dynatrace alongside the OTLP signals from the crew pods.

kubectl apply -f observability/collector/k8s/otel-logs-collector.yaml

# The operator creates a DaemonSet named otel-logs-collector-collector
kubectl -n observability rollout status daemonset/otel-logs-collector-collector --timeout=120s

Why two collectors? The gateway (otel-gateway, Deployment) receives OTLP pushed by the application pods — traces, metrics, and structured logs. The log collector (DaemonSet) pulls raw container stdout/stderr from the node filesystem, which catches logs from any pod regardless of whether it speaks OTLP. Both export to the same dynatrace-otlp secret and the same Dynatrace tenant.

8.4 Deploy the BMAD crew backend

# 1. Namespace
kubectl apply -f k8s/namespace.yaml

# 2. ConfigMap — substitute $OLLAMA_HOST at apply time
envsubst < k8s/configmap.yaml | kubectl apply -f -

# 3. Secret — inject your GitHub PAT directly (never commit the real value)
kubectl create secret generic bmad-crew-secrets \
  -n bmad-crew \
  --from-literal=GITHUB_PERSONAL_ACCESS_TOKEN="${GITHUB_PAT}" \
  --from-literal=OPENAI_API_KEY="" \
  --dry-run=client -o yaml | kubectl apply -f -

# 4. Storage, Deployment, Service, HPA
kubectl apply -f k8s/pvc.yaml
kubectl apply -f k8s/deployment.yaml
kubectl apply -f k8s/service.yaml
kubectl apply -f k8s/hpa.yaml

# 5. Wait for rollout
kubectl -n bmad-crew rollout status deploy/bmad-crew

Two things that bite in k8s:

  1. Egress to Ollama. Pods must reach http://${OLLAMA_HOST}:11434. Confirm routing from the cluster before you apply — if the model is unreachable, /kickoff hangs on the first LLM call.
  2. Memory persistence. CrewAI memory (LanceDB) is on local disk and ephemeral in a pod — the PVC mounts it at /app/.crewai. Only relevant if you enable memory=True.

8.5 Deploy the CopilotKit UI

The browser-based frontend. ui/k8s/configmap.yaml also contains ${OLLAMA_HOST} — use envsubst the same way:

# ConfigMap (substitutes $OLLAMA_HOST for the CopilotKit Ollama adapter)
envsubst < ui/k8s/configmap.yaml | kubectl apply -f -

# Deployment and Service
kubectl apply -f ui/k8s/deployment.yaml
kubectl apply -f ui/k8s/service.yaml
kubectl -n bmad-crew rollout status deploy/bmad-crew-ui

Traffic flow: Browser → /api/crew/* → Next.js rewrite → FastAPI pod (ClusterIP, internal). The FastAPI backend is never directly exposed to the internet — all external traffic enters through the UI's LoadBalancer service.


9. Run and verify

9.1 Open the CopilotKit UI

Get the LoadBalancer IP assigned to the UI service and open it in your browser:

# Wait for EXTERNAL-IP to appear (cloud cluster / MetalLB):
kubectl get svc bmad-crew-ui -n bmad-crew
# NAME            TYPE           CLUSTER-IP     EXTERNAL-IP      PORT(S)        AGE
# bmad-crew-ui    LoadBalancer   10.96.x.x      <EXTERNAL-IP>    80:31234/TCP   2m

# Grab the IP and open the browser:
BMAD_UI_IP=$(kubectl get svc bmad-crew-ui -n bmad-crew \
  -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
echo "CopilotKit UI → http://${BMAD_UI_IP}"
open "http://${BMAD_UI_IP}"        # macOS
# xdg-open "http://${BMAD_UI_IP}" # Linux

IP still <pending>? The cluster's LoadBalancer controller hasn't assigned an address yet. On a bare-metal cluster (no cloud LB) you need MetalLB. As a quick alternative, use port-forward:

kubectl port-forward svc/bmad-crew-ui -n bmad-crew 3000:80
# Open http://localhost:3000

Type a project brief in the chat sidebar, watch the Agent Timeline panel light up agent by agent (click any completed agent card to expand its full output), and read the QA report in the output panel once the pipeline finishes.

9.2 Headless / curl (no browser)

kubectl -n bmad-crew port-forward svc/bmad-crew 8080:80 &
curl -s localhost:8080/kickoff/async \
  -H 'content-type: application/json' \
  -d '{"project":"Status page","brief":"A public SLO status page..."}' | jq .
# returns {"run_id": "..."} — poll /run/{run_id}/result or stream /stream/{run_id}

/kickoff (sync) still works for quick tests; /kickoff/async + /stream/{run_id} is what the UI uses for live streaming.

9.3 Verify telemetry in Dynatrace

After triggering at least one crew run, confirm all three signal types landed in Dynatrace:

Distributed Traces

Open the Dynatrace UI → Distributed Traces and filter by service.name = bmad-crew. Each crew run produces one trace with this structure:

invoke_workflow BmadCrew           ← root span (full crew duration)
  create_agent <role>  ×8          ← agent initialisation
  invoke_agent <role>  ×8          ← one per BMAD agent
    chat qwen3.6       ×N          ← one per LLM call (carries token counts)
    execute_tool <name> ×M         ← GitHub tool calls (create_branch, etc.)
    mcp tools/call     ×P          ← MCP protocol calls

Token metrics (gen_ai.usage.input_tokens / gen_ai.usage.output_tokens) are attributes on every chat qwen3.6 span. A typical full pipeline consumes ≈ 130k input / 240k output tokens across 8 LLM calls.

Logs

Filter logs by service.name = bmad-crew. You should see agent lifecycle events (agent started, task completed) and any tool errors.

Dashboards

Import the pre-built dashboards from observability/dashboards/ in the repo:

  • CrewAI Efficiency — token consumption per agent, LLM latency
  • CrewAI Health — error rate, span counts, tool call breakdown
  • CrewAI + CopilotKit — end-to-end view combining UI and backend telemetry

If spans are missing: check that OTEL_EXPORTER_OTLP_ENDPOINT in the ConfigMap points at the Collector (http://otel-gateway-collector.observability.svc.cluster.local:4318), that the Collector pod is running, and that the Dynatrace export token has metrics.ingest + traces.ingest scopes. Walk the full pipeline in observability/README.md.


Troubleshooting

Symptom Cause / fix
crewai version ≠ 1.15.1 uv pip install 'crewai==1.15.1' — the tutorial is pinned
Hangs on first agent Ollama unreachable — check OLLAMA_BASE_URL and curl .../api/tags
OPENAI_API_KEY demanded You enabled memory=True without a local embedder — set the Ollama embedder (research §3.4)
Tool-calling / delegation flaky qwen tool-calling can wobble; fall back to sequential (drop --hierarchical)
No spans in Dynatrace OTEL_EXPORTER_OTLP_ENDPOINT unset/wrong, or Collector not forwarding
import crewai.project fails Version drift — verify the decorator import path for your CrewAI version
UI pod CrashLoopBackoff BMAD_CREW_URL wrong or Ollama unreachable from ui pod — check configmap and Ollama egress
Chat sidebar shows no response COPILOTKIT_ADAPTER=ollama but OLLAMA_BASE_URL unreachable from ui pod
Double crew run on kickoff kickoff_crew must be client-side only (useCopilotAction in page.tsx) — route.ts must have empty actions array
HITL never pauses BMAD_HUMAN_IN_LOOP=false in ConfigMap — set to true for interactive mode

Where to go next

  • Enable memory with a local Ollama embedder for cross-run recall (research §3.4).
  • Add custom tools (@tool / BaseTool) so the dev agent can read the repo or run tests.
  • Wrap the crew in a Flow (@start/@listen/@router) for retries and branching.
  • Contrast sequential vs hierarchical delegation on camera.