Skip to content

Latest commit

 

History

History
251 lines (188 loc) · 11.6 KB

File metadata and controls

251 lines (188 loc) · 11.6 KB

Autonomous Enterprise Operations Architecture

Last updated: 2026-05-18.

Purpose

Agentic Operations is the deployed control plane for a broader autonomous enterprise operations platform. SOC/ITSM remains the seed proof domain because it exercises alerts, tickets, logs, identity, approvals, remediation, evidence, and postmortems, but the architecture is intended to scale toward replacing or radically reducing enterprise IT/security/DevOps/service-desk labor with governed agents.

The control plane is designed to be:

  • work-intake agnostic: tickets, alerts, chat, email, schedules, CI/CD, direct operator prompts, and future business-operation triggers
  • ticketing-system and provider agnostic
  • agent-harness agnostic
  • local-model or cloud-model compatible
  • approval-gated for risky work
  • auditable enough for demos, troubleshooting, and future compliance review
  • modular enough to deploy into a new environment without rewriting the dashboard

Architectural doctrine: agent decisions belong to the harness; enforcement belongs to the platform. The dashboard should provide context and tools, then let Codex, Hermes, Claude Code, or a future harness decide how to proceed. Security, provider permissions, approvals, credential leases, audit, and unsafe action blocking are hard platform boundaries. See docs/AGENT_DECISION_MODEL.md.

The current implementation uses iTop as the reference ticketing provider, Hermes Agent as the preferred queue harness, Claude Code as a fallback harness, and the built-in AI proxy as the model gateway. None of these is treated as the permanent center of the architecture.

Runtime Components

Component Current implementation Replaceable boundary
Dashboard API FastAPI, raw asyncpg, PostgreSQL API routes should stay stable
Database PostgreSQL 16 PostgreSQL is mandatory
Frontend Vanilla HTML/CSS/JS Can be rebuilt as long as API contract remains
Ticket provider iTop plus local provider services/ticket_provider.py and services/provider_registry.py
Agent harness Hermes Agent by default; Claude Code and Codex fallback/selectable services/agent_harness.py
Model access Built-in ai-proxy at AGENT_LLM_BASE_URL Any OpenAI/Anthropic-compatible proxy/harness bridge
Approval system Dashboard change_requests table/API Can later sync to external CAB/change platforms
Memory/learning Skills, KB, workflows, postmortems Can later ingest external KB/tickets/docs
Email provider Mailcow reference stack plus optional API shim Exchange, Gmail, Proofpoint, Mimecast, or another email/security adapter
Enterprise providers Current SOC/IT reference modules Cloud, network, endpoint, SaaS, identity, code, storage, backup, and monitoring adapters

Canonical Data Model

The dashboard owns canonical records for enterprise work. External systems mirror into or out of those records.

Core tables:

  • tickets: canonical ticket record plus provider metadata.
  • ticket_notes: internal/user-visible notes from dashboard, agents, sync jobs, and future providers.
  • ticket_attachments: metadata for attachments. Ops Chat reference uploads are stored under runtime data/ops_chat_uploads; enterprise binary storage can be replaced by object storage/DMS.
  • agents: agent instance lifecycle.
  • agent_tasks: runnable unit of work, prompt, status, checkpoints, output, PID, work directory.
  • change_requests: approval gate for potentially destructive actions.
  • postmortems: structured learning after ticket completion.
  • agent_workflows: reusable automation/workflow blueprint records.
  • workflow_runs: execution records for reusable workflows.
  • knowledge_articles: local reusable documentation.
  • agent_skills: prompt-level reusable capabilities.
  • audit_log and event_log: durable action/event history.
  • tools and tool_checks: tool inventory and health checks.

As the platform broadens, the same model should expand around an enterprise ontology: users, groups, roles, org units, assets, apps, services, repos, environments, networks, data classifications, system owners, providers, policies, runbooks, and approval authorities. Those objects let agents reason about enterprise work beyond SOC tickets without hardcoding each customer's tooling.

Provider metadata on tickets:

  • provider: local, itop, later servicenow, jira, etc.
  • provider_ref: external ticket id/reference.
  • provider_class: external ticket class/type.
  • provider_url: external ticket URL when known.
  • provider_sync_status: synced, local_only, pending_create, create_failed, unknown.
  • provider_last_error: last provider sync/create failure.
  • provider_payload: raw provider payload or result for debugging.

Ticket Sync Flow

Inbound provider flow:

  1. Provider adapter discovers or fetches a ticket.
  2. Provider adapter upserts tickets.
  3. Provider metadata is recorded.
  4. Dashboard frontend shows the same canonical ticket shape regardless of provider.
  5. Agent context reads from /api/tickets/{id}/context.

Outbound dashboard flow:

  1. Dashboard creates a canonical ticket through POST /api/tickets.
  2. If sync_provider=false, ticket remains local-only.
  3. If sync_provider=true, ticket_service.create_ticket() calls provider_registry.create_ticket().
  4. Provider returns success, local_only, or error.
  5. Dashboard updates provider_sync_status, provider_last_error, and provider_payload.
  6. Existing tickets can be pushed manually with POST /api/tickets/{id}/push-provider.

iTop outbound creation is intentionally guarded. Incident/UserRequest creation requires:

  • ITOP_DEFAULT_ORG_ID
  • ITOP_DEFAULT_CALLER_ID

Without those values, the dashboard records create_failed instead of claiming a false sync.

Agent Harness Flow

  1. Operator assigns an agent or creates one from a prompt.
  2. API inserts an agents row.
  3. API inserts an agent_tasks row.
  4. Runner creates /app/agent_work/<agent_id>.
  5. Runner writes AGENTS.md, .claude/CLAUDE.md, .claude/settings.json, checkpoint.json, and output.log.
  6. Runner invokes the configured harness through services/agent_harness.py.
  7. Hermes currently runs with:
hermes chat -Q --provider <nous|dashboard-proxy> --model <model> --toolsets terminal,file --accept-hooks --max-turns 90 --source soc-dashboard --query "<prompt>"

Claude Code remains available with:

claude --allowedTools "Read,Write,Bash(curl *)" -p --settings <settings> --model <model> --permission-mode acceptEdits --no-session-persistence --output-format stream-json --verbose "<prompt>"

Codex remains available with:

codex exec --json --skip-git-repo-check --sandbox danger-full-access --model <model> --config 'approval_policy="never"' --config 'model_provider="agentic_proxy"' --config 'model_providers.agentic_proxy.base_url="<proxy>/v1"' --config 'model_providers.agentic_proxy.env_key="OPENAI_API_KEY"' --config 'model_providers.agentic_proxy.wire_api="responses"' "<prompt>"
  1. Runner streams stdout/stderr into output.log and mirrors tails into agent_tasks.output.
  2. task_tracker polls checkpoint.json and process state.
  3. If checkpoint status is done or completed, the task is completed and the harness process is terminated so local GPU work does not continue unnecessarily.

The runner contract must stay harness-agnostic and decision-preserving. Do not add app-side classifiers that decide business intent, ticket reuse, assignment, or workflow choice when the agent can make that decision from context. Narrow fallbacks are acceptable for idempotency, lost-message prevention, and safety recovery, but they must not become the normal operating model.

Wake, Restart, Stop

Wake:

  • If a queued/running task exists, refresh heartbeat and return that task.
  • If no active task exists, spawn a replacement using the latest stored prompt and task type.
  • It does not blindly toggle UI state.

Restart:

  • Stop active task if present.
  • Terminate old agent row.
  • Spawn a replacement for the same ticket/model/prompt.

Stop:

  • Terminate tracked subprocess when possible.
  • Mark task and agent stopped.
  • Record audit/event entries.

Approval Guardrail

Agents are instructed to request approval before any action that can alter an environment, data, accounts, mailboxes, firewall rules, blocklists, services, repositories, or production state.

The agent may decide that an action is needed, but the agent is not the approval authority. It must hit a real platform/provider barrier and request approval, access, or a credential lease. The system then enforces the decision through RBAC, approval APIs, provider permissions, and audit logs.

The approval API:

  • POST /api/changes/request
  • GET /api/changes/{id}/status
  • POST /api/changes/{id}/approve
  • POST /api/changes/{id}/reject
  • POST /api/changes/{id}/complete

The current lab can manually approve from the dashboard. Future deployments can attach this to iTop change workflows, ServiceNow change tasks, CAB rules, Keycloak roles, or external approval systems.

Learning Loop

Ticket work is intentionally separated from workflow creation.

Fast ticket path:

  1. Read full context.
  2. Resolve the task as quickly and safely as possible.
  3. Use approvals when required.
  4. Write notes/checkpoints.
  5. Finish.

Postmortem path:

  1. Review ticket, notes, logs, checkpoints, approvals, failures, and final result.
  2. Record what worked and what should improve.
  3. Propose skills/workflows/tests/guardrails.
  4. Mark for human review.
  5. Promote reviewed learning into reusable assets when appropriate: knowledge article, draft workflow, candidate skills, ticket note, and audit record. Promotion is idempotent per postmortem so repeated review updates the same assets rather than duplicating them.

Workflow-build path:

  1. Build reusable workflow blueprint.
  2. Define approval boundaries.
  3. Define tests and safe test environment assumptions.
  4. Persist workflow in draft/tested state.
  5. Stop before production activation until reviewed.

Fresh Deploy Versus Developed Environment

Fresh environment:

  • PostgreSQL starts from api/init_db.sql.
  • iTop can be disabled with ITOP_SYNC_ENABLED=false.
  • Local provider allows dashboard/agent workflows without any external ticketing system.
  • Agents can run when AGENT_LLM_BASE_URL, proxy routing, and the selected harness auth state are configured. Hermes uses OpenAI-compatible /v1/chat/completions; Claude Code uses Anthropic-compatible /v1/messages.

Developed environment:

  • Apply migrations in order.
  • Keep provider credentials in environment/vault only.
  • Sync provider tickets into canonical records.
  • Add provider-specific skills and KB articles.
  • Review and approve workflows before activation.

Email Provider Boundary

The dashboard should treat email as a provider capability, not as a hard dependency on Mailcow. The reference lab currently uses:

  • Mailcow for the open-source email stack.
  • Keycloak-Mailcow bridge scripts for provisioning and sync through direct MySQL.
  • Optional Mailcow HTTP API shim for read-only compatibility checks and future adapter-style reads.
  • Roundcube webmail as the deployable browser mailbox client and report-phish button surface for demos and email workflow tests.

The Mailcow HTTP shim and Roundcube sidecar are documented in docs/MAILCOW_API_SHIM.md. The API shim exposes only domain, mailbox, and alias inventory and intentionally omits password hashes. It should not be used as a generic write API. In production-style deployments, a customer email product should satisfy the same provider contract through its own adapter while tickets, approvals, audit logs, and agent context remain canonical in the dashboard.