Skip to content

Latest commit

 

History

History
620 lines (524 loc) · 35.7 KB

File metadata and controls

620 lines (524 loc) · 35.7 KB

Architecture

Both entry points — the in-process engine and the /v1 API — drive one Google ADK Workflow graph and shape its outcome into a Report. This page is the map of what runs between text in and report out. Read Concepts first if terms such as system model, lane, ground, or critic are unfamiliar.

The central split is simple: models extract facts and make security judgements; code performs checks with definite answers. The code does not make the whole analysis deterministic. It validates and constrains the probabilistic stages.

The pipeline

A static ADK Workflow with deterministic FunctionNode bookends around the model calls:

flowchart TD
    start([text in]) --> extract["extract<br/>(base)"]
    extract --> validate{{validate}}
    validate -- valid --> assert["assert<br/>(base, ANALYSIS_ASSERTIONS)"]
    validate -- invalid --> repair["repair<br/>(base)"]
    repair --> revalidate{{revalidate}}
    revalidate -- valid --> assert
    revalidate -- invalid --> reject([rejected])
    assert --> reread["reread<br/>(base, ANALYSIS_SOURCE_REVIEW)"]
    reread --> prepare[prepare]

    prepare --> analyze["lane agents, in parallel<br/>one per lane of each framework<br/>(strong)"]
    analyze --> merge["merge<br/>(per framework)"]
    merge --> critic["critic<br/>(per framework, strong)"]
    critic --> router{{route_review}}

    router -- accept --> assemble[assemble]
    router -- revise --> recritic["recritic<br/>(per framework, strong)"]
    recritic --> rereview{{rereview}}
    rereview -- accept --> assemble
    rereview -- revise --> failed([failed])
    assemble --> report([Report])

    classDef llm fill:#ede9fe,stroke:#7c3aed,stroke-width:1.5px,color:#2e1065
    classDef code fill:#e0f2fe,stroke:#0284c7,stroke-width:1.5px,color:#082f49
    classDef gate fill:#fef3c7,stroke:#d97706,stroke-width:1.5px,color:#451a03
    classDef good fill:#dcfce7,stroke:#16a34a,stroke-width:1.5px,color:#052e16
    classDef bad fill:#fee2e2,stroke:#dc2626,stroke-width:1.5px,color:#450a0a
    classDef io fill:#f1f5f9,stroke:#64748b,stroke-width:1.5px,color:#0f172a

    class extract,repair,assert,analyze,critic,recritic llm
    class prepare,merge,assemble code
    class validate,revalidate,router,rereview gate
    class report good
    class reject,failed bad
    class start io
Loading

Purple nodes are model calls. Everything else is a deterministic FunctionNode: blue ones do work, amber ones only choose an edge, and the rounded ends are the run's three outcomes. assert is in the graph only where the deployment sets ANALYSIS_ASSERTIONS; otherwise a valid model goes straight to prepare.

  • extract turns the untrusted text into a canonical system model (five DFD element types: external entity, process, data store, data flow, trust boundary). It writes that model in one of two transports, selected by the deployment — see The extraction transport.
  • validate is a mechanical gate. On the compact transport it expands the wire form into the same canonical model first, and an emission that is not a compact model at all is rejected rather than repaired. Failures route to repair (one bounded pass over the original text) and revalidate; a model that still fails, or is over the 150-element cap, ends as a rejection. Revalidate puts every element the issues did not name back as it was before it checks the whole, and the report's model_repair says which ones it had to.
  • prepare derives the per-analysis context, all of it a pure function of the validated model: the boundary crossings; the deterministic candidates for each lane; the domain packs this system earns; and the system model as the agents will see it — with source_excerpt, source_label and source_speaker stripped, so the only submitter words downstream of here are the ones a finding chose to quote.
  • lane agents (analyze_<framework>_<lane>) run in parallel — one per lane of each framework the job selected, which for STRIDE is its six categories. Each drafts claims in its own lane, and each claim cites at least one ground: a quote from the submitted text, an unknown attribute, or a boundary crossing. A candidate is never one of those — it is a structural lead code found, which an agent may investigate and reject, and which nothing downstream of the prompt reads.
  • merge joins one framework's drafts and runs the mechanical half of the fan-in: every reference resolves, no two lanes reused a claim ID, and every quote ground is matched against the bytes of the source it names. A refused quote is rewritten to the source's own nearest span where one is near enough (repaired_quotes). An unverifiable quote is marked and still renders; a claim where nothing verifies is dropped and marked as a dropped_claims entry. It also computes the per-lane coverage account over the drafts. Then that framework's critic rules on the drafts code did not settle, in one pass — verdicts, dedupe, severity calibration — spending judgement only on what code cannot check. A draft resting on an unknown is ruled needs-info in code before the critic reads, and when that leaves nothing to rule on, or the lanes drafted nothing, the critic is not called. Each package carries its own critic, blind to every other framework: two frameworks' subgraphs never touch between prepare and assemble, because a critic rules its own framework's claims against its own framework's question.
  • route_review runs the mechanical checks the assembler depends on. If the critic's output fails them, one bounded recritic re-ask runs; a second failure is a failed job, not a rejection.
  • assemble builds the final report.

Untrusted input is placed in fenced prompt sections that tell the model to treat it as data. This is an instruction-level defense, not a proof that prompt injection is impossible. Every model output is validated before code relies on it.

The extraction transport

extract produces the largest artifact any node here writes, and a measurable share of it is text the code throws away: an element ID is derived from the element's type and name, so every ID a model writes is overwritten, and every reference repeats it. The compact transport replaces each identifier with a short response-local ref, and a deterministic adapter expands the result into the same SystemModel before the validity gate runs.

Transport What extract writes Selected by
full a SystemModel the default
compact-v4 a compact wire form, expanded in code ANALYSIS_COMPACT_EXTRACTION

Nothing downstream of the gate can tell the two apart, which is the point and also the reason the report records which one ran, in execution.extraction_format. The adapter resolves references and copies fields; it decides no fact. repair writes a full SystemModel on either transport, so the one repair pass is never spent on a format conversion.

Whether the compact transport is worth running is a measurement, not a claim, and the measurement has been made four times. Five corpus sweeps per route put the saving at 3.7% of emitted tokens for compact-v4 — 27,462 against 26,435 over the thirteen cases, 1.44 standard deviations of the full route's own spread — against 3.3%, 4.8% and 4.5% for the three versions before it. uv run python -m evals.bench.deterministic transport prints 5.6% of characters, which is an upper bound.

compact-v4 is the first version whose safety fix raises the saving rather than spending it. Version 3 failed the gate on two flow refs the model derived from the flow's endpoints, so two flows between one pair collided; version 4 names a flow's ref after its label. A label slug runs 17.4 characters over the blessed corpus against 33.2 for an endpoint pair, worth a further 1.04% of the full emission — and it is also what the bench has priced all along, so part of the gap between 5.6% offline and 3.3% live was the bench pricing a ref the model was not writing.

The input side is close to free: the first live run measured 6,116 prompt tokens against the full route's 6,098 on the same case, because the compact schema is smaller by about what the delta prompt adds.

compact-v4 clears ten of the gate's eleven criteria, including duplicate-ref at 0 of 65 and first-pass validity tied with the full route at 0.938. The one it fails is the per-case veto, which the same run showed does not work: 11% of the full route's own readings fall below a floor built from four other sweeps of that route. So the route fails no criterion the gate can resolve, and it stays off on the argument rather than on the test — 3.7% of one node's output tokens, a second code path, and a latency this deployment cannot measure.

The route has failed its own predeclared gate three times — ADR 0035 records each application and what it cost. Latency is not in that gate and remains unmeasured: this deployment's pinned upstream has a per-call spread larger than the whole effect. The transport is off, and turning it off again is one variable and a restart.

Variable Effect
ANALYSIS_COMPACT_EXTRACTION Ask extract for the compact wire form. Off by default.

The extraction strategy

The transport above changes how extract spells one answer. A strategy changes what the reading node is asked for. ANALYSIS_FACTS_FIRST_EXTRACTION replaces extract with two nodes: facts reads the sources and writes a Source Fact Bundle — the things the text names, the interactions between them, one statement at a time, and the questions it leaves open, all under short handles the model invents — and resolve turns that bundle into the same SystemModel the validity gate reads and a catalog proposal prepare resolves. From the gate onward the graph is identical, which is what makes a comparison between the two strategies a comparison of reading order.

Strategy What the reading node writes Selected by
graph-first a SystemModel in one pass the default
facts-first a SourceFactBundle, resolved in code ANALYSIS_FACTS_FIRST_EXTRACTION
facts-split the same bundle, across two calls ANALYSIS_FACTS_SPLIT_EXTRACTION

The third exists because the first two differ in two ways at once: the reading order, and how many calls the work is spread over — graph-first runs extract and assert, facts-first runs one node that does both. A comparison between them cannot say which of those a difference belongs to. facts-split holds the order and splits the calls: inventory names the things, the interactions and the zones, and rows states the facts about them. Each call is read only for the half it owns, and a call that answered the other half is reported rather than merged in.

facts runs on the same tier row as extract, because it is the extraction stage under another order and the comparison needs both on one model. The report records which strategy ran, in execution.extraction_strategy.

The strategy carries its own assertion rows, so it is refused together with ANALYSIS_ASSERTIONS: a second pass over the model the bundle built would read one thing twice. It is refused with the compact transport for the same kind of reason — the bundle is not a model, so it has no compact wire form.

This is an experiment (#1003), disabled by default and unmeasured. No run has compared the two.

Variable Effect
ANALYSIS_FACTS_FIRST_EXTRACTION Read the sources facts-first. Off by default.

The assertion pass

ANALYSIS_ASSERTIONS puts one more base-tier call into every job, between the validity gate and prepare. The assert node reads the sources against the validated model and proposes one row per statement the sources make: a subject, a predicate from the service's registry, a value, a scope and the quote that says so. prepare resolves the proposal in code, once, and parks the record every later node reads. A settled row of a predicate the graph has no field for — a second factor stated absent, a credential stated shared — is then an entry in the evidence catalog every lane selects from, and the report embeds the catalog and the rows the resolver refused under assertions. See ADR 0034 for the contract and ADR 0036 for what a lane may cite.

An assert node that writes nothing fails the job, on the rule a silent lane fails it: a pass the deployment selected and a pass that never ran are not the same report. The pass is off by default, and a report built without it carries assertions: null and every other field it carried before.

Variable Effect
ANALYSIS_ASSERTIONS Run the assertion pass on every job. Off by default.

The source review

ANALYSIS_SOURCE_REVIEW puts a bounded repair pass between the catalog and prepare. reading renders the model and every recorded statement, reread reads the sources once more against them and proposes typed operations — add an element, add an interaction, add a statement, retract one, or mark a question open — and apply applies the batch or discards it whole. A correction is a retraction and an addition, because a statement's identity is computed from its own parts.

Nothing lands on a stale reading. Every operation names what it assumes about the artifacts it was written against, an operation that reaches a refused one is refused with it, and a batch whose result fails the validity or catalog gate is discarded entirely, with the artifacts the pass was shown left standing. The outcomes say which happened.

The pass reads a catalog, so it needs a route that produces one — either ANALYSIS_ASSERTIONS or ANALYSIS_FACTS_FIRST_EXTRACTION. One pass serves both, on purpose: it is the same three nodes over either head, so a comparison between them measures the extraction order and never two applicators.

This is an experiment (#1003), disabled by default and unmeasured.

Variable Effect
ANALYSIS_SOURCE_REVIEW Read the sources once more and patch what they support. Off by default.

The experiment is scored on the catalog, and no lane writes one, so the eval harness carries a head-only entry: run.py run --mode heads runs one arm's reading node, the validity gate, the bounded repair and whichever passes the deployment selects, resolves the catalog through the seam prepare resolves through, and stops. Measured against a recorded STRIDE sweep, running the lanes as well would cost roughly seventeen times as much for findings the endpoint never reads.

Models

Every LLM node runs on one of two model tiers — named for the job they do, not for any vendor's product line:

  • base — the workhorse: extract, repair and assert, which are transcription rather than judgement.
  • strong — judgement: every framework's lane agents, its critic, and the recritic.

config/model_tiers.toml maps nodes to tiers and each tier to a (vendor, model) pair. Deterministic FunctionNodes carry no model. The strong tier does most of a job's model work — the six-way category fan-out plus the critic and its re-ask.

The tiers choose their vendor independently, so they can run different vendors at the same time. A third tier, review, exists so criticism can be bound off the model it checks; the shipped node map points nothing at it. Every vendor is reached through one adapter (LiteLLM), and the LLM nodes share one adapter per bound tier — so the startup checks on credentials and decoding parameters run once per tier rather than once per node, and a selected tier nothing runs on costs no credential at all.

Provenance and certification

The implementation records three related values:

  • Served build — the model identifier the provider says actually answered a request, prefixed with its vendor (vertex_ai/gemini-2.5-pro-002). Not necessarily the one you asked for. Gemini is the worked example here because one had to be. It is also the one profiled family whose served build differs from the route you asked for, which is the distinction these three terms exist to draw. On Claude, both model fields hold the same string.
  • Sampling fingerprint — sha256 of the provider-qualified served model plus that tier's resolved decoding parameters. It identifies the generation setup, not everything that affected the output: the report records the input digest and instruction digest separately.
  • Blessed — a fingerprint recorded in config/blessed-fingerprints.toml because a measured, sanctioned run produced it. The list is this deployment's own; nothing about it ships from this repo.

Every completed report carries enough information to compare model selection, sampling, input identity, and instruction identity. It does not make a probabilistic result reproducible in the strict sense.

Field What it holds
NodeRun.requested_model The configured route — what was asked for (vertex_ai/gemini-2.5-pro).
NodeRun.model The served build — what actually answered (vertex_ai/gemini-2.5-pro-002).
NodeRun.instruction_sha256 The digest of the instructions the graph this node ran in carried.
NodeRun.execution_fingerprint The identity hash: both routes, the tier's decoding params, the instruction digest, and the build versions.
Report.execution The identity schema version, how far the served builds can be trusted, and those build versions.

The report records both model fields and compares neither directly. A served model change changes the execution fingerprint, which makes the run uncertified unless the deployment has blessed that fingerprint.

The identity binds the requested route as well as the served one. The served build is what the provider said answered, read off its own event stream, and nothing verifies it — each Report.nodes[].served_trust says so in the artifact rather than leaving a reader to assume better. Binding both routes is what stops the provider's word from selecting a blessed entry on its own: a translator that returns an approved build while the deployment asked for something cheaper presents a pair no manifest holds.

A prompt edit, a litellm bump or a service release moves every fingerprint. Each of those changes what a node can answer, so each re-baselines the manifest and runs read as uncertified until a sanctioned sweep blesses the new hashes.

The fingerprint is computed per node execution, not once at startup. If the served model changes during a run, different nodes can therefore carry different fingerprints. The vendor prefix is part of the hash because a served identifier alone carries no vendor—Vertex-hosted Claude and Anthropic-direct Claude can return the same model string.

config/blessed-fingerprints.toml records blessed fingerprints per tier, not per node. A fingerprint contains no node name, and critic and recritic run on the same tier, so they present a byte-identical hash; keying by node would call that one hash blessed under critic and unblessed under recritic, marking the first revise path in production uncertified on a technicality.

The list is deployment-local. This project can never ship a run that already counts as certified, because a repo-level blessing plus a local one could only resolve as one silently overriding the other. ANALYSIS_BLESSED_FINGERPRINTS chooses which single file is read — it does not layer a second one on top.

The service checks every completed job against the deployment's fingerprint manifest. This is a narrow operational attestation: it says whether the observed model-and-sampling fingerprints were approved. It does not judge the extracted model, findings, prompts, or input.

State Meaning Effect on GET /v1/jobs/{id}/report
certified Every observed fingerprint is blessed Served
uncertified At least one is not Served unless ANALYSIS_REQUIRE_CERTIFIED
unexercised A tier the graph declares presented no fingerprint at all Always withheld

The lists ship empty, so until you promote a measured baseline every run reads as uncertified. That is recorded, not fatal — a gate that fires before anyone knows the normal range just trains people to switch it off. Promoting one is python -m evals.harness.run promote <artifact> --yes, which derives the blessed fingerprints from the served builds the sweep observed rather than from anything typed in by hand — see TUNING.md. unexercised is different: every tier has a node that always runs, so it cannot happen on a run that produced a report at all. It is an internal assertion, and enforcing it costs nothing.

Withholding refuses the report; it never fails the job. A failed job carries no report at all, and the fingerprints that show what drifted live inside it. Nothing about certification appears in the job status view — it is operator-only.

A job that asks questions is certified in two parts. It pauses after its head, and the service certifies the nodes the head ran at that point. The questions route and the answers route refuse a waiting job under the same rules as a report. The job that the answers start runs the rest of the graph and carries the head's result. Its report is served only where both parts pass. A job that answers a finished report carries that report's result in the same way.

The two parts are two graphs, and the fingerprint of a node binds the digest of its whole graph's instructions. So an end-to-end sweep's fingerprints bless neither part. The paused head is the heads sweep's graph and the resumed run is the analysis sweep's graph, and tests/test_pause.py holds each pair to one digest. A deployment that serves jobs that ask questions promotes a heads sweep and an analysis sweep beside its end-to-end sweep.

Variable Effect
ANALYSIS_REQUIRE_CERTIFIED Withhold the report when the run is uncertified. Off by default.

Outcomes

The graph reaches one of three states, surfaced identically by both entry points:

  • completed — a Report.
  • rejected — the input failed the validity gate; carries the ValidationIssues.
  • failed (raises) — an internal error. No partial report is ever produced: an empty Tampering section means "looked, found nothing", never "the Tampering agent errored".

That last guarantee is enforced, not assumed. An LLM node whose completion is truncated writes no output key at all — ADK saves one only from a final event carrying text — so "the agent errored" and "the agent found nothing" arrive as an absent key and an empty one. merge_drafts distinguishes them and fails the job on the first, naming the lanes and the knob; validate_extraction does the same one node earlier. Read as equivalent, a truncated agent would delete a sixth of the analysis and finish green, because the critic rules what it is handed and by_category omits a lane with no threats rather than carrying a zero.

Resilience

Retry, timeout and the per-job deadline are configured in config/resilience.toml and attached to the adapter itself, so a retry is invisible to the graph and the report's nodes array is unchanged by one. A per-request timeout turns a hang into an error the retry can act on. Three attempts by default.

The retry loop sits above the provider seam described below, and the translator sits beneath it. That is where the loop's two bounds are expressible: one token bucket shared by the whole process, and jitter that spreads out lanes which failed at the same instant. Neither is something a single call can do for itself.

Concurrency and isolation

Concurrent analyses are independent. Each analyze() call runs one job in its own ADK session, created fresh with a unique session id and seeded with only that job's own input text; the graph's per-run data lives entirely in that session's state, read back by the same session id. The engine, runner, and compiled pipeline hold no per-run state — they carry only read-only configuration (the node→model map, the loaded prompts and skills), so one engine is safe to share across every call.

  • Isolation is per session, not per caller. Two analyses submitted at the same time by the same caller still get separate sessions, so they can never read or overwrite each other's state.
  • Within a single analysis, the lane agents run in parallel but each writes to its own category-keyed slot in the session, so the parallel branches don't clobber one another before the merge.
  • Untrusted input stays contained. Because a job's text lives only in its own session (as fenced data — see The pipeline), a prompt-injection attempt in one submission cannot reach another running analysis.

Containment is not resistance, and the two are measured differently. Fencing is structural and deterministic: every caller byte sits inside a marker sized to its own content, so a submission cannot close the block it is in and continue in instruction position. evals/adversarial/ carries a source built to try exactly that, and CI asserts it fails.

What fencing does not do is stop a model reading ignore all previous instructions from inside a fence and deciding to obey it. That is semantic, it is a property of a model and a prompt set rather than of this code, and it is measured by scoring a report against what the injection asked for — deterministic grading, no model judge. That half needs a live model and has never run; see evals/adversarial/README.md for the bar and the residual risk.

The consequence of a failure is bounded by what a model here can reach: no model in this service holds any tool or host authority. Every LLM node returns structured text that a deterministic FunctionNode validates, so a model talked into something is talked into producing a bad report, not into acting.

This guarantee assumes the intended concurrency model: async calls on a single event loop. The default InMemorySessionService is an in-process store — safe for cooperative async concurrency with distinct session ids, but not a thread-safe store to share across OS threads. Scaling across processes keeps analyses isolated (nothing is shared), but then each worker has its own in-memory job and session state, so a job must be routed to the worker that holds it — which is what the persistent backends below are for.

The model translator runs in this process

Every model reaches its provider through ADK's LiteLlm and, beneath that, litellm — ADR 0015 records why that is one substrate rather than a swappable adapter. Both run in the service process, with the service's authority. LiteLLM holds the provider credentials, opens the network connections, and is the only code between a node's request and a provider's answer.

A provider seam sits between the graph and that translator: analysis_service.provider defines one bounded request that goes out, one result that comes back, and a failure as a value rather than as an exception object. Everything the service asks of a provider crosses it.

The seam is not containment and does not claim to be. There is one implementation and it runs the translator in this process, under this process's credentials. What the seam buys is that the boundary is written down: a request carries no credential and nothing executable, so an implementation that ran elsewhere has a contract to meet rather than a dependency's type graph to reason about. Containment is the deployment work below, and this repository still ships none of it.

State the consequence plainly: a compromise of that dependency is a compromise of this application. A malicious release, or arbitrary code execution inside it, would reach the process's environment — including every provider credential this deployment declared — its filesystem, and its outbound network. Nothing in this repository contains that, because containment is deployment work and this repository ships no deployment packaging: no image, no container definition, no egress policy. An operator running this service in production owns that boundary. What it should look like is set out in #502.

What the repository does bound is the seam — the set of values this service hands the translator, and where each comes from:

What crosses Where it comes from
the model route Vendor.route() — a registry entry, in code
the credential kwargs the vendor's own table, read from declared ANALYSIS_* variables
the decoding params sampling.toml plus an explicit env allowlist
num_retries=0 a literal

No provider endpoint crosses it at all. Nothing sets api_base, base_url or api_version anywhere in the package, so a request goes where the vendor's own client sends it. custom_llm_provider is set — the build-time capability probe has to name a provider — and only ever from a Vendor.

That matters because prompts, submitted sources, corpus text and model output all flow through this process. An adapter that accepted an address from any of them would be an SSRF and endpoint-substitution path. tests/test_translator_seam.py fails if one appears, if a decoding param could express one, or if a new value starts crossing the seam.

No TLS setting crosses it either. LiteLLM resolves certificate verification from a call kwarg, then from SSL_VERIFY in the process environment, then from its own litellm.ssl_verify. The package sets none of the four kwargs that reach that resolution, and neither sampling.toml nor the env allowlist can express one. tests/test_translator_seam.py also drives LiteLLM's own resolver, because the environment is a path no kwarg lint can see.

Only a declared credential authenticates a run. LiteLLM reads a key out of the process environment on its own, so a machine that ever ran another tool against a provider carries a key this deployment never declared. Vendor.credential_kwargs reads the ANALYSIS_* variable instead and fails the build closed when it is absent, so no adapter that could fall back is ever built. tests/test_transport_conformance.py holds both halves against the bytes: the declared key beats an ambient one, and an ambient one alone builds nothing.

Dependency versions are pinned exactly (pyproject.toml) and hashed at install (uv.lock), and the installed version of every distribution between a node and its provider is inside each run's execution identity — so a litellm bump moves every fingerprint rather than silently reusing a blessing taken before it.

Seams

The pipeline is reached through interfaces, so backends are swappable and the whole graph runs offline against scripted models:

Seam Interface Default Status
Pipeline execution PipelineRunner AdkPipelineRunner (real graph) / StubPipelineRunner (tests) Complete
Job persistence JobStore InMemoryJobStore (memory) Backend selected by ANALYSIS_JOB_STORE via a fail-closed registry; only the non-durable memory backend ships — a durable one is a new registry entry
ADK sessions BaseSessionService InMemorySessionService In-memory only; a session_service_uri backend is unwired

The in-memory defaults are enough to get a report in process. Choosing a backend is already wired for the JobStore (ANALYSIS_JOB_STORE, which stops startup on an unset or unknown value rather than quietly falling back). Still out of scope for the current work: a durable JobStore implementation, a session backend, deployment packaging (container, Cloud Run), and observability. The interfaces and the selection seam are in place for all of them.

Where the code lives

Module Responsibility
analysis_service.engine In-process Engine facade.
analysis_service.api The /v1 FastAPI app.
analysis_service.jobs Job lifecycle, JobStore, PipelineRunner seams.
analysis_service.sources Source: the untrusted text a job is built from, the per-deployment bounds both entry points enforce, and the fenced render that is the only way caller bytes reach a model (OWASP LLM01).
analysis_service.deployment One installation's config, resolved once: the files, the graph they configure, its runner and its certification gate.
analysis_service.pipeline AdkPipelineRunner: one job's identity, input digest and certification around a Graph Run.
analysis_service.execution Drives a built graph, stamps each node execution, and turns the finished run into its report or its rejection. Shared by the service and the eval harness.
analysis_service.graph Topology and node functions.
analysis_service.system_model Canonical model + validity helpers.
analysis_service.analysis Deterministic traversal of a validated model: flows, reachability, paths, unknown controls. No security claims.
analysis_service.candidates The rule table. Structural conditions an agent should investigate — leads, never findings, never evidence.
analysis_service.domains Which skills/domains/ packs a model earns, decided from its own technology fields.
analysis_service.coverage Per-lane accounting: what each agent was offered, and what its drafts cite.
analysis_service.claims The neutral Claim, its grounds, the proposal and ruling wrappers, the marks, the severity model and the FrameworkAnalysis block.
analysis_service.report The Report envelope: the job, the input, the node runs and the blocks.
analysis_service.frameworks The framework-package contract, its registry and its deployment gate.
analysis_service.validation The mechanical validity gate.
analysis_service.fan_in One framework's lane batches merged into the drafts its critic reads: evidence resolution, the whole-set checks, and every mark the service records about a draft, behind one call.
analysis_service.critic The mechanical checks around the critic's ruling — the ones no model should be asked to perform — and the view a critic reads.
analysis_service.skills / .prompts / .markdown_loader Skill/prompt loading and composition.
analysis_service.model_tiers / .sampling / .resilience Config loaders.
analysis_service.vendors The vendor registry: each vendor's router prefix, credential mode, and model-name rules.
analysis_service.binding Builds one adapter per tier from (vendor, model, sampling, resilience), and the NodeBinding the graph binds onto its LLM nodes.
analysis_service.model_gate The startup check that asks the provider library whether a tier's parameters are actually supported.
analysis_service.certification Compares a run's fingerprints against the deployment's blessed manifest.
analysis_service.auth Bearer-token (OIDC JWT) verification.
analysis_service.errors ConfigError, the base every fail-closed config loader raises — so a caller can handle "this deployment cannot run" without enumerating which knob was wrong.