Both entry points — the in-process engine and the
/v1 API — drive one Google ADK Workflow graph and shape its
outcome into a Report. This page is the map of what
runs between text in and report out. Read Concepts first if terms
such as system model, lane, ground, or critic are unfamiliar.
The central split is simple: models extract facts and make security judgements; code performs checks with definite answers. The code does not make the whole analysis deterministic. It validates and constrains the probabilistic stages.
A static ADK Workflow with deterministic FunctionNode bookends around the
model calls:
flowchart TD
start([text in]) --> extract["extract<br/>(base)"]
extract --> validate{{validate}}
validate -- valid --> assert["assert<br/>(base, ANALYSIS_ASSERTIONS)"]
validate -- invalid --> repair["repair<br/>(base)"]
repair --> revalidate{{revalidate}}
revalidate -- valid --> assert
revalidate -- invalid --> reject([rejected])
assert --> reread["reread<br/>(base, ANALYSIS_SOURCE_REVIEW)"]
reread --> prepare[prepare]
prepare --> analyze["lane agents, in parallel<br/>one per lane of each framework<br/>(strong)"]
analyze --> merge["merge<br/>(per framework)"]
merge --> critic["critic<br/>(per framework, strong)"]
critic --> router{{route_review}}
router -- accept --> assemble[assemble]
router -- revise --> recritic["recritic<br/>(per framework, strong)"]
recritic --> rereview{{rereview}}
rereview -- accept --> assemble
rereview -- revise --> failed([failed])
assemble --> report([Report])
classDef llm fill:#ede9fe,stroke:#7c3aed,stroke-width:1.5px,color:#2e1065
classDef code fill:#e0f2fe,stroke:#0284c7,stroke-width:1.5px,color:#082f49
classDef gate fill:#fef3c7,stroke:#d97706,stroke-width:1.5px,color:#451a03
classDef good fill:#dcfce7,stroke:#16a34a,stroke-width:1.5px,color:#052e16
classDef bad fill:#fee2e2,stroke:#dc2626,stroke-width:1.5px,color:#450a0a
classDef io fill:#f1f5f9,stroke:#64748b,stroke-width:1.5px,color:#0f172a
class extract,repair,assert,analyze,critic,recritic llm
class prepare,merge,assemble code
class validate,revalidate,router,rereview gate
class report good
class reject,failed bad
class start io
Purple nodes are model calls. Everything else is a deterministic FunctionNode:
blue ones do work, amber ones only choose an edge, and the rounded ends are the
run's three outcomes. assert is in the graph only where the deployment sets
ANALYSIS_ASSERTIONS; otherwise a valid model goes straight to prepare.
- extract turns the untrusted text into a canonical system model (five DFD element types: external entity, process, data store, data flow, trust boundary). It writes that model in one of two transports, selected by the deployment — see The extraction transport.
- validate is a mechanical gate. On the compact transport it expands the
wire form into the same canonical model first, and an emission that is not a
compact model at all is rejected rather than repaired. Failures route to repair (one bounded
pass over the original text) and revalidate; a model that still fails, or is
over the 150-element cap, ends as a rejection.
Revalidate puts every element the issues did not name back as it was before
it checks the whole, and the report's
model_repairsays which ones it had to. - prepare derives the per-analysis context, all of it a pure function of the
validated model: the boundary crossings; the deterministic candidates for
each lane; the domain packs this system earns; and the system model as the
agents will see it — with
source_excerpt,source_labelandsource_speakerstripped, so the only submitter words downstream of here are the ones a finding chose to quote. - lane agents (
analyze_<framework>_<lane>) run in parallel — one per lane of each framework the job selected, which for STRIDE is its six categories. Each drafts claims in its own lane, and each claim cites at least one ground: a quote from the submitted text, anunknownattribute, or a boundary crossing. A candidate is never one of those — it is a structural lead code found, which an agent may investigate and reject, and which nothing downstream of the prompt reads. - merge joins one framework's drafts and runs the mechanical half of the
fan-in: every reference resolves, no two lanes reused a claim ID, and every
quote ground is matched against the bytes of the source it names. A refused
quote is rewritten to the source's own nearest span where one is near enough
(
repaired_quotes). An unverifiable quote is marked and still renders; a claim where nothing verifies is dropped and marked as adropped_claimsentry. It also computes the per-lanecoverageaccount over the drafts. Then that framework's critic rules on the drafts code did not settle, in one pass — verdicts, dedupe, severity calibration — spending judgement only on what code cannot check. A draft resting on anunknownis ruledneeds-infoin code before the critic reads, and when that leaves nothing to rule on, or the lanes drafted nothing, the critic is not called. Each package carries its own critic, blind to every other framework: two frameworks' subgraphs never touch betweenprepareandassemble, because a critic rules its own framework's claims against its own framework's question. - route_review runs the mechanical checks the assembler depends on. If the
critic's output fails them, one bounded recritic re-ask runs; a second
failure is a
failedjob, not a rejection. - assemble builds the final report.
Untrusted input is placed in fenced prompt sections that tell the model to treat it as data. This is an instruction-level defense, not a proof that prompt injection is impossible. Every model output is validated before code relies on it.
extract produces the largest artifact any node here writes, and a measurable
share of it is text the code throws away: an element ID is derived from the
element's type and name, so every ID a model writes is overwritten, and every
reference repeats it. The compact transport replaces each identifier with a
short response-local ref, and a deterministic adapter expands the result into
the same SystemModel before the validity gate runs.
| Transport | What extract writes |
Selected by |
|---|---|---|
| full | a SystemModel |
the default |
| compact-v4 | a compact wire form, expanded in code | ANALYSIS_COMPACT_EXTRACTION |
Nothing downstream of the gate can tell the two apart, which is the point and
also the reason the report records which one ran, in
execution.extraction_format. The adapter resolves
references and copies fields; it decides no fact. repair writes a full
SystemModel on either transport, so the one repair pass is never spent on a
format conversion.
Whether the compact transport is worth running is a measurement, not a claim,
and the measurement has been made four times. Five corpus sweeps per route put
the saving at 3.7% of emitted tokens for compact-v4 — 27,462 against
26,435 over the thirteen cases, 1.44 standard deviations of the full route's own
spread — against 3.3%, 4.8% and 4.5% for the three versions before it. uv run python -m evals.bench.deterministic transport prints 5.6% of characters,
which is an upper bound.
compact-v4 is the first version whose safety fix raises the saving rather than
spending it. Version 3 failed the gate on two flow refs the model derived from
the flow's endpoints, so two flows between one pair collided; version 4 names a
flow's ref after its label. A label slug runs 17.4 characters over the blessed
corpus against 33.2 for an endpoint pair, worth a further 1.04% of the full
emission — and it is also what the bench has priced all along, so part of the
gap between 5.6% offline and 3.3% live was the bench pricing a ref the model was
not writing.
The input side is close to free: the first live run measured 6,116 prompt tokens against the full route's 6,098 on the same case, because the compact schema is smaller by about what the delta prompt adds.
compact-v4 clears ten of the gate's eleven criteria, including
duplicate-ref at 0 of 65 and first-pass validity tied with the full route at
0.938. The one it fails is the per-case veto, which the same run showed does not
work: 11% of the full route's own readings fall below a floor built from four
other sweeps of that route. So the route fails no criterion the gate can
resolve, and it stays off on the argument rather than on the test — 3.7% of one
node's output tokens, a second code path, and a latency this deployment cannot
measure.
The route has failed its own predeclared gate three times — ADR 0035 records each application and what it cost. Latency is not in that gate and remains unmeasured: this deployment's pinned upstream has a per-call spread larger than the whole effect. The transport is off, and turning it off again is one variable and a restart.
| Variable | Effect |
|---|---|
ANALYSIS_COMPACT_EXTRACTION |
Ask extract for the compact wire form. Off by default. |
The transport above changes how extract spells one answer. A strategy
changes what the reading node is asked for. ANALYSIS_FACTS_FIRST_EXTRACTION
replaces extract with two nodes: facts reads the sources and writes a
Source Fact Bundle — the things the text names, the interactions between
them, one statement at a time, and the questions it leaves open, all under
short handles the model invents — and resolve turns that bundle into the
same SystemModel the validity gate reads and a catalog proposal prepare
resolves. From the gate onward the graph is identical, which is what makes a
comparison between the two strategies a comparison of reading order.
| Strategy | What the reading node writes | Selected by |
|---|---|---|
| graph-first | a SystemModel in one pass |
the default |
| facts-first | a SourceFactBundle, resolved in code |
ANALYSIS_FACTS_FIRST_EXTRACTION |
| facts-split | the same bundle, across two calls | ANALYSIS_FACTS_SPLIT_EXTRACTION |
The third exists because the first two differ in two ways at once: the
reading order, and how many calls the work is spread over — graph-first runs
extract and assert, facts-first runs one node that does both. A comparison
between them cannot say which of those a difference belongs to. facts-split
holds the order and splits the calls: inventory names the things, the
interactions and the zones, and rows states the facts about them. Each call is
read only for the half it owns, and a call that answered the other half is
reported rather than merged in.
facts runs on the same tier row as extract, because it is the extraction
stage under another order and the comparison needs both on one model. The
report records which strategy ran, in
execution.extraction_strategy.
The strategy carries its own assertion rows, so it is refused together with
ANALYSIS_ASSERTIONS: a second pass over the model the bundle built would
read one thing twice. It is refused with the compact transport for the same
kind of reason — the bundle is not a model, so it has no compact wire form.
This is an experiment (#1003), disabled by default and unmeasured. No run has compared the two.
| Variable | Effect |
|---|---|
ANALYSIS_FACTS_FIRST_EXTRACTION |
Read the sources facts-first. Off by default. |
ANALYSIS_ASSERTIONS puts one more base-tier call into every job, between
the validity gate and prepare. The assert node reads the sources against
the validated model and proposes one row per statement the sources make: a
subject, a predicate from the service's registry, a value, a scope and the
quote that says so. prepare resolves the proposal in code, once, and parks
the record every later node reads. A settled row of a predicate the graph has
no field for — a second factor stated absent, a credential stated shared — is
then an entry in the evidence catalog every lane selects from, and the report
embeds the catalog and the rows the resolver refused under assertions. See
ADR 0034 for
the contract and
ADR 0036 for
what a lane may cite.
An assert node that writes nothing fails the job, on the rule a silent lane
fails it: a pass the deployment selected and a pass that never ran are not the
same report. The pass is off by default, and a report built without it carries
assertions: null and every other field it carried before.
| Variable | Effect |
|---|---|
ANALYSIS_ASSERTIONS |
Run the assertion pass on every job. Off by default. |
ANALYSIS_SOURCE_REVIEW puts a bounded repair pass between the catalog and
prepare. reading renders the model and every recorded statement, reread
reads the sources once more against them and proposes typed operations — add an
element, add an interaction, add a statement, retract one, or mark a question
open — and apply applies the batch or discards it whole. A correction is a
retraction and an addition, because a statement's identity is computed from its
own parts.
Nothing lands on a stale reading. Every operation names what it assumes about the artifacts it was written against, an operation that reaches a refused one is refused with it, and a batch whose result fails the validity or catalog gate is discarded entirely, with the artifacts the pass was shown left standing. The outcomes say which happened.
The pass reads a catalog, so it needs a route that produces one — either
ANALYSIS_ASSERTIONS or ANALYSIS_FACTS_FIRST_EXTRACTION. One pass serves
both, on purpose: it is the same three nodes over either head, so a comparison
between them measures the extraction order and never two applicators.
This is an experiment (#1003), disabled by default and unmeasured.
| Variable | Effect |
|---|---|
ANALYSIS_SOURCE_REVIEW |
Read the sources once more and patch what they support. Off by default. |
The experiment is scored on the catalog, and no lane writes one, so the eval
harness carries a head-only entry: run.py run --mode heads runs one arm's
reading node, the validity gate, the bounded repair and whichever passes the
deployment selects, resolves the catalog through the seam prepare resolves
through, and stops. Measured against a recorded STRIDE sweep, running the lanes
as well would cost roughly seventeen times as much for findings the endpoint
never reads.
Every LLM node runs on one of two model tiers — named for the job they do, not for any vendor's product line:
base— the workhorse:extract,repairandassert, which are transcription rather than judgement.strong— judgement: every framework's lane agents, itscritic, and therecritic.
config/model_tiers.toml maps nodes to tiers and each tier
to a (vendor, model) pair. Deterministic FunctionNodes carry no model. The
strong tier does most of a job's model work — the six-way category fan-out
plus the critic and its re-ask.
The tiers choose their vendor independently, so they can run different
vendors at the same time. A third tier, review, exists so criticism can be
bound off the model it checks; the shipped node map points nothing at it. Every
vendor is reached through one adapter (LiteLLM), and the LLM nodes share one
adapter per bound tier — so the startup checks on credentials and decoding
parameters run once per tier rather than once per node, and a selected tier
nothing runs on costs no credential at all.
The implementation records three related values:
- Served build — the model identifier the provider says actually answered a
request, prefixed with its vendor (
vertex_ai/gemini-2.5-pro-002). Not necessarily the one you asked for. Gemini is the worked example here because one had to be. It is also the one profiled family whose served build differs from the route you asked for, which is the distinction these three terms exist to draw. On Claude, both model fields hold the same string. - Sampling fingerprint —
sha256of the provider-qualified served model plus that tier's resolved decoding parameters. It identifies the generation setup, not everything that affected the output: the report records the input digest and instruction digest separately. - Blessed — a fingerprint recorded in
config/blessed-fingerprints.tomlbecause a measured, sanctioned run produced it. The list is this deployment's own; nothing about it ships from this repo.
Every completed report carries enough information to compare model selection, sampling, input identity, and instruction identity. It does not make a probabilistic result reproducible in the strict sense.
| Field | What it holds |
|---|---|
NodeRun.requested_model |
The configured route — what was asked for (vertex_ai/gemini-2.5-pro). |
NodeRun.model |
The served build — what actually answered (vertex_ai/gemini-2.5-pro-002). |
NodeRun.instruction_sha256 |
The digest of the instructions the graph this node ran in carried. |
NodeRun.execution_fingerprint |
The identity hash: both routes, the tier's decoding params, the instruction digest, and the build versions. |
Report.execution |
The identity schema version, how far the served builds can be trusted, and those build versions. |
The report records both model fields and compares neither directly. A served model change changes the execution fingerprint, which makes the run uncertified unless the deployment has blessed that fingerprint.
The identity binds the requested route as well as the served one. The served
build is what the provider said answered, read off its own event stream, and
nothing verifies it — each Report.nodes[].served_trust says so in the
artifact rather than leaving a reader to assume better. Binding both routes is
what stops the provider's word from selecting a blessed entry on its own: a
translator that returns an approved build while the deployment asked for
something cheaper presents a pair no manifest holds.
A prompt edit, a litellm bump or a service release moves every fingerprint.
Each of those changes what a node can answer, so each re-baselines the manifest
and runs read as uncertified until a sanctioned sweep blesses the new hashes.
The fingerprint is computed per node execution, not once at startup. If the served model changes during a run, different nodes can therefore carry different fingerprints. The vendor prefix is part of the hash because a served identifier alone carries no vendor—Vertex-hosted Claude and Anthropic-direct Claude can return the same model string.
config/blessed-fingerprints.toml records blessed fingerprints per tier,
not per node. A fingerprint contains no node name, and critic and recritic
run on the same tier, so they present a byte-identical hash; keying by node
would call that one hash blessed under critic and unblessed under recritic,
marking the first revise path in production uncertified on a technicality.
The list is deployment-local. This project can never ship a run that already
counts as certified, because a repo-level blessing plus a local one could only
resolve as one silently overriding the other.
ANALYSIS_BLESSED_FINGERPRINTS chooses which single file is read — it does not
layer a second one on top.
The service checks every completed job against the deployment's fingerprint manifest. This is a narrow operational attestation: it says whether the observed model-and-sampling fingerprints were approved. It does not judge the extracted model, findings, prompts, or input.
| State | Meaning | Effect on GET /v1/jobs/{id}/report |
|---|---|---|
| certified | Every observed fingerprint is blessed | Served |
| uncertified | At least one is not | Served unless ANALYSIS_REQUIRE_CERTIFIED |
| unexercised | A tier the graph declares presented no fingerprint at all | Always withheld |
The lists ship empty, so until you promote a measured baseline every run
reads as uncertified. That is recorded, not fatal — a gate that fires before
anyone knows the normal range just trains people to switch it off. Promoting one
is python -m evals.harness.run promote <artifact> --yes, which derives the
blessed fingerprints from the served builds the sweep observed rather than from
anything typed in by hand — see
TUNING.md.
unexercised is different: every tier has a node that always runs, so it cannot
happen on a run that produced a report at all. It is an internal assertion, and
enforcing it costs nothing.
Withholding refuses the report; it never fails the job. A failed job carries no report at all, and the fingerprints that show what drifted live inside it. Nothing about certification appears in the job status view — it is operator-only.
A job that asks questions is certified in two parts. It pauses after its head, and the service certifies the nodes the head ran at that point. The questions route and the answers route refuse a waiting job under the same rules as a report. The job that the answers start runs the rest of the graph and carries the head's result. Its report is served only where both parts pass. A job that answers a finished report carries that report's result in the same way.
The two parts are two graphs, and the fingerprint of a node binds the digest of
its whole graph's instructions. So an end-to-end sweep's fingerprints bless
neither part. The paused head is the heads sweep's graph and the resumed run
is the analysis sweep's graph, and tests/test_pause.py holds each pair to
one digest. A deployment that serves jobs that ask questions promotes a heads
sweep and an analysis sweep beside its end-to-end sweep.
| Variable | Effect |
|---|---|
ANALYSIS_REQUIRE_CERTIFIED |
Withhold the report when the run is uncertified. Off by default. |
The graph reaches one of three states, surfaced identically by both entry points:
- completed — a
Report. - rejected — the input failed the validity gate; carries the
ValidationIssues. - failed (raises) — an internal error. No partial report is ever produced: an empty Tampering section means "looked, found nothing", never "the Tampering agent errored".
That last guarantee is enforced, not assumed. An LLM node whose completion is
truncated writes no output key at all — ADK saves one only from a final event
carrying text — so "the agent errored" and "the agent found nothing" arrive
as an absent key and an empty one. merge_drafts distinguishes them and
fails the job on the first, naming the lanes and the knob; validate_extraction
does the same one node earlier. Read as equivalent, a truncated agent would
delete a sixth of the analysis and finish green, because the critic rules what
it is handed and by_category omits a lane with no threats rather than
carrying a zero.
Retry, timeout and the per-job deadline are configured in config/resilience.toml
and attached to the adapter itself, so a retry is invisible to the graph and the
report's nodes array is unchanged by one. A per-request timeout turns a hang
into an error the retry can act on. Three attempts by default.
The retry loop sits above the provider seam described below, and the translator sits beneath it. That is where the loop's two bounds are expressible: one token bucket shared by the whole process, and jitter that spreads out lanes which failed at the same instant. Neither is something a single call can do for itself.
Concurrent analyses are independent. Each analyze() call runs one job in its
own ADK session, created fresh with a unique session id and seeded with only that
job's own input text; the graph's per-run data lives entirely in that session's
state, read back by the same session id. The engine, runner, and compiled
pipeline hold no per-run state — they carry only read-only configuration (the
node→model map, the loaded prompts and skills), so one engine is safe to share
across every call.
- Isolation is per session, not per caller. Two analyses submitted at the same time by the same caller still get separate sessions, so they can never read or overwrite each other's state.
- Within a single analysis, the lane agents run in parallel but each writes to its own category-keyed slot in the session, so the parallel branches don't clobber one another before the merge.
- Untrusted input stays contained. Because a job's text lives only in its own session (as fenced data — see The pipeline), a prompt-injection attempt in one submission cannot reach another running analysis.
Containment is not resistance, and the two are measured differently. Fencing
is structural and deterministic: every caller byte sits inside a marker sized to
its own content, so a submission cannot close the block it is in and continue in
instruction position. evals/adversarial/ carries a source built to try exactly
that, and CI asserts it fails.
What fencing does not do is stop a model reading ignore all previous instructions from inside a fence and deciding to obey it. That is semantic, it
is a property of a model and a prompt set rather than of this code, and it is
measured by scoring a report against what the injection asked for — deterministic
grading, no model judge. That half needs a live model and has never run; see
evals/adversarial/README.md for the bar and the residual risk.
The consequence of a failure is bounded by what a model here can reach: no
model in this service holds any tool or host authority. Every LLM node returns
structured text that a deterministic FunctionNode validates, so a model talked
into something is talked into producing a bad report, not into acting.
This guarantee assumes the intended concurrency model: async calls on a single
event loop. The default InMemorySessionService is an in-process store — safe
for cooperative async concurrency with distinct session ids, but not a
thread-safe store to share across OS threads. Scaling across processes keeps
analyses isolated (nothing is shared), but then each worker has its own in-memory
job and session state, so a job must be routed to the worker that holds it —
which is what the persistent backends below are for.
Every model reaches its provider through ADK's LiteLlm and, beneath that,
litellm — ADR 0015 records
why that is one substrate rather than a swappable adapter. Both run in the
service process, with the service's authority. LiteLLM holds the provider
credentials, opens the network connections, and is the only code between a
node's request and a provider's answer.
A provider seam sits between the graph and that translator:
analysis_service.provider defines one bounded request that goes out, one
result that comes back, and a failure as a value rather than as an exception
object. Everything the service asks of a provider crosses it.
The seam is not containment and does not claim to be. There is one implementation and it runs the translator in this process, under this process's credentials. What the seam buys is that the boundary is written down: a request carries no credential and nothing executable, so an implementation that ran elsewhere has a contract to meet rather than a dependency's type graph to reason about. Containment is the deployment work below, and this repository still ships none of it.
State the consequence plainly: a compromise of that dependency is a compromise of this application. A malicious release, or arbitrary code execution inside it, would reach the process's environment — including every provider credential this deployment declared — its filesystem, and its outbound network. Nothing in this repository contains that, because containment is deployment work and this repository ships no deployment packaging: no image, no container definition, no egress policy. An operator running this service in production owns that boundary. What it should look like is set out in #502.
What the repository does bound is the seam — the set of values this service hands the translator, and where each comes from:
| What crosses | Where it comes from |
|---|---|
| the model route | Vendor.route() — a registry entry, in code |
| the credential kwargs | the vendor's own table, read from declared ANALYSIS_* variables |
| the decoding params | sampling.toml plus an explicit env allowlist |
num_retries=0 |
a literal |
No provider endpoint crosses it at all. Nothing sets api_base, base_url
or api_version anywhere in the package, so a request goes where the vendor's
own client sends it. custom_llm_provider is set — the build-time capability
probe has to name a provider — and only ever from a Vendor.
That matters because prompts, submitted sources, corpus text and model output
all flow through this process. An adapter that accepted an address from any of
them would be an SSRF and endpoint-substitution path.
tests/test_translator_seam.py fails if one appears, if a decoding param could
express one, or if a new value starts crossing the seam.
No TLS setting crosses it either. LiteLLM resolves certificate verification
from a call kwarg, then from SSL_VERIFY in the process environment, then from
its own litellm.ssl_verify. The package sets none of the four kwargs that
reach that resolution, and neither sampling.toml nor the env allowlist can
express one. tests/test_translator_seam.py also drives LiteLLM's own
resolver, because the environment is a path no kwarg lint can see.
Only a declared credential authenticates a run. LiteLLM reads a key out of
the process environment on its own, so a machine that ever ran another tool
against a provider carries a key this deployment never declared.
Vendor.credential_kwargs reads the ANALYSIS_* variable instead and fails the
build closed when it is absent, so no adapter that could fall back is ever
built. tests/test_transport_conformance.py holds both halves against the
bytes: the declared key beats an ambient one, and an ambient one alone builds
nothing.
Dependency versions are pinned exactly (pyproject.toml) and hashed at install
(uv.lock), and the installed version of every distribution between a node and
its provider is inside each run's
execution identity — so a litellm bump moves
every fingerprint rather than silently reusing a blessing taken before it.
The pipeline is reached through interfaces, so backends are swappable and the whole graph runs offline against scripted models:
| Seam | Interface | Default | Status |
|---|---|---|---|
| Pipeline execution | PipelineRunner |
AdkPipelineRunner (real graph) / StubPipelineRunner (tests) |
Complete |
| Job persistence | JobStore |
InMemoryJobStore (memory) |
Backend selected by ANALYSIS_JOB_STORE via a fail-closed registry; only the non-durable memory backend ships — a durable one is a new registry entry |
| ADK sessions | BaseSessionService |
InMemorySessionService |
In-memory only; a session_service_uri backend is unwired |
The in-memory defaults are enough to get a report in process. Choosing a backend
is already wired for the JobStore (ANALYSIS_JOB_STORE, which stops startup on
an unset or unknown value rather than quietly falling back). Still out of scope
for the current work: a durable JobStore implementation, a session backend,
deployment packaging (container, Cloud Run), and observability. The interfaces
and the selection seam are in place for all of them.
| Module | Responsibility |
|---|---|
analysis_service.engine |
In-process Engine facade. |
analysis_service.api |
The /v1 FastAPI app. |
analysis_service.jobs |
Job lifecycle, JobStore, PipelineRunner seams. |
analysis_service.sources |
Source: the untrusted text a job is built from, the per-deployment bounds both entry points enforce, and the fenced render that is the only way caller bytes reach a model (OWASP LLM01). |
analysis_service.deployment |
One installation's config, resolved once: the files, the graph they configure, its runner and its certification gate. |
analysis_service.pipeline |
AdkPipelineRunner: one job's identity, input digest and certification around a Graph Run. |
analysis_service.execution |
Drives a built graph, stamps each node execution, and turns the finished run into its report or its rejection. Shared by the service and the eval harness. |
analysis_service.graph |
Topology and node functions. |
analysis_service.system_model |
Canonical model + validity helpers. |
analysis_service.analysis |
Deterministic traversal of a validated model: flows, reachability, paths, unknown controls. No security claims. |
analysis_service.candidates |
The rule table. Structural conditions an agent should investigate — leads, never findings, never evidence. |
analysis_service.domains |
Which skills/domains/ packs a model earns, decided from its own technology fields. |
analysis_service.coverage |
Per-lane accounting: what each agent was offered, and what its drafts cite. |
analysis_service.claims |
The neutral Claim, its grounds, the proposal and ruling wrappers, the marks, the severity model and the FrameworkAnalysis block. |
analysis_service.report |
The Report envelope: the job, the input, the node runs and the blocks. |
analysis_service.frameworks |
The framework-package contract, its registry and its deployment gate. |
analysis_service.validation |
The mechanical validity gate. |
analysis_service.fan_in |
One framework's lane batches merged into the drafts its critic reads: evidence resolution, the whole-set checks, and every mark the service records about a draft, behind one call. |
analysis_service.critic |
The mechanical checks around the critic's ruling — the ones no model should be asked to perform — and the view a critic reads. |
analysis_service.skills / .prompts / .markdown_loader |
Skill/prompt loading and composition. |
analysis_service.model_tiers / .sampling / .resilience |
Config loaders. |
analysis_service.vendors |
The vendor registry: each vendor's router prefix, credential mode, and model-name rules. |
analysis_service.binding |
Builds one adapter per tier from (vendor, model, sampling, resilience), and the NodeBinding the graph binds onto its LLM nodes. |
analysis_service.model_gate |
The startup check that asks the provider library whether a tier's parameters are actually supported. |
analysis_service.certification |
Compares a run's fingerprints against the deployment's blessed manifest. |
analysis_service.auth |
Bearer-token (OIDC JWT) verification. |
analysis_service.errors |
ConfigError, the base every fail-closed config loader raises — so a caller can handle "this deployment cannot run" without enumerating which knob was wrong. |