Skip to content

[integration][observability] Add OpenTelemetry GenAI exporter for Agent Traces (Phase 1) - #1160

Open
Zhuoxi2000 wants to merge 7 commits into
apache:mainfrom
Zhuoxi2000:otel-exporter-phase1-main
Open

Zhuoxi2000 wants to merge 7 commits into
apache:mainfrom
Zhuoxi2000:otel-exporter-phase1-main

Conversation

@Zhuoxi2000

@Zhuoxi2000 Zhuoxi2000 commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Linked issue: #970

Purpose of change

An Agent Trace Event Log (event-log.trace.enabled: true) from a Java or Python agent can now be exported to any OTLP backend as OpenTelemetry GenAI traces:

EventLogOTelExporter [--endpoint URL] [--protocol grpc|http/protobuf] [--service-name NAME] <file-or-dir>...

This is a new optional module, flink-agents-integrations-observability-otel, and is not bundled into dist. The goal is to make the execution hierarchy already recorded in Agent Trace visible in standard tracing tools without adding runtime overhead.

Runtime flow

  1. exportFiles expands directories to their events-*.log files and streams JSON records into TraceRecord (unknown fields are ignored).
  2. AgentTraceSpans.assemble keeps lifecycle records with executionId, inputRunId, and a parseable timestamp, grouped by execution and input run.
  3. Each input run becomes a trace with a synthesized invoke_agent root. Each execution becomes a span under its parentExecutionId, or under the root if no parent is present, based on entityType.
  4. Spans are passed to the OTLP SpanExporter in batches of at most 512 (--batch-size). Diagnostics are logged and returned in ExportSummary.

Key decisions

  • Batch conversion from the Event Log, rather than an in-job exporter: no runtime overhead, and the same input format works for both Java and Python. Continuous consumption is out of scope.
  • Deterministic IDs (SHA-256 of inputRunId / executionId): re-exporting the same records produces the same IDs; delivery remains at-least-once.
  • Incomplete executions use status UNSET, not ERROR: a missing terminal record may come from a crash, dropped write, or recovery, and the log alone cannot distinguish them.
  • gen_ai.request.model comes from entityMetadata.model (entityName is the ChatModel resource). gen_ai.provider.name is left unset because the record does not contain it.
  • execute_tool spans remain INTERNAL regardless of the tool transport.
  • businessKey is the keyed-stream key, a conversation only in chat pipelines, so it is always flink_agents.business_key and becomes gen_ai.conversation.id only when enabled.

Behavioral Semantics

Interaction decisions

started is used as the start record when present, otherwise created; created is never treated as a terminal record.

Records for one execution Span Status Diagnostic
start + finished start → terminal UNSET none
start + failed start → terminal ERROR + error.type none
start only (incl. created only) zero-length, incomplete=true UNSET INCOMPLETE_EXECUTION
terminal only, not reused zero-length, incomplete=true from terminal MISSING_START
reused only zero-length, status=reused UNSET none

Behavioral contracts

  1. One trace per inputRunId, rooted at invoke_agent {agentName} (INTERNAL), with gen_ai.agent.name and, when present, flink_agents.business_key; gen_ai.conversation.id = businessKey only with --business-key-as-conversation-id.
  2. An execution span's parent is its parentExecutionId span; otherwise it is attached to the run root.
  3. llm → chat {model} (CLIENT): gen_ai.request.model from entityMetadata.model, and gen_ai.usage.*_tokens from promptTokens / completionTokens when the terminal record carries them (the built-in reporters do not yet). Without a model, the span is named chat and no request model is set.
  4. tool → execute_tool {name} (INTERNAL): gen_ai.tool.name; gen_ai.tool.call.id = externalId, otherwise toolCallId; gen_ai.tool.type = function, extension (remote_function, mcp), or unset (model_built_in). The raw value is kept in flink_agents.tool.type. gen_ai.agent.name is set when the tool's records carry agentName.
  5. action → action {name}; parser → parse {name} with gen_ai.operation.name=parse; both are INTERNAL.
  6. A failed execution has ERROR status with the recorded error message, and error.type = the recorded error type, falling back to problemCategory.
  7. Execution spans carry flink_agents.{input_run_id, execution_id, entity_type, entity_name, execution.status}, plus flink_agents.business_key when present.
  8. Exporting the same records again, in any order, produces the same spans and IDs.
  9. Non-lifecycle records, and records missing executionId, inputRunId, or timestamp, are skipped; a record whose timestamp does not parse is skipped with a MALFORMED_RECORD diagnostic.

Failure behavior

  • An undecodable record produces a MALFORMED_RECORD diagnostic naming the file. Reading continues, but a broken stream or 1,000 malformed records stop that file.
  • A batch that fails or takes over 30 s throws ExportFailedException with the number of spans the earlier batches delivered; there is no retry, and re-running is safe.
  • Unsupported --protocol throws IllegalArgumentException at build time.
  • CLI: a flag without a value, a non-positive --batch-size, or no input, prints usage and exits 2; a missing input path throws NoSuchFileException.

Tests

Contract Tests
1 AgentTraceSpansTest.testRunBecomesSingleTraceWithRootSpan, testBusinessKeyAsConversationIdIsOptIn
2 testParenting, testDiagnosticsForIncompleteAndMissingStart
3 testGenAiAttributes, testChatWithoutModel
4 testGenAiAttributes, testToolTypeMapping, testToolSpanCarriesAgentName
5 testParserMapping, testParenting
6 testGenAiAttributes, testErrorTypeFallsBackToProblemCategory, testTruncatedErrorMessage
7 testCorrelationAttributes
8 testDeterministicIds, testRecordOrderDoesNotMatter
9 testNonLifecycleRecordsIgnored, testInvalidTimestampIsSkippedWithDiagnostic, EventLogOTelExporterTest.testExportJsonlFile
Interaction table testCreatedStartedTerminal, testCreatedOnly, testCreatedThenFailedWithoutStart, testIncompleteExecution, testDiagnosticsForIncompleteAndMissingStart, testReusedExecution
Malformed input, protocol, directories testMalformedRecordDiagnostic, testUnsupportedProtocol, testDirectoryDiscovery
Batched export testExportsInBoundedBatches, testFailedBatchReportsPartialProgress, testRejectsNonPositiveBatchSize

Span mapping is covered for each entity type and lifecycle case; input handling uses real files.

Not verified:

  • A real OTLP endpoint: tests use an in-memory SpanExporter, so transport, the receiver's size limit, and the 30 s per-batch timeout are not covered (a failing batch is).
  • The 1,000-malformed cap, broken-stream abort, and CLI exit codes (manual only).
  • A Python-produced log (same format, no fixture).
  • Very large logs: one invocation currently holds all records in memory.
Implementation invariants and supporting evidence
  • Test fixtures match the shape written by EventLogRecordJsonSerializer: entityMetadata is an object, status / problemCategory are top-level fields, and errorType / errorMessage are inside eventAttributes. A failed-tool record and an LLM record produced by the real serializer (EventLogRecord + ExecutionLifecycleEvents) assemble into the same spans as the fixtures.
  • OTelIds: trace id = first 16 bytes of SHA-256(inputRunId), span id = first 8 bytes of SHA-256(executionId); the run root hashes inputRunId with a domain suffix so it cannot collide with an execution span id; an all-zero id is adjusted to remain valid.
  • Attribute keys are pinned as literals rather than taken from the incubating semconv artifact, so the exported schema changes only with a Flink Agents release.
  • OpenTelemetry artifacts are aligned through opentelemetry-bom (all resolve to 1.51.0).
  • Verified locally, rebased onto main after [api] Introduce multimodal content blocks #1060 and [hotfix][ci] Pull MinIO for Milvus tests from Bitnami's legacy image #1162: module tests 28/28, spotless:check, and verify for the module and its upstream modules.

API

No existing API, record format, or runtime behavior changes. dist is unchanged, so nothing is added to a job's classpath.

New public entry points: EventLogOTelExporter (builder and main), AgentTraceSpans, and TraceRecord.

Existing trace-enabled Event Logs are consumed as-is. The GenAI conventions are still at development stability, so the exported attribute set is pinned per release.

Documentation

  • doc-needed
  • doc-not-needed
  • doc-included

Was this patch authored or co-authored using generative AI tooling?

  • Yes
  • No

If yes, include a Generated-by: <tool name and version> (<model name and version>) line, for example Generated-by: Claude Code 2.1.226 (Claude Opus 4.6), in the commit message so it reaches Git history. Repeat the same line here for reviewer visibility. See the ASF generative tooling guidance.

Generated-by: Claude Code 2.1.259 (Claude Fable 5, Claude Opus 5.5)

@github-actions github-actions Bot added doc-label-missing The Bot applies this label either because none or multiple labels were provided. fixVersion/0.4.0 priority/major Default priority of the PR or issue. labels Sep 25, 2026
@Zhuoxi2000 Zhuoxi2000 changed the title Otel exporter phase1 main [integration][observability] Add OpenTelemetry GenAI exporter for Agent Traces (Phase 1) Sep 25, 2026
@github-actions github-actions Bot added doc-label-missing The Bot applies this label either because none or multiple labels were provided. and removed doc-label-missing The Bot applies this label either because none or multiple labels were provided. labels Sep 25, 2026

@sangkyoonnam sangkyoonnam left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking the suggestions in. I pulled the branch, built the module and ran its tests (18 + 4 pass), and checked the span shapes against semantic-conventions-genai at e57c543b: the root invoke_agent INTERNAL span, chat as CLIENT, execute_tool as INTERNAL with gen_ai.tool.call.id and gen_ai.tool.type, and error.type carrying the recorded error type (which the built-in reporters derive from the root-cause class) all match the internal-agent and tool span tables. gen_ai.provider.name is Required on the chat span and stays unset because the record doesn't carry it; the docs say so, and that's the one gap I'd keep stated rather than fixed here. None of this blocks. Three things worth a look below (the usage mapping, request size, and the conversation id), plus two nits.

private ExportSummary export(List<TraceRecord> records, List<ConverterDiagnostic> diagnostics) {
List<SpanData> spans = assembler.assemble(records, diagnostics);
if (!spans.isEmpty()) {
CompletableResultCode result = spanExporter.export(spans);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This exports every assembled span in one call. I checked the pinned OpenTelemetry 1.51.0 gRPC and HTTP exporters: neither splits the collection, each marshals it into a single request. The OTel Collector's default gRPC receive limit is 4 MiB (grpc-go's default unless max_recv_msg_size_mib is set), so an oversized export is rejected as a whole. Could exports use bounded batches and report partial progress if a later batch fails? The SDK's 512-span BatchSpanProcessor default is one reference point, though a span count alone doesn't guarantee a request stays under the byte limit. If batching is added, should the 30 s wait at L170 apply per batch or to the whole run?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 36e9717: 512-span batches, 30 s per batch, and failures report delivered spans.

}

private static long epochNanos(String isoTimestamp) {
Instant instant = Instant.parse(isoTimestamp);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: an otherwise valid lifecycle record with an invalid timestamp throws DateTimeParseException through RunAccumulator.accept and aborts the whole export, while a JSON decoding failure becomes a MALFORMED_RECORD diagnostic. EventContext writes Instant.now().toString() on the built-in path, but the other constructor and the timestamped report methods take the string as given. Could the timestamp be validated before the record enters either accumulator, with a diagnostic and skip instead? A test with one invalid timestamp followed by a valid record would cover it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done: an invalid timestamp now yields a MALFORMED_RECORD diagnostic and skips that record.

if (eventAttributes == null) {
return;
}
Object prompt = eventAttributes.get("promptTokens");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

executionFinished() produces no usage attributes: the Java and Python chat models keep token counts in the response's extraArgs and in metric counters (BaseChatModelSetup.java:181), and nothing writes promptTokens / completionTokens into a terminal record's eventAttributes. The branches here are exercised only by the test fixtures. The docs already defer richer usage metrics, so this is a clarification rather than a request to grow the PR: could the mapping docs say that logs from the current built-in reporters don't supply the terminal attributes this converter needs for gen_ai.usage.*, so "when recorded" doesn't read as "usually"?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch; the docs now say built-in reporters don't write these token attributes yet.

// the framework running the tool, whatever transport the tool itself uses.
kind = SpanKind.INTERNAL;
attributes.put(GEN_AI_OPERATION_NAME, "execute_tool");
attributes.put(GEN_AI_TOOL_NAME, entityName);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: the execute_tool table lists gen_ai.agent.name as conditionally required when applicable, and TraceRecord exposes agentName, but it's only set on the root span. Could tool spans carry it when the record has one? That keeps a tool span self-describing when a backend shows it without its root.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done: tool spans now carry gen_ai.agent.name whenever their records have one.

}
attributes.put(FA_ENTITY_NAME, entityName);
if (any.getBusinessKey() != null) {
attributes.put(GEN_AI_CONVERSATION_ID, any.getBusinessKey());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

businessKey is the keyed-stream key rendered as text. It is a conversation in a chat pipeline and an order or device id elsewhere, while the convention asks for a conversation identifier the library actually has. Would you carry it as flink_agents.business_key unconditionally and map it to gen_ai.conversation.id behind an option? Same for the root mapping at L163.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed; it's always flink_agents.business_key now, and conversation id is opt-in.

@Zhuoxi2000
Zhuoxi2000 force-pushed the otel-exporter-phase1-main branch from 224b58d to 36e9717 Compare September 28, 2026 03:55
@github-actions github-actions Bot added doc-label-missing The Bot applies this label either because none or multiple labels were provided. and removed doc-label-missing The Bot applies this label either because none or multiple labels were provided. labels Sep 28, 2026
@sangkyoonnam

Copy link
Copy Markdown
Contributor

All five are addressed in 36e9717. I rebuilt the module at that commit and its 28 tests pass. Exports now go out in bounded batches and stop at the first failed or timed-out one, reporting the spans from the batches that succeeded before it, which answers my batching point. Nothing more from my side.

@Zhuoxi2000

Copy link
Copy Markdown
Contributor Author

Thanks for the careful review; @wenjin272, could you take a look when you have time?

@wenjin272

Copy link
Copy Markdown
Contributor

Thanks for the PR. Since @joeyutong designed most of Flink Agents’ observability, I think he would be best placed to review this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

doc-label-missing The Bot applies this label either because none or multiple labels were provided. fixVersion/0.4.0 priority/major Default priority of the PR or issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants