Skip to content

[api][java][python] Refactor ChatMessage and separate model invocation results - #1185

Open
wenjin272 wants to merge 3 commits into
apache:mainfrom
wenjin272:codex/chat-message-refactor
Open

wenjin272 wants to merge 3 commits into
apache:mainfrom
wenjin272:codex/chat-message-refactor

Conversation

@wenjin272

@wenjin272 wenjin272 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Linked issue: Closes #1056

Purpose of change

Outcome and runtime flow

Represent conversation content as ordered typed blocks in Java and Python. Chat models return ChatResult, containing an assistant ChatMessage, model/response identifiers, token usage, finish reason and result metadata. This separates conversation history from per-invocation output previously mixed into extra_args and untyped tool-call maps.

ChatRequestEvent -> chat action -> provider adapter -> ChatResult. The action reads usage and finish reason, stores only the assistant message in history, and dispatches typed tool calls. Tool execution returns ToolResponse; history records its model-facing content as ToolResultBlock under the same call ID. Structured output and routing information remain on ChatResponseEvent.

Key decisions and review order

The three commits group core contracts/execution, provider adaptations, and downstream consumers/bridges/fixtures. Review the first two for design and protocol behavior; apply the complete series for repository buildability.

Keep ToolResponse separate from ToolResultBlock: execution status and timing are distinct from model-facing conversation content. Keep ordinary mutable metadata/input maps without deep-freezing arbitrary values. Stateless provider conversions use local static utilities, with protocol field constants.

Behavioral Semantics

Interaction decisions

Input / condition Behavior
SYSTEM / USER / ASSISTANT SYSTEM accepts text; USER accepts text/media; ASSISTANT accepts text/media/reasoning/tool calls.
TOOL message Exactly one ToolResultBlock; its content accepts text/media. Each provider checks its supported media.
Assistant text plus reasoning/tool calls Block order survives serialization; text and tool-call properties project only matching top-level blocks.
Provider reasoning with continuation metadata Anthropic and Bedrock retain signed/redacted content; Gemini retains thought signatures; OpenAI Responses retains native reasoning items. Ollama stores reasoning but does not replay it.
Structured output plus chat history Parsed output is an event attribute; the original assistant message remains in history.
Java/Python boundary or restored event Convert the same typed message/result shape, including nested blocks and tool responses.

Contracts and failure behavior

  • Tool calls have explicit IDs, names and input maps. Duplicate call IDs within a message and invalid role/block combinations fail validation.
  • Tool results reference the same call ID; provider IDs are retained when available and adapters generate IDs when absent. Gemini-generated IDs are not echoed as native IDs.
  • Usage distinguishes unknown from zero. Metrics read typed usage; finish reasons remain strings, with existing canonical limit/filter mappings used by the execution guard.
  • Invalid model output, unsupported media and structured-output parsing errors follow validation/action failure paths; service errors retain the existing retry policy and exhausted calls emit failed response events.
  • Event Log sanitizes media payloads/URLs while retaining reasoning, tool input and metadata; transport/state serialization preserves full content. Mutable nested maps can be shared and must not be treated as deep snapshots.

Tests

Contract Coverage
Role validation, ordered blocks, result envelope, usage and finish reasons Java ChatMessageTest/ChatMessageSerializationTest; Python test_chat_message.py/test_chat_result.py
Original message retained with structured output; typed metrics and tool lifecycle Java ChatModelActionTest/ToolCallActionTest; corresponding Python action and token metric tests
Provider mapping and reasoning continuation Anthropic, Bedrock, Gemini, OpenAI and Ollama provider tests, including serialization-before-replay cases
Ollama user images survive the new model Java OllamaMultimodalTest and Python test_ollama_multimodal.py
Failed events/tools and cross-language wire shape ChatResponseEventTest, CrossLanguageEventSnapshotTest and Python snapshot tests
Bridge/state/log consumers JavaResourceAdapterTest, ActionStateSerdeTest, FileEventLoggerTest and Python runtime conversion tests

After rebasing onto main (99103da67): full Java reactor compilation/install succeeded; Java non-E2E tests passed (2627 passed, 44 skipped), excluding FlussActionStateStoreIntegrationTest because its embedded service previously blocked during initialization. Python non-integration tests: 1707 passed, 14 skipped. Anthropic parsing tests also passed against CI SDK version 1.11.0 (96 tests); the mocked Tongyi call-ID round trip and both MCP input-action paths passed (3 tests). Prompt-driven example action regressions cover three agents. Mock chat MiniCluster E2E: 2 passed. Spotless and Ruff checks for changed files passed.

Not verified: live model services, exhaustive parity across providers, and upgrade/recovery from the old message wire format. This PR intentionally breaks that format.

API

Breaking Java/Python change: chat methods return ChatResult; replace extra_args/extraArgs and untyped tool_calls with typed blocks and scoped metadata. History still consists of ChatMessage; text convenience factories remain available. Use event structured-output accessors for parsed results. Existing serialized messages/events/checkpoints require migration; no compatibility layer is provided.

Documentation

  • doc-needed
  • doc-not-needed
  • doc-included

Was this patch authored or co-authored using generative AI tooling?

  • Yes
  • No

Generated-by: Codex 0.153.4 (GPT-6)

Introduce ordered content blocks and ChatResult wrapping an assistant ChatMessage
in Java and Python. Keep model result metadata separate from conversation
content, structured output on ChatResponseEvent, and one tool call ID through
tool execution and history. Update the execution flow and direct contract tests.

This is the core design review commit. No backward compatibility with the
previous chat API is retained.

The series is grouped by review responsibility; apply all three commits
together for full-repository buildability.

Refs apache#1056

Generated-by: Codex 0.153.4 (GPT-6)
Co-authored-by: Codex <codex@openai.com>
@github-actions github-actions Bot added doc-included Your PR already contains the necessary documentation updates. fixVersion/0.4.0 priority/major Default priority of the PR or issue. labels Oct 1, 2026
wenjin272 and others added 2 commits October 1, 2026 16:10
Migrate Ollama, OpenAI, Azure, Anthropic, Gemini, Bedrock, Watsonx, Tongyi and
related adapters to the new chat contracts. Preserve provider-native reasoning
signatures and encrypted content for replay, and map typed tool calls/results.

Keep stateless SDK conversions in provider utility methods, with local protocol
constants. Include provider parsing, request and reasoning replay tests.

The series is grouped by review responsibility; apply all three commits
together for full-repository buildability.

Refs apache#1056

Generated-by: Codex 0.153.4 (GPT-6)
Co-authored-by: Codex <codex@openai.com>
Update Java/Python resource bridges, ReAct and routing consumers, prompts,
tools, memory and event logging to the new chat contracts. Migrate downstream
tests, cross-language event snapshots, examples, YAML consumers and docs.

These changes follow the core and provider contracts in the preceding commits.

The series is grouped by review responsibility; apply all three commits
together for full-repository buildability.

Refs apache#1056

Generated-by: Codex 0.153.4 (GPT-6)
Co-authored-by: Codex <codex@openai.com>
@wenjin272
wenjin272 force-pushed the codex/chat-message-refactor branch from e0d4f8a to 4f26d55 Compare October 1, 2026 08:10
@github-actions github-actions Bot added doc-included Your PR already contains the necessary documentation updates. and removed doc-included Your PR already contains the necessary documentation updates. labels Oct 1, 2026
@wenjin272
wenjin272 marked this pull request as ready for review October 1, 2026 12:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

doc-included Your PR already contains the necessary documentation updates. fixVersion/0.4.0 priority/major Default priority of the PR or issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Tech Debt][API][ChatMessage] Review ChatMessage responsibilities and data model

1 participant