Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
1787c1e
fix: enforce tool approval regardless of call order in mixed batches
CorgiBoyG Sep 4, 2026
3ac0f83
fix: preserve unknown-call termination after approval resume
CorgiBoyG Sep 4, 2026
13fb5e9
fix: keep per-call classification inside approval-pausing batches
CorgiBoyG Sep 5, 2026
43da2d4
docs: correct batch-classification precedence comment to match fail-c…
CorgiBoyG Sep 5, 2026
4b6d453
refactor: route both declaration-only pause surfaces through a shared…
CorgiBoyG Sep 5, 2026
253140a
fix: preserve mixed pause order
CorgiBoyG Sep 6, 2026
8b1a889
test: satisfy mixed-batch typing checks
CorgiBoyG Sep 10, 2026
1861068
fix: preserve host-owned calls on approval replay
CorgiBoyG Sep 10, 2026
1b7ad00
fix: align host pause accounting after replay
CorgiBoyG Sep 11, 2026
c2ce1b1
fix: persist mixed pause batches atomically
CorgiBoyG Sep 11, 2026
1ba8738
fix: close mixed pause replay gaps
CorgiBoyG Sep 11, 2026
f1a5a75
fix: make stateless pause matching order-independent
CorgiBoyG Sep 11, 2026
e84c464
fix: harden mixed pause response correlation
CorgiBoyG Sep 11, 2026
8bcf2a4
fix: reject ambiguous mixed approval replays
CorgiBoyG Sep 11, 2026
50246f5
fix: preserve stateless mixed response order
CorgiBoyG Sep 11, 2026
d691822
fix: reject conflicting stateless approval replays
CorgiBoyG Sep 11, 2026
80d0d4d
fix: prioritize explicit approval for host tools
CorgiBoyG Sep 15, 2026
7ec256c
fix: enforce approval for additional host tools
CorgiBoyG Sep 15, 2026
83df045
fix: preserve budget for replacement approvals
CorgiBoyG Sep 15, 2026
c3bab94
fix: avoid double counting A2UI handoffs
CorgiBoyG Sep 15, 2026
6795b50
fix: preserve mixed batch outbox on invalidation
CorgiBoyG Sep 16, 2026
a3031c8
fix: preserve outbox after host pause cancellation
CorgiBoyG Sep 16, 2026
c42e41f
fix: avoid replaying published outbox updates
CorgiBoyG Sep 16, 2026
e51f4cc
fix: persist host-only mixed batch outbox
CorgiBoyG Sep 16, 2026
e896c37
Merge branch 'main' into fix/mixed-batch-classification-precedence
eavanvalkenburg Sep 16, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -54,8 +54,12 @@ usability is the portable default while invalidation is provider-specific protoc
`FunctionInvocationLayer` remains independent of `finish_reason`: a newly completed actionable call proceeds when
argument preparation and schema validation succeed. A provider that knows partial response output was invalidated
raises `ResponseInvalidatedException`; any local function calls from that response must not execute. The layer
abandons that current iteration, clears request budget state, restores the last valid continuation, avoids successful
response persistence and local function side effects, and re-raises the same exception.
abandons that current iteration, ordinarily clears request budget state, restores the last valid continuation, avoids
successful response persistence and local function side effects from the invalid response, and re-raises the same exception.
If the invalidated call was instead delivering results from an already completed mixed approval/Host batch, the layer
retains that serializable provider outbox and its charged budget. A retry replays the stored Host and local results
without recovering approval authority or executing the local side effect again; only a successful provider response
clears the outbox.

Anthropic applies the signal only to local actionable `tool_use` blocks. A valid stream has closed local blocks, a
terminal `stop_reason` of `tool_use`, and `message_stop`. Non-tool terminal reasons, an open block at `message_stop`,
Expand Down
26 changes: 20 additions & 6 deletions docs/specs/004-python-function-calling-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,7 @@ Code-reading landmarks:
- `_get_response_with_function_invocation(...)` owns non-streaming aggregation.
- `_stream_response_with_function_invocation(...)` owns streamed emission/finalization.
- `_resolve_approval_responses(...)` handles only inbound approval decisions.
- `_stage_pending_pause_batch_responses(...)` durably assembles mixed approval and Host-owned responses.
- `_process_model_function_calls(...)` handles only calls from a completed model response.
- `_try_execute_function_calls(...)` decides approval/declaration/execution behavior for a batch.
- `_replace_approval_contents_with_results(...)` is the occurrence-aware approval transcript normalizer.
Expand Down Expand Up @@ -187,6 +188,15 @@ sequenceDiagram
The terminal result is caller-visible in both modes. The private normalized message copy is model-visible. The
original caller input and earlier response remain unchanged.

When one model batch mixes local approval requests with declaration-only or `additional_tools` calls, all of those
pauses form one ordered barrier. The session stores immutable request snapshots plus any partial responses. Until
every occurrence has a response, later calls return without executing an approved tool or invoking the model,
including runs that contain only unrelated input. Once complete, the layer reconstructs the response batch in the
model's original call order. Correlation uses each content occurrence id rather than assuming provider `call_id`
values are unique. Workflow `request_info` responses preserve that occurrence id when they are handed back to the
agent. The stored representation must survive an `AgentSession.to_dict()` / `AgentSession.from_dict()` round trip,
and streaming and non-streaming paths must have equivalent behavior.

### Reasoning-bound function-call groups

Some hosted services bind reasoning content or an opaque reasoning signature to the function call that follows it.
Expand Down Expand Up @@ -355,11 +365,14 @@ that manually replay messages own the equivalent rule: do not resend an approval
report another reason, including `length`, while still returning a complete and service-usable call.
- A provider adapter that knows its protocol invalidated or cancelled partial response output raises
`ResponseInvalidatedException`; any local function calls already exposed from that response must not execute. The
function invocation layer then abandons only that current model/tool iteration, clears its request budget state,
restores the last valid service continuation, and re-raises
the same exception. It emits no new approval request or function result, invokes no function middleware or tool
body, makes no second model request or dangling-call settlement request, and does not success-persist that service
response. Approval decisions resolved before the invalidated provider call retain their existing semantics.
function invocation layer then abandons only that current model/tool iteration, restores the last valid service
continuation, and re-raises the same exception. It emits no new approval request or function result, invokes no
function middleware or tool body for calls from the invalid response, makes no second model request or dangling-call
settlement request, and does not success-persist that service response. Ordinarily it clears the request budget.
When the failed request was delivering a completed mixed approval/Host batch, however, the serializable pending
batch, approval authority, ordered Host and local results, and already-charged budget remain in a provider outbox.
A later run replays that outbox without executing the approved function again, and clears it only after a provider
response succeeds.
Streaming providers may yield local function-call deltas before discovering invalidation; direct stream consumers
must treat the exception as an explicit instruction to discard those calls and never execute them. Caller
cancellation remains cancellation rather than becoming this provider signal.
Expand Down Expand Up @@ -598,6 +611,7 @@ that manually replay messages own the equivalent rule: do not resend an approval

| Scenario | Required invariant | Primary regression test |
|---|---|---|
| Mixed approval and Host pause batch | Approval decisions and declaration-only or `additional_tools` results from one model batch form one ordered barrier. Partial replies persist in serializable session state without executing a tool or calling the model; after every occurrence is answered, the batch is submitted in original model order. A provider-invalidated submission retains a serializable two-phase outbox, approval authority, results, continuation, and charged budget for execution-free retry, including at `max_function_calls=1`; provider success clears it. Correlation is occurrence-aware when provider `call_id` values repeat, with equivalent streaming, non-streaming, and workflow behavior. | `packages/core/tests/core/test_function_invocation_logic.py::test_mixed_approval_host_batch_stages_partial_responses_in_original_order`, `test_completed_mixed_batch_replays_persisted_outbox_after_provider_invalidation`, `packages/core/tests/workflow/test_agent_executor_tool_calls.py::test_agent_executor_mixed_pause_batch_waits_and_preserves_occurrence_order` |
| Safe and approval-required calls in one batch | Hidden safe calls replay only with the matching visible approval. | `packages/core/tests/core/test_harness_tool_approval.py::test_mixed_batch_hides_already_approved_request_until_approval_replay` |
| Restored approval state | Serialized `ToolApprovalState` restores mixed-batch behavior. | `test_mixed_batch_accepts_restored_tool_approval_state` |
| Unrelated turn before approval | Hidden calls do not execute on an unrelated turn. | `test_hidden_mixed_batch_requests_do_not_replay_on_unrelated_turn` |
Expand Down Expand Up @@ -638,7 +652,7 @@ that manually replay messages own the equivalent rule: do not resend an approval
| Middleware failure batch cancellation | A fatal signal fails the whole parallel batch: in-flight sibling tool invocations are cancelled and awaited before the failure propagates. Cancellation is cooperative — an async sibling stops at its next suspension point; a synchronous tool body already executing in a worker thread cannot be interrupted and may complete its side effects, but its result is discarded and never reaches the transcript, the model, or history, and failure propagation is not delayed behind it. | `TestMiddlewareFailure::test_failure_cancels_concurrent_sibling_tool`, `test_failure_with_sync_sibling_discards_late_result` |
| Middleware failure on a service-managed conversation | The continuation state is already persisted when the batch fails, so before propagating, the loop settles the hosted thread: one error `function_result` per dangling call, sent with `tool_choice="none"` in one extra request; the persisted continuation advances to the settlement response (required for response-ID continuations, a no-op for conversation-object ids) and the settlement response is otherwise discarded; a settlement failure never masks the abort. Without a service-managed conversation no extra request is made. | `TestMiddlewareFailure::test_failure_settles_dangling_calls_on_service_conversation`, `test_failure_settles_service_conversation_streaming`, `test_failure_settlement_advances_response_id_continuation`, `test_failure_without_service_conversation_makes_no_settlement_request` |
| Middleware failure during approved-tool replay | A fatal abort while the approval-resolution phase replays an approved tool escapes loudly (never absorbed into a rejection result), the tool's original — already service-persisted — call is settled the same way, and the continuation advances; both response modes. | `TestMiddlewareFailure::test_failure_during_approved_replay_settles_and_escapes`, `test_failure_during_approved_replay_streaming` |
| Provider-invalidated partial response output | `ResponseInvalidatedException` propagates unchanged in both response modes after clearing request budget state and restoring the last valid continuation. Partial streamed call deltas may remain visible, but no new approval, function middleware, tool body, result, follow-up model call, or settlement occurs. Anthropic requires every local call block to close, terminal `stop_reason="tool_use"`, and `message_stop`; non-tool terminal reasons, open blocks, missing `message_stop`, and non-cancellation stream errors after local call start invalidate the call. Hosted/server-only calls and caller cancellation remain unaffected. | `packages/core/tests/core/test_function_invocation_logic.py::test_response_invalidation_short_circuits_current_iteration`, `test_invalidated_final_no_tool_response_preserves_prior_continuation`, `packages/anthropic/tests/test_anthropic_client.py::test_non_streaming_local_tool_call_with_invalidating_stop_reason_raises`, `test_valid_streaming_local_tool_call_executes_and_continues`, `test_streaming_local_tool_call_invalid_terminal_sequences_raise`, `test_streaming_provider_error_after_local_call_is_wrapped_as_invalidation`, `test_streaming_cancellation_after_local_call_is_not_wrapped`, `test_streaming_hosted_tool_pause_turn_is_not_invalidated` |
| Provider-invalidated partial response output | `ResponseInvalidatedException` propagates unchanged in both response modes after restoring the last valid continuation. It clears ordinary request budget state, but preserves a completed mixed-batch provider outbox and its charged budget for execution-free retry. Partial streamed call deltas may remain visible, but no new approval, function middleware, tool body, result, follow-up model call, or settlement occurs for calls from the invalid response. Anthropic requires every local call block to close, terminal `stop_reason="tool_use"`, and `message_stop`; non-tool terminal reasons, open blocks, missing `message_stop`, and non-cancellation stream errors after local call start invalidate the call. Hosted/server-only calls and caller cancellation remain unaffected. | `packages/core/tests/core/test_function_invocation_logic.py::test_response_invalidation_short_circuits_current_iteration`, `test_invalidated_final_no_tool_response_preserves_prior_continuation`, `test_completed_mixed_batch_replays_persisted_outbox_after_provider_invalidation`, `packages/anthropic/tests/test_anthropic_client.py::test_non_streaming_local_tool_call_with_invalidating_stop_reason_raises`, `test_valid_streaming_local_tool_call_executes_and_continues`, `test_streaming_local_tool_call_invalid_terminal_sequences_raise`, `test_streaming_provider_error_after_local_call_is_wrapped_as_invalidation`, `test_streaming_cancellation_after_local_call_is_not_wrapped`, `test_streaming_hosted_tool_pause_turn_is_not_invalidated` |
| Maximum iterations | No orphan calls; a final no-tool response or deterministic fallback is returned. | `test_max_iterations_limit`, `test_max_iterations_no_orphaned_function_calls`, `test_max_iterations_makes_final_toolchoice_none_call`, `test_max_iterations_blank_final_fallback_synthesizes_message`, streaming equivalents |
| Maximum function calls | Parallel overshoot is bounded after the batch; every executed result group counts even without a `function_result`; blank final responses get fallback content. | `test_max_function_calls_limits_parallel_invocations`, `test_max_function_calls_single_calls_per_iteration`, `test_user_input_request_multiple_contents_propagate`, `test_approval_resume_user_input_counts_toward_function_call_budget`, `test_max_function_calls_blank_final_fallback_synthesizes_message`, streaming equivalent |
| Provider tool content after an active limit | Locally actionable calls and local approval requests returned despite `tool_choice="none"` are removed in both response modes. Provider-executed informational call/result pairs, hosted approval requests, and metadata-only streaming updates remain visible; fallback text never replaces retained transcript content. | `test_function_invocation_limit_drops_unexecutable_tool_content`, `test_streaming_function_invocation_limit_drops_unexecutable_tool_content`, `test_streaming_function_invocation_limit_preserves_metadata_after_tool_content_is_dropped`, `test_function_invocation_limit_preserves_provider_executed_tool_pair`, `test_streaming_function_invocation_limit_preserves_provider_executed_tool_pair`, `test_function_invocation_limit_appends_fallback_after_provider_executed_tool_pair`, `test_streaming_function_invocation_limit_appends_fallback_after_provider_executed_tool_pair`, `test_function_invocation_limit_preserves_hosted_approval_request`, `test_streaming_function_invocation_limit_preserves_hosted_approval_request` |
Expand Down
16 changes: 15 additions & 1 deletion python/packages/ag-ui/agent_framework_ag_ui/_a2ui/_agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -179,6 +179,15 @@ def record_adapter_calls(self, count: int) -> None:
raise ValueError("Adapter function-call count cannot be negative.")
self._state["total_function_calls"] = self.calls_used + count

def offset_core_handoffs(self, count: int) -> None:
"""Remove core's charge for private handoffs before charging actual renders."""
if count < 0:
raise ValueError("Core handoff count cannot be negative.")
adjusted_count = self.calls_used - count
if adjusted_count < 0:
raise RuntimeError("Core handoff count exceeds the shared function-call count.")
self._state["total_function_calls"] = adjusted_count


def _generate_tool_schema() -> dict[str, Any]:
"""JSON schema for the planner-facing generate_a2ui tool (string args)."""
Expand Down Expand Up @@ -720,7 +729,7 @@ async def _execute_server_tools(
invocation_kwargs = run_kwargs.get("function_invocation_kwargs")
custom_args = dict(cast(Mapping[str, Any], invocation_kwargs)) if invocation_kwargs is not None else {}
try:
groups, should_terminate = await _try_execute_function_call_groups(
groups, should_terminate, _ = await _try_execute_function_call_groups(
custom_args=custom_args,
function_calls=server_calls,
tools=tools,
Expand Down Expand Up @@ -955,6 +964,11 @@ async def _run_streaming(
core_result_ids = {result.call_id for result in core_results}
# Completed handoffs carry the invocation arguments after function middleware.
generate_calls: list[Content] = list(handoff_calls.values()) if core_execution else []
if core_execution:
# Core correctly records the private handoff as executed. It is only an
# adapter control seam, though, so replace that charge with one for each
# render the adapter actually starts below.
execution_budget.offset_core_handoffs(len(generate_calls))
server_calls: list[Content] = []
client_calls: list[Content] = []
for cid in call_order:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1950,7 +1950,7 @@ async def forward_hosted_decision(approval: Content = approval) -> list[Content]

async def execute_local_call(approval: Content = approval, call_id: str = call_id) -> list[Content]:
try:
result_groups, _ = await _try_execute_function_call_groups(
result_groups, _, _ = await _try_execute_function_call_groups(
custom_args=tool_kwargs,
function_calls=[approval],
tools=tools,
Expand Down
32 changes: 26 additions & 6 deletions python/packages/ag-ui/tests/ag_ui/test_a2ui.py
Original file line number Diff line number Diff line change
Expand Up @@ -1092,19 +1092,34 @@ def search(query: str, context: FunctionInvocationContext) -> str:


@pytest.mark.parametrize(
("max_calls", "expected_actions", "expected_surfaces"),
[(1, ["search"], 0), (3, ["search", "search"], 1)],
("max_calls", "expected_actions", "expected_surfaces", "expected_tool_choices", "expected_execution_trace"),
[
(1, ["search"], 0, ["auto"], ["search", "generate_a2ui"]),
(3, ["search", "search"], 1, ["auto", "auto"], ["search", "generate_a2ui"] * 2),
],
)
async def test_core_a2ui_shares_call_budget_with_surface_generation(
streaming_chat_client_stub, max_calls: int, expected_actions: list[str], expected_surfaces: int
streaming_chat_client_stub,
max_calls: int,
expected_actions: list[str],
expected_surfaces: int,
expected_tool_choices: list[str | None],
expected_execution_trace: list[str],
) -> None:
from collections.abc import AsyncIterator
from collections.abc import AsyncIterator, Awaitable, Callable

from agent_framework import Agent, FunctionTool
from agent_framework import Agent, FunctionInvocationContext, FunctionMiddleware, FunctionTool

executed: list[str] = []
execution_trace: list[str] = []
tool_choices: list[str | None] = []
round_number = 0

class RecordExecution(FunctionMiddleware):
async def process(self, context: FunctionInvocationContext, call_next: Callable[[], Awaitable[None]]) -> None:
execution_trace.append(context.function.name)
await call_next()

def search() -> str:
executed.append("search")
return "Search completed"
Expand All @@ -1113,7 +1128,9 @@ async def stream_fn(
messages: list[Message], options: dict[str, Any], **kwargs: Any
) -> AsyncIterator[ChatResponseUpdate]:
nonlocal round_number
if options.get("tool_choice") == "none":
tool_choice = options.get("tool_choice")
tool_choices.append(tool_choice)
if tool_choice == "none":
yield ChatResponseUpdate(role="assistant", contents=[Content.from_text("Done.")])
return
round_number += 1
Expand All @@ -1131,9 +1148,12 @@ async def stream_fn(
agent = Agent(client=client, tools=[FunctionTool(name="search", description="Search", func=search)])
assert isinstance(client, FunctionInvocationLayer)
client.function_invocation_configuration["max_function_calls"] = max_calls
client.function_middleware.append(RecordExecution())
kinds = await _drive(enable_a2ui(agent, _RenderSub()))

assert executed == expected_actions
assert tool_choices == expected_tool_choices
assert execution_trace == expected_execution_trace
assert len([kind for kind in kinds if kind[0] == "result" and kind[1].startswith("generate-")]) == expected_surfaces
assert len([kind for kind in kinds if kind[0] == "result" and kind[1].startswith("search-")]) == len(
expected_actions
Expand Down
Loading
Loading