Repository navigation
ClientSession Error Handling #1401
Description
Activity
- addedquestionFurther information is requestedFurther information is requested
on Sep 25, 2025 - addedbugSomething isn't workingSomething isn't workingready for workEnough information for someone to start working onEnough information for someone to start working onP1Significant bug affecting many users, highly requested featureSignificant bug affecting many users, highly requested featureand removedquestionFurther information is requestedFurther information is requested
on Oct 3, 2025 Thanks for this report - this seems like something we should fix to raise the exception for more predictable behavior.
This is a great issue to surface — from a broader perspective, the SDK could benefit from a more consistent error taxonomy and cross-layer behavior, not just in
ClientSession.A few related pain points:
- Different layers (transports, sessions, FastMCP, low-level server) use different patterns:
- some raise exceptions directly,
- others wrap in
ToolErroror similar, - others convert to
ErrorDatawith limited structure.
- It’s hard for clients (and LLM agents) to distinguish:
- transient / retryable errors (network, timeouts),
- validation errors (bad input),
- permission / auth problems,
- “tool blew up internally” bugs.
It might be worth expanding this issue (or creating a related one) to define:
-
A small error taxonomy, e.g.:
TransientError– retryableValidationError– bad input, non-retryablePermissionError– authn/authz issuesNotFoundError, etc.
and how they map into
ErrorData.code/data. -
A consistent propagation rule:
- At protocol boundaries, always return an
ErrorDataresponse rather than raising. - Preserve the original exception type and message in a structured way in
ErrorData.data. - Ensure clients can reliably inspect the error category and decide whether to retry.
- At protocol boundaries, always return an
-
A minimal error code registry:
- A short list of documented error codes with clear semantics and retry guidance.
If maintainers are interested, I’d be happy to help sketch a concrete proposal for an error taxonomy and how it would map across server, client, and transports, building on the work you’re doing here for
ClientSession.- Different layers (transports, sessions, FastMCP, low-level server) use different patterns:
Additional related issue: #1789
The reproduction in #1789 reveals there are actually two related bugs in the streamable HTTP transport:
Bug 1 (original issue): In
_handle_sse_response(), timeout exceptions are caught at line 388 but not propagated to the waiting request whenlast_event_idisNone.Bug 2 (from #1789): In
handle_get_stream(), theattemptcounter is reset to0whenever the SSE stream ends "normally" - but httpx read timeouts cause a graceful stream end, not an exception. This means the reconnection loop runs forever instead of respectingMAX_RECONNECTION_ATTEMPTS. The same issue exists in_handle_reconnection()which recursively calls itself withattempt=0.The fix likely needs to:
- Propagate timeout errors to the response stream (for bug 1)
- Track whether any events were received before the stream ended - if none were received, treat it as a timeout/error rather than a normal server close (for bug 2)
Hi @ivanbelenky , thank you for the fix. I was curious when we can expect the PR to be merged and released?
@mehtarac the PR is waiting for review, I am not a maintainer so I wouldn't know.
14 remaining items
- added a commit that references this issue
on Aug 23, 2026 I've implemented a fix for this issue in PR #3369. The fix adds a dedicated callback to and classes, allowing users to properly handle transport-level exceptions (timeouts, connection errors) instead of having them silently swallowed.
Summary of the fix:
- New callback receives transport exceptions directly
- Falls back to for backwards compatibility
- Default now logs transport exceptions at ERROR level
- All 5,744 existing tests pass + 2 new tests added
The PR was auto-closed due to the assignment requirement. I'm happy to discuss the approach or make adjustments — just let me know if a maintainer can assign this issue so the PR can reopen.
- added a commit that references this issue
on Aug 25, 2026 Hi! I'd like to work on this issue.
I've investigated and have a fix ready for review: when the read stream yields an exception, the dispatcher now fails all pending
send_raw_requestwaiters with that exception (re-raised as-is, preserving the original type likehttpx.ReadTimeout) instead of parking them until their timeout elapses. Theon_stream_exceptionobserver tee is unchanged, and the dispatcher keeps serving once the stream recovers.Happy to adjust if you'd prefer a different approach.
- added a commit that references this issue
on Aug 31, 2026 Confirmed on current
main:_default_message_handlerstill onlycheckpoint()s, whileIncomingMessageisServerNotification | Exception. Transport faults (e.g.httpx.ReadTimeoutfromsse_read_timeout) therefore reach the handler and disappear;_deliver_stream_exceptionnever logs because nothing is raised.Minimal fix: re-raise
Exceptionitems in the default handler so the existing spawn-and-log path surfaces them at ERROR. Custommessage_handlers are unchanged. Failing in-flightsend_requestwaiters is a separate follow-up.Confirmed on current
main(2026-09-02):_default_message_handlerstill onlyawait anyio.lowlevel.checkpoint()whileIncomingMessageisServerNotification | Exception. Transport faults such ashttpx.ReadTimeoutfromsse_read_timeoutstill reach the handler and vanish;_deliver_stream_exceptiononly logs if the handler raises.Propose: re-raise
Exceptionitems in the default handler only, so the existing spawn-and-log path surfaces them at ERROR. Custommessage_handlers unchanged. Not opening a PR until this issue is assigned.Confirmed on current
main(working tree from zipball snapshot).Remaining gap
_default_message_handlerstill onlyawait anyio.lowlevel.checkpoint(), so transportExceptionitems that reachmessage_handlerare discarded;_deliver_stream_exception's ERROR log path never fires unless a custom handler raises.JSONRPCDispatcher._dispatchobservesExceptionitems but does not fail in-flightsend_raw_requestwaiters, so callers can sit until their own timeout (the hang reported here) even when the transport already failed.
Proposed minimal fix (two coherent pieces)
- Default handler: re-raise
Exceptionitems; notifications unchanged. Mirror the same behavior inClient._evicting_message_handlerwhen no user handler is set. - Dispatcher: on a read-stream
Exception,_fail_pending(exc)so waiters are woken with the original exception type (e.g.httpx.ReadTimeout), then continue toon_stream_exception/message_handler._fan_out_closedbecomes a thin wrapper around_fail_pending(CONNECTION_CLOSED).
Custom
message_handlercallables stay untouched. Not opening a PR until this issue is assigned (per CONTRIBUTING).Tests prepared locally
test_default_message_handler_raises_on_transport_exceptiontest_transport_exception_in_stream_logs_at_error_leveltest_send_raw_request_raises_transport_exception_yielded_mid_await(asyncio + trio)
Happy to adjust scope (handler-only vs handler+pending) if maintainers prefer a smaller first PR.
Disclosure: drafted with AI assistance; I personally reviewed the approach against
session.py/jsonrpc_dispatcher.pyand the existing closed PRs on this issue.Hi! I’d like to contribute to this issue.
My understanding is that an exception received by
ClientSessionreaches the default message handler but is not propagated to the pending tool-call task. With a Streamable HTTP read timeout, that can leave the caller waiting indefinitely instead of receiving an exception it can handle.My proposed first step is to reproduce this with a deliberately slow local tool, trace the exception path through
ClientSessionand its pending request handling, and then propose the smallest backwards-compatible fix plus a regression test. I’ll keep the change scoped and will not alter public API behaviour without confirming the intended semantics here.Does that approach match the maintainers’ intent, and may I be assigned to the issue?
Disclosure: I used AI to help orient myself in the codebase; I will personally reproduce, test, and own any change I submit.
@VAIBHAVVVV76 Thanks for jumping in. I left a gap analysis earlier on this thread (default handler swallowing transport Exception items, plus pending send_raw_request waiters not being failed when the read stream yields an Exception).
I am holding off on a PR until a maintainer assigns the issue (per CONTRIBUTING). Happy to collaborate on the smallest backwards-compatible fix + regression tests once it is assigned - whether that is handler-only first, or handler + failing pending waiters together.
Happy to sync so we do not duplicate work.
Confirmed on current
main(f1b6589): two gaps still match the original hang.1. Default handler swallows transport faults.
IncomingMessageisServerNotification | Exception, and_deliver_stream_exceptiononly logs when the handler raises._default_message_handlerstill only doesawait anyio.lowlevel.checkpoint(), so faults that reach it disappear and the ERROR path never fires.2. Pending requests keep waiting.
JSONRPCDispatcher._dispatchobservesExceptionitems viaon_stream_exceptionbut does not wake in-flightsend_raw_requestwaiters. EOF already fans outCONNECTION_CLOSED; a mid-stream transport fault (e.g.httpx.ReadTimeoutfromsse_read_timeout) does not, socall_tool/send_requestcan hang until the caller's own timeout.Proposed minimal fix (backwards-compatible):
- Re-raise
Exceptionitems in_default_message_handleronly (custom handlers unchanged) so the existing spawn-and-log path surfaces them at ERROR. - On a read-stream
Exception, also fail pending waiters withCONNECTION_CLOSEDwhose message carries the original fault (same shape as the EOF fan-out), then still callon_stream_exception. - Regression tests for both paths.
I use this SDK in agent/MCP client work (Steeve-Crypto / Haquant) and hit the same silent-timeout class of bug. Implementing against this approach and opening a PR with
Fixes #1401+ tests. Happy to adjust if maintainers want handler-only first.Disclosure: written with AI assistance; I reviewed the analysis against current
mainand will own the PR.- Re-raise
Branch pushed with handler + dispatcher fix and regression tests (4 pytest cases passed locally): https://github.com/FOWEPJF255/python-sdk/tree/fix/1401-client-session-error-handling
Holding off on opening a PR until this issue is assigned (per CONTRIBUTING). Happy to open
Fixes #1401once a maintainer assigns.
Question
I work on the Strands SDK, and we have an integration with MCP where we invoke tools through an MCP client.
Our design spins up a thread, creates an asyncio event loop, and schedules a task to create a session with an MCP server using the python-sdk mcp ClientSession: https://github.com/strands-agents/sdk-python/blob/main/src/strands/tools/mcp/mcp_client.py#L391
Then, when an LLM decides to call a tool, or multiple tools, we schedule additional tasks on this thread's event loop to call the mcp server: https://github.com/strands-agents/sdk-python/blob/main/src/strands/tools/mcp/mcp_client.py#L324-L330
We ran into an issue where when the mcp
streamablehttp_clientsse_read_timeout was lower than the time it took for a tool to return, the tool invocation tasks would hang due to an exception that was not propagated out of the ClientSession. We see an exception message (included in the additional context section), but the stack trace never enters our code, it stops in the mcp code.After investigating further, I found that this exception is sent through the MCP ClientSession
message_handler, and in the default implementation of this message handler, an exception is never raised:python-sdk/src/mcp/client/session.py
Lines 57 to 60 in 71889d7
To work around this issue, I have introduced my own message_handler to raise an exception passed to it: strands-agents/harness-sdk#922
I wanted to know why these exceptions are not raised in the MCP ClientSession? This silent failure took a long time to debug, and I want able to find any documentation on this behavior. As a user of the ClientSession, I would expect these exceptions to be raised so that my code can handle them.
Additional Context
Error stack trace: