Skip to content

ClientSession Error Handling #1401

Description

@Unshure

Question

I work on the Strands SDK, and we have an integration with MCP where we invoke tools through an MCP client.

Our design spins up a thread, creates an asyncio event loop, and schedules a task to create a session with an MCP server using the python-sdk mcp ClientSession: https://github.com/strands-agents/sdk-python/blob/main/src/strands/tools/mcp/mcp_client.py#L391

Then, when an LLM decides to call a tool, or multiple tools, we schedule additional tasks on this thread's event loop to call the mcp server: https://github.com/strands-agents/sdk-python/blob/main/src/strands/tools/mcp/mcp_client.py#L324-L330

We ran into an issue where when the mcp streamablehttp_client sse_read_timeout was lower than the time it took for a tool to return, the tool invocation tasks would hang due to an exception that was not propagated out of the ClientSession. We see an exception message (included in the additional context section), but the stack trace never enters our code, it stops in the mcp code.

After investigating further, I found that this exception is sent through the MCP ClientSession message_handler, and in the default implementation of this message handler, an exception is never raised:

async def _default_message_handler(
message: RequestResponder[types.ServerRequest, types.ClientResult] | types.ServerNotification | Exception,
) -> None:
await anyio.lowlevel.checkpoint()

To work around this issue, I have introduced my own message_handler to raise an exception passed to it: strands-agents/harness-sdk#922

I wanted to know why these exceptions are not raised in the MCP ClientSession? This silent failure took a long time to debug, and I want able to find any documentation on this behavior. As a user of the ClientSession, I would expect these exceptions to be raised so that my code can handle them.

Additional Context

Error stack trace:

Error reading SSE stream:
Traceback (most recent call last):
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_transports/default.py", line 101, in map_httpcore_exceptions
    yield
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_transports/default.py", line 271, in __aiter__
    async for part in self._httpcore_stream:
        yield part
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_async/connection_pool.py", line 407, in __aiter__
    raise exc from None
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_async/connection_pool.py", line 403, in __aiter__
    async for part in self._stream:
        yield part
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_async/http11.py", line 342, in __aiter__
    raise exc
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_async/http11.py", line 334, in __aiter__
    async for chunk in self._connection._receive_response_body(**kwargs):
        yield chunk
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_async/http11.py", line 203, in _receive_response_body
    event = await self._receive_event(timeout=timeout)
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_async/http11.py", line 217, in _receive_event
    data = await self._network_stream.read(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        self.READ_NUM_BYTES, timeout=timeout
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    )
    ^
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_backends/anyio.py", line 32, in read
    with map_exceptions(exc_map):
         ~~~~~~~~~~~~~~^^^^^^^^^
  File "/opt/homebrew/Cellar/python@3.13/3.13.5/Frameworks/Python.framework/Versions/3.13/lib/python3.13/contextlib.py", line 162, in __exit__
    self.gen.throw(value)
    ~~~~~~~~~~~~~~^^^^^^^
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpcore/_exceptions.py", line 14, in map_exceptions
    raise to_exc(exc) from exc
httpcore.ReadTimeout

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/mcp/client/streamable_http.py", line 326, in _handle_sse_response
    async for sse in event_source.aiter_sse():
    ...<10 lines>...
            break
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx_sse/_api.py", line 42, in aiter_sse
    async for line in lines:
    ...<3 lines>...
            yield sse
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_models.py", line 1031, in aiter_lines
    async for text in self.aiter_text():
        for line in decoder.decode(text):
            yield line
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_models.py", line 1018, in aiter_text
    async for byte_content in self.aiter_bytes():
    ...<2 lines>...
            yield chunk
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_models.py", line 997, in aiter_bytes
    async for raw_bytes in self.aiter_raw():
    ...<2 lines>...
            yield chunk
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_models.py", line 1055, in aiter_raw
    async for raw_stream_bytes in self.stream:
    ...<2 lines>...
            yield chunk
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_client.py", line 176, in __aiter__
    async for chunk in self._stream:
        yield chunk
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_transports/default.py", line 270, in __aiter__
    with map_httpcore_exceptions():
         ~~~~~~~~~~~~~~~~~~~~~~~^^
  File "/opt/homebrew/Cellar/python@3.13/3.13.5/Frameworks/Python.framework/Versions/3.13/lib/python3.13/contextlib.py", line 162, in __exit__
    self.gen.throw(value)
    ~~~~~~~~~~~~~~^^^^^^^
  File "/Users/ncclegg/Library/Application Support/hatch/env/virtual/strands-agents/X7vsTQrp/strands-agents/lib/python3.13/site-packages/httpx/_transports/default.py", line 118, in map_httpcore_exceptions
    raise mapped_exc(message) from exc
httpx.ReadTimeout

Activity

  1. added
    bugSomething isn't working
    ready for workEnough information for someone to start working on
    P1Significant bug affecting many users, highly requested feature
    and removed
    questionFurther information is requested
    on Oct 3, 2025
  2. felixweinberger commented on Oct 3, 2025

    @felixweinberger
    Contributor

    Thanks for this report - this seems like something we should fix to raise the exception for more predictable behavior.

  3. dgenio commented on Nov 28, 2025

    @dgenio

    This is a great issue to surface — from a broader perspective, the SDK could benefit from a more consistent error taxonomy and cross-layer behavior, not just in ClientSession.

    A few related pain points:

    • Different layers (transports, sessions, FastMCP, low-level server) use different patterns:
      • some raise exceptions directly,
      • others wrap in ToolError or similar,
      • others convert to ErrorData with limited structure.
    • It’s hard for clients (and LLM agents) to distinguish:
      • transient / retryable errors (network, timeouts),
      • validation errors (bad input),
      • permission / auth problems,
      • “tool blew up internally” bugs.

    It might be worth expanding this issue (or creating a related one) to define:

    1. A small error taxonomy, e.g.:

      • TransientError – retryable
      • ValidationError – bad input, non-retryable
      • PermissionError – authn/authz issues
      • NotFoundError, etc.

      and how they map into ErrorData.code / data.

    2. A consistent propagation rule:

      • At protocol boundaries, always return an ErrorData response rather than raising.
      • Preserve the original exception type and message in a structured way in ErrorData.data.
      • Ensure clients can reliably inspect the error category and decide whether to retry.
    3. A minimal error code registry:

      • A short list of documented error codes with clear semantics and retry guidance.

    If maintainers are interested, I’d be happy to help sketch a concrete proposal for an error taxonomy and how it would map across server, client, and transports, building on the work you’re doing here for ClientSession.

  4. maxisbey commented on Dec 16, 2025

    @maxisbey
    Contributor

    Additional related issue: #1789

    The reproduction in #1789 reveals there are actually two related bugs in the streamable HTTP transport:

    Bug 1 (original issue): In _handle_sse_response(), timeout exceptions are caught at line 388 but not propagated to the waiting request when last_event_id is None.

    Bug 2 (from #1789): In handle_get_stream(), the attempt counter is reset to 0 whenever the SSE stream ends "normally" - but httpx read timeouts cause a graceful stream end, not an exception. This means the reconnection loop runs forever instead of respecting MAX_RECONNECTION_ATTEMPTS. The same issue exists in _handle_reconnection() which recursively calls itself with attempt=0.

    The fix likely needs to:

    1. Propagate timeout errors to the response stream (for bug 1)
    2. Track whether any events were received before the stream ended - if none were received, treat it as a timeout/error rather than a normal server close (for bug 2)

    AI Disclaimer

  5. ivanbelenky commented on Dec 23, 2025

    @ivanbelenky

    hey @maxisbey just wanted to mention that Bug 1 has been addressed here.

  6. mehtarac commented on Dec 30, 2025

    @mehtarac

    Hi @ivanbelenky , thank you for the fix. I was curious when we can expect the PR to be merged and released?

  7. ivanbelenky commented on Dec 30, 2025

    @ivanbelenky

    @mehtarac the PR is waiting for review, I am not a maintainer so I wouldn't know.

  8. 14 remaining items

  9. DSHCorrectover commented on Aug 23, 2026

    @DSHCorrectover
  10. added a commit that references this issue on Aug 23, 2026
    80cd21e
  11. mukktinaadh commented on Aug 23, 2026

    @mukktinaadh

    I've implemented a fix for this issue in PR #3369. The fix adds a dedicated callback to and classes, allowing users to properly handle transport-level exceptions (timeouts, connection errors) instead of having them silently swallowed.

    Summary of the fix:

    • New callback receives transport exceptions directly
    • Falls back to for backwards compatibility
    • Default now logs transport exceptions at ERROR level
    • All 5,744 existing tests pass + 2 new tests added

    The PR was auto-closed due to the assignment requirement. I'm happy to discuss the approach or make adjustments — just let me know if a maintainer can assign this issue so the PR can reopen.

  12. chaucerj commented on Aug 27, 2026

    @chaucerj

    Hi! I'd like to work on this issue.

    I've investigated and have a fix ready for review: when the read stream yields an exception, the dispatcher now fails all pending send_raw_request waiters with that exception (re-raised as-is, preserving the original type like httpx.ReadTimeout) instead of parking them until their timeout elapses. The on_stream_exception observer tee is unchanged, and the dispatcher keeps serving once the stream recovers.

    Happy to adjust if you'd prefer a different approach.

  13. AnnasMazhar commented on Sep 1, 2026

    @AnnasMazhar

    Confirmed on current main: _default_message_handler still only checkpoint()s, while IncomingMessage is ServerNotification | Exception. Transport faults (e.g. httpx.ReadTimeout from sse_read_timeout) therefore reach the handler and disappear; _deliver_stream_exception never logs because nothing is raised.

    Minimal fix: re-raise Exception items in the default handler so the existing spawn-and-log path surfaces them at ERROR. Custom message_handlers are unchanged. Failing in-flight send_request waiters is a separate follow-up.

  14. RanaPriyansh commented on Sep 2, 2026

    @RanaPriyansh

    Confirmed on current main (2026-09-02): _default_message_handler still only await anyio.lowlevel.checkpoint() while IncomingMessage is ServerNotification | Exception. Transport faults such as httpx.ReadTimeout from sse_read_timeout still reach the handler and vanish; _deliver_stream_exception only logs if the handler raises.

    Propose: re-raise Exception items in the default handler only, so the existing spawn-and-log path surfaces them at ERROR. Custom message_handlers unchanged. Not opening a PR until this issue is assigned.

  15. FOWEPJF255 commented on Sep 7, 2026

    @FOWEPJF255

    Confirmed on current main (working tree from zipball snapshot).

    Remaining gap

    1. _default_message_handler still only await anyio.lowlevel.checkpoint(), so transport Exception items that reach message_handler are discarded; _deliver_stream_exception's ERROR log path never fires unless a custom handler raises.
    2. JSONRPCDispatcher._dispatch observes Exception items but does not fail in-flight send_raw_request waiters, so callers can sit until their own timeout (the hang reported here) even when the transport already failed.

    Proposed minimal fix (two coherent pieces)

    1. Default handler: re-raise Exception items; notifications unchanged. Mirror the same behavior in Client._evicting_message_handler when no user handler is set.
    2. Dispatcher: on a read-stream Exception, _fail_pending(exc) so waiters are woken with the original exception type (e.g. httpx.ReadTimeout), then continue to on_stream_exception / message_handler. _fan_out_closed becomes a thin wrapper around _fail_pending(CONNECTION_CLOSED).

    Custom message_handler callables stay untouched. Not opening a PR until this issue is assigned (per CONTRIBUTING).

    Tests prepared locally

    • test_default_message_handler_raises_on_transport_exception
    • test_transport_exception_in_stream_logs_at_error_level
    • test_send_raw_request_raises_transport_exception_yielded_mid_await (asyncio + trio)

    Happy to adjust scope (handler-only vs handler+pending) if maintainers prefer a smaller first PR.

    Disclosure: drafted with AI assistance; I personally reviewed the approach against session.py / jsonrpc_dispatcher.py and the existing closed PRs on this issue.

  16. VAIBHAVVVV76 commented on Sep 8, 2026

    @VAIBHAVVVV76

    Hi! I’d like to contribute to this issue.

    My understanding is that an exception received by ClientSession reaches the default message handler but is not propagated to the pending tool-call task. With a Streamable HTTP read timeout, that can leave the caller waiting indefinitely instead of receiving an exception it can handle.

    My proposed first step is to reproduce this with a deliberately slow local tool, trace the exception path through ClientSession and its pending request handling, and then propose the smallest backwards-compatible fix plus a regression test. I’ll keep the change scoped and will not alter public API behaviour without confirming the intended semantics here.

    Does that approach match the maintainers’ intent, and may I be assigned to the issue?

    Disclosure: I used AI to help orient myself in the codebase; I will personally reproduce, test, and own any change I submit.

  17. FOWEPJF255 commented on Sep 8, 2026

    @FOWEPJF255

    @VAIBHAVVVV76 Thanks for jumping in. I left a gap analysis earlier on this thread (default handler swallowing transport Exception items, plus pending send_raw_request waiters not being failed when the read stream yields an Exception).

    I am holding off on a PR until a maintainer assigns the issue (per CONTRIBUTING). Happy to collaborate on the smallest backwards-compatible fix + regression tests once it is assigned - whether that is handler-only first, or handler + failing pending waiters together.

    Happy to sync so we do not duplicate work.

  18. Steeve-Crypto commented on Sep 24, 2026

    @Steeve-Crypto

    Confirmed on current main (f1b6589): two gaps still match the original hang.

    1. Default handler swallows transport faults. IncomingMessage is ServerNotification | Exception, and _deliver_stream_exception only logs when the handler raises. _default_message_handler still only does await anyio.lowlevel.checkpoint(), so faults that reach it disappear and the ERROR path never fires.

    2. Pending requests keep waiting. JSONRPCDispatcher._dispatch observes Exception items via on_stream_exception but does not wake in-flight send_raw_request waiters. EOF already fans out CONNECTION_CLOSED; a mid-stream transport fault (e.g. httpx.ReadTimeout from sse_read_timeout) does not, so call_tool / send_request can hang until the caller's own timeout.

    Proposed minimal fix (backwards-compatible):

    • Re-raise Exception items in _default_message_handler only (custom handlers unchanged) so the existing spawn-and-log path surfaces them at ERROR.
    • On a read-stream Exception, also fail pending waiters with CONNECTION_CLOSED whose message carries the original fault (same shape as the EOF fan-out), then still call on_stream_exception.
    • Regression tests for both paths.

    I use this SDK in agent/MCP client work (Steeve-Crypto / Haquant) and hit the same silent-timeout class of bug. Implementing against this approach and opening a PR with Fixes #1401 + tests. Happy to adjust if maintainers want handler-only first.

    Disclosure: written with AI assistance; I reviewed the analysis against current main and will own the PR.

  19. FOWEPJF255 commented on Sep 30, 2026

    @FOWEPJF255

    Branch pushed with handler + dispatcher fix and regression tests (4 pytest cases passed locally): https://github.com/FOWEPJF255/python-sdk/tree/fix/1401-client-session-error-handling

    Holding off on opening a PR until this issue is assigned (per CONTRIBUTING). Happy to open Fixes #1401 once a maintainer assigns.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1Significant bug affecting many users, highly requested featurebugSomething isn't workinggood first issueGood for newcomersready for workEnough information for someone to start working on

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions