Skip to content

Truncate junitxml failure and error message attributes - #15147

Closed
Cherith1222 wants to merge 2 commits into
pytest-dev:mainfrom
Cherith1222:fix/12223-junitxml-message-bound
Closed

Cherith1222 wants to merge 2 commits into
pytest-dev:mainfrom
Cherith1222:fix/12223-junitxml-message-bound

Conversation

@Cherith1222

@Cherith1222 Cherith1222 commented Oct 7, 2026 •

Copy link
Copy Markdown

Summary

--junitxml wrote the full assertion message into the message attribute of <failure> and <error>. With -vv, that attribute can be tens of thousands of characters. The attribute is now shortened from the original text, before XML escaping: whole lines are kept up to a 500-character budget. A single line longer than the budget is cut at 500 characters and then escaped, so an escape such as #x1B is not split. When anything is dropped, the attribute ends with a note that the full report is in the element text. The element body still contains the full longrepr.

Fixes #12223.

The commit already includes changelog/12223.bugfix.rst. AUTHORS was not changed.

Tests

On Windows, Python 3.13.3, in this checkout, via the project virtualenv:

  • Before this revision, python -m pytest testing/test_junitxml.py -k message_is_truncated -p no:cacheprovider was 3 passed, 140 deselected, and -k failure_verbose_message was 2 passed, 141 deselected.
  • After this revision: -k TestShortenFailureMessage was 4 passed; -k message_is_truncated was 3 passed; -k failure_message_cuts_on_line_boundary was 1 passed; -k raises_match_message_is_not_shortened was 1 passed; -k failure_verbose_message was 2 passed.
  • ruff 0.16.9 format and check passed on src/_pytest/junitxml.py and testing/test_junitxml.py.

Not run: the rest of the pytest suite, mypy, and Python 3.11.

AI assistance

Drafted with Cursor. The account owner authorized opening this pull request. A human has not reviewed the diff. There is no Co-authored-by, Reviewed-by, or Signed-off-by trailer.

Checklist

  • Include new tests or update existing tests when applicable.
  • Allow maintainers to push and squash when merging my commits.
  • Add text like Fixes #12223 to the PR description. The commit message already says the same.
  • Changelog fragment: changelog/12223.bugfix.rst.
  • AUTHORS was not updated.
  • AI assistance is disclosed in this description and in the commit message. A Co-authored-by trailer was not added.

Truncate the escaped message attribute of <failure> and <error> elements to 1000 characters plus '...', keeping the full report in the element body. Fixes pytest-dev#12223.

Drafted with AI assistance (Cursor); not yet reviewed by the account owner.
@psf-chronographer psf-chronographer Bot added the bot:chronographer:provided (automation) changelog entry is part of PR label Oct 7, 2026

Copy link
Copy Markdown
Member

🤖 Written by Claude Opus 5.5 via Claude Code for the pytest maintainers; I prompted it, it did the work, I read it.

Thanks for picking this up. The goal is right: a short message attribute, with the full report in the element text, which already carries it. I'd suggest changing how the attribute is shortened.

Why the current approach falls short

  • It cuts at 1000 chars with a bare .... A reader of the attribute has no hint that the full report is in the element text.
  • It cuts after bin_xml_escape, so it can split an escape like #x1B. Shorten the raw message first, then escape it.
  • This isn't only a -vv problem. A plain raise ValueError(5000 * "z") produces a 5,018-char attribute at default verbosity.

What common failures look like on main. I measured message for 16 typical failures: bare, == and in asserts on int/str/dict/list/approx, pytest.fail, pytest.raises (no raise and bad match), KeyError, AttributeError, a multiline exception, ExceptionGroup, and a fixture setup error.

  • At -q every one is ≤ 283 chars and ≤ 7 lines.
  • At -vv only container diffs grow: dict 326, list(range(200)) 4,426, and the #12223 example about 25k.
  • Keeping only the first line is too aggressive. For pytest.raises(match=...), line 1 is just Regex pattern did not match., and the expected regex and actual message are on lines 2–3. String diffs and multiline exceptions lose their useful part the same way.

Suggestion: keep whole lines up to a budget. If anything was dropped, end with a pointer line:

_MESSAGE_BUDGET = 300  # or ~500

def _shorten_message(msg: str) -> str:
    lines = msg.splitlines() or [""]
    kept: list[str] = []
    size = 0
    for line in lines:
        if kept and size + len(line) + 1 > _MESSAGE_BUDGET:
            break
        kept.append(line)
        size += len(line) + 1
    out = "\n".join(kept)
    cut = len(out) > _MESSAGE_BUDGET
    if cut:
        out = out[:_MESSAGE_BUDGET] + "..."
    rest = len(lines) - len(kept)
    if not cut and not rest:
        return out
    note = f"+{rest} more lines" if rest else "truncated"
    return f"{out}\n[{note}; full report in element text]"

It would be applied as bin_xml_escape(_shorten_message(message)) in append_failure and append_error.

With a budget of 300, all 16 cases stay byte-identical at -q. At -vv only the dict and list diffs get cut, for example 4,426 → 350 chars ending in [+207 more lines; full report in element text]. A budget around 500 would keep the small -vv dict intact too. The common case doesn't change, and the pointer only shows up when something was actually dropped.

On tests: it would be good to cover the line-boundary cut, the hard cut of a single overlong first line, and an unchanged short multiline message (e.g. the pytest.raises(match=...) case).


Generated by Claude Code

Keep whole lines of the raw failure and error text up to 500 characters, then escape. When text is dropped, note that the full report is in the element text.

AI-assisted (Cursor). Not yet reviewed by a human.
@Cherith1222

Copy link
Copy Markdown
Author

The raw message is shortened before bin_xml_escape. Whole lines are kept up to a 500-character budget. A single line longer than that budget is cut at 500 characters, and only then is the text escaped. When anything is dropped, the attribute ends with a note that the full report is in the element text: [+N more lines; full report in element text], or [truncated; full report in element text] when the first line itself was cut.

I used 500 rather than 300 so a small -vv container diff can still fit. A message that already fits is left unchanged, including a short multiline pytest.raises(match=...) failure.

This revision was drafted with Cursor. A human has not line-reviewed the new diff.

Copy link
Copy Markdown
Member

🤖 Written by Claude Opus 5.5 via Claude Code for the pytest maintainers; I prompted it, it did the work, I read it.

Closing in favor of #15150. That PR is a clean reimplementation of the line-budget approach for #12223, and its CI is green.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bot:chronographer:provided (automation) changelog entry is part of PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

junit-xml output attribute can be too big to handle with -vv

2 participants