Skip to content

fix(rst): a section's text no longer reads an unwritten byte - #2583

Merged
DeusData merged 1 commit into
mainfrom
fix/rst-body-cut
Oct 10, 2026
Merged

DeusData merged 1 commit into
mainfrom
fix/rst-body-cut

Conversation

@DeusData

Copy link
Copy Markdown
Owner

rst_body cuts a reST section's text at 500 bytes. It did so by setting a stop value one past the end and then backing off from out[RST_BODY_MAX] while that byte looked like a UTF-8 continuation byte. That byte was never written, which had two effects:

  • Run-to-run variance. Whether the text ended at 499 or 500 bytes depended on whatever the arena memory held.
  • Garbage bytes. When a multibyte character didn't fit, bytes that were never written could end up in the text.

Fix: only whole characters are written, and the text ends at the last one written.

Measured

Test. rst_section_text_cut pins the cut at the boundary for three cases: a word that ends exactly at byte 500, a two-byte character that would cross it, and a word cut at a character. The words are spread over short paragraphs on purpose.

  • ASan lane: the unwritten byte happens to read as zero in this test process, so the old code passes the test too.
  • MSan lane (docker compose -f test-infrastructure/docker-compose.yml run --rm test-msan doc_links_rst, local arm64): with the old rst_body, the test aborts with MemorySanitizer: stack-overflow … nested bug in the same thread. That happened three runs out of three, always at the same pc. With the fix, the suite passes: 9 passed, nothing reported.

Possibly related. The stack overflow appears only when the code reads the unwritten byte, so it looks like MSan failing while it reports that read. The five suites scripts/msan.sh excludes for "stack-overflow under instrumentation" may hide real uninitialized reads the same way. This PR doesn't change those exclusions.

rst_body cut a reST section's text at 500 bytes by setting a stop value
past the end, then backing off from out[RST_BODY_MAX] while that byte
looked like a UTF-8 continuation byte. That byte was never written, so
where the text ended (499 or 500 bytes) depended on what the arena memory
held, and when a multibyte character did not fit, bytes never written
could end up in the text.

Only whole characters are written now, and the text ends at the last one
written. The same django tree indexed twice gave doc_link_candidates rows
that differed in about 250 of 11,300 (through the section text's TF-IDF);
with this change two runs are identical, and so are the 6,594 sections.

The new test pins the cut at the boundary: a word that fills the 500 bytes,
a two-byte character that would cross them, a word cut at a character.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData
DeusData merged commit 83edbd3 into main Oct 10, 2026
28 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant