Repository navigation
Accept normalized UTF-8 encoding declarations with a BOM - #1371
Open
harbinresearcher wants to merge 1 commit into
Open
harbinresearcher wants to merge 1 commit into
harbinresearcher wants to merge 1 commit into
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
parse_encodingrejects a UTF-8 BOM when the encoding declaration usesUTF-8,utf_8,UTF_8orutf-8-sig, although Python accepts these declarations. Consequently, message extraction fails before tokenization on valid source files.Normalize case and underscores only for the BOM compatibility check, accepting
utf-8andutf-8-*as Python does. Keep the existingutf-8return value, preserve the original declaration in error messages and leave non-BOM behavior unchanged. General codec lookup is intentionally avoided: Python rejects BOM declarations such asutf8andcp65001even though they identify UTF-8 codecs.Tests cover both declaration lines, valid aliases, incompatible declarations and stream-position restoration on success and failure. Each accepted example is first compiled by Python to check that it is valid source.
Validation on Windows Python 3.12.14:
extract_pythonsmoke check now extracts a message from a BOM-prefixed source usingUTF_8.Other Python versions and operating systems have not been tested locally.
AI assistance: this is an autonomous, user-authorized Codex contribution. Codex investigated, implemented and tested the change; a separate agent reviewed it and reran the targeted tests. No manual human review or production incident is claimed.