FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

gh-158445: Reject invalid UCS4 in PyUnicode_FromKindAndData() by vstinner · Pull Request #158447 · python/cpython · GitHub

Repository navigation

gh-158445: Reject invalid UCS4 in PyUnicode_FromKindAndData() - #158447

Closed
vstinner wants to merge 1 commit into
python:mainfrom
vstinner:from_ucs4
Closed

vstinner wants to merge 1 commit into
python:mainfrom
vstinner:from_ucs4

Conversation

vstinner commented Sep 29, 2026 •
edited by bedevere-app Bot
Loading

Copy link
Copy Markdown
Member

PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and PyUnicodeWriter_WriteUCS4() now raise an exception if a character is not in the [U+0000; U+10ffff] range, instead of creating an invalid str object.

  • Add _testinternalcapi._Py_MAX_UNICODE.
  • Add unicode_invalid_character() helper function.

PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and
PyUnicodeWriter_WriteUCS4() now raise an exception if a character is
not in the [U+0000; U+10ffff] range, instead of creating an invalid
str object.

* Add _testinternalcapi._Py_MAX_UNICODE.
* Add unicode_invalid_character() helper function.

Copy link
Copy Markdown

Documentation build overview

📚 cpython-previews | 🛠️ Build #34835623 | 📁 Comparing da5cb33 against main (115c297)

  🔍 Preview build  

2 files changed
± c-api/unicode.html
± whatsnew/changelog.html

Copy link
Copy Markdown
Member Author

@serhiy-storchaka: Would you mind to review this change? See the issue for the rationale.

The change makes the two functions a little bit slower, but also makes them safer. It should not be possible to create an invalid string in Python.

In the wild, I mostly saw invalid characters when debugging CPython. For example, PyUnicode_New(size, 0x10ffff) creates a UCS-4 buffer filled with the byte pattern 0xff which creates invalid characters \Uffffffff on purpose: to detect usage of uninitialized characters.

The other case where I saw invalid characters was on Solaris with wchar_t* strings (Py_UCS4 strings in practice). The _Py_DecodeNonUnicodeWchar() function was added to fix these characters.

Copy link
Copy Markdown
Member

No, I do not think it is worth to slow down this function. If you need an additional check -- use the UTF32 decoder.

Copy link
Copy Markdown
Member Author

@serhiy-storchaka:

No, I do not think it is worth to slow down this function. If you need an additional check -- use the UTF32 decoder.

It's a little bit surprising that only 2 functions of the C API ignores invalid characters. But you have a point with performance.

I wrote PR gh-158502 to document the undefined behavior, only detect invalid characters in debug mode (raise SystemError), and add tests on the behavior in release and debug mode.

Copy link
Copy Markdown
Member Author

Rejected in favor of #158502.

vstinner closed this Sep 30, 2026
vstinner deleted the from_ucs4 branch September 30, 2026 14:51
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants


Back | FazBrowse Home | New Git URL