| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and PyUnicodeWriter_WriteUCS4() now raise an exception if a character is not in the [U+0000; U+10ffff] range, instead of creating an invalid str object. * Add _testinternalcapi._Py_MAX_UNICODE. * Add unicode_invalid_character() helper function.
Documentation build overview2 files changed ± c-api/unicode.html ± whatsnew/changelog.html |
Sorry, something went wrong.
|
@serhiy-storchaka: Would you mind to review this change? See the issue for the rationale. The change makes the two functions a little bit slower, but also makes them safer. It should not be possible to create an invalid string in Python. In the wild, I mostly saw invalid characters when debugging CPython. For example, PyUnicode_New(size, 0x10ffff) creates a UCS-4 buffer filled with the byte pattern 0xff which creates invalid characters \Uffffffff on purpose: to detect usage of uninitialized characters. The other case where I saw invalid characters was on Solaris with wchar_t* strings (Py_UCS4 strings in practice). The _Py_DecodeNonUnicodeWchar() function was added to fix these characters. |
Sorry, something went wrong.
|
No, I do not think it is worth to slow down this function. If you need an additional check -- use the UTF32 decoder. |
Sorry, something went wrong.
It's a little bit surprising that only 2 functions of the C API ignores invalid characters. But you have a point with performance. I wrote PR gh-158502 to document the undefined behavior, only detect invalid characters in debug mode (raise SystemError), and add tests on the behavior in release and debug mode. |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and PyUnicodeWriter_WriteUCS4() now raise an exception if a character is not in the [U+0000; U+10ffff] range, instead of creating an invalid str object.