| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
When an encode error handler returned a position before the error, _PyUnicode_EncodeUTF16() reserved one code unit for each character to be re-encoded, but non-BMP characters take two (a surrogate pair). This underestimated the buffer size and caused a heap buffer overflow.
| Back | FazBrowse Home | New Git URL |
Fixes #158925
If an encode error handler returns a position before the error, _PyUnicode_EncodeUTF16() grows the buffer by one code unit for each character it will encode again (moreunits += pos - newpos). In a UCS-4 string a non-BMP character takes two units, so ucs4lib_utf16_encode() ends up writing past the end of the buffer.
Fix: in that case reserve two units per character, same as the UTF-8 encoder does for backward jumps (max_char_size * (startpos - newpos)). Whatever isn't used gets trimmed by PyBytesWriter_FinishWithPointer(). UCS-2 strings and forward jumps don't need it, and UTF-32 is always 4 bytes per character so it's fine.
Tested on a debug build: the reproducer from the issue and the new test both crash without the patch and pass with it (also with -R 3:3). The test checks the UTF-8, UTF-16 and UTF-32 encoders. test_codeccallbacks, test_codecs, test_str, test_bytes and test_capi.test_unicode pass.
The same code is in 3.10, so this should be backported.