| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
Well. In fact, the issue is broader: no only _bootstrap_python is affected, any python program is affected since the Modules/getpath.c code is used by all Python executables. |
Sorry, something went wrong.
|
I rebased and updated the PR to clarify that this issue affects the Python path configuration (sys.path creation). |
Sorry, something went wrong.
|
Sadly, Modules/getpath.c is not a regular extension module, it cannot be loaded in test_getpath to easily write unit tests. There are getpath_methods which are injected inside a namespace (dict) by funcs_to_dict() function. It may be interesting to convert it to a regular extension module (_getpath?). |
Sorry, something went wrong.
| const char *path; | ||
| if (!PyArg_ParseTuple(args, "s", &path)) { | ||
| PyObject *path; | ||
| if (!PyArg_ParseTuple(args, "U", &path)) { |
There was a problem hiding this comment.
BTW, I would use METH_O and PyArg_Parse() in these functions, but this is another issue.
Why cannot they be implemented in Python?
Sorry, something went wrong.
There was a problem hiding this comment.
BTW, I would use METH_O and PyArg_Parse() in these functions, but this is another issue.
I tried to minimize the changes.
Why cannot they be implemented in Python?
Ask @zooba who designed this. Maybe it can be changed?
Sorry, something went wrong.
There was a problem hiding this comment.
Perf, mostly. These trivial ones probably could be, but don't fall into the trap of trying to port the full ntpath/posixpath implementations into getpath - we don't have a lot of the functionality needed to handle those at this stage (e.g. no codecs, no os module).
Sorry, something went wrong.
Fix the Python path configuration used to initialized sys.path at Python startup. Paths are no longer encoded to UTF-8/strict to avoid encoding errors if it contains surrogate characters (bytes paths are decoded with the surrogateescape error handler). getpath_basename() and getpath_dirname() functions no longer encode the path to UTF-8/strict, but work directly on Unicode strings. These functions now use PyUnicode_FindChar() and PyUnicode_Substring() on the Unicode path, rather than strrchr() on the encoded bytes string.
I rephrased the NEWS entry to omit function names. Is it better? I only named functions in the commit message. |
Sorry, something went wrong.
|
Although this fixes the issue, this is a bit fragile since it can break if any other function were to use utf8 handler for encoding. This PR avoids the case, can you add a comment that utf8 should be avoided here? |
Sorry, something went wrong.
|
Thanks @vstinner for the PR 🌮🎉.. I'm working now to backport this PR to: 3.11. |
Sorry, something went wrong.
|
Sorry @vstinner, I had trouble checking out the 3.11 backport branch. |
Sorry, something went wrong.
Yes, a regression can be introduced again tomorrow. Well, we can fix it in this case :-)
I'm not sure about the intent of a comment explaining that UTF-8 should not be used, since the modified functions now use Unicode (no encode/decode). |
Sorry, something went wrong.
|
Thanks @vstinner for the PR 🌮🎉.. I'm working now to backport this PR to: 3.11. |
Sorry, something went wrong.
|
GH-97677 is a backport of this pull request to the 3.11 branch. |
Sorry, something went wrong.
…H-97645) Fix the Python path configuration used to initialized sys.path at Python startup. Paths are no longer encoded to UTF-8/strict to avoid encoding errors if it contains surrogate characters (bytes paths are decoded with the surrogateescape error handler). getpath_basename() and getpath_dirname() functions no longer encode the path to UTF-8/strict, but work directly on Unicode strings. These functions now use PyUnicode_FindChar() and PyUnicode_Substring() on the Unicode path, rather than strrchr() on the encoded bytes string. (cherry picked from commit 9f2f1dd) Co-authored-by: Victor Stinner <vstinner@python.org>
Fix the Python path configuration used to initialized sys.path at Python startup. Paths are no longer encoded to UTF-8/strict to avoid encoding errors if it contains surrogate characters (bytes paths are decoded with the surrogateescape error handler). getpath_basename() and getpath_dirname() functions no longer encode the path to UTF-8/strict, but work directly on Unicode strings. These functions now use PyUnicode_FindChar() and PyUnicode_Substring() on the Unicode path, rather than strrchr() on the encoded bytes string. (cherry picked from commit 9f2f1dd) Co-authored-by: Victor Stinner <vstinner@python.org>
…97645) Fix the Python path configuration used to initialized sys.path at Python startup. Paths are no longer encoded to UTF-8/strict to avoid encoding errors if it contains surrogate characters (bytes paths are decoded with the surrogateescape error handler). getpath_basename() and getpath_dirname() functions no longer encode the path to UTF-8/strict, but work directly on Unicode strings. These functions now use PyUnicode_FindChar() and PyUnicode_Substring() on the Unicode path, rather than strrchr() on the encoded bytes string.
| Back | FazBrowse Home | New Git URL |
Fix the Python path configuration used to initialized sys.path at
Python startup. getpath_basename() and getpath_dirname() functions no
longer encode the path to UTF-8/strict to avoid encoding errors if it
contains surrogate characters (created by decoding a bytes path with
the surrogateescape error handler).
The functions now use PyUnicode_FindChar() and PyUnicode_Substring()
on the Unicode path, rather than strrchr() on the encoded bytes
string.