FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Optimize UTF-8 decoding to UCS1 · Issue #158931 · python/cpython · GitHub

Repository navigation

Optimize UTF-8 decoding to UCS1 #158931

Description

Feature or enhancement

Is it expected that decoding 2-byte UTF-8 sequences to UCS1 is no faster than decoding them to UCS2? I thought UCS1 would have an advantage because the resulting string uses half as much memory and requires half as many bytes to be written.

$ python -VV
Python 3.15.0rc3 (tags/v3.15.0rc3:8a8eb0b, Oct  2 2026, 19:16:27) [MSC v.1951 64 bit (AMD64)]
$ python -m timeit -s 's = ("\x80" * 65_536).encode()' 's.decode()'  
5000 loops, best of 5: 98.8 usec per loop
$ python -m timeit -s 's = ("\u0100" * 65_536).encode()' 's.decode()'
5000 loops, best of 5: 98.7 usec per loop

For comparison, the corresponding encoding operations do show a difference:

$ python -m timeit -s 's = "\x80" * 65_536' 's.encode()'
5000 loops, best of 5: 69.1 usec per loop
$ python -m timeit -s 's = "\u0100" * 65_536' 's.encode()'           
1000 loops, best of 5: 99.9 usec per loop

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)pendingThe issue will be closed if no feedback is providedperformancePerformance or resource usagetype-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL