Feature or enhancement
Is it expected that decoding 2-byte UTF-8 sequences to UCS1 is no faster than decoding them to UCS2? I thought UCS1 would have an advantage because the resulting string uses half as much memory and requires half as many bytes to be written.
$ python -VV
Python 3.15.0rc3 (tags/v3.15.0rc3:8a8eb0b, Oct 2 2026, 19:16:27) [MSC v.1951 64 bit (AMD64)]
$ python -m timeit -s 's = ("\x80" * 65_536).encode()' 's.decode()'
5000 loops, best of 5: 98.8 usec per loop
$ python -m timeit -s 's = ("\u0100" * 65_536).encode()' 's.decode()'
5000 loops, best of 5: 98.7 usec per loop
For comparison, the corresponding encoding operations do show a difference:
$ python -m timeit -s 's = "\x80" * 65_536' 's.encode()'
5000 loops, best of 5: 69.1 usec per loop
$ python -m timeit -s 's = "\u0100" * 65_536' 's.encode()'
1000 loops, best of 5: 99.9 usec per loop
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Reactions are currently unavailable
Feature or enhancement
Is it expected that decoding 2-byte UTF-8 sequences to UCS1 is no faster than decoding them to UCS2? I thought UCS1 would have an advantage because the resulting string uses half as much memory and requires half as many bytes to be written.
For comparison, the corresponding encoding operations do show a difference:
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response