FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

C API: Rework and enhance the `PyBytesWriter` implementation · Issue #158585 · python/cpython · GitHub

Repository navigation

C API: Rework and enhance the PyBytesWriter implementation #158585

Activity

  1. added 3 commits that reference this issue on Oct 2, 2026
  2. vstinner commented on Oct 2, 2026

    MemberAuthor

    Currently in the main branch, there are 87 calls to PyBytesWriter_Create(). Only 15 resize the writer later:

    • _PyBytes_FormatEx(): call PyBytesWriter_GrowAndUpdatePointer()
    • _PyBytes_FromIterator(): call _PyBytesWriter_ResizeAndUpdatePointer()
    • _PyUnicode_EncodeCharmap(): call PyBytesWriter_Resize()
    • _PyUnicode_EncodeIconv(): call PyBytesWriter_Grow()
    • _PyUnicode_EncodeUTF16(): call PyBytesWriter_GrowAndUpdatePointer()
    • _PyUnicode_EncodeUTF32(): call PyBytesWriter_GrowAndUpdatePointer()
    • _io.FileIO.readall(): call PyBytesWriter_WriteBytes() multiple times
    • _io._IOBase.readline(): call PyBytesWriter_WriteBytes() multiple times
    • _io._RawIOBase.readall(): call PyBytesWriter_WriteBytes() multiple times
    • assemble.c (3 writers): call PyBytesWriter_Resize()
    • codeobject.c: remove_column_info(): call PyBytesWriter_Resize(res, PyBytesWriter_GetSize(res) * 2)
    • encode_code_page_errors(): call PyBytesWriter_GrowAndUpdatePointer()
    • raw_unicode_escape(): call PyBytesWriter_GrowAndUpdatePointer()
    • unicode_encode_ucs1(): call PyBytesWriter_GrowAndUpdatePointer()
    • unicode_encode_utf8(): call PyBytesWriter_GrowAndUpdatePointer()
  3. added 4 commits that reference this issue on Oct 2, 2026
  4. vstinner commented on Oct 3, 2026

    MemberAuthor

    I ran benchmarks on the following code to measure the PyBytesWriter overhead over PyBytes_FromStringAndSize(NULL, size).

    I didn't find any obvious way to optimize PyBytesWriter.

    Benchmark on:

        const Py_ssize_t size = 3;
        PyBytesWriter *writer = PyBytesWriter_Create(size);
        if (writer == NULL) {
            return NULL;
        }
    
        char *str = PyBytesWriter_GetData(writer);
        memset(str, 'x', size);
    
        return PyBytesWriter_Finish(writer);

    I tried to add a freelist to bytes for sizes in range [0; 255] (bytes): see draft PR #158657.

    Benchmark with size=3 bytes:

    • Main branch: 29.8 ns +- 0.3 ns
    • Add bytes freelist: 28.1 ns +- 0.0 ns
    • Remove PyBytesWriter freelist: 37.8 ns +- 1.0 ns

    Benchmark with size=64 bytes:

    • Main branch: 31.0 ns +- 0.9 ns
    • No writer small buffer: 30.2 ns +- 0.3 ns
    • Add bytes freelist: 27.8 ns +- 0.5 ns
    • Add bytes freelist, no writer small buffer: 25.0 ns +- 0.2 ns
    • Remove PyBytesWriter freelist: 40.4 ns +- 4.4 ns
    • Remove PyBytesWriter freelist, no writer small buffer: 37.5 ns +- 0.4 ns

    Benchmark with size=255 bytes:

    • Main branch: 35.5 ns +- 0.9 ns
    • Add bytes freelist: 31.6 ns +- 1.7 ns
    • Add bytes freelist, no writer small buffer: 27.6 ns +- 0.2 ns
    • Remove PyBytesWriter freelist: 50.1 ns +- 0.7 ns

    Benchmark with size=300 bytes (don't use writer small buffer):

    • Main branch: 33.5 ns +- 0.5 ns
    • Remove PyBytesWriter freelist: 40.8 ns +- 0.6 ns

    Notes:

    • Removing the PyBytesWriter freelist makes all benchmarks slower. So the freelist is more efficient than always calling PyMem_Malloc(). _Py_freelists_GET() costs a few nanoseconds, but it's acceptable.
    • The speed up of adding a freelist to bytes is not very obvious to me. It only saves 3.1 ns on size=64 bytes (31.0 ns => 27.8 ns). On size=3, it only saves 1.7 ns. Adding a bytes freelist can increase Python memory usage.
    • Changing byteswriter_resize() allocation strategy to not use the small buffer saves 0.8 ns with size=64. Using the writer small buffer or not doesn't seem to really impact the performance. There is no major difference on performance.
  5. added 3 commits that reference this issue on Oct 3, 2026
  6. 1 remaining item

  7. added 7 commits that reference this issue on Oct 3, 2026
  8. vstinner commented on Oct 5, 2026

    MemberAuthor

    I ran #158665 benchmark (create the string b'abc') to compare the 3.15 branch and the (current) main branches:

    Mean +- std dev: [py315] 41.6 ns +- 1.0 ns -> [main] 32.5 ns +- 0.6 ns: 1.28x faster

    Oh nice, the PyBytesWriter overhead is now way smaller on the main branch! The main branch is 9.1 ns faster.

  9. added 3 commits that reference this issue on Oct 5, 2026
  10. vstinner commented on Oct 10, 2026

    MemberAuthor

    I wrote an article on this work (and PyUnicodeWriter work): https://vstinner.github.io/optimize-pybyteswriter-pyunicodewriter-implementation.html.

  11. added 2 commits that reference this issue on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)topic-C-APItype-refactorCode refactoring (with no changes in behavior)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL