FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

base64.decode: linebreaks are not ignored · Issue #76672 · python/cpython · GitHub

Repository navigation

base64.decode: linebreaks are not ignored #76672

Description

BPO 32491
Nosy @gpshead, @bitdancer, @vadmium

Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

Show more details

GitHub fields:

assignee = None
closed_at = None
created_at = <Date 2018-01-03.23:35:25.802>
labels = ['3.7', 'type-bug', 'library']
title = 'base64.decode: linebreaks are not ignored'
updated_at = <Date 2018-01-04.03:44:16.618>
user = 'https://github.com/gpshead'

bugs.python.org fields:

activity = <Date 2018-01-04.03:44:16.618>
actor = 'martin.panter'
assignee = 'none'
closed = False
closed_date = None
closer = None
components = ['Library (Lib)']
creation = <Date 2018-01-03.23:35:25.802>
creator = 'gregory.p.smith'
dependencies = []
files = []
hgrepos = []
issue_num = 32491
keywords = []
message_count = 3.0
messages = ['309449', '309451', '309454']
nosy_count = 3.0
nosy_names = ['gregory.p.smith', 'r.david.murray', 'martin.panter']
pr_nums = []
priority = 'normal'
resolution = None
stage = None
status = 'open'
superseder = None
type = 'behavior'
url = 'https://bugs.python.org/issue32491'
versions = ['Python 3.6', 'Python 3.7']

Activity

  1. gpshead commented on Jan 3, 2018

    MemberAuthor

    I've tried reading various RFCs around Base64 encoding, but I couldn't make the ends meet. Yet there is an inconsistency between base64.decodebytes() and base64.decode() in that how they handle linebreaks that were used to collate the encoded text. Below is an example of what I'm talking about:

    >>> import base64
    >>> foo = base64.encodebytes(b'123456789')
    >>> foo
    b'MTIzNDU2Nzg5\n'
    >>> foo = b'MTIzND\n' + b'U2Nzg5\n'
    >>> foo
    b'MTIzND\nU2Nzg5\n'
    >>> base64.decodebytes(foo)
    b'123456789'
    >>> from io import BytesIO
    >>> bytes_in = BytesIO(foo)
    >>> bytes_out = BytesIO()
    >>> bytes_in.seek(0)
    0
    >>> base64.decode(bytes_in, bytes_out)
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
      File "/somewhere/lib/python3.6/base64.py", line 512, in decode
        s = binascii.a2b_base64(line)
    binascii.Error: Incorrect padding
    >>> bytes_in = BytesIO(base64.encodebytes(b'123456789'))
    >>> bytes_in.seek(0)
    0
    >>> base64.decode(bytes_in, bytes_out)
    >>> bytes_out.getvalue()
    b'123456789'

    Obviously, I'd expect encodebytes() and encode both to either accept or to reject the same input.

    Thanks.

    Oleg

    via Oleg Sivokon on python-dev (who was having trouble getting bugs.python.org account creation to work)

  2. added
    stdlibStandard Library Python modules in the Lib/ directory
    type-bugAn unexpected behavior, bug, or error
    on Jan 3, 2018
  3. bitdancer commented on Jan 4, 2018

    Member

    This reduces to the following:

    >>> from binascii import a2b_base64 as f
    >>> f(b'MTIzND\nU2Nzg5\n')
    b'123456789'
    >>> f(b'MTIzND\n')
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
    binascii.Error: Incorrect padding

    That is, decode does its decoding line by line, whereas decodebytes passes the entire object to a2b_base64 as a single entity. Apparently a2b_base64 looks at the padding for the entirety of what it is given, which I believe is in accordance with the RFC. This means that decode is fundamentally broken per the RFC, and there is no obvious way to fix it without adding an incremental decoder to binascii. And an incremental decoder probably belongs in codecs (assuming we ever resolved the transcode interface issue, I can't actually remember...).

    Note that it will work as long as an "integral" number of base64 encoding units are in each line.

  4. vadmium commented on Jan 4, 2018

    Member

    I wrote an incremental base-64 decoder for the "codecs" module in bpo-27799, which you could use. It just does some preprocessing using a regular expression to pick four-character chunks before passing the data to a2b_base64. Or maybe implementing it properly in the "binascii" module is better.

    Quickly reading RFC 2045, I saw it says "All line breaks or other characters not found in Table 1 [64 alphabet characters plus padding character] must be ignored by decoding software." So this is a real bug, although I think a base-64 encoder that triggers it would be rare.

  5. transferred this issue fromon Apr 10, 2022
  6. serhiy-storchaka commented on Mar 31, 2026

    Member

    One way -- implement incremental decoder. But incremental decoders, as they are designed in Python, have fatal flaw -- they are vulnerable to the quadratic complexity attack. For example, if we have incomplete group on one line, followed by ignorable characters, and all other lines contain only ignorable characters, we will have quadratic complexity. The incremental decoder returns position before the incomplete group. After reading new line, it will be added to the buffer, and the decoder will start from the start of the incomplete group. After scanning to the end of the buffer, it will return the same position, because the group is still incomplete. And this will continue with the next line and growing buffer.

    So, there is a great pitfall if implement and use an incremental decoder. Also, implementing incremental decoders will significantly increase the API in base64 and binascii modules.

    Other way -- reuse the binascii.Incomplete exception. It was only used in the BinHex 4 (HQX) decoder, and after removing it remains unused. We can make it a subclass of binascii.Error for compatibility and start raising for incomplete groups. Not limited by the incremental decoder interface, we can encode the current state (the number and value of alphabetic characters in the incomplete group, the number of padding character), so the next chunk of input can be processed without accumulating in the unlimitedly growing buffer. And this will not grow the API. We should also add the new final parameter (true by default). If it is false, then Incomplete will raise for padded group, because the pad character is ignored if not at the end of the input.

    This will also allow to implement the correct incremental decoder, although it will still be vulnerable to the quadratic complexity attack.

  7. gpshead commented on Mar 31, 2026

    MemberAuthor

    i'm not sure if this issue is important - it was filed for someone based on a mailing list conversation but it hasn't come up since. my inclination is to just close it as not planned. it can be re-opened if there is a need.

    In practice I don't believe people insert extraneous characters between sets of 4 that decode to 3 bytes for base64 encoded data. And if that has been done with line breaks... the workaround anyone would've done is to just split and join first.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.7 (EOL)end of lifestdlibStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL