Proposed new feature or change:
Follow-up to Sebastian’s review comment on #32757.
NumPy currently infers Unicode for mixed bytes/text inputs:
>>> import numpy as np
>>> np.asarray([b"a", "b"])
array(['a', 'b'], dtype='<U1')
>>> np.concatenate(([b"a", "b"], ["c"]))
array(['a', 'b', 'c'], dtype='<U1')
We should add an opt-in mode (exposed to NumPy Python internals via _array_converter) that detects these mixtures during dtype discovery, before coercion happens and we lose track of the input dtypes. Shared C conversion machinery would let concatenate detect mixtures within its inputs as well as between them.
The same detection could support deprecating implicit Unicode inference for mixed bytes/text inputs. These should eventually produce object arrays that preserve the original values. Explicit conversions such as dtype="U" would remain supported.
This has come up in the past, in particular:
- #15327 proposed deprecating implicit string/non-string promotion, including in concatenate.
- #19078 described downstream difficulties with promotion warnings: callers wanted numeric inference for numeric inputs and object arrays for mixed inputs.
- #19101 explored an opt-in object-fallback mode during dtype discovery; it was not merged.
Reactions are currently unavailable
Proposed new feature or change:
Follow-up to Sebastian’s review comment on #32757.
NumPy currently infers Unicode for mixed bytes/text inputs:
We should add an opt-in mode (exposed to NumPy Python internals via _array_converter) that detects these mixtures during dtype discovery, before coercion happens and we lose track of the input dtypes. Shared C conversion machinery would let concatenate detect mixtures within its inputs as well as between them.
The same detection could support deprecating implicit Unicode inference for mixed bytes/text inputs. These should eventually produce object arrays that preserve the original values. Explicit conversions such as dtype="U" would remain supported.
This has come up in the past, in particular: