FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

gh-157710: Soft deprecate C API modifying str objects (#157711) · python/cpython@84b0669 · GitHub

Repository navigation

Commit 84b0669

Browse files
gh-157710: Soft deprecate C API modifying str objects (#157711)
Soft deprecate PyUnicode_New(), PyUnicode_CopyCharacters(), PyUnicode_Fill(), PyUnicode_Resize(), PyUnicode_WRITE() and PyUnicode_WriteChar() functions. Use the PyUnicodeWriter API instead. * Add tests on PyUnicode_New() and PyUnicode_Resize(). Check that the result is either a mutable string or the empty string singleton. * Check that PyUnicode_Fill(), PyUnicode_CopyCharacters() and PyUnicode_WriteChar() fail to modify a string with 2 references. Co-authored-by: Peter Bierma <zintensitydev@gmail.com>
1 parent 182f323 commit 84b0669

6 files changed

Lines changed: 138 additions & 22 deletions

File tree

‎Doc/c-api/unicode.rst‎

Lines changed: 40 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -168,8 +168,14 @@ access to internal read-only data of Unicode objects:
168168
The function performs no checks for any of its requirements,
169169
and is intended for usage in loops.
170170
171+
While :class:`str` objects are usually immutable in Python, this special C API allows
172+
mutating a fresh :class:`str` object if the string has not been "used" yet.
173+
171174
.. versionadded:: 3.3
172175
176+
.. soft-deprecated:: next
177+
Use the :c:type:`PyUnicodeWriter` API instead.
178+
173179
174180
.. c:function:: Py_UCS4 PyUnicode_READ(int kind, void *data, Py_ssize_t index)
175181
@@ -407,9 +413,15 @@ APIs:
407413
using the :c:type:`PyUnicodeWriter` API, or one of the ``PyUnicode_From*``
408414
functions below.
409415
416+
While :class:`str` objects are usually immutable in Python, this special C API
417+
returns a :class:`str` object that can be mutated, except if *size* is zero, in which
418+
case it returns the immutable empty string constant.
410419
411420
.. versionadded:: 3.3
412421
422+
.. soft-deprecated:: next
423+
Use the :c:type:`PyUnicodeWriter` API instead.
424+
413425
414426
.. c:function:: PyObject* PyUnicode_FromKindAndData(int kind, const void *buffer, \
415427
Py_ssize_t size)
@@ -754,11 +766,16 @@ APIs:
754766
possible. Returns ``-1`` and sets an exception on error, otherwise returns
755767
the number of copied characters.
756768
757-
The string must not have been “used” yet.
769+
While :class:`str` objects are usually immutable in Python, this special C API allows
770+
mutating a fresh :class:`str` object if the string has not been "used" yet.
771+
758772
See :c:func:`PyUnicode_New` for details.
759773
760774
.. versionadded:: 3.3
761775
776+
.. soft-deprecated:: next
777+
Use the :c:type:`PyUnicodeWriter` API instead.
778+
762779
763780
.. c:function:: int PyUnicode_Resize(PyObject **unicode, Py_ssize_t length);
764781
@@ -774,6 +791,14 @@ APIs:
774791
The function doesn't check string content, the result may not be a
775792
string in canonical representation.
776793
794+
While :class:`str` objects are usually immutable in Python, this special C API
795+
can resize a :class:`str` object in-place if the string has not been "used" yet.
796+
It returns a :class:`str` object which can be mutated, except if *size* is zero, in
797+
which case it returns the immutable empty string constant.
798+
799+
.. soft-deprecated:: next
800+
Use the :c:type:`PyUnicodeWriter` API instead.
801+
777802
778803
.. c:function:: Py_ssize_t PyUnicode_Fill(PyObject *unicode, Py_ssize_t start, \
779804
Py_ssize_t length, Py_UCS4 fill_char)
@@ -784,14 +809,19 @@ APIs:
784809
Fail if *fill_char* is bigger than the string maximum character, or if the
785810
string has more than 1 reference.
786811
787-
The string must not have been “used” yet.
788-
See :c:func:`PyUnicode_New` for details.
789-
790812
Return the number of written characters, or return ``-1`` and raise an
791813
exception on error.
792814
815+
While :class:`str` objects are usually immutable in Python, this special C API allows
816+
mutating a fresh :class:`str` object if the string has not been "used" yet.
817+
818+
See :c:func:`PyUnicode_New` for details.
819+
793820
.. versionadded:: 3.3
794821
822+
.. soft-deprecated:: next
823+
Use the :c:type:`PyUnicodeWriter` API instead.
824+
795825
796826
.. c:function:: int PyUnicode_WriteChar(PyObject *unicode, Py_ssize_t index, \
797827
Py_UCS4 character)
@@ -804,11 +834,16 @@ APIs:
804834
See :c:func:`PyUnicode_WRITE` for a version that skips these checks,
805835
making them your responsibility.
806836
807-
The string must not have been “used” yet.
837+
While :class:`str` objects are usually immutable in Python, this special C API allows
838+
mutating a fresh :class:`str` object if the string has not been "used" yet.
839+
808840
See :c:func:`PyUnicode_New` for details.
809841
810842
.. versionadded:: 3.3
811843
844+
.. soft-deprecated:: next
845+
Use the :c:type:`PyUnicodeWriter` API instead.
846+
812847
813848
.. c:function:: Py_UCS4 PyUnicode_ReadChar(PyObject *unicode, Py_ssize_t index)
814849

‎Doc/whatsnew/3.16.rst‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1224,6 +1224,13 @@ Deprecated C APIs
12241224
:c:func:`PyModule_GetFilenameObject` instead is still recommended.
12251225
(Contributed by Victor Stinner in :gh:`154757`.)
12261226

1227+
* Soft deprecate functions modifying Unicode strings:
1228+
:c:func:`PyUnicode_New`, :c:func:`PyUnicode_CopyCharacters`,
1229+
:c:func:`PyUnicode_Fill`, :c:func:`PyUnicode_Resize`,
1230+
:c:func:`PyUnicode_WRITE` and :c:func:`PyUnicode_WriteChar`.
1231+
Use the safer :c:type:`PyUnicodeWriter` API instead.
1232+
(Contributed by Victor Stinner in :gh:`157710`.)
1233+
12271234
.. Add C API deprecations above alphabetically, not here at the end.
12281235
12291236
.. include:: ../deprecations/c-api-pending-removal-in-3.18.rst

‎Lib/test/test_capi/test_unicode.py‎

Lines changed: 38 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -25,6 +25,7 @@
2525
# Maximum invalid character which fits into 32-bit Py_UCS4
2626
MAX_INVALID_CHAR = 0xFFFF_FFFF
2727
NULL = None
28+
USED_STR_ERROR = 'Cannot modify a string currently used'
2829

2930
class Str(str):
3031
pass
@@ -76,9 +77,27 @@ def test_checkexact(self):
7677
# Test PyUnicode_CheckExact()
7778
self._test_check(_testlimitedcapi.unicode_checkexact, exact=True)
7879

80+
def assert_is_mutable(self, result, refcnt):
81+
# Check that result is a "mutable" Unicode string
82+
self.assertEqual(refcnt, 1)
83+
self.assertFalse(sys._is_immortal(result))
84+
85+
def assert_is_empty_singleton(self, result):
86+
# Check that result is the empty string singleton
87+
self.assertEqual(result, '')
88+
self.assertTrue(sys._is_immortal(result))
89+
7990
def test_new(self):
8091
"""Test PyUnicode_New()"""
81-
new = _testcapi.unicode_new
92+
_unicode_new = _testcapi.unicode_new
93+
94+
def new(size, maxchar):
95+
result = _unicode_new(size, maxchar)
96+
if size != 0:
97+
self.assert_is_mutable(result, sys.getrefcount(result))
98+
else:
99+
self.assert_is_empty_singleton(result)
100+
return result
82101

83102
for maxchar in 0, 0x61, 0xa1, 0x4f60, 0x1f600, 0x10ffff:
84103
self.assertEqual(new(0, maxchar), '')
@@ -123,6 +142,10 @@ def test_fill(self):
123142
self.assertEqual(fill(to, start, length, fill_char),
124143
(expected, filled))
125144

145+
# A string with 2 references cannot be modified
146+
with self.assertRaisesRegex(SystemError, USED_STR_ERROR):
147+
fill('abc', 0, 3, ord('x'), incref=True)
148+
126149
s = strings[0]
127150
self.assertRaises(IndexError, fill, s, -1, 0, 0x78)
128151
self.assertRaises(IndexError, fill, s, PY_SSIZE_T_MIN, 0, 0x78)
@@ -162,7 +185,12 @@ def _test_writechar(self, writechar, *, check):
162185

163186
def test_writechar(self):
164187
"""Test PyUnicode_WriteChar()"""
165-
self._test_writechar(_testlimitedcapi.unicode_writechar, check=True)
188+
writechar = _testlimitedcapi.unicode_writechar
189+
self._test_writechar(writechar, check=True)
190+
191+
# A string with 2 references cannot be modified
192+
with self.assertRaisesRegex(SystemError, USED_STR_ERROR):
193+
writechar('abc', 1, ord('x'), incref=True)
166194

167195
def test_write_macro(self):
168196
"""Test PyUnicode_WRITE()"""
@@ -187,18 +215,16 @@ def resize(s, length, new=True, compute_hash=False):
187215
self.assertFalse(is_new_obj)
188216
elif length == 0:
189217
# Get the empty Unicode string
190-
self.assertEqual(result, '')
191-
self.assertTrue(sys._is_immortal(result))
218+
self.assert_is_empty_singleton(result)
192219
self.assertTrue(is_new_obj)
193220
elif (not new) or compute_hash:
194221
# Get a fresh copy
195-
self.assertEqual(refcnt, 1)
222+
self.assert_is_mutable(result, refcnt)
196223
self.assertTrue(is_new_obj)
197-
self.assertFalse(sys._is_immortal(result))
198224
else:
199225
# In-size replace can return the same address, or not.
200226
# So 'is_new_obj' cannot be tested.
201-
self.assertFalse(sys._is_immortal(result))
227+
self.assert_is_mutable(result, refcnt)
202228

203229
return result
204230

@@ -1793,6 +1819,11 @@ def test_copycharacters(self):
17931819
self.assertRaises(SystemError, unicode_copycharacters, s, 0, s, 0, PY_SSIZE_T_MIN)
17941820
self.assertRaises(SystemError, unicode_copycharacters, s, 0, b'', 0, 0)
17951821
self.assertRaises(SystemError, unicode_copycharacters, s, 0, [], 0, 0)
1822+
1823+
# A string with 2 references cannot be modified
1824+
with self.assertRaisesRegex(SystemError, USED_STR_ERROR):
1825+
unicode_copycharacters('abc', 0, 'abc', 0, 1, incref=True)
1826+
17961827
# CRASHES unicode_copycharacters(s, 0, NULL, 0, 0)
17971828
# TODO: Test PyUnicode_CopyCharacters() with non-unicode and
17981829
# non-modifiable unicode as "to".
Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
Soft deprecate functions modifying Unicode strings:
2+
:c:func:`PyUnicode_New`, :c:func:`PyUnicode_CopyCharacters`,
3+
:c:func:`PyUnicode_Fill`, :c:func:`PyUnicode_Resize`,
4+
:c:func:`PyUnicode_WRITE` and :c:func:`PyUnicode_WriteChar`. Use
5+
the :c:type:`PyUnicodeWriter` API instead. Patch by Victor Stinner.

‎Modules/_testcapi/unicode.c‎

Lines changed: 33 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -59,13 +59,20 @@ unicode_copy(PyObject *unicode)
5959

6060
/* Test PyUnicode_Fill() */
6161
static PyObject *
62-
unicode_fill(PyObject *self, PyObject *args)
62+
unicode_fill(PyObject *self, PyObject *args, PyObject *kwargs)
6363
{
64+
static char *kwlist[] = {"to", "start", "length", "fill_char",
65+
"incref", NULL};
6466
PyObject *to, *to_copy;
6567
Py_ssize_t start, length, filled;
6668
unsigned int fill_char;
69+
int incref = 0;
6770

68-
if (!PyArg_ParseTuple(args, "OnnI", &to, &start, &length, &fill_char)) {
71+
if (!PyArg_ParseTupleAndKeywords(args, kwargs,
72+
"OnnI|p", kwlist,
73+
&to, &start, &length,
74+
&fill_char, &incref))
75+
{
6976
return NULL;
7077
}
7178

@@ -74,7 +81,14 @@ unicode_fill(PyObject *self, PyObject *args)
7481
return NULL;
7582
}
7683

84+
if (incref) {
85+
Py_INCREF(to_copy);
86+
}
7787
filled = PyUnicode_Fill(to_copy, start, length, (Py_UCS4)fill_char);
88+
if (incref) {
89+
Py_DECREF(to_copy);
90+
}
91+
7892
if (filled == -1 && PyErr_Occurred()) {
7993
Py_DECREF(to_copy);
8094
return NULL;
@@ -190,13 +204,18 @@ unicode_asutf8(PyObject *self, PyObject *args)
190204

191205
/* Test PyUnicode_CopyCharacters() */
192206
static PyObject *
193-
unicode_copycharacters(PyObject *self, PyObject *args)
207+
unicode_copycharacters(PyObject *self, PyObject *args, PyObject *kwargs)
194208
{
209+
static char *kwlist[] = {"to", "to_start", "from", "from_start",
210+
"howmany", "incref", NULL};
195211
PyObject *from, *to, *to_copy;
196212
Py_ssize_t from_start, to_start, how_many, copied;
213+
int incref = 0;
197214

198-
if (!PyArg_ParseTuple(args, "UnOnn", &to, &to_start,
199-
&from, &from_start, &how_many)) {
215+
if (!PyArg_ParseTupleAndKeywords(args, kwargs,
216+
"UnOnn|p", kwlist,
217+
&to, &to_start, &from, &from_start,
218+
&how_many, &incref)) {
200219
return NULL;
201220
}
202221

@@ -210,8 +229,15 @@ unicode_copycharacters(PyObject *self, PyObject *args)
210229
return NULL;
211230
}
212231

232+
if (incref) {
233+
Py_INCREF(to_copy);
234+
}
213235
copied = PyUnicode_CopyCharacters(to_copy, to_start, from,
214236
from_start, how_many);
237+
if (incref) {
238+
Py_DECREF(to_copy);
239+
}
240+
215241
if (copied == -1 && PyErr_Occurred()) {
216242
Py_DECREF(to_copy);
217243
return NULL;
@@ -856,12 +882,12 @@ static PyType_Spec Writer_spec = {
856882

857883
static PyMethodDef TestMethods[] = {
858884
{"unicode_new", unicode_new, METH_VARARGS},
859-
{"unicode_fill", unicode_fill, METH_VARARGS},
885+
{"unicode_fill", _PyCFunction_CAST(unicode_fill), METH_VARARGS | METH_KEYWORDS},
860886
{"unicode_fromkindanddata", unicode_fromkindanddata, METH_VARARGS},
861887
{"unicode_asucs4", unicode_asucs4, METH_VARARGS},
862888
{"unicode_asucs4copy", unicode_asucs4copy, METH_VARARGS},
863889
{"unicode_asutf8", unicode_asutf8, METH_VARARGS},
864-
{"unicode_copycharacters", unicode_copycharacters, METH_VARARGS},
890+
{"unicode_copycharacters", _PyCFunction_CAST(unicode_copycharacters), METH_VARARGS | METH_KEYWORDS},
865891
{"unicode_GET_CACHED_HASH", unicode_GET_CACHED_HASH, METH_O},
866892
{"test_py_identifier", test_py_identifier, METH_NOARGS},
867893
{"corrupt_unicode", corrupt_unicode, METH_VARARGS},

‎Modules/_testlimitedcapi/unicode.c‎

Lines changed: 15 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -140,14 +140,19 @@ unicode_copy(PyObject *unicode)
140140

141141
/* Test PyUnicode_WriteChar() */
142142
static PyObject *
143-
unicode_writechar(PyObject *self, PyObject *args)
143+
unicode_writechar(PyObject *self, PyObject *args, PyObject *kwargs)
144144
{
145+
static char *kwlist[] = {"to", "index", "character", "incref", NULL};
145146
PyObject *to, *to_copy;
146147
Py_ssize_t index;
147148
unsigned int character;
148149
int result;
150+
int incref = 0;
149151

150-
if (!PyArg_ParseTuple(args, "OnI", &to, &index, &character)) {
152+
if (!PyArg_ParseTupleAndKeywords(args, kwargs,
153+
"OnI|p", kwlist,
154+
&to, &index, &character, &incref))
155+
{
151156
return NULL;
152157
}
153158

@@ -156,7 +161,14 @@ unicode_writechar(PyObject *self, PyObject *args)
156161
return NULL;
157162
}
158163

164+
if (incref) {
165+
Py_INCREF(to_copy);
166+
}
159167
result = PyUnicode_WriteChar(to_copy, index, (Py_UCS4)character);
168+
if (incref) {
169+
Py_DECREF(to_copy);
170+
}
171+
160172
if (result == -1 && PyErr_Occurred()) {
161173
Py_DECREF(to_copy);
162174
return NULL;
@@ -1915,7 +1927,7 @@ static PyMethodDef TestMethods[] = {
19151927
test_unicode_compare_with_ascii, METH_NOARGS},
19161928
{"test_string_from_format", test_string_from_format, METH_NOARGS},
19171929
{"test_widechar", test_widechar, METH_NOARGS},
1918-
{"unicode_writechar", unicode_writechar, METH_VARARGS},
1930+
{"unicode_writechar", _PyCFunction_CAST(unicode_writechar), METH_VARARGS | METH_KEYWORDS},
19191931
{"unicode_resize", unicode_resize, METH_VARARGS},
19201932
{"unicode_resize_null", unicode_resize_null, METH_VARARGS},
19211933
{"unicode_append", unicode_append, METH_VARARGS},

0 commit comments

Comments
 (0)

Back | FazBrowse Home | New Git URL