FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

gh-119592: gh-152967: Fix ProcessPoolExecutor stranding submitted work when a max_tasks_per_child worker exits by gpshead · Pull Request #152978 · python/cpython · GitHub

/ cpython Public

gh-119592: gh-152967: Fix ProcessPoolExecutor stranding submitted work when a max_tasks_per_child worker exits - #152978

Merged
gpshead merged 1 commit into
python:mainfrom
gpshead:fix-gh-119592-worker-replacement
Jul 8, 2026
Merged

gh-119592: gh-152967: Fix ProcessPoolExecutor stranding submitted work when a max_tasks_per_child worker exits#152978
gpshead merged 1 commit into
python:mainfrom
gpshead:fix-gh-119592-worker-replacement

Conversation

gpshead commented Jul 3, 2026
edited
Loading

Copy link
Copy Markdown
Member

Fix ProcessPoolExecutor stranding submitted work when a max_tasks_per_child worker exits. Worker replacement went through the executor object: the manager thread read executor attributes that shutdown(wait=False) clears concurrently, and could not replace workers at all once the executor was garbage collected. A worker exiting at its max_tasks_per_child limit in those states left the remaining submitted work permanently unexecuted and hung interpreter exit; the racing case could crash the manager thread.

Replace workers from the executor manager thread using its own state plus configuration read through the live executor weakref, which shutdown() never clears:


Drafted and investigated entirely by Claude Fable 5 based on the issues. edit: Looks to be in good shape. Undrafting. (I originally put this up as a draft to better iterate on review to see what shape this should take and how feasible backporting this further as a bugfix could be).

…n a max_tasks_per_child worker exits

Worker replacement went through the executor object: the manager thread
read executor attributes that shutdown(wait=False) clears concurrently,
and could not replace workers at all once the executor was garbage
collected. A worker exiting at its max_tasks_per_child limit in those
states left the remaining submitted work permanently unexecuted and hung
interpreter exit; the racing case could crash the manager thread.

Replace workers from the executor manager thread using its own state
plus configuration read through the live executor weakref, which
shutdown() never clears:

- After shutdown(wait=False) with the executor still referenced, a
  replacement is spawned and the remaining work is executed as
  documented.
- Once the executor has been garbage collected (pythongh-152967), or a
  replacement worker cannot be started and no workers remain, the
  remaining futures now fail with BrokenProcessPool instead of never
  resolving.
- A new _force_shutting_down flag stops both spawn paths from starting
  workers that would escape terminate_workers()/kill_workers().

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
self._join_executor_internals(broken=True)

def terminate_broken(self, cause):
def terminate_broken(self, cause, bpe_message=None):

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Choose a reason Spam Abuse Off Topic Outdated Duplicate Resolved Low Quality

While this might look like a public API change in the diff... it's on the _ExecutorManagerThread internal use only class. Fine to backport.

gpshead marked this pull request as ready for review July 4, 2026 05:04
gpshead added needs backport to 3.14 bugs and security fixes 🔨 test-with-buildbots Test PR w/ buildbots; report in status section labels Jul 4, 2026

Copy link
Copy Markdown

🤖 New build scheduled with the buildbot fleet by @gpshead for commit fd234c9 🤖

Results will be shown at:

https://buildbot.python.org/all/#/grid?branch=refs%2Fpull%2F152978%2Fmerge

If you want to schedule another build, you need to add the 🔨 test-with-buildbots label again.

bedevere-bot removed the 🔨 test-with-buildbots Test PR w/ buildbots; report in status section label Jul 8, 2026
gpshead merged commit 0c6422f into python:main Jul 8, 2026
138 of 143 checks passed

Copy link
Copy Markdown

Thanks @gpshead for the PR 🌮🎉.. I'm working now to backport this PR to: 3.14, 3.15.
🐍🍒⛏🤖

Copy link
Copy Markdown

Sorry, @gpshead, I could not cleanly backport this to 3.14 due to a conflict.
Please backport using cherry_picker on command line.

cherry_picker 0c6422ff6a13ae309493fb7a358cb35d7ea959c8 3.14

bedevere-app Bot commented Jul 8, 2026

Copy link
Copy Markdown

GH-153363 is a backport of this pull request to the 3.15 branch.

bedevere-app Bot removed the needs backport to 3.15 pre-release feature fixes, bugs and security fixes label Jul 8, 2026
gpshead added a commit that referenced this pull request Jul 8, 2026
…ted work when a max_tasks_per_child worker exits (GH-152978) (#153363)

gh-119592: gh-152967: Fix ProcessPoolExecutor stranding submitted work when a max_tasks_per_child worker exits (GH-152978)

gh-119592: Fix ProcessPoolExecutor stranding submitted work when a max_tasks_per_child worker exits

Worker replacement went through the executor object: the manager thread
read executor attributes that shutdown(wait=False) clears concurrently,
and could not replace workers at all once the executor was garbage
collected. A worker exiting at its max_tasks_per_child limit in those
states left the remaining submitted work permanently unexecuted and hung
interpreter exit; the racing case could crash the manager thread.

Replace workers from the executor manager thread using its own state
plus configuration read through the live executor weakref, which
shutdown() never clears:

- After shutdown(wait=False) with the executor still referenced, a
  replacement is spawned and the remaining work is executed as
  documented.
- Once the executor has been garbage collected (gh-152967), or a
  replacement worker cannot be started and no workers remain, the
  remaining futures now fail with BrokenProcessPool instead of never
  resolving.
- A new _force_shutting_down flag stops both spawn paths from starting
  workers that would escape terminate_workers()/kill_workers().
(cherry picked from commit 0c6422f)



Reviewed-multiple-times-by: Gregory P. Smith

Co-authored-by: Gregory P. Smith <68491+gpshead@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

Copy link
Copy Markdown
Member

Please don't forget about backports.

bedevere-app Bot commented Jul 29, 2026

Copy link
Copy Markdown

GH-154873 is a backport of this pull request to the 3.14 branch.

bedevere-app Bot removed the needs backport to 3.14 bugs and security fixes label Jul 29, 2026
hugovk added a commit that referenced this pull request Jul 30, 2026
…ted work when a max_tasks_per_child worker exits (GH-152978) (#154873)

Co-authored-by: Gregory P. Smith <68491+gpshead@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants


Back | FazBrowse Home | New Git URL