| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
anyio's lease-counted runner creates a fresh event loop for every async test when only function-scoped async fixtures exist. On Windows each loop's self-pipe is an emulated loopback-TCP socketpair; churning ~2,700 of those across a ~30s run transiently exhausts kernel socket buffers and flakes CI with WinError 10055 raised from asyncio.new_event_loop() before an arbitrary test's body starts, plus a collateral "unclosed event loop" failure pinned on the next test by filterwarnings=error. Hold a module-scoped autouse runner lease so each xdist worker reuses one loop per module. Measured on windows-latest at the flaked commit: 2,737 -> ~380 socketpair calls per run, peak TIME_WAIT 4,164 -> 400, identical test outcomes on both platforms. Modules that parametrize anyio_backend or call trio.run directly shadow the lease with a sync no-op: a module-scoped lease cannot depend on the function-scoped parameter (ScopeMismatch), and a held asyncio loop's wakeup fd breaks direct trio runs on Windows.
There was a problem hiding this comment.
I didn't find any bugs — the fixture and the six shadow modules line up exactly with every module that parametrizes anyio_backend or calls trio.run directly. Deferring to a human because this establishes a new suite-wide testing convention (module-shared event loops + the shadow-fixture requirement in AGENTS.md) that maintainers should sign off on.
Extended reasoning...This PR is test-infrastructure only: a new module-scoped, autouse _module_runner_lease fixture in tests/conftest.py (plus a session-scoped anyio_backend), sync no-op shadows of that fixture in six test modules, and a documentation update in AGENTS.md describing the new convention. No files under src/ are touched.
I cross-checked that the six modules given shadow fixtures are exactly the set of test modules that parametrize anyio_backend or call trio.run directly — no module is missed, and no shadow is superfluous. A forgotten shadow in a future parametrizing module fails loudly with ScopeMismatch, as the PR description and AGENTS.md note. The bug-hunting pass found no issues.
None. The change does not touch runtime code, authentication, transport, or any security-sensitive path; it only alters how the test suite allocates event loops.
The blast radius is limited to CI/test behavior: worst case is new test flakiness or order-dependence within a module, which CI would surface rather than shipping to users. However, the PR changes the isolation semantics of every async test in the suite (loop shared per module) and introduces a new convention contributors must follow (the shadow requirement, one of whose failure modes is only a Windows-specific flake per the direct-trio case). That is a project-wide testing-convention decision that maintainers should deliberately accept, so I'm not shadow-approving despite finding no defects.
The author reports full-suite runs on both windows-latest and ubuntu-latest at the flaked SHA with identical outcomes and measured socketpair-churn reduction, plus an audit for loop-scoped state dependence. Those claims can't be independently re-verified here, but the design and documentation are internally consistent.
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Share one event loop per test module instead of creating a fresh one for every async test, to stop a Windows CI flake caused by event-loop churn.
Motivation and Context
test (3.12, locked, windows-latest) flaked with two failures that are really one event:
The churn behind it: anyio's runner is lease-counted, so with only function-scoped async fixtures every async test creates and destroys its own event loop (anyio docs, agronholm/anyio#686). This suite has ~2,740 async tests and runs them in ~30s across 4 xdist workers on Windows CI — hundreds of loopback socketpair create/close cycles per second, each parking a connection in TIME_WAIT for ~120s. WinError 10055 is a transient kernel buffer failure under that churn; each loop creation is a lottery ticket and a Windows job currently buys ~2,740 of them (measured: exactly one socketpair() call per async test).
The fix holds a module-scoped anyio runner lease (autouse fixture in tests/conftest.py), so each xdist worker reuses one loop per module. Measured on the flaked SHA with a socket.socketpair counter: 2,737 → ~380 socketpair calls per run on windows-latest, peak TIME_WAIT 4,164 → 400, with identical test outcomes on both windows-latest and ubuntu-latest.
Six modules parametrize anyio_backend (trio + MockClock, SelectorEventLoop) and shadow the lease fixture with a sync no-op: a module-scoped lease can't depend on the function-scoped parameter (pytest fails with ScopeMismatch — loud, so a future module that forgets the shadow can't misbehave silently), the held runner wouldn't match the requested backend anyway, and tests/client/test_stdio.py's direct trio.run(...) calls collide with a lingering asyncio loop's wakeup fd on Windows. When these modules run, the previous module's runner has already been torn down, so their behavior is unchanged.
Session scope would cut churn further (~10 socketpairs/run) but was rejected: it breaks all anyio_backend-parametrized tests with ScopeMismatch and the direct-trio tests on Windows via the wakeup-fd collision (both verified empirically).
How Has This Been Tested?
Breaking Changes
None — test infrastructure only. New convention for contributors: a test module that parametrizes anyio_backend must shadow _module_runner_lease (documented in the fixture's docstring in tests/conftest.py; forgetting it fails loudly at setup).
Types of changes
Checklist
Additional context
Tests in the same module now share loop-scoped state. An audit found no bare asyncio.create_task, set_event_loop, exception-handler, or executor usage anywhere in tests/ or src/mcp/, so nothing depends on per-test loop isolation today; a future test leaking a background task would surface as order-dependence within its own module.
The residual ~380 loop creations (the six opt-out modules plus one loop per module per worker) keep a small tail of exposure. If the flake ever reappears, the next increment is moving the ~12 parametrized/direct-trio tests into small sibling modules so the two big opted-out files (test_jsonrpc_dispatcher.py, test_streamable_http_modern.py) rejoin the shared loop.
AI Disclaimer