| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
|
Tagging subscribers to this area: @dotnet/ncl |
Sorry, something went wrong.
|
Reliability problem for services, we should fix it in both 9.0.x and 8.0.x. Marking for servicing. |
Sorry, something went wrong.
|
Servicing approved by Steve Carroll on Jan 8th via email. |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Backport of #110744 to release/9.0-staging
Fixes #110598
/cc @MihaZupan
Customer Impact
HttpClient is often used for server-to-server communication. If a service experiences an outage, requests to that service will fail as expected, but HttpClient is expected to eventually recover once the service becomes available again.
Due to a race condition, HttpClient's connection pool may become stuck, preventing new connections to the other server from being established.
To recover, users must restart the process in the service that didn't have an outage in the first place.
We've heard from two internal services (one of them reported in #110598) where this issue led to prolonged recovery times after an outage in Azure.
Regression
Yes - introduced in .NET 6 (where we rewrote HTTP connection pool).
Testing
Added a targeted test (also into CI) that reliably reproduces the stuck connection pool state.
Further manual validation was performed to make sure the problem is fully addressed.
Risk
Low.
There are logically 3 places where we use some state, and 2 of them were already using a lock. This change is limited in scope to effectively update the 3rd place to do the same.