| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
The Ubuntu ThreadSanitizer job on Linux was hitting the 30-minute timeout ~45-75% of runs in a silent hang (no TSan report, no output progress). Reproduced locally in WSL Ubuntu 24.04 with the exact same packages as CI. Root cause ---------- JSC's concurrent garbage collector on Linux suspends each mutator thread at GC safepoints using a SIGUSR1-based protocol: Collector Thread Mutator Thread N ---------------- ---------------- pthread_kill(N, SIGUSR1) -> (signal handler runs, sem_post) sem_wait(sem) <- (handler returns) Under ThreadSanitizer, signal delivery is intercepted and serialized. When the mutator is inside an instrumented section, TSan defers the SIGUSR1 handler indefinitely. The Collector Thread's sem_wait then blocks forever, hanging the whole process. Confirmed with a gdb capture of a hung inferior: Thread 33 "ollector Thread": BabylonJS#5 __interceptor_sem_wait BabylonJS#6 WTF::Thread::suspend(WTF::ThreadSuspendLocker const&) BabylonJS#7-BabylonJS#21 [JSC GC stop-the-world path] Thread 3 "UnitTests": <pending SIGUSR1> (never delivered) macOS JSC uses Mach thread_suspend() rather than Unix signals, which is why the macOS TSan job has been passing in ~2.7 min the whole time. Fix --- Set JSC_useConcurrentGC=0 for the Ubuntu_ThreadSanitizer job only. This removes the dedicated Collector Thread; GC runs on the mutator without any cross-thread signaling. Also revert the 30-min timeout bump — with the hang fixed the Linux TSan job should finish in roughly the same time as macOS TSan (~3 min). Verification ------------ - Default (concurrent GC on) : 9-11 hangs per 20 runs - JSC_useConcurrentGC=0 : 0 hangs per 30 runs [Created by Copilot on behalf of @bghgary] Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Captured a full backtrace via lldb -k (see prior CI commit). The smoking gun: frame BabylonJS#5: ___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED frame BabylonJS#6: napi_delete_reference at hermes_napi_reference.cpp:113 frame BabylonJS#7: Napi::Reference<...>::~Reference at napi-inl.h:3262 frame BabylonJS#8: Napi::ObjectReference::~ObjectReference frame BabylonJS#10: Babylon::Polyfills::Internal::URL::~URL at URL.h:10 frame BabylonJS#12: Napi::ObjectWrap<...URL>::FinalizeCallback at napi-inl.h:4963 frame BabylonJS#13: napi_env__::shutdown at hermes_napi.cpp:214 frame BabylonJS#16: hermes::vm::Runtime::~Runtime Root cause: Hermes's napi_env__::shutdown() iterates refListHead_ and delete ref; one at a time. It only sets ref->deletionPending_ on the *current* ref before its finalize_cb fires. If the finalizer transitively destroys a node-addon-api wrapper (Napi::Reference / Napi::ObjectReference) whose underlying napi_ref was already deleted earlier in the same loop, napi_delete_reference reads ref->deletionPending_ from freed memory and proceeds to delete ref again -> double-free. The exact path: URL (an ObjectWrap subclass) has a Napi::ObjectReference member m_searchParamsReference. addReference prepends to the linked list, so m_searchParamsReference's ref is processed BEFORE URL's wrap ref. When URL's wrap finalizer runs delete this, ~URL destroys m_searchParamsReference, whose destructor calls napi_delete_reference on the already-freed sibling ref. macOS libmalloc detects this (malloc: *** error for object 0x...: pointer being freed was not allocated -> SIGABRT). Linux glibc and Windows CRT happen to miss it. Fix: PATCH Hermes shutdown() to mark ALL refs deletionPending in a pre-pass BEFORE iterating. Apply as a FetchContent PATCH_COMMAND via the new ApplyPatchIfNeeded.cmake helper (idempotent — uses git apply --check --reverse to detect already-applied state across reconfigure). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
| Back | FazBrowse Home | New Git URL |
No description provided.