FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Don't attach dbus-python to the GLib main loop (fixes random SIGABRT in dbus_watch_handle) by jslay88 · Pull Request #634 · StreamController/StreamController · GitHub

Don't attach dbus-python to the GLib main loop (fixes random SIGABRT in dbus_watch_handle) - #634

Merged
Core447 merged 5 commits into
StreamController:mainfrom
jslay88:fix-dbus-mainloop-crash
Aug 6, 2026
Merged

Core447 merged 5 commits into
StreamController:mainfrom
jslay88:fix-dbus-mainloop-crash

Conversation

jslay88 commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

Problem

StreamController dies with SIGABRT (exit 134) after anywhere from hours to days of uptime. Nothing lands in logs.log because the abort happens in native code. The journal has it:

dbus[3]: arguments to dbus_watch_handle() were incorrect, assertion "watch != NULL" failed in file ../dbus/dbus-watch.c line 738.
This is normally a bug in some application using the D-Bus library.
/app/bin/launch.sh: line 3:     3 Aborted                 (core dumped) python3 /app/bin/StreamController/main.py "$@"
app-com.core447.StreamController@autostart.service: Main process exited, code=exited, status=134/n/a
app-flatpak-com.core447.StreamController-2233223863.scope: Consumed 1h 5min 6.185s CPU time over 5d 4h 38min 11.318s wall clock time

Stack from the core (1.5.0-beta.15 flatpak, KDE/Wayland):

#3  libdbus-1.so.3.38.3            <- assertion + abort
#4  libdbus-1.so.3.38.3
#5  libdbus-1.so.3.38.3            <- dbus_watch_handle
#6  _dbus_glib_bindings...so       <- io_handler_dispatch
#7  libglib-2.0.so.0               <- g_main_dispatch
#8  libglib-2.0.so.0
#9  libglib-2.0.so.0
#10 libgio-2.0.so.0
#11 libffi.so.8

Cause

dbus-python's GLib integration is not thread safe, which is longstanding and documented by the dbus maintainer ("dbus-glib ... makes no attempt to be thread-safe"). In dbus-gmain.c:

static void
watch_toggled (DBusWatch *watch, void *data)
{
  /* Because we just exit on OOM, enable/disable is
   * no different from add/remove
   */
  if (dbus_watch_get_enabled (watch))
    add_watch (watch, data);
  else
    remove_watch (watch, data);
}

libdbus toggles the write watch on essentially every message send, on whichever thread made the call, so both paths run off the main thread. They mutate an unlocked GSList on the connection (cs->ios = g_slist_prepend (...) in connection_setup_add_watch, handler->cs->ios = g_slist_remove (...) in io_handler_destroy_source). Meanwhile the main loop dispatches those same sources, and io_handler_dispatch has no NULL check:

  /* Note that we don't touch the handler after this, because
   * dbus may have disabled the watch and thus killed the
   * handler.
   */
  dbus_watch_handle (handler->watch, dbus_condition);

io_handler_watch_freed is what sets handler->watch = NULL.

main.py called DBusGMainLoop(set_as_default=True), which attaches every dbus.SessionBus() created afterwards, including the shared singleton plugins get. Plugins then call D-Bus from the per-deck tick_actions threads via on_tick(), roughly once a second, forever. MediaPlugin and Battery both do. That is a continuous race against the main loop, and eventually it loses.

Note this is not the older dbus_threads_init_default() problem. libdbus has been thread-safe by default since 1.8 and the runtime ships 1.16.2, so threads_init() is a no-op here. The unsynchronised state is dbus-python's own GLib glue.

Fix

Stop attaching any dbus-python connection to the main loop. Without the attachment there are no GSources, io_handler_dispatch never runs, and the assertion is unreachable, which fixes it for plugin code too without plugins needing to change.

dbus-python is still used for blocking method calls (quit_running, make_api_calls, GnomeExtensions), which never needed a main loop. The four places that needed signals move to GDBus, which is thread safe and already used in this codebase (src/backend/trayicon.py uses Gio.DBusProxy / Gio.bus_get_sync, src/api.py uses dasbus).

  • main.py: drop DBusGMainLoop(set_as_default=True) and the unused import dbus.service
  • LockScreenDetector: add subscribe_to_screen_saver() on top of Gio.DBusConnection.signal_subscribe
  • Detectors/{KDE,Gnome,Cinnamon}: use it instead of add_signal_receiver, and stop re-installing the main loop from the LockScreenManager setup thread
  • WindowGrabber/Integrations/Gnome: port to Gio.DBusProxy and its g-signal
  • tests/: stdlib regression guard, no new dependency

Compatibility

add_signal_receiver is the only dbus-python API that breaks without a default main loop (blocking calls are unaffected, verified in the flatpak). I checked all 59 plugins in the official store by cloning each and scanning for add_signal_receiver, connect_to_signal, DBusGMainLoop, dbus.mainloop, reply_handler, dbus.service.Object and set_default_main_loop.

Zero hits. Only three touch dbus-python at all, and all three make blocking calls only, so they keep working unchanged:

Plugin Bus Usage Called from
MediaPlugin session MPRIS get_object + Interface().Get() on_tick
Battery system UPower get_object + Interface() on_tick
GnomeWindowCalls session Gnome Shell Extensions get_object elsewhere

None use pydbus, dasbus, sdbus or jeepney either. Two of those three are currently exposed to this crash and are fixed by this change.

The only remaining exposure would be an out-of-store plugin receiving dbus-python signals. It would now get a clear RuntimeError from dbus-python instead of racing, and it was already crash-prone. Happy to add an opt-in helper for that case if you would prefer one.

Testing

Attached gdb to a process running the old configuration and broke on dbus_watch_handle. The live backtrace matches the crash core frame for frame:

#0  dbus_watch_handle ()                        libdbus-1.so.3
#1  ??                                          _dbus_glib_bindings...so
#2  g_main_dispatch ()                          libglib-2.0.so.0
#3  g_main_context_iterate_unlocked.isra ()     libglib-2.0.so.0
#4  g_main_loop_run ()                          libglib-2.0.so.0
#5  ffi_call_unix64 ()                          libffi.so.8
rdi            0x55c0448f77c0

rdi is the DBusWatch *. Valid there, 0 in the crash.

With the patch the library backing frame #1 is never loaded, so that frame cannot exist:

$ grep -c _dbus_glib_bindings /proc/<pid>/maps
5     # 1.5.0-beta.15 as shipped
0     # this branch

Ran the patched app on KDE/Wayland. Decks, pages, plugins and the dasbus API all load normally, and dbus-monitor shows the lock screen detector registering a match rule identical to the old add_signal_receiver one:

type='signal',interface='org.freedesktop.ScreenSaver',member='ActiveChanged',path='/org/freedesktop/ScreenSaver'

python3 -m unittest discover -s tests -t . passes, and fails on 129bdb5 flagging all four signal sites.

Two things I could not cover. I never got the abort to fire on demand: the stress tool ran 15 minutes and about 9M calls without hitting it, which is consistent with a race that took 5d 4h in the wild, so tests/manual_dbus_thread_stress.py is a manual tool rather than a test. And since I am on KDE, the Gnome window grabber port and the Gnome/Cinnamon detectors are reviewed but not exercised at runtime. Would appreciate a second pair of eyes from someone on GNOME or Cinnamon.

jslay88 added 2 commits July 30, 2026 00:26
dbus-python's GLib integration (dbus-gmain) is not thread safe. watch_toggled
implements watch enable/disable as a full GSource add/remove, and both paths
mutate an unlocked GSList on the connection from whichever thread sent the
message. The main loop dispatches those same sources, and io_handler_dispatch
calls dbus_watch_handle(handler->watch) with no NULL check. Once anything
touches D-Bus off the main thread the process eventually aborts:

  dbus[3]: arguments to dbus_watch_handle() were incorrect, assertion
  "watch != NULL" failed in file ../dbus/dbus-watch.c line 738.

The deck tick threads reach D-Bus through plugins (MediaPlugin and Battery both
call it from on_tick), so this shows up as a random SIGABRT after hours or days
of uptime, with nothing in the logs because the abort is in native code.

Setting DBusGMainLoop as the default attaches every dbus.SessionBus() in the
process, including the shared one plugins get. Without it, no GSources are
created and the crash is unreachable.

* main.py: drop DBusGMainLoop(set_as_default=True) and the unused dbus.service
  import. dbus-python is now only used for blocking method calls, which never
  needed a main loop.
* LockScreenDetector: add subscribe_to_screen_saver() built on GDBus
* Detectors/{KDE,Gnome,Cinnamon}: use it instead of add_signal_receiver, and
  stop re-installing the main loop from the LockScreenManager setup thread
* WindowGrabber/Integrations/Gnome: port to Gio.DBusProxy

All 59 plugins in the official store were checked for APIs that need the main
loop (add_signal_receiver, connect_to_signal, reply_handler, dbus.service).
None use any of them. The three that use dbus-python at all (MediaPlugin,
GnomeWindowCalls, Battery) only make blocking calls and keep working unchanged.
Static check over main.py, src/ and GtkHelper/ for dbus-python APIs that only
work on a connection attached to a GLib main loop (DBusGMainLoop, dbus.mainloop,
set_default_main_loop, add_signal_receiver, connect_to_signal, reply_handler,
dbus.service). Those are what make a background D-Bus call able to abort the
process in dbus_watch_handle(), so blocking method calls are all that is left
allowed. Fails on 129bdb5, passes now.

stdlib unittest, no new dependency:

    python3 -m unittest discover -s tests -t .

Also adds tests/manual_dbus_thread_stress.py, which reproduces the shape of the
race (attached shared connection, main loop dispatching its watches, background
threads sending large messages). Kept out of the automated suite because hitting
it depends on winning the race, it ran for 15 minutes and 9M calls here without
aborting, while the real crash took just over 5 days of uptime.
jslay88 and others added 3 commits July 30, 2026 00:47
Logind.py was added by StreamController#632 after this branch was cut, so it wasn't
covered by the earlier audit here: it still called
dbus.mainloop.glib.DBusGMainLoop(set_as_default=True) and
connect_to_signal() for the Lock/Unlock signals, on the
LockScreenManager setup thread. That reinstalls the exact non-thread-safe
main loop this branch removes, and LockScreenManager tries logind before
falling back to the DE-specific detectors, so it's the active path on
most systemd-based distros.

Switch the Lock/Unlock subscriptions to Gio.DBusConnection.signal_subscribe
on the system bus, same pattern already used for the screen-saver
detectors. Session-path resolution stays on dbus-python since it's a
blocking call and works fine without an attached main loop.
Core447 merged commit 4cecd76 into StreamController:main Aug 6, 2026

Core447 commented Aug 6, 2026

Copy link
Copy Markdown
Member

Thank you! Great work!!

Core447 added a commit that referenced this pull request Aug 6, 2026
)

* Enhance BetterDeck context management and add image hashing for MediaPlayer tasks

* Fix: Improve UI image handling in ControllerTouchScreen for better performance

* Fix(OnboardingWindow): Problems with Gnome extension hint

* Chore(DonationDialog): Remove option to comission custom plugins

* Feat(Settings): Add option to disable label scrolling

* Chore: Remove debug scroll print

* Chore(ActionChooser): Header not aligned properly

* Fix(locales): Wrong semicolons lead to weird parsing sometimes

* Fix: use correct method name for page lookup in --change-state (#570)

on_change_state in app.py and the api_state_requests handler in
DeckController.py both call get_best_page_path_match_from_name, which
does not exist on PageManagerBackend. The correct method is
find_matching_page_path, matching on_change_page and api_page_requests.

This causes --change-state CLI commands to fail silently (AttributeError
swallowed by GTK's action handler) when SC is running, and crash at
startup when SC is not running.

* Refactor: Enhance image handling and thumbnail generation in ScreenBar for improved performance

* Refactor: Optimize image handling and update scheduling in ScreenBar and IconSelector for improved performance

* Bump to 1.5.0-beta.14

* Remove pyproject.toml

* Update changelog

* Refactor: Adjust image sizing logic in ScreenBarImage for better responsiveness based on touchscreen dimensions

* feat(hyprland): use IPC socket events instead of polling for window changes

Replace the 200ms polling loop that spawns 'hyprctl activewindow -j' (via
flatpak-spawn + xdg-dbus-proxy when running in Flatpak) with a direct
connection to Hyprland's socket2 IPC event stream.

The old approach spawns ~5 processes/second even when no window changes occur,
causing 10-30% CPU usage on the flatpak-session-helper and xdg-dbus-proxy
(see #433 and #457). The new approach uses a blocking Unix socket read with
zero CPU when idle.

Changes:
- Listen for 'activewindow>>' events on Hyprland's .socket2.sock
- Auto-detect socket path via HYPRLAND_INSTANCE_SIGNATURE env var
- Fallback to legacy polling if socket is unavailable
- Add --filesystem=xdg-run/hypr:ro to Flatpak manifest for socket access
- Reconnect automatically if the socket is closed (e.g. compositor restart)

Fixes #433 (partially — eliminates the subprocess-spawning component)
Fixes #457 (eliminates the flatpak-spawn/xdg-dbus-proxy tight loop)

* perf: throttle media player to 2 FPS when no animated content

The MediaPlayerThread runs at 30 FPS unconditionally, iterating over all
keys/dials every ~33ms even when there is no video background, no key
videos, and no scrolling labels. On a 15-key Stream Deck MK.2, this
burns ~20% CPU in pure Python overhead for no visible benefit.

Add dynamic FPS: detect whether any animated content (video, scroll
labels) or pending image tasks exist. When idle, throttle to 2 FPS
(500ms sleep). Immediately return to 30 FPS when animation starts or
tasks are queued.

This preserves full responsiveness for video/animation while reducing
idle CPU from ~20% to ~1%.

* fix: reduce idle CPU usage from ~25% to ~3% (#579)

* fix: reduce KDE window grabber CPU usage by caching window id

Only fetch window name and class when the active window id changes.
Previously, every poll cycle (200ms) spawned 3 subprocesses via
flatpak-spawn/kdotool even when the active window hadn't changed.
Now only 1 subprocess runs per cycle in idle, with the other 2 only
triggered on actual window change.

* fix: cache get_has_scroll_labels to avoid repeated font rendering

get_has_scroll_labels() calls getbbox() on every label to measure text
width. This was called 30 times per second (MediaPlayer FPS) for every
key on the deck, causing ~1350 font rendering calls/sec in idle.

Cache the result and invalidate when labels actually change
(set_page_label, set_action_label, clear_labels).

* fix: reduce idle CPU usage with adaptive FPS and render caching

Three optimizations that together reduce idle CPU from ~25% to ~3%:

1. Adaptive MediaPlayer FPS: drop from 30fps to 2fps when no animated
   content (video/GIF) is active and no pending render tasks exist.
   Ramps back to 30fps instantly when needed.

2. Skip redundant key ticks: only iterate keys in the 30fps loop when
   at least one key has video/GIF content. Checked once per second.

3. Label dirty-check: set_action_label compares all label properties
   before triggering a full re-render. Plugins like CPUTemp that set
   the same label text every tick no longer cause unnecessary renders.

4. Image hash guard in ControllerKey.update(): hash the rendered image
   and skip native format conversion + deck write when unchanged.

---------

Co-authored-by: Stefan Toczek <39330537+stefantoczek@users.noreply.github.com>

* Bump metainfo to beta 14

* Update kofi supporter list

* Update .gitignore

* Automatically fetch github contributors

* Feat: Add uninstall button to plugin settings page

* feat(hyprland): use IPC socket events instead of polling for window changes

Replace the 200ms polling loop that spawns 'hyprctl activewindow -j' (via
flatpak-spawn + xdg-dbus-proxy when running in Flatpak) with a direct
connection to Hyprland's socket2 IPC event stream.

The old approach spawns ~5 processes/second even when no window changes occur,
causing 10-30% CPU usage on the flatpak-session-helper and xdg-dbus-proxy
(see #433 and #457). The new approach uses a blocking Unix socket read with
zero CPU when idle.

Changes:
- Listen for 'activewindow>>' events on Hyprland's .socket2.sock
- Auto-detect socket path via HYPRLAND_INSTANCE_SIGNATURE env var
- Fallback to legacy polling if socket is unavailable
- Add --filesystem=xdg-run/hypr:ro to Flatpak manifest for socket access
- Reconnect automatically if the socket is closed (e.g. compositor restart)

Fixes #433 (partially — eliminates the subprocess-spawning component)
Fixes #457 (eliminates the flatpak-spawn/xdg-dbus-proxy tight loop)

* Add based API for #548 (and other future applicaations) (#551)

The main component of this change is adding api.py, a dasbus based dbus API server with the following features:
* icon pack enumeration
* page enumeration
* addpage(name, json)
* removepage
* streamdeck enumeration
* export the datadir location (for eventual use adding other metadata/backgrounds on-the-fly)

To support this feature a small number of changes were needed elsewhere:
* change add_page so it can throw if the page already exists (the threeish usages
elsewhere in the app now have a try/catch to handle that)
* The flatpak yaml had to be updated to add dasbus (and since the tool is now newer it emitted slightly different yaml)

Co-authored-by: Core447 <100139110+Core447@users.noreply.github.com>

* Improve page swich speeds

* Bump runtime version to 50

* Bump to beta.15

* Fix syntax error in metainfo

* add logind detection for Niri (#632)

* Adding Restart option in tray.py and fixing Open button in Data Path (#617)

* Update tray.py

added an option to Quit and Restart

* Update Settings.py

Fixed the "Open" button for Settings > Developer > Data Path. Will open user default file manager at set location.

* Update Settings.py

updated for beta.15

* Update Settings.py

updated for beta.15

* Delete src/windows/Settings/Settings.py

* Add files via upload

* Fix semaphore leak (#598)

* fix: restore detached subprocess launch in HelperMethods.run_command

Agent-Logs-Url: https://github.com/gensyn/StreamController/sessions/62cf1117-382f-42a9-a4c1-934f159d8374

Co-authored-by: gensyn <36128035+gensyn@users.noreply.github.com>

* Initial plan

* Fix open_web command injection: use argv list instead of shell=True

Agent-Logs-Url: https://github.com/gensyn/StreamController/sessions/d37f03d2-b8d2-4224-bec5-825a7da82f80

Co-authored-by: gensyn <36128035+gensyn@users.noreply.github.com>

* Removed run_command and updated open_web

* Added try/except to open_web

* Re-introduced run_command

* Re-introduced run_command

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>

* feat(Store): Add Git Branch Selection for Plugins in Store UI (#540)

* feat(Store): Add Git branch override functionality for plugins

- Implement methods to get, set, and remove Git branch overrides for plugins in StoreBackend.
- Enhance InfoPage and SourceGroup to allow users to select between store version and specific Git branches.
- Update localization files to support new UI elements related to Git branch selection.
- Ensure proper handling of plugin data and UI updates when switching between sources.

* feat(Store): Enhance plugin source selection with Git tags support

* Make plugin info page srollable

* Chore(InfoPage): Hide plugin branch selector for custom plugins

* Chore(InfoPage): Remove apply and status rows

* Feat(InfoPage): Add branch caching to avoid rate limiting

---------

Co-authored-by: Core447 <100139110+Core447@users.noreply.github.com>

* [Bug Fix]: Stabilize resume by keeping controllers alive during transient HID errors (#526)

* Fix resume stability by keeping controller alive

* Gate touchscreen transient-error handling behind beta-resume-mode

Address review feedback: only skip the removal/reconnect fallback
when "Use new resume mode (beta)" is enabled, mirroring the existing
behavior in MediaPlayerSetImageTask. Keeps the ability to reset a
genuinely stuck deck when beta-resume mode is off, and raises the
failure threshold from 5 to 10 as suggested in review. Also fixes
n_failed_in_row being shared across all decks instead of tracked
per serial number.

---------

Co-authored-by: Core447 <100139110+Core447@users.noreply.github.com>

* Hold Gio.Application while keep-running is enabled (silent idle-quit in background mode) (#626)

* Hold Gio.Application while running in background mode

With keep-running enabled, closing the main window only hides it.
Gio.Application's reference counting can then decide the app is idle
and quit the process silently: the deck goes dark with no traceback
and no log line. Holding the application while keep-running is active
(and releasing it in on_quit) pins the lifecycle to the explicit quit
path instead.

Related to the lifecycle rework in #558; this is the minimal bugfix
for current main.

* Sync Gio.Application hold with keep-running when toggled at runtime

The hold taken in __init__ only reflected the setting at startup. If
keep-running was toggled afterwards (via the settings page or the
KeepRunningDialog shown on first close), the hold state never updated,
so the original silent idle-quit could still happen after enabling
keep-running mid-session, or the app could stay held longer than
intended after disabling it. Centralize the hold/release logic in
set_keep_running_hold() and call it from both places the setting can
change at runtime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Core447 <100139110+Core447@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: repair page import and reload callbacks (#633)

* fix: Add proper GIF support with transparency and per-frame timing (#595)

* fix: Add proper GIF support with transparency and per-frame timing

GIFs were previously decoded using OpenCV (cv2), which does not support
GIF transparency or animation metadata. This caused two issues:
- Transparent GIFs rendered with a white background
- All GIF frames played at a fixed FPS regardless of per-frame delays,
  causing animations to play at the wrong speed or appear cut short

Changes:
- key_video_cache.py: Detect GIFs by file extension and decode them
  using PIL instead of cv2. Convert each frame to RGBA to preserve
  the alpha channel. Cache GIF frames as PNG instead of JPEG to retain
  transparency. Read per-frame delays from GIF metadata and store them
  in frame_delays for use during playback.
- KeyVideo.py: Add _get_next_gif_frame() which advances playback based
  on per-frame delay in milliseconds rather than a fixed FPS value.
  This correctly handles GIFs with variable frame timing (e.g. a long
  first frame followed by short animation frames).

* perf: Non-blocking GIF loading with static preview, delay caching and deduplication

Loading pages with multiple animated GIFs caused multi-second freezes because
all frames were decoded synchronously on the calling thread.

- Render frame 0 synchronously in __init__ so every button shows a static
  preview immediately; all remaining work runs in a background daemon thread
- Save per-frame delays to delays.json next to the cached frames so the slow
  PIL seek-loop only runs once; subsequent loads read the JSON instantly
- Add VideoFrameCache.get_or_create() registry so the same GIF+size is only
  loaded once even when multiple buttons reference the same file; use it in
  InputVideo to avoid redundant work on page switches

* fix: Prevent cache-write race condition (segfault) and improve GIF timing accuracy

key_video_cache.py: Replace the single global registry lock with a
per-key initialization lock. The previous approach ran __init__ (including
Image.open and LANCZOS resize) outside the global lock, allowing two threads
to create instances for the same GIF simultaneously and write to the same
cache files concurrently — corrupting PNGs and causing segfaults in PIL.
The per-key lock ensures only one thread initialises a given (path, size)
at a time without blocking unrelated keys.

KeyVideo.py: Replace tick-based GIF timing with time.perf_counter(). The
previous implementation accumulated elapsed time in media player ticks,
which caused playback to drift when the media player was under CPU load.
Also replace the single if with a while loop so multiple frames are skipped
when catch-up is needed, keeping the animation in sync with real time.

* fix: include file path in video cache key to avoid collisions

_fast_cache_key() hashed only mtime+size, and the on-disk cache dir was
keyed on that hash alone (no path). Two different files sharing size and
mtime (e.g. assets bulk-extracted from a zip with identical timestamps)
would silently read/write each other's cached frames. Folding the
absolute path into the hash keeps the fast stat()-only lookup while
requiring an actual path+size+mtime match.

* fix: free video caches once no button references them anymore

_registry was a plain dict, so every distinct (path, size) VideoFrameCache
ever shown stayed fully decoded in memory for the life of the app, even
after the button using it was gone. Using a WeakValueDictionary lets an
entry disappear once nothing holds a strong reference to it, matching the
intended "shared while in use, freed when unused" behavior instead of
leaking indefinitely.

* fix: preserve transparency and correct colors for GIF backgrounds

background_video_cache.py decoded all backgrounds via cv2.VideoCapture,
which loses alpha and mangles colors for GIFs. Decode .gif backgrounds
with PIL instead, keeping the existing tick-driven frame advancement.

BackgroundVideo.crop_key_image_from_deck_sized_image (DeckController.py)
is the method that actually runs at playback time (it overrides the base
class's version), and it pasted every background frame onto an opaque
RGB key image before this, dropping alpha regardless of how the frame was
decoded. Return the cropped RGBA segment directly instead, matching how
BackgroundImage already handles static image backgrounds.

---------

Co-authored-by: Core447 <100139110+Core447@users.noreply.github.com>

* Fix GTK calls from non-main threads (SIGSEGV class of #566) (#625)

Three remaining sites make GTK4 widget calls from background threads,
which corrupts GTK state and crashes with a silent SIGSEGV and no
Python traceback, the same failure class fixed for MediaPlayerThread
in #566:

- DeckManager.remove_controller(): called from
  FlatpakDeckDisconnectThread / udev callbacks; deck_stack.remove_page()
  now goes through GLib.idle_add, mirroring how
  add_newly_connected_deck() already handles add_page().
- MainWindow.show_info_toast(): callers include background threads
  (e.g. update_assets in main.py); the Adw.Toast is now constructed and
  added on the main loop.
- MissingRow.hide_install_error(): runs on a threading.Timer thread;
  the four widget calls are now dispatched via GLib.idle_add.

* Don't attach dbus-python to the GLib main loop (fixes random SIGABRT in dbus_watch_handle) (#634)

* Don't attach dbus-python to the GLib main loop

dbus-python's GLib integration (dbus-gmain) is not thread safe. watch_toggled
implements watch enable/disable as a full GSource add/remove, and both paths
mutate an unlocked GSList on the connection from whichever thread sent the
message. The main loop dispatches those same sources, and io_handler_dispatch
calls dbus_watch_handle(handler->watch) with no NULL check. Once anything
touches D-Bus off the main thread the process eventually aborts:

  dbus[3]: arguments to dbus_watch_handle() were incorrect, assertion
  "watch != NULL" failed in file ../dbus/dbus-watch.c line 738.

The deck tick threads reach D-Bus through plugins (MediaPlugin and Battery both
call it from on_tick), so this shows up as a random SIGABRT after hours or days
of uptime, with nothing in the logs because the abort is in native code.

Setting DBusGMainLoop as the default attaches every dbus.SessionBus() in the
process, including the shared one plugins get. Without it, no GSources are
created and the crash is unreachable.

* main.py: drop DBusGMainLoop(set_as_default=True) and the unused dbus.service
  import. dbus-python is now only used for blocking method calls, which never
  needed a main loop.
* LockScreenDetector: add subscribe_to_screen_saver() built on GDBus
* Detectors/{KDE,Gnome,Cinnamon}: use it instead of add_signal_receiver, and
  stop re-installing the main loop from the LockScreenManager setup thread
* WindowGrabber/Integrations/Gnome: port to Gio.DBusProxy

All 59 plugins in the official store were checked for APIs that need the main
loop (add_signal_receiver, connect_to_signal, reply_handler, dbus.service).
None use any of them. The three that use dbus-python at all (MediaPlugin,
GnomeWindowCalls, Battery) only make blocking calls and keep working unchanged.

* Add regression guard for the dbus-python main loop crash

Static check over main.py, src/ and GtkHelper/ for dbus-python APIs that only
work on a connection attached to a GLib main loop (DBusGMainLoop, dbus.mainloop,
set_default_main_loop, add_signal_receiver, connect_to_signal, reply_handler,
dbus.service). Those are what make a background D-Bus call able to abort the
process in dbus_watch_handle(), so blocking method calls are all that is left
allowed. Fails on 129bdb5, passes now.

stdlib unittest, no new dependency:

    python3 -m unittest discover -s tests -t .

Also adds tests/manual_dbus_thread_stress.py, which reproduces the shape of the
race (attached shared connection, main loop dispatching its watches, background
threads sending large messages). Kept out of the automated suite because hitting
it depends on winning the race, it ran for 15 minutes and 9M calls here without
aborting, while the real crash took just over 5 days of uptime.

* End the files touched by this branch with a trailing newline

* Port LogindLockScreenDetector off the dbus-python main loop too

Logind.py was added by #632 after this branch was cut, so it wasn't
covered by the earlier audit here: it still called
dbus.mainloop.glib.DBusGMainLoop(set_as_default=True) and
connect_to_signal() for the Lock/Unlock signals, on the
LockScreenManager setup thread. That reinstalls the exact non-thread-safe
main loop this branch removes, and LockScreenManager tries logind before
falling back to the DE-specific detectors, so it's the active path on
most systemd-based distros.

Switch the Lock/Unlock subscriptions to Gio.DBusConnection.signal_subscribe
on the system bus, same pattern already used for the screen-saver
detectors. Session-path resolution stays on dbus-python since it's a
blocking call and works fine without an attached main loop.

---------

Co-authored-by: Core447 <100139110+Core447@users.noreply.github.com>

---------

Co-authored-by: Core447 <100139110+Core447@users.noreply.github.com>
Co-authored-by: Ty Smith <ty@tysmith.me>
Co-authored-by: Luc LETOFFE <lucletoffe@hey.com>
Co-authored-by: oneandonlyno1 <39330537+oneandonlyno1@users.noreply.github.com>
Co-authored-by: Stefan Toczek <39330537+stefantoczek@users.noreply.github.com>
Co-authored-by: geeksville <kevinh@geeksville.com>
Co-authored-by: Nicolò Maria Semprini <nicosemp@gmail.com>
Co-authored-by: tsunpot <154901805+Mattdubs9699@users.noreply.github.com>
Co-authored-by: gensyn <36128035+gensyn@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Pavel Savinov <getjump0@gmail.com>
Co-authored-by: Stevo <79025479+CupOfOwls@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Pieter de Villiers <pieter.devilliers@gmail.com>
Co-authored-by: David F <david.frason1@gmail.com>
Co-authored-by: jslay88 <justin.slay@gmail.com>
pull Bot pushed a commit to mattmattox/StreamController that referenced this pull request Aug 12, 2026
ScreenSaver.show() blanked DeckController.inputs and let init_inputs() refill
it one input type at a time. MediaPlayerThread.run() and tick_actions() read
that dict without a lock, so they could see it without the input type they were
about to index. Neither loop caught exceptions, so the KeyError ended the thread
for good and the deck stayed frozen until the app was restarted - which is what
the resume from the lock screen looks like from the outside.

Reproduced on a fake deck with no artificial timing help: MediaPlayerThread died
with KeyError: Input.Dial and tick_actions with ValueError: Operation on closed
image after 53-87 lock/unlock cycles.

* init_inputs(): build the new dict with all input types present and publish it
  in a single assignment. ScreenSaver.show() and set_rotation() no longer blank
  self.inputs first, since init_inputs() replaces it wholesale now.
* MediaPlayerThread.run() and tick_actions(): log and continue instead of
  letting one bad tick take the thread down permanently.
* tests/manual_screensaver_race.py drives the lock/unlock cycle against a fake
  deck. Manual like tests/manual_dbus_thread_stress.py, since it depends on
  winning a race and needs a data dir with a page:

      python3 tests/manual_screensaver_race.py --data data

  683, 595 and 575 cycles with both threads still alive after the fix.

Note this is not the SIGABRT/segfault class from the same report - that came
from dbus-python being attached to the GLib main loop and was fixed in StreamController#634.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants


Back | FazBrowse Home | New Git URL