| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
|
🤖 New build scheduled with the buildbot fleet by @diegorusso for commit ac018d6 🤖 Results will be shown at: https://buildbot.python.org/all/#/grid?branch=refs%2Fpull%2F146071%2Fmerge If you want to schedule another build, you need to add the 🔨 test-with-buildbots label again. |
Sorry, something went wrong.
|
I have some questions about the EH frame generation and how it applies to the different code regions. Looking at jit_record_code, it's called in two places:
Both end up calling _PyJitUnwind_GdbRegisterCode, which builds the same EH frame via _PyJitUnwind_BuildEhFrame. The EH frame in elf_init_ehframe describes a specific prologue/epilogue sequence. On x86_64 for example: push %rbp (1 byte) mov %rsp, %rbp (3 bytes) call *%rcx (2 bytes) pop %rbp (1 byte) ret I understand how this is correct for jit_shim. Looking at Tools/jit/shim.c, it's a normal C function that calls into the executor: _Py_CODEUNIT *
_JIT_ENTRY(...) {
jit_func_preserve_none jitted = (jit_func_preserve_none)exec->jit_code;
return jitted(exec, frame, stack_pointer, tstate, ...);
}The compiler will emit exactly the prologue/epilogue the EH frame describes. But I don't understand how the same EH frame is correct for jit_executor. The executor code region is a concatenation of many stencils, each compiled from Tools/jit/template.c with __attribute__((preserve_none)), chaining together via __attribute__((musttail)) tail calls. These stencils don't have the push rbp / mov rsp,rbp prologue that the EH frame describes. They use a completely different calling convention. The FDE covers the full code_size + trampolines.size range but the CFI instructions only describe ~7 bytes of prologue/epilogue. DWARF will apply the last rule (CFA = RSP + 8 on x86_64) to all remaining addresses in the range. I don't understand why that rule would be correct at arbitrary points within the stencil code. Is it guaranteed that preserve_none stencils never modify RSP? Or is there something else going on that makes this work? The test (test_jit.py) sets a breakpoint at id(42) which hits in the interpreter, not in the middle of a stencil. So the test verifies that the symbols appear in GDB's backtrace, but I don't think it exercises unwinding from an arbitrary point within the executor code region. Could we add a test that triggers unwinding from inside JIT code (e.g., via a signal or Ctrl+C while executing JIT code)? Am I missing something about how the stencils interact with the stack, or is the EH frame intentionally approximate for the executor region? |
Sorry, something went wrong.
There was a problem hiding this comment.
A bunch of questions I have from reading the code so far
Sorry, something went wrong.
| struct jit_code_entry *first_entry; | ||
| }; | ||
|
|
||
| static volatile struct jit_descriptor __jit_debug_descriptor = { |
There was a problem hiding this comment.
Should these be non-static? The GDB JIT interface spec says GDB locates __jit_debug_descriptor and __jit_debug_register_code by name in the symbol table. With static linkage they would be invisible in .dynsym on stripped builds and when CPython is loaded as a shared library via dlopen. Am I missing something, or would this silently break in release/packaged builds where .symtab is stripped?
Maybe also worth adding __attribute__((used)) to prevent the linker from eliding them?
Sorry, something went wrong.
There was a problem hiding this comment.
Yes, you are right. Instead of removing the static I've exported with the macro Py_EXPORTED_SYMBOL
Sorry, something went wrong.
| id(42) | ||
| return | ||
|
|
||
| warming_up = True |
There was a problem hiding this comment.
Could this loop hang? When warming_up=True, the call passes warming_up_caller=True which returns immediately at line 8, so the recursive body never actually executes. If the JIT does not activate via some other path, would this not spin forever until the timeout kills it? Should there be a max iteration count as a safety net?
Also, line 16 uses bitwise & instead of and. Was that intentional? It means is_active() is always evaluated even when is_enabled() is False.
Sorry, something went wrong.
There was a problem hiding this comment.
I've simplified the test, the loop is not more controlled and deterministic.
Sorry, something went wrong.
| return; | ||
| } | ||
| _PyJitUnwind_GdbRegisterCode( | ||
| code_addr, (unsigned int)code_size, entry, filename); |
There was a problem hiding this comment.
code_size comes in as size_t but gets cast to unsigned int here. I know JIT regions will not be 4GB, but should the API just take size_t throughout for consistency?
Sorry, something went wrong.
There was a problem hiding this comment.
This is now done.
Sorry, something went wrong.
What this change synthesises for jit_executor is one unwind description for the executor as a whole, not compiler-emitted per-stencil CFI. Because the stencils are musttail-chained, the jumps between stencils do not add extra native call frames. The unwind job here is just to recover the caller of the executor frame. We don't want to describe each stencil as its own frame. When GDB stops at a PC inside py::jit_executor:<jit>:
On AArch64, for most of the covered executor range, the synthetic CFI says:
Good catch for the testing gap. I’ve now added a new test that breaks inside the jit executor. It sill breaks at the builtin_id but GDB then finishes out through the C helper frames until the selected frame is py::jit_executor:<jit> (thanks to some GDB-python scripting), single-steps twice inside the executor, and only then runs bt. |
Sorry, something went wrong.
|
@diegorusso @pablogsal I think I may have come up with a solution that works. EDIT: I think I gdb doesn't only use backtrace. So we're still stuck. Sorry for the noise! Background info (skip if not interested):
The current issue:
The solution:
This should work with backtrace from execinfo.h. without issues. It should even work with gdb step debugging/bt, with the exception that at the function prologue, it might be broken. However the most important thing is that we unbreak all C extension code that uses backtrace! Also, backtrace should be fast as our DWARF would be tiny and simple. TLDR: frame pointers = eh_frame is simple. |
Sorry, something went wrong.
@diegorusso I have to say that I am tremendously confused here. If GDB or backtrace() stops at an arbitrary PC inside py::jit_executor:<jit>, the unwind info for that exact PC should let the unwinder reconstruct the caller frame (py::jit_shim:<jit>) and then continue into _PyEval_*. So the real question is not “does the FDE cover the address range?” and it is not “do the stencils form one logical frame?”. The real question is: does the CFI row that applies at that PC actually describe the machine state there? That is the part I do not think has been explained. I agree with the narrow musttail point: tail-chaining the stencils means you do not accumulate one native call frame per stencil. Fine. But that only tells us that we want to unwind the executor as one logical frame. It does not tell us that one fixed synthetic unwind recipe is valid everywhere inside the executor blob. And that is exactly where I think the argument goes off the rails. jit_executor is not one ordinary C function with one stable prologue/epilogue. It is a concatenation of many preserve_none stencils, glued together with musttail. For a single synthetic FDE to be correct across the whole region, there has to be some invariant that says “for any PC in executor code, the CFA and saved return state look like this”. I do not see that invariant stated anywhere, and the current explanation seems to jump from “musttail” straight to “the unwind is correct”, which are not the same thing. A concrete x86_64 example of why this seems wrong to me: with the same sort of flags used for executor stencils, a preserve_none + musttail function can compile to something as trivial as jmp callee or, if it needs temporary stack space / spills, something more like subq $24, %rsp ... addq $24, %rsp jmp callee In the first case there is no %rbp frame at all. In the second case the CFA is temporarily %rsp-relative and changes inside the body. So I do not understand how one synthetic %rbp-based description for the entire covered executor range is supposed to be generally correct. For jit_shim I can at least see the intended story, because it is one ordinary non-tail C function that calls into JIT code. For jit_executor, I still do not see what makes the unwind recipe valid for arbitrary PCs inside the blob. Also, I rebuilt the branch locally and tried the exact “finish to py::jit_executor:<jit>, step twice, then bt” flow. On x86_64 I still get: #0 py::jit_executor:<jit> () #1 ?? () ... Backtrace stopped: previous frame inner to this frame (corrupt stack?) So this is not just a theoretical concern for me. I still do not understand why the model being described here is supposed to work.I am of course not objecting to the goal. I am saying I still do not see the correctness argument. If the claim is that this is actually a correct unwind description for jit_executor as a whole, then I think what is missing from the discussion is the key invariant: what exactly is guaranteed to be true about the CFA / saved FP / saved return address at an arbitrary PC inside executor code that makes this one synthetic FDE valid? |
Sorry, something went wrong.
Thanks, thanks for the comment. I regenerated the x86_64 and AArch64 stencils after the recent frame-pointer changes. What we have today is that shim gets a real frame-pointer prologue, but the executor stencils still are not uniformly rbp/x29-framed, so I don’t think the current generated code is enough to justify a single executor-wide CFA = rbp + 16 / x29 + const rule for arbitrary PCs in the blob. The current implementation is still one synthetic executor-wide FDE. The unwinder uses the current PC to select that FDE and apply its CFI to recover the caller frame. That works where the actual machine state matches the synthetic rule at the stop PC, but it is still approximate executor-wide unwind metadata, not exact per-stencil CFI. Separately, once this PR lands, wiring up libgcc-backed backtrace should be fairly easy. We already synthesise .eh_frame; the remaining work is to call the appropriate __register_frame* and deregistration API for that blob so the unwinder can see it. |
Sorry, something went wrong.
Ok, I think now I understand. After re-checking the generated stencils I agree the current explanation was too bold. musttail only establishes the narrow point that the stencil-to-stencil transitions do not accumulate one native call frame per stencil and it does not by itself establish the stronger property needed for unwinding: that for an arbitrary PC inside jit_executor, the CFA and saved return state always have a shape described by one executor-wide FDE. That stronger property is the missing invariant here. After looking again at the regenerated x86_64 and AArch64 stencils, I don't think we have that invariant today:
I cannot justify the current synthetic executor-wide FDE as being correct for arbitrary PCs in the executor blob. The new test I added is still useful, but it proves something narrower: that the synthetic FDE works for the exercised in-executor stop. It does not prove that the same CFI is exact for every interior PC in the region (like you did in your example) I think the real options are:
The current implementation does not yet have the invariant needed to justify one executor-wide FDE for jit_executor but at the same time I don't really like the suggestions above. Let me think about it |
Sorry, something went wrong.
The current generation reserve the rbp. So all current stencils assume an rbp. Do you think it would fix it if we emitted our own prologue for the very first JIT executor uop ie (push %rbp; movq %rsp, %rbp) , and teardown (popq %rbp) at all rets ? I have a working branch that does that. FWIW, it can be done quite easily using the assembly manipulator we have in the JIT. Will that make it appropiate rbp/x29-framed?
Unfortunately, it seems you're right here. I dug around libgcc a little more and that's the only interface I see that intercepts _Unwind_Find_FDE. The function is public but undocumented, which is annoying. I'm just shocked that libgcc does not seem to use frame pointers as a fallback for x86_64 or AArch64 when I looked around it. |
Sorry, something went wrong.
Not all of them. See _SET_IP family. But you can see others as well. On AArch64 if we reserve the frame pointer, it will be barely touched (just a few uops set it). If we don't reserve it, then we have the standard prologue/epilogue for the majority of the uops. I'm not entirely sure your statement is true. |
Sorry, something went wrong.
Huh that's surprising! On x86_64, the current main produces code that doesn't touch rbp at all (from manual inspection at least). I wonder why it's different on AArch64, thanks for reporting back. |
Sorry, something went wrong.
Oh sorry I'm wrong, not main, my branch, I had to pass the usual: "-fno-omit-frame-pointer",
"-mno-omit-leaf-frame-pointer",
to it to get things like that. |
Sorry, something went wrong.
Accept either JIT backtrace shape in test_jit.py: - py::jit:executor -> _PyEval_* - py::jit:executor -> _PyJIT_Entry -> _PyEval_* This avoids baking in architecture-specific unwind details while still checking that GDB gets out of the JIT region and back into the eval loop.
Documentation build overview100 files changed · + 1 added · ± 99 modified + Added ± Modified |
Sorry, something went wrong.
There was a problem hiding this comment.
I've just tried this out and it works nicely (on linux AArch64). Here's an example:
Hitting ctrl-C while running this under gdb:
def count(seq):
t = 0
for i in seq:
t += math.sqrt(1.0)Gives me this backtrace:
#0 PyFloat_AsDouble (op=0xfffff7840770) at Objects/floatobject.c:253 #1 0x0000fffff7779074 in math_1 (err_msg=0xfffff777bf38 "expected a nonnegative input, got %s", can_overflow=0, func=<optimized out>, arg=<optimized out>) at ./Modules/mathmodule.c:792 #2 math_sqrt (self=<optimized out>, args=<optimized out>) at ./Modules/mathmodule.c:1161 #3 0x0000fffff7fe72f0 in py::jit:executor () #4 0x0000aaaaaab57960 in _PyEval_EvalFrameDefault (tstate=0xaaaaab31b3c0, frame=0xfffffffffffffffd, throwflag=-1424409936, throwflag@entry=0) at Python/generated_cases.c.h:5941 #5 0x0000aaaaaad3a96c in _PyEval_EvalFrame (throwflag=0, frame=0xfffff7fea020, tstate=0xaaaaab1c7de8 <_PyRuntime+344816>) at ./Include/internal/pycore_ceval.h:122 ...
and I can use the up, down, step and continue commands as I would expect.
Stepping into the jitted code even works (sort of):
Single stepping until exit from function py::jit:executor, which has no line number information.
Sorry, something went wrong.
Generate per-target JIT unwind constants from the compiled shim EH frame, emit dispatcher headers for the unwind info, and update the GDB JIT unwind tests.
| #else | ||
| # error "Unsupported target architecture" | ||
| #endif | ||
| DWRF_U8(DWRF_CFA_def_cfa); |
There was a problem hiding this comment.
This is really elegant!
Sorry, something went wrong.
There was a problem hiding this comment.
Thank you very much, Diego. This PR has been quite a Sisyphean effort, and the result looks very solid.
The DWARF writing is very elegant, in particular deriving the unwind info from the shim CFI instead of hard-coding another copy of the rules. Excellent job ❤️ 👌
I found and fixed two small issues before landing:
I’m landing this now.
Great work man! 🚀
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
The PR adds the support to GDB for unwinding JIT frames by emitting eh frames.
It reuses part of the existent infrastructure for the perf_jit from @pablogsal.
This is part of the overall plan laid out here: #126910 (comment)
The output in GDB looks like:
Program received signal SIGINT, Interrupt. 0x0000fffff7fb50f8 in py::jit_entry:<jit> () (gdb) bt #0 0x0000fffff7fb50f8 in py::jit_entry:<jit> () #2 0x0000aaaaaad5e314 in _PyEval_EvalFrameDefault (tstate=0xfffff7fb80f0, frame=0xfffff774bab0, throwflag=6, throwflag@entry=0) at ../../Python/generated_cases.c.h:5711 #3 0x0000aaaaaad61350 in _PyEval_EvalFrame (tstate=0xaaaaab1d57b0 <_PyRuntime+344632>, frame=0xfffff7fb8020, throwflag=0) at ../../Include/internal/pycore_ceval.h:122 ...