| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Both used np.dot/np.cross on 3-vectors, whose generic dispatch overhead - built for arbitrary shapes/broadcasting - dominates cost at this size, same pattern as isR/trnorm/tr2adjoint (rai-opensource#213, rai-opensource#214). qqmul: replaced with explicit scalar Hamilton-product arithmetic, bit-identical to the prior result (verified to ~1 ULP over 2000 random trials). qvmul: replaced the q * pure(v) * conj(q) sandwich (two full Hamilton products, each wasting work on a zero scalar part, via qqmul/qpure/ qconj with their own getvector re-validation) with the closed-form rotation identity v' = v + 2s(w x v) + 2 w x (w x v) for q = (s, w). This is a different, well-known equivalent formula rather than a re-expression of the same one, so it is not bit-identical - verified to ~1e-14 absolute over 2000 random trials (v scaled to magnitude ~10), well within floating-point noise. qqmul ~10x faster (15.0us -> 1.4us isolated), qvmul ~24x faster (33.4us -> 1.4us isolated). End to end: Q1 * v drops from ~33us to ~4us (~9x), Q1 * Q2 from ~19us to ~10us (~2x - qqmul itself is no longer the bottleneck there; the remainder is UnitQuaternion construction overhead, in particular qunit()'s np.linalg.norm/np.r_ calls, which is a separate, un-addressed candidate for a future fix in the same spirit). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
⚠️ Please install the Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Summary
Test plan
🤖 Generated with Claude Code