| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
…mul_7_6, _reduce_big_sum_7
It would be good to profile both versions on a few different architectures first. |
Sorry, something went wrong.
|
Apparently, this slows these two cases down on a Mac. Anyway, I realize the reason why two arrays is faster than one only works when the fixed multiplier has much less than 64 bits (50 bits in this case). It should be possible to rearrange the computation somewhat though. |
Sorry, something went wrong.
|
Update: this does not slow down these two cases in a Mac (it's neutral, probably because Mac has too many registers). Is mostly human mistake (check out an old build, forget to --enable-assembly, forget to clean up autogenerated src/config.h[.in] because the build system does not automatically do that) before--- fft_small 1 thread --- 20.00: 2.381 2.157 3.412 20.10: 2.132 2.117 2.159 20.20: 2.262 2.246 2.280 20.30: 2.221 2.201 2.245 20.40: 2.330 2.313 2.339 20.50: 2.289 2.274 2.303 20.60: 2.245 2.215 2.277 20.70: 2.234 2.201 2.268 20.80: 2.201 2.161 2.233 20.90: 2.236 2.221 2.259 precomp: avg 0.328439, max 10.001241 21.00: 2.234 2.204 2.270 21.10: 2.191 2.149 2.217 21.20: 2.343 2.297 2.374 21.30: 2.314 2.282 2.335 21.40: 2.341 2.322 2.355 21.50: 2.300 2.283 2.319 21.60: 2.261 2.223 2.292 21.70: 2.247 2.205 2.275 21.80: 2.211 2.168 2.238 21.90: 2.253 2.217 2.267 precomp: avg 0.095625, max 0.433954 22.00: 2.240 2.202 2.258 22.10: 2.205 2.171 2.222 22.20: 2.340 2.328 2.349 22.30: 2.305 2.280 2.321 22.40: 2.356 2.287 2.423 22.50: 2.324 2.248 2.399 22.60: 2.313 2.296 2.328 22.70: 2.297 2.291 2.303 22.80: 2.251 2.237 2.258 22.90: 2.285 2.275 2.291 precomp: avg 0.053925, max 0.303884 23.00: 2.290 2.267 2.335 23.10: 2.251 2.231 2.279 23.20: 2.367 2.356 2.371 23.30: 2.320 2.313 2.325 23.40: 2.444 2.410 2.473 23.50: 2.390 2.385 2.395 23.60: 2.335 2.321 2.344 23.70: 2.320 2.310 2.328 23.80: 2.318 2.288 2.383 23.90: 2.317 2.300 2.324 precomp: avg 0.032803, max 0.220024 24.00: 2.327 2.313 2.348 24.10: 2.283 2.270 2.317 24.20: 2.409 2.394 2.432 24.30: 2.363 2.355 2.367 24.40: 2.480 2.460 2.530 24.50: 2.442 2.423 2.508 24.60: 2.464 2.411 2.498 24.70: 2.422 2.377 2.445 24.80: 2.375 2.338 2.450 24.90: 2.441 2.395 2.482 precomp: avg 0.015412, max 0.218362 25.00: 2.377 2.361 2.390 25.10: 2.340 2.312 2.373 25.20: 2.462 2.447 2.473 25.30: 2.427 2.417 2.435 25.40: 2.581 2.557 2.601 25.50: 2.499 2.482 2.521 25.60: 2.533 2.500 2.557 25.70: 2.498 2.480 2.511 25.80: 2.429 2.407 2.444 25.90: 2.471 2.435 2.491 precomp: avg 0.016391, max 0.143278 --- fft_small 1 thread --- 20.00: 2.416 2.159 3.615 20.10: 2.133 2.123 2.142 20.20: 2.265 2.250 2.285 20.30: 2.229 2.203 2.272 20.40: 2.351 2.332 2.369 20.50: 2.315 2.295 2.332 20.60: 2.244 2.215 2.273 20.70: 2.238 2.203 2.270 20.80: 2.196 2.157 2.232 20.90: 2.241 2.208 2.268 precomp: avg 0.336676, max 10.766291 21.00: 2.232 2.205 2.268 21.10: 2.192 2.155 2.230 21.20: 2.346 2.283 2.383 21.30: 2.317 2.285 2.337 21.40: 2.369 2.345 2.399 21.50: 2.331 2.309 2.353 21.60: 2.266 2.230 2.292 21.70: 2.248 2.206 2.274 21.80: 2.214 2.173 2.241 21.90: 2.254 2.215 2.263 precomp: avg 0.098097, max 0.420648 22.00: 2.241 2.220 2.256 22.10: 2.205 2.179 2.225 22.20: 2.335 2.321 2.348 22.30: 2.306 2.297 2.320 22.40: 2.394 2.318 2.471 22.50: 2.362 2.285 2.445 22.60: 2.315 2.300 2.326 22.70: 2.287 2.274 2.295 22.80: 2.254 2.244 2.260 22.90: 2.286 2.277 2.292 precomp: avg 0.052319, max 0.269028 23.00: 2.277 2.270 2.284 23.10: 2.231 2.228 2.240 23.20: 2.361 2.346 2.369 23.30: 2.322 2.316 2.327 23.40: 2.469 2.458 2.476 23.50: 2.437 2.434 2.440 23.60: 2.351 2.333 2.359 23.70: 2.319 2.304 2.327 23.80: 2.287 2.278 2.292 23.90: 2.327 2.311 2.340 precomp: avg 0.028490, max 0.223473 24.00: 2.305 2.294 2.313 24.10: 2.271 2.263 2.281 24.20: 2.388 2.375 2.399 24.30: 2.359 2.354 2.367 24.40: 2.494 2.487 2.505 24.50: 2.472 2.460 2.499 24.60: 2.440 2.401 2.472 24.70: 2.400 2.360 2.426 24.80: 2.350 2.330 2.370 24.90: 2.410 2.383 2.422 precomp: avg 0.019453, max 0.219641 25.00: 2.367 2.358 2.376 25.10: 2.297 2.266 2.326 25.20: 2.454 2.443 2.466 25.30: 2.410 2.403 2.417 25.40: 2.562 2.547 2.571 25.50: 2.521 2.519 2.523 25.60: 2.522 2.492 2.538 25.70: 2.484 2.473 2.494 25.80: 2.424 2.409 2.442 25.90: 2.470 2.435 2.497 precomp: avg 0.015258, max 0.158447 That said, it's true that umul_ppmm without --enable-assembly (which split the 64-bit limb into two 32-bit limbs) is much slower than _madd which uses int128. Should be safe to have an implementation that check __SIZEOF_INT128__ (or __GNUC__ like crt_helpers.h, clang depends this too despite the name) |
Sorry, something went wrong.
|
This makes flint_mpn_mul 4% slower on my machine (Zen 3). |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Currently, in order to reduce dependency chain length, functions in crt_helpers.h use two parallel arrays r and t to represent a multi-limb values.
Unfortunately, this means a N-limb value needs 2N registers. In the largest case, this leads to massive slowdown, likely because of register spilling to memory.
This PR address that by changing the code of the largest case to use just N registers.
Speedup can be verified by running ./build/fft_small/profile/p-mul before and after the change, the .40 and .50 rows are ~ 2x faster, and is essentially tied with #2692 . (Probably because of added instruction level parallelism.)
Since the change here is obviously much simpler than #2692 , this is to be preferred.
Points for discussion: is it worth switching everything to use this one-accumulator version? It might lead to slower performance because of the need to propagate carry.
Benchmark result
`crt-register-reduce` (this PR)Benchmark instruction