…e index type
tf::distance computed (end - beg + step - 1) / step in the index type.
When the span of the range does not fit in that type, the signed sum
overflowed (undefined behaviour) and the unsigned sum wrapped, so
for_each_index, for_each_by_index and reduce_by_index silently ran zero
iterations, for example for_each_index(0u, UINT_MAX, 1u << 20) or
for_each_index(0, INT_MAX, 1 << 20). The signed path also divided by
zero for an empty range with a zero step, such as for_each_index(5, 5, 0),
and did not compile for int8_t and int16_t.
Compute the span in the unsigned type of the same width, which always
holds it, and return 0 for an empty or zero-step range.
IndexRange::unravel and the N-dimensional box builders had the same
problem one step later: the exclusive end of the last partition is one
step past the last index and wrapped, which dropped the last partition
(4095 of 4096 iterations with the default partitioner on 4 workers). Add
tf::index_at, which returns that end as the range's own end when it is
not representable. for_each_index also no longer advances its index past
the last iteration, which overflowed a signed index near INT_MAX.
Anyone who calls for_each_index, for_each_by_index or reduce_by_index over a range whose span does not fit in the index type (for example for_each_index(0u, UINT_MAX, 1u << 20) to walk a 32-bit space in 1 MiB blocks) got zero iterations and no error, and for_each_index(5, 5, 0) crashed with SIGFPE.
tf::distance computed (end - beg + step - 1) / step in the index type, which went wrong in three ways:
Two more places had the same overflow one step later:
The fix computes the span in std::make_unsigned_t<T>, which always holds it, and returns 0 for an empty or zero-step range. A new helper tf::index_at returns the index at a position, or the range's own end when that position is not representable in T; unravel and the box builders use it. The for_each_index loops no longer step past the last index. Results for ranges that already worked are unchanged.
Tests added to unittests/test_for_each.cpp:
Run locally (g++ 13, C++20, -O1, Linux):
No scheduling code changed, so there is no measurable performance impact: a few integer operations per range, and one fewer index increment per chunk.