af_device_info documents a recommended minimum size of 64 bytes for
d_name (docs/details/device.dox, device_func_prop), and the CPU and
OpenCL backends write at most 64. The CUDA backend wrote up to 256 and
its sanitize loop then read d_name[256], touching 257 bytes of a buffer
callers were told to size at 64.
Clamp the write to 64 and bound the lookahead loop at 63 so it stays
inside the documented buffer.
Confirmed with AddressSanitizer against a 64-byte heap allocation:
before, "heap-buffer-overflow ... WRITE of size 84"; after, clean.
Fixes #3712
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012SRueobKykUocB1SRdHsEE
The CUDA backend wrote up to 256 bytes into d_name and then scanned 256 bytes sanitizing it, while af_device_info documents 64 and the CPU and OpenCL backends respect that, so conforming callers overflowed. Write and scan within 64 bytes and stop at the terminator. Fixes #3712.