Every frame header and payload read called prepareRead, which returned a
cleanup function to clear the timeout and translate close or cancellation
errors. That function captured the context, connection, and the address of
the caller's named error result. Because prepareRead returned the function,
the closure outlived its stack frame and Go allocated its captured state on
the heap for every frame read.
Move the cleanup logic to a normal finishRead method and defer a direct
method call instead. This preserves timeout cleanup and error translation
without returning a closure. Compiler escape analysis with -gcflags=-m=2
confirms that the old function literal escaped and forced the named error
result in readFrameHeader and readFramePayload onto the heap; neither escape
remains after this change.
The results below compare parent d099e16 with this commit on an Apple
M4 Max using GOMAXPROCS=1 and benchstat over 10 samples. A temporary
in-package harness, not included in this commit, repeatedly called one
internal frame read. Header reads parse a minimal frame header; payload reads
copy 512 bytes from a buffered repeating reader. The background cases use
context.Background, while the cancelable cases use an uncanceled
context.WithCancel.
Removing the escaping closure eliminates 2 allocations and 64 bytes from
every frame read: one allocation for the closure environment and one for the
named error result retained by that closure. Header reads with a background
context improve from 170.5 to 118.5 ns/op, 216 to 152 B/op, and 6 to
4 allocs/op. Header reads with a cancelable context improve from 211.1 to
158.7 ns/op with the same allocation reduction. Payload reads remove the
same fixed overhead; they remain slower because the benchmark also copies
512 bytes.
Both context types benefit because this commit does not change timeout
registration. It only removes cleanup allocations made after every
prepareRead call. An interleaved 12-sample
BenchmarkConn/disabledCompress run, which exercises complete message reads,
improves from 3.875 to 3.657 us/op (-5.63%), 42 to 32 allocs/op (-23.81%),
and 1536 to 1216 B/op (-20.83%).
Replace the per-frame cleanup closure returned by prepareRead with a normal finishRead method.
The closure captured the connection, context, and caller’s named error result. Escape analysis showed that both the closure and error result moved to the heap on every frame header and payload read. The new method preserves timeout cleanup and error mapping without those allocations.
This is admittedly a micro optimization but I decided to hunt out any allocation wins I can get in the direct request path for a project I'm working on, and assessed that the fixes were maintainable and easy to understand.
AI disclosure: I used AI to help find and fix this, but fully understand the problem myself and reviewed the code. The final shape contains manual adaptations, too.
Benchmarks
Each frame read removes 2 allocations and 64 bytes:
The focused frame benchmarks used a temporary in-package harness and are not included in this change. I didn't think you'd want them in the repo cause they're such a micro optimization.