| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
This package is not in the latest version of its module.
Go to latest Published: Mar 30, 2026 License: Apache-2.0The Go module system was introduced in Go 1.11 and is the official dependency management solution for Go.
Redistributable licenses place minimal restrictions on how software can be used, modified, and redistributed.
Modules with tagged versions give importers more predictable builds.
When a project reaches major version v1 it is considered stable.
Package float16 implements both 16-bit floating point data types: - Float16: IEEE 754-2008 half-precision (1 sign, 5 exponent, 10 mantissa bits) - BFloat16: "Brain Floating Point" format (1 sign, 8 exponent, 7 mantissa bits)
This implementation provides conversion between both types and other floating-point types (float32 and float64) with support for various rounding modes and error handling.
The float16 type supports all IEEE 754-2008 special values:
When converting to higher-precision types (float32/float64), subnormal float16 values are preserved. However, when converting back from higher-precision types to float16, subnormal values may be rounded to the nearest representable normal float16 value. This behavior is consistent with many hardware implementations that handle subnormals in a similar way for performance reasons.
The following rounding modes are supported for conversions:
Conversion functions with a ConversionMode parameter can return errors for:
See: http://en.wikipedia.org/wiki/Half-precision_floating-point_format
const ( BFloat16SignMask = 0x8000 // 0b1000000000000000 - Sign bit mask BFloat16ExponentMask = 0x7F80 // 0b0111111110000000 - Exponent bits mask BFloat16MantissaMask = 0x007F // 0b0000000001111111 - Mantissa bits mask BFloat16MantissaLen = 7 // Number of mantissa bits BFloat16ExponentLen = 8 // Number of exponent bits // Exponent bias and limits for BFloat16 // bias = 2^(exponent_bits-1) - 1 = 2^7 - 1 = 127 (same as Float32) BFloat16ExponentBias = 127 // Bias for 8-bit exponent BFloat16ExponentMax = 255 // Maximum exponent value BFloat16ExponentMin = 0 // Minimum exponent value // Normalized exponent range BFloat16ExponentNormalMin = 1 // Minimum normalized exponent BFloat16ExponentNormalMax = 254 // Maximum normalized exponent (infinity at 255) // Special exponent values BFloat16ExponentZero = 0 // Zero and subnormal numbers BFloat16ExponentInfinity = 255 // Infinity and NaN )
BFloat16 format constants
const ( Version = "1.0.0" VersionMajor = 1 VersionMinor = 0 VersionPatch = 0 )
Package version information
const ( SignMask = 0x8000 // 0b1000000000000000 - Sign bit mask ExponentMask = 0x7C00 // 0b0111110000000000 - Exponent bits mask MantissaMask = 0x03FF // 0b0000001111111111 - Mantissa bits mask MantissaLen = 10 // Number of mantissa bits ExponentLen = 5 // Number of exponent bits // Exponent bias and limits for IEEE 754 half-precision // bias = 2^(exponent_bits-1) - 1 = 2^4 - 1 = 15 ExponentBias = 15 // Bias for 5-bit exponent ExponentMax = 31 // Maximum exponent value (11111 binary) ExponentMin = 0 // Minimum exponent value // Normalized exponent range ExponentNormalMin = 1 // Minimum normalized exponent ExponentNormalMax = 30 // Maximum normalized exponent (infinity at 31) // Float32 constants for conversion Float32ExponentBias = 127 // IEEE 754 single precision bias Float32ExponentLen = 8 // Float32 exponent bits Float32MantissaLen = 23 // Float32 mantissa bits // Special exponent values ExponentZero = 0 // Zero and subnormal numbers ExponentInfinity = 31 // Infinity and NaN )
IEEE 754 half-precision format constants
var ( DefaultArithmeticMode = ModeIEEEArithmetic DefaultRounding = DefaultRoundingMode )
Global arithmetic settings
var ( BFloat16Zero = BFloat16PositiveZero BFloat16One = BFloat16FromFloat32(1.0) BFloat16Two = BFloat16FromFloat32(2.0) BFloat16Half = BFloat16FromFloat32(0.5) BFloat16E = BFloat16FromFloat32(float32(math.E)) BFloat16Pi = BFloat16FromFloat32(float32(math.Pi)) BFloat16Sqrt2 = BFloat16FromFloat32(float32(math.Sqrt2)) )
Convenience constants for common BFloat16 values
var ( DefaultConversionMode ConversionMode = ModeIEEE DefaultRoundingMode RoundingMode = RoundNearestEven )
var ( // Common integer values Zero16 = PositiveZero One16 = FromFloat32(1.0) Two16 = FromFloat32(2.0) Three16 = FromFloat32(3.0) Four16 = FromFloat32(4.0) Five16 = FromFloat32(5.0) Ten16 = FromFloat32(10.0) // Common fractional values Half16 = FromFloat32(0.5) Quarter16 = FromFloat32(0.25) Third16 = FromFloat32(1.0 / 3.0) // Special mathematical values NaN16 = QuietNaN PosInf = PositiveInfinity NegInf = NegativeInfinity // Commonly used constants Deg2Rad = FromFloat32(float32(math.Pi / 180.0)) // Degrees to radians Rad2Deg = FromFloat32(float32(180.0 / math.Pi)) // Radians to degrees )
Constants for common values
var ( E = FromFloat32(float32(math.E)) // Euler's number Pi = FromFloat32(float32(math.Pi)) // Pi Phi = FromFloat32(float32(math.Phi)) // Golden ratio Sqrt2 = FromFloat32(float32(math.Sqrt2)) // Square root of 2 SqrtE = FromFloat32(float32(math.SqrtE)) // Square root of E SqrtPi = FromFloat32(float32(math.SqrtPi)) // Square root of Pi SqrtPhi = FromFloat32(float32(math.SqrtPhi)) // Square root of Phi Ln2 = FromFloat32(float32(math.Ln2)) // Natural logarithm of 2 Log2E = FromFloat32(float32(math.Log2E)) // Base-2 logarithm of E Ln10 = FromFloat32(float32(math.Ln10)) // Natural logarithm of 10 Log10E = FromFloat32(float32(math.Log10E)) // Base-10 logarithm of E )
Mathematical constants as Float16 values
BFloat16Equal returns true if a equals b
BFloat16Greater returns true if a > b
BFloat16GreaterEqual returns true if a >= b
BFloat16Less returns true if a < b
BFloat16LessEqual returns true if a <= b
BFloat16ToSlice32 converts a slice of BFloat16 values to float32
BFloat16ToSlice64 converts a slice of BFloat16 values to float64
func Configure(cfg *Config)
Configure applies the given configuration to the package
func DebugInfo() map[string]interface{}
DebugInfo returns debugging information about the package state
func GetBenchmarkOperations() map[string]BenchmarkOperation
GetBenchmarkOperations returns a map of operations suitable for benchmarking
func GetMemoryUsage() int
GetMemoryUsage returns the current memory usage of the package in bytes
IsInf reports whether f is an infinity, according to sign If sign > 0, IsInf reports whether f is positive infinity If sign < 0, IsInf reports whether f is negative infinity If sign == 0, IsInf reports whether f is either infinity
IsNormal reports whether f is a normal number (not zero, subnormal, infinite, or NaN)
IsSubnormal reports whether f is a subnormal number
ValidateSliceLength checks if two slices have the same length
type ArithmeticMode int
ArithmeticMode defines the precision/performance trade-off for arithmetic operations
const ( // ModeIEEE provides full IEEE 754 compliance with proper rounding ModeIEEEArithmetic ArithmeticMode = iota // ModeFastArithmetic optimizes for speed, may sacrifice some precision ModeFastArithmetic // ModeExactArithmetic provides exact results when possible, errors on precision loss ModeExactArithmetic )
type BFloat16 uint16
BFloat16 represents a 16-bit "Brain Floating Point" format value Used by Google Brain, TensorFlow, and various ML frameworks Format: 1 sign bit, 8 exponent bits, 7 mantissa bits
const ( BFloat16PositiveZero BFloat16 = 0x0000 // +0.0 BFloat16NegativeZero BFloat16 = 0x8000 // -0.0 BFloat16PositiveInfinity BFloat16 = 0x7F80 // +∞ BFloat16NegativeInfinity BFloat16 = 0xFF80 // -∞ BFloat16QuietNaN BFloat16 = 0x7FC0 // Quiet NaN BFloat16SignalingNaN BFloat16 = 0x7F81 // Signaling NaN // Largest finite values BFloat16MaxValue BFloat16 = 0x7F7F // Largest positive normal BFloat16MinValue BFloat16 = 0xFF7F // Largest negative normal (most negative) BFloat16SmallestPos BFloat16 = 0x0080 // Smallest positive normal BFloat16SmallestNeg BFloat16 = 0x8080 // Smallest negative normal // Smallest subnormal values BFloat16SmallestPosSubnormal BFloat16 = 0x0001 // Smallest positive subnormal BFloat16SmallestNegSubnormal BFloat16 = 0x8001 // Smallest negative subnormal )
Special BFloat16 values
BFloat16Abs returns the absolute value of b
BFloat16Add adds two BFloat16 values
BFloat16AddSlice performs element-wise addition of two BFloat16 slices
func BFloat16AddWithMode(a, b BFloat16, mode ArithmeticMode, rounding RoundingMode) (BFloat16, error)
BFloat16AddWithMode performs addition with specified arithmetic and rounding modes.
BFloat16Cos returns the cosine of b (in radians).
BFloat16Div divides two BFloat16 values
BFloat16DivSlice performs element-wise division of two BFloat16 slices
func BFloat16DivWithMode(a, b BFloat16, mode ArithmeticMode, rounding RoundingMode) (BFloat16, error)
BFloat16DivWithMode performs division with specified arithmetic and rounding modes.
BFloat16FMA computes a fused multiply-add (a*b + c) for BFloat16 values. This is a stub that returns an error; a full implementation is planned for a future phase.
BFloat16FastSigmoid computes an approximate sigmoid using a rational polynomial. Uses the approximation: sigmoid(x) ≈ 0.5 + 0.5 * x / (1 + |x|) which avoids exp() entirely.
BFloat16FastTanh computes an approximate tanh using a rational polynomial. Uses the approximation: tanh(x) ≈ x*(27 + x*x) / (27 + 9*x*x) which is a Padé approximant accurate to within ~0.004 for |x| < 3.
FromBits creates a BFloat16 from its bit representation
BFloat16FromFloat16 converts a Float16 to BFloat16
FromFloat32 converts a float32 to BFloat16 using round-to-nearest-even BFloat16 is essentially a truncated float32, so conversion is straightforward
func BFloat16FromFloat32WithMode(f32 float32, convMode ConversionMode, roundMode RoundingMode) (BFloat16, error)
BFloat16FromFloat32WithMode converts a float32 to BFloat16 with specified conversion and rounding modes.
func BFloat16FromFloat32WithRounding(f float32, mode RoundingMode) BFloat16
BFloat16FromFloat32WithRounding converts a float32 to BFloat16 with the specified rounding mode.
FromFloat64 converts a float64 to BFloat16
func BFloat16FromFloat64WithMode(f64 float64, convMode ConversionMode, roundMode RoundingMode) (BFloat16, error)
BFloat16FromFloat64WithMode converts a float64 to BFloat16 with specified conversion and rounding modes.
func BFloat16FromFloat64WithRounding(f float64, mode RoundingMode) BFloat16
BFloat16FromFloat64WithRounding converts a float64 to BFloat16 with the specified rounding mode.
BFloat16FromSlice32 converts a slice of float32 values to BFloat16
BFloat16FromSlice64 converts a slice of float64 values to BFloat16
BFloat16FromString parses a string into a BFloat16 value. It handles special values (NaN, Inf) and numeric strings.
BFloat16Log returns the natural logarithm of b.
BFloat16Log2 returns the base-2 logarithm of b.
BFloat16Max returns the larger of a or b
BFloat16Min returns the smaller of a or b
BFloat16Mul multiplies two BFloat16 values
BFloat16MulSlice performs element-wise multiplication of two BFloat16 slices
func BFloat16MulWithMode(a, b BFloat16, mode ArithmeticMode, rounding RoundingMode) (BFloat16, error)
BFloat16MulWithMode performs multiplication with specified arithmetic and rounding modes.
BFloat16Neg returns the negation of b
BFloat16ScaleSlice multiplies each element in the slice by a scalar
BFloat16Sigmoid returns 1 / (1 + exp(-b)).
BFloat16Sin returns the sine of b (in radians).
BFloat16Sqrt returns the square root of the BFloat16 value.
BFloat16Sub subtracts two BFloat16 values
BFloat16SubSlice performs element-wise subtraction of two BFloat16 slices
func BFloat16SubWithMode(a, b BFloat16, mode ArithmeticMode, rounding RoundingMode) (BFloat16, error)
BFloat16SubWithMode performs subtraction with specified arithmetic and rounding modes.
BFloat16SumSlice returns the sum of all elements in the slice
BFloat16Tanh returns the hyperbolic tangent of b.
func (b BFloat16) Class() FloatClass
Class returns the IEEE 754 classification of the BFloat16 value
CopySign returns a value with the magnitude of f and the sign of s
Format implements fmt.Formatter, supporting %e, %f, %g, %E, %F, %G, %v, and %s verbs.
GoString returns a Go syntax representation of the BFloat16 value.
IsSubnormal reports whether b is a subnormal number
IsZero returns true if the BFloat16 is zero (positive or negative)
MarshalBinary implements encoding.BinaryMarshaler. The encoding is 2 bytes in little-endian order.
MarshalJSON implements json.Marshaler.
UnmarshalBinary implements encoding.BinaryUnmarshaler. The encoding is 2 bytes in little-endian order.
UnmarshalJSON implements json.Unmarshaler.
BFloat16Error provides detailed error information for bfloat16 operations
func (e *BFloat16Error) Error() string
BenchmarkOperation represents a benchmarkable operation
type Config struct {
DefaultConversionMode ConversionMode
DefaultRoundingMode RoundingMode
DefaultArithmeticMode ArithmeticMode
EnableFastMath bool // Package float16 implements the 16-bit floating point data type (IEEE 754-2008).
}
Package configuration
func DefaultConfig() *Config
DefaultConfig returns the default package configuration
type ConversionMode int
ConversionMode controls error reporting behavior for conversions
const ( // ModeIEEE performs IEEE-style conversion, saturating to Inf/0 with no errors ModeIEEE ConversionMode = iota // ModeStrict reports errors for NaN, Inf, overflow, and underflow ModeStrict )
type ErrorCode int
ErrorCode represents specific error categories for float16 operations
type Float16 uint16
Float16 represents a 16-bit IEEE 754 half-precision floating-point value
const ( PositiveZero Float16 = 0x0000 // +0.0 NegativeZero Float16 = 0x8000 // -0.0 PositiveInfinity Float16 = 0x7C00 // +∞ NegativeInfinity Float16 = 0xFC00 // -∞ // Largest finite values MaxValue Float16 = 0x7BFF // Largest positive finite value (~65504) MinValue Float16 = 0xFBFF // Largest negative finite value (~-65504) // Smallest normalized positive value SmallestNormal Float16 = 0x0400 // 2^-14 ≈ 6.103515625e-05 // Largest subnormal value LargestSubnormal Float16 = 0x03FF // (1023/1024) * 2^-14 ≈ 6.097555161e-05 // Smallest positive subnormal value SmallestSubnormal Float16 = 0x0001 // 2^-24 ≈ 5.960464478e-08 // Common NaN representations QuietNaN Float16 = 0x7E00 // Quiet NaN (most significant mantissa bit set) SignalingNaN Float16 = 0x7D00 // Signaling NaN NegativeQNaN Float16 = 0xFE00 // Negative quiet NaN )
Special values following IEEE 754 half-precision standard
func AddWithMode(a, b Float16, mode ArithmeticMode, rounding RoundingMode) (Float16, error)
AddWithMode performs addition with specified arithmetic and rounding modes
func DivWithMode(a, b Float16, mode ArithmeticMode, rounding RoundingMode) (Float16, error)
DivWithMode performs division with specified arithmetic and rounding modes
DotProduct computes the dot product of two Float16 slices
Float16FromBFloat16 converts a BFloat16 to Float16
Frexp breaks f into a normalized fraction and an integral power of two It returns frac and exp satisfying f == frac × 2^exp, with the absolute value of frac in the interval [0.5, 1) or zero
FromFloat32 converts a float32 value to a Float16 value. It handles special cases like NaN, infinities, and zeros. The conversion follows IEEE 754-2008 rules for half-precision.
func FromFloat32WithRounding(f32 float32, mode RoundingMode) Float16
FromFloat32WithRounding converts a float32 to Float16 using the provided rounding mode. It mirrors fromFloat32New but respects the explicit rounding mode instead of always rounding to nearest-even.
FromFloat64 converts a float64 value to a Float16 value. It handles special cases like NaN, infinities, and zeros.
func FromFloat64WithMode(f64 float64, convMode ConversionMode, roundMode RoundingMode) (Float16, error)
FromFloat64WithMode converts a float64 to Float16 with specified conversion and rounding modes
FromSlice64 converts a slice of float64 to a slice of Float16
Inf returns a Float16 infinity value If sign >= 0, returns positive infinity If sign < 0, returns negative infinity
Modf returns integer and fractional floating-point numbers that sum to f Both values have the same sign as f
func MulWithMode(a, b Float16, mode ArithmeticMode, rounding RoundingMode) (Float16, error)
MulWithMode performs multiplication with specified arithmetic and rounding modes
NextAfter returns the next representable Float16 value after f in the direction of g
Parse converts a string to a Float16 value This is a simplified implementation for testing
ParseFloat converts a string to a Float16 value. The precision parameter is ignored for Float16. It returns the Float16 value and an error if the string cannot be parsed.
RoundToEven returns the nearest integer value to f, rounding ties to even
ScaleSlice multiplies each element in the slice by a scalar
func SubWithMode(a, b Float16, mode ArithmeticMode, rounding RoundingMode) (Float16, error)
SubWithMode performs subtraction with specified arithmetic and rounding modes
ToFloat16 converts a float64 to a Float16 value. This is a convenience wrapper used in tests and utilities.
ToSlice16 converts a slice of float32 to a slice of Float16. This is a convenience wrapper used in tests and utilities.
func ToSlice16WithMode(s []float32, convMode ConversionMode, roundMode RoundingMode) ([]Float16, []error)
ToSlice16WithMode converts a slice of float32 to Float16 with specified modes
VectorAdd performs vectorized addition (placeholder for future SIMD implementation)
VectorMul performs vectorized multiplication (placeholder for future SIMD implementation)
func (f Float16) Class() FloatClass
Class returns the IEEE 754 classification of the value
IsFinite returns true if the Float16 value is finite (not infinity or NaN)
IsInf returns true if the Float16 value represents infinity If sign > 0, returns true only for positive infinity If sign < 0, returns true only for negative infinity If sign == 0, returns true for either infinity
IsNormal returns true if the Float16 value is normalized (not zero, subnormal, infinite, or NaN)
IsSubnormal returns true if the Float16 value is subnormal (denormalized)
IsZero returns true if the Float16 value represents zero (positive or negative)
Sign returns the sign of the Float16 value: 1 for positive, -1 for negative, 0 for zero
ToBFloat16 converts a Float16 to BFloat16
ToFloat32 converts a Float16 value to a float32 value. It handles special cases like NaN, infinities, and zeros.
ToFloat64 converts a Float16 value to a float64 value. It handles special cases like NaN, infinities, and zeros.
Float16Error provides detailed error information for float16 operations
func (e *Float16Error) Error() string
type FloatClass int
FloatClass enumerates the IEEE 754 classification of a Float16 value
const ( ClassPositiveZero FloatClass = iota ClassNegativeZero ClassPositiveSubnormal ClassNegativeSubnormal ClassPositiveNormal ClassNegativeNormal ClassPositiveInfinity ClassNegativeInfinity ClassQuietNaN ClassSignalingNaN )
type RoundingMode int
RoundingMode controls how results are rounded during conversion/arithmetic
const ( // Round to nearest, ties to even RoundNearestEven RoundingMode = iota // Round toward zero (truncate) RoundTowardZero // Round toward +Inf RoundTowardPositive // Round toward -Inf RoundTowardNegative // Round to nearest, ties away from zero RoundNearestAway )
SliceStats computes basic statistics for a Float16 slice
func ComputeSliceStats(s []Float16) SliceStats
ComputeSliceStats calculates statistics for a Float16 slice
| Back | FazBrowse Home | New Git URL |