| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
This package is not in the latest version of its module.
Go to latest Published: Mar 30, 2026 License: Apache-2.0The Go module system was introduced in Go 1.11 and is the official dependency management solution for Go.
Redistributable licenses place minimal restrictions on how software can be used, modified, and redistributed.
Modules with tagged versions give importers more predictable builds.
When a project reaches major version v1 it is considered stable.
const ( Version = "2.0.0" VersionMajor = 2 VersionMinor = 0 VersionPatch = 0 )
Version information
const ( SignMask = 0b10000000 // 0x80 - Sign bit mask ExponentMask = 0b01111000 // 0x78 - Exponent bits mask MantissaMask = 0b00000111 // 0x07 - Mantissa bits mask MantissaLen = 3 // Number of mantissa bits // Exponent bias and limits // See https://en.wikipedia.org/wiki/Exponent_bias // bias = 2^(|exponent|-1) - 1 ExponentBias = 7 // Bias for 4-bit exponent ExponentMax = 15 // Maximum exponent value ExponentMin = -7 // Minimum exponent value // Float32 constants for conversion Float32Bias = 127 // IEEE 754 single precision bias // Special values PositiveZero Float8 = 0x00 NegativeZero Float8 = 0x80 PositiveInfinity Float8 = 0x78 // IEEE 754 E4M3FN: S.1111.000 = 0.1111.000₂ NegativeInfinity Float8 = 0xF8 // IEEE 754 E4M3FN: S.1111.000 = 1.1111.000₂ NaN Float8 = 0x7F // IEEE 754 E4M3FN: S.1111.111 (0x7F or 0xFF) MaxValue Float8 = 0x7E // Largest finite positive value MinValue Float8 = 0xFE // Largest finite negative value SmallestPositive Float8 = 0x01 // Smallest positive normalized value )
Bit masks and constants for Float8 format
var ( E = ToFloat8(2.718281828459045) // Euler's number Pi = ToFloat8(3.141592653589793) // Pi Phi = ToFloat8(1.618033988749895) // Golden ratio Sqrt2 = ToFloat8(1.4142135623730951) // Square root of 2 SqrtE = ToFloat8(1.6487212707001282) // Square root of E SqrtPi = ToFloat8(1.7724538509055159) // Square root of Pi Ln2 = ToFloat8(0.6931471805599453) // Natural logarithm of 2 Log2E = ToFloat8(1.4426950408889634) // Base-2 logarithm of E Ln10 = ToFloat8(2.302585092994046) // Natural logarithm of 10 Log10E = ToFloat8(0.4342944819032518) // Base-10 logarithm of E )
Constants as Float8 values
var ( ErrOverflow = &Float8Error{Op: "convert", Msg: "value too large for float8"} ErrUnderflow = &Float8Error{Op: "convert", Msg: "value too small for float8"} ErrNaN = &Float8Error{Op: "convert", Msg: "NaN not representable in float8"} )
Common error instances
var DefaultArithmeticMode = ArithmeticAuto
Global arithmetic mode
var DefaultConversionMode = ModeDefault
Global conversion mode (can be changed for different behavior)
func Configure(config *Config)
Configure applies the given configuration to the package
func DebugInfo() map[string]interface{}
DebugInfo returns debugging information about the package state
func DisableFastArithmetic()
DisableFastArithmetic disables lookup tables and uses algorithmic operations
func DisableFastConversion()
DisableFastConversion disables lookup table and uses algorithmic conversion
func EnableFastArithmetic()
EnableFastArithmetic enables lookup tables for arithmetic operations
func EnableFastConversion()
EnableFastConversion enables lookup table for ToFloat32 conversion
func GetMemoryUsage() int
GetMemoryUsage returns the current memory usage of lookup tables in bytes
ToSlice32 converts a slice of Float8 to float32 with optimized performance.
This function is optimized for batch conversion of Float8 values to float32. It handles all special values correctly, including negative zero, infinity, and NaN.
Parameters:
Returns:
Note: The conversion from Float8 to float32 is always exact since Float8 is a subset of float32. For large slices, consider using a pool of []float32 to reduce allocations.
type ArithmeticMode int
ArithmeticMode defines which implementation to use for arithmetic operations
const ( // ArithmeticAuto chooses the best implementation automatically ArithmeticAuto ArithmeticMode = iota // ArithmeticAlgorithmic forces algorithmic implementation ArithmeticAlgorithmic // ArithmeticLookup forces lookup table implementation (if available) ArithmeticLookup )
type Config struct {
EnableFastArithmetic bool
EnableFastConversion bool
DefaultMode ConversionMode
ArithmeticMode ArithmeticMode
}
Config holds package configuration options
func DefaultConfig() *Config
DefaultConfig returns the default package configuration
type ConversionMode int
ConversionMode defines how conversions handle edge cases
const ( // ModeDefault uses standard IEEE 754 rounding behavior ModeDefault ConversionMode = iota // ModeStrict returns errors for overflow/underflow ModeStrict // ModeFast uses lookup tables when available (default for arithmetic) ModeFast )
type Float8 uint8
Float8 represents an 8-bit floating-point number using the IEEE 754 FP8 E4M3FN format. This format is commonly used in machine learning for reduced-precision arithmetic.
Bit layout:
Special values:
This implementation follows the E4M3FN variant which has no infinities and two NaNs.
Add returns the sum of the operands a and b.
This is a convenience function that calls AddWithMode with DefaultArithmeticMode. For more control over the arithmetic behavior, use AddWithMode directly.
Special cases:
Add(+0, ±0) = +0 Add(-0, -0) = -0 Add(±Inf, ∓Inf) = NaN Add(NaN, x) = NaN Add(x, NaN) = NaN
For finite numbers, the result is rounded to the nearest representable Float8 value using the current rounding mode (typically round-to-nearest-even).
AddSlice performs element-wise addition of two Float8 slices.
This function adds corresponding elements of the input slices and returns a new slice with the results. The input slices must have the same length; otherwise, the function will panic.
Parameters:
Returns:
Panics:
Example:
a := []Float8{1.0, 2.0, 3.0}
b := []Float8{4.0, 5.0, 6.0}
result := AddSlice(a, b) // Returns [5.0, 7.0, 9.0]
func AddWithMode(a, b Float8, mode ArithmeticMode) Float8
AddWithMode returns the sum of the operands a and b using the specified arithmetic mode.
The arithmetic mode determines how the addition is performed:
Special cases are handled according to IEEE 754 rules:
For finite numbers, the result is rounded to the nearest representable Float8 value. If the exact result is exactly halfway between two representable values, it is rounded to the value with an even least significant bit (round-to-nearest-even).
Ceil returns the least integer value greater than or equal to f.
Special cases are:
Ceil(±0) = ±0 Ceil(±Inf) = ±Inf Ceil(NaN) = NaN
For finite x, the result is the least integer value ≥ x. The result is exact (no rounding occurs).
Cos returns the cosine of f (in radians).
Special cases are:
Cos(±0) = 1 Cos(±Inf) = NaN Cos(NaN) = NaN
For finite x, the result is the cosine of x in the range [-1, 1]. The result is rounded to the nearest representable Float8 value.
Div returns the quotient a/b of the operands a and b.
This is a convenience function that calls DivWithMode with DefaultArithmeticMode. For more control over the arithmetic behavior, use DivWithMode directly.
Special cases:
Div(±0, ±0) = NaN Div(±Inf, ±Inf) = NaN Div(x, ±0) = ±Inf for x finite and not zero (sign obeys rule for signs) Div(±Inf, y) = ±Inf for y finite and not zero (sign obeys rule for signs) Div(x, y) = NaN if x or y is NaN
The sign of the result follows the standard sign rules for division. For finite numbers, the result is rounded to the nearest representable Float8 value. Division by zero results in ±Inf with the sign determined by the rule of signs.
func DivWithMode(a, b Float8, mode ArithmeticMode) Float8
DivWithMode performs division with specified arithmetic mode
Floor returns the greatest integer value less than or equal to f.
Special cases are:
Floor(±0) = ±0 Floor(±Inf) = ±Inf Floor(NaN) = NaN
For finite x, the result is the greatest integer value ≤ x. The result is exact (no rounding occurs).
Fmod returns the floating-point remainder of x/y.
The result has the same sign as x and magnitude less than the magnitude of y.
Special cases are:
Fmod(±0, y) = ±0 for y != 0 Fmod(±Inf, y) = NaN Fmod(x, 0) = NaN Fmod(NaN, y) = NaN Fmod(x, NaN) = NaN Fmod(x, ±Inf) = x for x not infinite
For finite x and y (y ≠ 0), the result is x - n*y where n is the integer nearest to x/y. If two integers are equally near, the even one is chosen. The result is rounded to the nearest representable Float8 value.
FromFloat64 converts a float64 to Float8 (with potential precision loss)
Log returns the natural logarithm of f.
Special cases are:
Log(+Inf) = +Inf Log(0) = -Inf Log(x < 0) = NaN Log(NaN) = NaN
For finite x > 0, the result is the natural logarithm of x. The result is rounded to the nearest representable Float8 value.
Max returns the larger of two Float8 values. If either value is NaN, returns NaN. Max(+Inf, x) returns +Inf Max(-Inf, x) returns x (if x is finite or +Inf) Max(x, +Inf) returns +Inf Max(x, -Inf) returns x (if x is finite or +Inf)
Min returns the smaller of two Float8 values. If either value is NaN, returns NaN. Min(+Inf, x) returns x (if x is finite or -Inf) Min(-Inf, x) returns -Inf Min(x, +Inf) returns x (if x is finite or -Inf) Min(x, -Inf) returns -Inf
Mul returns the product of the operands a and b.
This is a convenience function that calls MulWithMode with DefaultArithmeticMode. For more control over the arithmetic behavior, use MulWithMode directly.
Special cases:
Mul(±0, ±Inf) = NaN Mul(±Inf, ±0) = NaN Mul(±0, ±0) = ±0 (sign obeys the rule for signs of zero products) Mul(±0, y) = ±0 for y finite and not zero Mul(±Inf, y) = ±Inf for y finite and not zero Mul(x, y) = NaN if x or y is NaN
The sign of the result follows the standard sign rules for multiplication. For finite numbers, the result is rounded to the nearest representable Float8 value.
MulSlice performs element-wise multiplication of two Float8 slices.
This function multiplies corresponding elements of the input slices and returns a new slice with the results. The input slices must have the same length; otherwise, the function will panic.
Parameters:
Returns:
Panics:
Example:
a := []Float8{1.0, 2.0, 3.0}
b := []Float8{4.0, 5.0, 6.0}
result := MulSlice(a, b) // Returns [4.0, 10.0, 18.0]
func MulWithMode(a, b Float8, mode ArithmeticMode) Float8
MulWithMode performs multiplication with specified arithmetic mode
Pow returns f raised to the power of exp.
Special cases are:
Pow(±0, exp) = ±0 for exp > 0 Pow(±0, exp) = +Inf for exp < 0 Pow(1, exp) = 1 for any exp (even NaN) Pow(f, 0) = 1 for any f (including NaN, +Inf, -Inf) Pow(f, 1) = f for any f Pow(NaN, exp) = NaN Pow(f, NaN) = NaN Pow(±0, -Inf) = +Inf Pow(±0, +Inf) = +0 Pow(+Inf, exp) = +Inf for exp > 0 Pow(+Inf, exp) = +0 for exp < 0 Pow(-Inf, exp) = -0 for exp a negative odd integer Pow(-Inf, exp) = +0 for exp a negative non-odd integer Pow(-Inf, exp) = -Inf for exp a positive odd integer Pow(-Inf, exp) = +Inf for exp a positive non-odd integer Pow(-1, ±Inf) = 1 Pow(f, +Inf) = +Inf for |f| > 1 Pow(f, -Inf) = +0 for |f| > 1 Pow(f, +Inf) = +0 for |f| < 1 Pow(f, -Inf) = +Inf for |f| < 1
The result is rounded to the nearest representable Float8 value.
Round returns the nearest integer value to f, rounding ties to even.
Special cases are:
Round(±0) = ±0 Round(±Inf) = ±Inf Round(NaN) = NaN
For finite x, the result is the nearest integer to x. Ties are rounded to the nearest even integer. The result is exact (no rounding occurs).
ScaleSlice multiplies each element in the slice by a scalar
Sin returns the sine of f (in radians).
Special cases are:
Sin(±0) = ±0 Sin(±Inf) = NaN Sin(NaN) = NaN
For finite x, the result is the sine of x in the range [-1, 1]. The result is rounded to the nearest representable Float8 value.
Sqrt returns the square root of the Float8 value.
Special cases are:
Sqrt(+0) = +0 Sqrt(-0) = -0 Sqrt(+Inf) = +Inf Sqrt(x) = NaN if x < 0 (including -Inf) Sqrt(NaN) = NaN
For finite x ≥ 0, the result is the greatest Float8 value y such that y² ≤ x. The result is rounded to the nearest representable Float8 value.
Sub returns the difference of a-b, i.e., the result of subtracting b from a.
This is a convenience function that calls SubWithMode with DefaultArithmeticMode. For more control over the arithmetic behavior, use SubWithMode directly.
Special cases:
Sub(+0, +0) = +0 Sub(+0, -0) = +0 Sub(-0, +0) = -0 Sub(-0, -0) = +0 Sub(±Inf, ±Inf) = NaN Sub(NaN, x) = NaN Sub(x, NaN) = NaN
For finite numbers, the result is rounded to the nearest representable Float8 value.
func SubWithMode(a, b Float8, mode ArithmeticMode) Float8
SubWithMode performs subtraction with specified arithmetic mode
SumSlice returns the sum of all elements in the slice.
This function computes the sum of all Float8 values in the input slice. If the slice is empty, it returns PositiveZero.
The summation is performed using the standard addition rules for Float8, including proper handling of special values (NaN, Inf, etc.).
Parameters:
Returns:
Example:
s := []Float8{1.0, 2.0, 3.0, 4.0}
sum := SumSlice(s) // Returns 10.0
Tan returns the tangent of f (in radians).
Special cases are:
Tan(±0) = ±0 Tan(±Inf) = NaN Tan(NaN) = NaN
For finite x, the result is the tangent of x. The result is rounded to the nearest representable Float8 value. Note that the result may be extremely large or small for inputs near (2n+1)π/2.
ToFloat8 converts a float32 value to Float8 format using the default conversion mode.
This is a convenience function that calls ToFloat8WithMode with DefaultConversionMode. For more control over the conversion process, use ToFloat8WithMode directly.
Special cases:
For finite numbers, the conversion may lose precision or result in overflow/underflow. The default mode handles these cases by saturating to the maximum/minimum representable values.
func ToFloat8WithMode(f32 float32, mode ConversionMode) (Float8, error)
ToFloat8WithMode converts a float32 to Float8 with the specified conversion mode.
The conversion mode determines how edge cases are handled:
Special cases are handled as follows:
For finite numbers, the conversion follows these steps:
Returns the converted Float8 value and an error if the conversion fails in strict mode.
ToSlice8 converts a slice of float32 to Float8 with optimized performance.
This function is optimized for batch conversion of float32 values to Float8. It handles special values correctly, including negative zero, infinity, and NaN.
Parameters:
Returns:
Note: This function preserves negative zero by checking the sign bit of zero values. For large slices, consider using a pool of []Float8 to reduce allocations.
Trunc returns the integer value of f with any fractional part removed.
Special cases are:
Trunc(±0) = ±0 Trunc(±Inf) = ±Inf Trunc(NaN) = NaN
For finite x, the result is the integer part of x with the sign of x. This is equivalent to rounding toward zero. The result is exact (no rounding occurs).
Abs returns the absolute value of f.
Special cases are:
Abs(±Inf) = +Inf Abs(NaN) = NaN Abs(±0) = +0
For all other values, Abs clears the sign bit to return a positive number.
IsFinite reports whether f is a finite value (not infinite and not NaN).
A Float8 value is finite if its exponent is not all 1s (0x0F). This includes both normal numbers (with an implicit leading 1 bit) and subnormal numbers (with an implicit leading 0 bit).
Returns:
IsInf reports whether f is an infinity, either positive or negative.
In the E4M3FN format, infinity values have all exponent bits set (0x78 for +Inf, 0xF8 for -Inf) and a zero mantissa. This is different from the standard IEEE 754 format used in float32/float64.
Returns:
IsNaN reports whether f is a "not-a-number" (NaN) value.
In the E4M3FN format, NaN is represented with all exponent bits set (0x0F) and all mantissa bits set (0x07). This results in two possible NaN values: 0x7F (positive NaN) and 0xFF (negative NaN).
Returns:
IsNormal returns true if the Float8 is a normal (non-zero, non-infinite) number
IsZero reports whether f represents the floating-point value zero (either positive or negative).
According to IEEE 754, both +0 and -0 are considered zero, though they may have different bit patterns and behave differently in certain operations (like 1/+0 = +Inf, 1/-0 = -Inf).
Returns:
Sign returns the sign of the Float8 value.
The return values are:
Note that negative zero is treated as zero (returns 0), following the IEEE 754 standard where +0 and -0 compare as equal. However, they can be distinguished using bitwise operations or by examining the sign bit directly.
For NaN values, Sign returns 0, consistent with math/big.Float's behavior.
ToFloat32 converts a Float8 value to float32.
This conversion is always exact since Float8 is a subset of float32. Special values are preserved:
The conversion uses a fast path for common values and falls back to algorithmic conversion for other values.
type Float8Error struct {
Op string // Operation that caused the error
Value float32 // Input value that caused the error (if applicable)
Msg string // Error message
}
Float8Error represents errors that can occur during Float8 operations
func (e *Float8Error) Error() string
| Back | FazBrowse Home | New Git URL |