Skip to content

Ensure that std::float16_t/std::bfloat16_t support is exact #300

Description

@lemire

In release 8.0.0, we support float16_t and bfloat16_t (thanks @dalle). We have reasonable testing and our code is based on an implementation publicly available since GCC 13 (thanks @jakubjelinek for providing support).

However issues remain:

  1. At least one issue was identified. It is a minor issue so we still released, but it should be fixed. With float16_t, we have that he smallest value (subnormal) that can be represented using float16 is 2**-24 Consider 5.9604644775390625E-8 which is exactly 2**-25. This value is exactly midpoint between the float16 0 and the smallest float16 value. It should be zero (with rounding to even) but it is not. GCC shares this issue and it is quite minor. But there may be other similar issues and this requires investigation. A cause of the issue is that our subnormal code assumes that there cannot be short strings requiring round-to-even: and that is a mathematically proven assumption for 32-bit and 64-bit floats. However that is not true in general.
  2. More generally, we need to go through both the float16_t and bfloat16_t and prove (mathematically) that all parameters are correct and optimal.

Activity

  1. youdie006 commented on Sep 14, 2026

    @youdie006

    On part 1 — "there may be other similar issues and this requires investigation" — here is an exhaustive answer for both 16-bit types.

    Method: for every adjacent pair of finite bit patterns, take the exact midpoint, print it with enough decimal places to be exact, parse it, and compare the raw bit pattern against ties-to-even. Built with g++ -std=c++23 -O2 against a8a02f7.

    type midpoint pairs swept wrong
    float16_t 31,743 9
    bfloat16_t 32,639 0
    float (subnormal range, control) 4,000 0

    The nine are consecutive and all subnormal — pairs 0x0000/0x0001, 0x0002/0x0003, … 0x0010/0x0011, i.e. (2k+1)·2⁻²⁵ for even k ≤ 16. 5.9604644775390625E-8 from your report is the first of them. Each returns the odd neighbour where ties-to-even wants the even one; ec is errc() in every case, so nothing signals it.

    They stop at 0x0010/0x0011 for a reason worth recording: that midpoint needs exactly 19 significant digits, and the next one (0x0012/0x0013) needs 20. Twenty overflows the 19-digit fast path and falls through to the exact big-integer path, which rounds correctly. So the defect is bounded precisely by the fast path's digit limit, not by the subnormal range.

    bfloat16_t is clean because its subnormal midpoints all need far more than 19 digits.

    Cross-checked against numpy float16, which returns the even neighbour for all nine.

    I have not attempted part 2 (the proof that every parameter is correct and optimal), and I am not opening a PR — you have this filed as known and deferred, so this is only the enumeration you asked for.


    Disclosure: I used Claude (an AI assistant) for the sweep. The numbers above are from a harness I ran and cross-checked, not a tool's opinion.

  2. changed the title [-]Insure that std::float16_t/std::bfloat16_t support is exact[/-] [+]Ensure that std::float16_t/std::bfloat16_t support is exact[/+] on Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions