15. Floating Point Arithmetic: Issues and Limitations

Floating-point numbers are represented in computer hardware as base 2 (binary) fractions. For example, the decimal fraction 0.125 has value 1/10 + 2/100 + 5/1000, and in the same way the binary fraction 0.001 has value 0/2 + 0/4 + 1/8. These two fractions have identical values, the only real difference being that the first is written in base 10 fractional notation, and the second in base 2.

Leider lassen sich die meisten Dezimalbrüche nicht exakt als Binärbrüche darstellen. Dies hat zur Folge, dass die von dir eingegebenen Dezimal-Gleitkommazahlen im Allgemeinen nur annähernd durch die tatsächlich im Rechner gespeicherten Binär-Gleitkommazahlen wiedergegeben werden.

Das Problem lässt sich zunächst im Dezimalsystem leichter verstehen. Betrachten wir den Bruch 1/3. Man kann ihn als Bruch im Dezimalsystem näherungsweise darstellen:

0.3

oder besser noch:

0.33

oder besser noch:

0.333

und so weiter. Ganz gleich, wie viele Stellen man notieren möchte, das Ergebnis wird niemals genau 1/3 sein, sondern eine immer bessere Annäherung an 1/3 darstellen.

Ebenso lässt sich der Dezimalwert 0,1 – ganz gleich, wie viele Ziffern im Zweier-System man verwenden möchte – nicht exakt als Bruch im Zweier-System darstellen. Im Zweier-System ist 1/10 der sich unendlich wiederholende Bruch

0.0001100110011001100110011001100110011001100110011...

Hält man bei einer beliebigen endlichen Anzahl von Bits an, erhält man eine Näherung. Auf den meisten heutigen Rechnern werden Gleitkommazahlen mithilfe eines binären Bruchs approximiert, wobei der Zähler aus den ersten 53 Bits (beginnend mit dem höchstwertigen Bit) besteht und der Nenner eine Zweierpotenz ist. Im Fall von 1/10 lautet der Binärbruch 3602879701896397 / 2 ** 55, was nahe am tatsächlichen Wert von 1/10 liegt, diesem aber nicht exakt entspricht.

Many users are not aware of the approximation because of the way values are displayed. Python only prints a decimal approximation to the true decimal value of the binary approximation stored by the machine. On most machines, if Python were to print the true decimal value of the binary approximation stored for 0.1, it would have to display

>>> 0.1
0.1000000000000000055511151231257827021181583404541015625

That is more digits than most people find useful, so Python keeps the number of digits manageable by displaying a rounded value instead

>>> 1 / 10
0.1

Denke daran: Auch wenn das ausgegebene Ergebnis genau wie der Wert 1/10 aussieht, ist der tatsächlich gespeicherte Wert der nächstgelegene darstellbare Binärbruch.

Interessanterweise gibt es viele verschiedene Dezimalzahlen, die denselben nächstliegenden binären Näherungsbruch haben. Beispielsweise werden die Zahlen 0.1, 0.10000000000000001 und 0.1000000000000000055511151231257827021181583404541015625 alle durch 3602879701896397 / 2 ** 55 approximiert. Da alle diese Dezimalwerte dieselbe Approximation haben, könnte jeder einzelne von ihnen angezeigt werden, ohne dass die Invariante eval(repr(x)) == x verloren geht.

Früher wählten die Python-Eingabeaufforderung und die integrierte Funktion repr() die Zahl mit 17 signifikanten Stellen aus: 0.10000000000000001. Ab Python 3.1 ist Python (auf den meisten Systemen) nun in der Lage, die kürzeste dieser Zahlen auszuwählen und einfach 0.1 anzuzeigen.

Note that this is in the very nature of binary floating-point: this is not a bug in Python, and it is not a bug in your code either. You’ll see the same kind of thing in all languages that support your hardware’s floating-point arithmetic (although some languages may not display the difference by default, or in all output modes).

For more pleasant output, you may wish to use string formatting to produce a limited number of significant digits:

>>> format(math.pi, '.12g')  # give 12 significant digits
'3.14159265359'

>>> format(math.pi, '.2f')   # give 2 digits after the point
'3.14'

>>> repr(math.pi)
'3.141592653589793'

Man muss sich bewusst machen, dass es sich hierbei im wahrsten Sinne des Wortes um eine Illusion handelt: Man rundet lediglich die Anzeige des tatsächlichen Maschinenwerts.

One illusion may beget another. For example, since 0.1 is not exactly 1/10, summing three values of 0.1 may not yield exactly 0.3, either:

>>> .1 + .1 + .1 == .3
False

Also, since the 0.1 cannot get any closer to the exact value of 1/10 and 0.3 cannot get any closer to the exact value of 3/10, then pre-rounding with round() function cannot help:

>>> round(.1, 1) + round(.1, 1) + round(.1, 1) == round(.3, 1)
False

Though the numbers cannot be made closer to their intended exact values, the round() function can be useful for post-rounding so that results with inexact values become comparable to one another:

>>> round(.1 + .1 + .1, 10) == round(.3, 10)
True

Binary floating-point arithmetic holds many surprises like this. The problem with „0.1“ is explained in precise detail below, in the „Representation Error“ section. See Examples of Floating Point Problems for a pleasant summary of how binary floating-point works and the kinds of problems commonly encountered in practice. Also see The Perils of Floating Point for a more complete account of other common surprises.

As that says near the end, „there are no easy answers.“ Still, don’t be unduly wary of floating-point! The errors in Python float operations are inherited from the floating-point hardware, and on most machines are on the order of no more than 1 part in 2**53 per operation. That’s more than adequate for most tasks, but you do need to keep in mind that it’s not decimal arithmetic and that every float operation can suffer a new rounding error.

Es gibt zwar Ausnahmefälle, doch bei den meisten alltäglichen Anwendungen der Gleitkommaarithmetik erhältst du letztendlich das erwartete Ergebnis, wenn du die Anzeige deiner Endergebnisse einfach auf die gewünschte Anzahl von Dezimalstellen rundest. str() reicht in der Regel aus; für eine feinere Steuerung findest du die Formatbezeichner der Methode str.format() unter Format String Syntax.

Für Anwendungsfälle, die eine exakte Dezimaldarstellung erfordern, solltest du das Modul decimal verwenden, das eine Dezimalarithmetik implementiert, die für Buchhaltungsanwendungen und Anwendungen mit hoher Genauigkeit geeignet ist.

Eine weitere Form der exakten Arithmetik wird vom Modul fractions unterstützt, das eine auf rationalen Zahlen basierende Arithmetik implementiert (sodass Zahlen wie 1/3 exakt dargestellt werden können).

Wenn du häufig mit Gleitkommaoperationen arbeitest, solltest du das NumPy-Paket sowie viele andere Pakete für mathematische und statistische Operationen ansehen, die vom SciPy-Projekt bereitgestellt werden. Siehe <https://scipy.org>.

Python provides tools that may help on those rare occasions when you really do want to know the exact value of a float. The float.as_integer_ratio() method expresses the value of a float as a fraction:

>>> x = 3.14159
>>> x.as_integer_ratio()
(3537115888337719, 1125899906842624)

Since the ratio is exact, it can be used to losslessly recreate the original value:

>>> x == 3537115888337719 / 1125899906842624
True

The float.hex() method expresses a float in hexadecimal (base 16), again giving the exact value stored by your computer:

>>> x.hex()
'0x1.921f9f01b866ep+1'

This precise hexadecimal representation can be used to reconstruct the float value exactly:

>>> x == float.fromhex('0x1.921f9f01b866ep+1')
True

Da die Darstellung exakt ist, eignet sie sich gut für die zuverlässige Übertragung von Werten zwischen verschiedenen Python-Versionen (Plattformunabhängigkeit) und den Datenaustausch mit anderen Sprachen, die dasselbe Format unterstützen (wie beispielsweise Java und C99).

Another helpful tool is the math.fsum() function which helps mitigate loss-of-precision during summation. It tracks „lost digits“ as values are added onto a running total. That can make a difference in overall accuracy so that the errors do not accumulate to the point where they affect the final total:

>>> sum([0.1] * 10) == 1.0
False
>>> math.fsum([0.1] * 10) == 1.0
True

15.1. Darstellungsfehler

In diesem Abschnitt wird das Beispiel „0.1“ ausführlich erläutert und gezeigt, wie du Fälle wie diesen selbst genau analysieren kannst. Grundkenntnisse über die binäre Gleitkommadarstellung werden vorausgesetzt.

Darstellungsfehler bezieht sich auf die Tatsache, dass einige (eigentlich die meisten) Dezimalbrüche nicht exakt als Binärbrüche (Basis 2) dargestellt werden können. Dies ist der Hauptgrund dafür, dass Python (oder Perl, C, C++, Java, Fortran und viele andere) oft nicht genau die Dezimalzahl anzeigen, die man erwartet.

Warum ist das so? 1/10 lässt sich nicht exakt als binärer Bruch darstellen. Seit mindestens dem Jahr 2000 verwenden fast alle Rechner die binäre IEEE-754-Gleitkommaarithmetik, und fast alle Plattformen ordnen Python-Float-Werte den IEEE-754-binary64-Werten mit „doppelter Genauigkeit“ zu. IEEE-754-binary64-Werte enthalten 53 Bits an Genauigkeit, daher versucht der Computer bei der Eingabe, 0,1 in den nächstgelegenen Bruch der Form J/2**N umzuwandeln, wobei J eine Ganzzahl mit genau 53 Bits ist. Umformulierung

1 / 10 ~= J / (2**N)

als

J ~= 2**N / 10

and recalling that J has exactly 53 bits (is >= 2**52 but < 2**53), the best value for N is 56:

>>> 2**52 <=  2**56 // 10  < 2**53
True

That is, 56 is the only value for N that leaves J with exactly 53 bits. The best possible value for J is then that quotient rounded:

>>> q, r = divmod(2**56, 10)
>>> r
6

Since the remainder is more than half of 10, the best approximation is obtained by rounding up:

>>> q+1
7205759403792794

Daher lautet die bestmögliche Annäherung an 1/10 in IEEE 754-Doppelnauigkeit:

7205759403792794 / 2 ** 56

Durch Division von Zähler und Nenner durch zwei lässt sich der Bruch wie folgt vereinfachen:

3602879701896397 / 2 ** 55

Beachte, dass dieser Wert, da wir aufgerundet haben, tatsächlich etwas größer als 1/10 ist; hätten wir nicht aufgerundet, wäre der Quotient etwas kleiner als 1/10 gewesen. Aber auf keinen Fall kann er genau 1/10 betragen!

Der Computer „sieht“ also niemals 1/10: Was er sieht, ist genau der oben angegebene Bruch, die bestmögliche IEEE-754-Doppelgenauigkeits-Annäherung, die er erzielen kann:

>>> 0.1 * 2 ** 55
3602879701896397.0

If we multiply that fraction by 10**55, we can see the value out to 55 decimal digits:

>>> 3602879701896397 * 10 ** 55 // 2 ** 55
1000000000000000055511151231257827021181583404541015625

meaning that the exact number stored in the computer is equal to the decimal value 0.1000000000000000055511151231257827021181583404541015625. Instead of displaying the full decimal value, many languages (including older versions of Python), round the result to 17 significant digits:

>>> format(0.1, '.17f')
'0.10000000000000001'

The fractions and decimal modules make these calculations easy:

>>> from decimal import Decimal
>>> from fractions import Fraction

>>> Fraction.from_float(0.1)
Fraction(3602879701896397, 36028797018963968)

>>> (0.1).as_integer_ratio()
(3602879701896397, 36028797018963968)

>>> Decimal.from_float(0.1)
Decimal('0.1000000000000000055511151231257827021181583404541015625')

>>> format(Decimal.from_float(0.1), '.17')
'0.10000000000000001'