Third in a series — one report per release candidate through GA. Previous editions: preview 7, build 26366.102, preview 7, build 26376.106 (frozen as published).
Author stance: independent, self-appointed decimal128 specialist. Not affiliated with the .NET team. Findings are anchored to IEEE 754-2019 and GDAS (Cowlishaw) and cross-checked with an independent implementation that uses BID implementation (decimal128-csharp-bid) plus Intel libbid.
1. Framing
This edition reviews SDK 11.0.100-rc.1.26401.101 (daily build 26401.101, 2026-08-01). Prior editions are frozen as published; every prior finding is retested here and marked FIXED, IMPROVED, or UNCHANGED.
Given Microsoft’s stance on BID in-memory representation, I have cloned my decimal128-csharp implementation and modified it to use BID. The resulting decimal128-csharp-bid provides a basis for direct functionality and performance comparison of .NET 11 Decimal128, libbid, and decimal128-csharp-bid.
This edition is deliberately shorter on prose and centered on performance (section 7). The conformance ledger (section 5) records verdicts and moves on.
2. Executive Summary
- BID in-memory representation is frozen. Microsoft has definitively stated that they will not consider an alternate in-memory representation, accepting the additional decoding cost that must take place for each operation.
Decimal128will not ship IEEE 754 conformant, and at this point it cannot. Two required clauses are unimplemented: rounding-direction attributes (clause 4.3 — only roundTiesToEven of the five exists) and status flags (clause 7 — no surface at all). Neither is a design preference the standard leaves open, and neither is the kind of change an RC accepts: both are API surface, and GA locks surface in. The type that ships in November will be numerically correct within the subset it implements, and outside conformance with respect to those two clauses (sections 5.1, 5.2).- Clause 5.12 is closed: the quantum-preserving ToString fix has landed.
dotnet/runtime PR #131422 reached the RC 1 line in this build, and the finding is now
FIXED.
Decimal128.MaxValue.ToString()has gone from 6,145 characters to 41;1E+7no longer prints as10000000;0E+2no longer collapses to0; a signed zero survivesF2as-0.00. This closes the standing recommendation carried since the 26376.106 edition, and with it the last gap that could be closed inside this release (section 5.3). - totalOrder stands conformant.
TotalOrderIeee754Comparer<Decimal128>re-verified on this build; the total-order vectors pass with the rest (section 5.5). - The conformance sweep is GREEN and wider: 54,820 vectors run, 0 fail — up 921 on the 26380.103 sweep, entirely from ToString vectors that now execute instead of being skipped. Nothing regressed to make room for them.
- Both machines were re-measured on this build, so the arm64 and x86-64 tables in section 7 are a genuine cross-arch pair rather than two builds photographed apart (section 3).
- Performance has stopped moving. The previous edition’s headline was divide up to 4× faster; this one has no such headline, and that is the finding. Across all 52 comparable cells the Decimal128 build-over-build geomean is 1.009 on arm64 and 1.011 on x86-64 — measured independently on an M3 Pro and an i9-9880H, two microarchitectures agreeing to two decimal places that nothing happened. Every pathological band called out in the previous editions — far-quantum add/sub, wide multiply, exact/power-of-ten divide, FMA — ships here unchanged (section 7).
3. Scope & Methodology
- Version under test: .NET 11 RC 1 line, SDK 11.0.100-rc.1.26401.101.
- Harness — one deliberate change this edition: the standards-anchored rosetta
dispatch is otherwise as before, including the total-order operations driven through
TotalOrderIeee754Comparer(the comparer is .NET’s clause 5.10 vehicle;CompareToremains a value order by design), with every bit-level verdict flowing through Microsoft’s own clause 5.5.2DecodeBinary/EncodeBinary. What changed is that the to-scientific-string family is no longer excluded:toSci/tosci/apply/d128_toStringnow execute againstToString("R")and are checked the way §5.12 actually specifies — exact spelling passes, and anything else is reparsed and compared bit-for-bit, which separates quantum not recovered (a failure) from quantum recovered, spelled differently (permitted). Engineering notation stays excluded on a narrower reason: .NET has no exponent-multiple-of-3 format. This widens the sweep by 921 vectors (section 5.3). - Comparison cohort — changed since the 26376.106 edition:
System.Numerics.Decimal128(system under test), Intel libbid, and decimal128-csharp-bid — my own C# implementation, BID128-native in memory, arithmetic on undecoded operands, non-canonical inputs handled in every operation. All three implementations therefore compute directly on the IEEE 754 BID128 encoding.System.Decimalis retired: it is a 96-bit, 28-digit type with no Inf/NaN/subnormal story, and its presence in prior tables answered a question (“what does the incumbent cost?”) this series has now answered twice. The port column is not comparable to prior editions’decimal128-csharpcolumn (different implementation); the Decimal128 (.NET 11) column is comparable build-over-build (identical corpus and methodology). - Benchmark pairing: the Decimal128 (.NET 11) and decimal128-csharp-bid rows come
from the same emit run under the same rc.1 SDK, one process per cell (tier-1/PGO
code is per-process and training-order-dependent). Both arches were re-measured for
this edition — arm64 run
Rcsbid9, x86-64 runxRcsbid9, both on 26401.101 against an unchanged engine — which also restores a property the series lost when the two machines drifted onto different SDKs: the arm64 and x86-64 tables below are the same build. Corpus and bands are identical to the 26376.106 edition (sign-split add/sub datasets; mul/div/FMA bands unchanged). - One caveat on the libbid column, stated plainly: it was not re-measured for
this edition. libbid is a C library, runtime-independent, and an SDK bump cannot move
it — but the 26380.103 measurements made a point of re-running libbid in the same session
on x86-64, because that machine drifts at day scale. Those x86 libbid rows are now
cross-session against
xRcsbid9. On arm64 the libbid rows come from the standing cross-port C arm, as before, where the M3’s day-scale repeatability makes this sound. Treat x86-64 libbid ratios as good to roughly ±10%, not to the digit. - Reproducibility: decimal128-csharp-bid is the engine under
github.com/migueldecimal128/decimal128-csharp-bid, commit64e891a. The engine is held at that commit across every build compared here, deliberately: holding it fixed is what makes the build-over-build movement attributable to the runtime.
4. What They Got Right
- Everything from the prior editions stands: numerically correct required arithmetic,
canonical BID encoding, honest subnormals, truly-fused
FusedMultiplyAdd(re-verified on this build), correctly-roundedSqrtandRootN. - totalOrder was finished properly, and finished before GA — exactly the window the 26376.106 edition argued for, since after GA the cohort-collapsing behavior would have become observable behavior someone depends on. Runtime PRs #131087/#131205 gave the generic comparer real decimal cohort ordering plus fast paths for the built-in decimal types. It re-verifies clean on this build.
ToStringnow preserves the quantum, and was fixed thoroughly rather than narrowly (PR #131422, section 5.3). The obvious patch would have been to special-case the unrepresentable values; instead theG/Rpath gained a principled rule — emit scientific notation when the quantum exponent requires it or when it is more compact — reusing the same heuristic and the sameFormatExponentthe binary types use, soDecimal128now formats stylistically likedoublerather than like a special case. A newNumberBufferKind.DecimalIeee754separates these types fromSystem.Decimal, which has no negative zero, so signed zero survives every specifier and formatting rounds ties-to-even per §5.12.1. That is more work than the finding demanded, and it is the kind of fix that does not need revisiting.
5. IEEE 754 Compliance — prior findings retested
| # | Finding (clause) | 26366.102 | 26376.106 | 26380.103 | 26401.101 |
|---|---|---|---|---|---|
| 5.1 | Rounding attributes / double-rounding (4.3, 5.1) | ❌ | ❌ | ❌ | UNCHANGED |
| 5.2 | Status flags absent (7) | ❌ | ❌ | ❌ | UNCHANGED |
| 5.3 | Quantum-preserving ToString / string conversion (5.12) |
❌ | ❌ | ❌ | ✅ FIXED |
| 5.4 | fusedMultiplyAdd (5.4.1) | ❌ | ✅ FIXED | ✅ STANDS | STANDS (re-verified fused) |
| 5.5 | totalOrder (5.10) | ❌ | ⚠️ partial | ✅ FIXED | STANDS |
| 5.6 | Clause 9.2 ops faithful, not correctly rounded | — | ⚠️ | ⚠️ | UNCHANGED |
| 5.7 | Log/Log10 catastrophic loss near 1 | — | ❌ | ❌ | UNCHANGED |
Every row above was re-run against this build; none is carried forward on paper. The sweep now runs 54,820 vectors with 0 failures, up 921 on the 26380.103 sweep — every one of those 921 a ToString vector that this series had been skipping since its first edition and now executes (section 5.3). Skips fall to 17,731 accordingly.
5.1 Rounding — UNCHANGED, and this one is not optional
Arithmetic rounds tiesToEven only; there is no rounding-attribute surface; and composing a directed result out of already-rounded arithmetic double-rounds, so the workaround does not recover the missing behavior.
I have described this in earlier editions as a long-horizon item, which understated it. IEEE 754-2019 does not offer rounding-direction attributes as a feature an implementation may decline: clause 4.3 requires them to be provided, and for a decimal format that means roundTiesToEven and roundTiesToAway, roundTowardPositive, roundTowardNegative and roundTowardZero. Supplying one of the five is not a subset of conformance; it is non-conformance with respect to that clause. The type is numerically correct within the one attribute it implements — which is what the 54,820-vector sweep establishes, since the corpus skips 8,360 vectors precisely because they specify a rounding direction this type cannot express.
5.2 Flags — UNCHANGED, and also not optional
No status-flag surface, so the harness still runs value-only. Clause 7 requires an implementation to provide status flags for the five exceptions — invalid, divideByZero, overflow, underflow, inexact — and their absence is the second respect in which this type does not conform. The practical cost is the one the first edition described: with no flag channel, inexactness and underflow are unobservable, so a caller cannot detect that a result was rounded, and my own harness cannot verify a whole class of behavior independently of the returned value.
5.3 String conversion (clause 5.12) — FIXED
dotnet/runtime PR #131422 — “Ensure that Decimal32/64/128 ToString roundtrips and preserves the cohort” — merged 2026-07-30 and has reached this build. The finding that opened this series is closed.
The defect was structural, not cosmetic. The G/R path went through
FormatGeneral(..., suppressScientific: true), which emits fixed-point unconditionally;
a positive quantum exponent has no fixed-point spelling, so the output reparsed as a
different member of the cohort. FormatGeneralAndRoundTripDecimalIeee754 now emits
scientific notation when it is required (quantum exponent > 0) or more compact
(adjusted exponent < −4), the latter being the same G heuristic the binary types use.
Two supporting fixes rode along: DecimalIeee754ToNumber no longer clamps Scale to 0
for a zero coefficient (which had made 0E+2 indistinguishable from 0), and a new
NumberBufferKind.DecimalIeee754 separates these types from System.Decimal, which has
no negative zero — so a signed zero now survives every format specifier, and formatting
rounds ties-to-even as §5.12.1 requires.
Measured on this build, every symptom this series has carried since the first edition is gone:
| 26380.103 | 26401.101 | |
|---|---|---|
MaxValue.ToString() |
6,145 characters | 41 characters |
1E+7 |
10000000 |
1E+07 |
0E+2 |
0 |
0E+02 |
(-0).ToString("F2") |
0.00 |
-0.00 |
| reparse recovers quantum | no | yes |
Verifying this required a harness change, and the reason is worth recording. The corpus carried 2,003 ToString vectors behind a static exclusion — “MS ToString is not GDAS-conformant formatting” — asserted once when the harness was written and re-asserted by every sweep since without being retested. A green run therefore said nothing whatever about ToString, and would have kept reporting this finding as live indefinitely. An exclusion is an untested assertion, and it does not expire when the thing it describes is fixed. Driving the family for real:
- 921 vectors now run and pass with the exact GDAS spelling — vectors this series had been skipping since its first edition.
- 193 recover the quantum but spell the exponent
1E+07where GDAS writes1E+7. .NET routes through the sharedFormatExponentwithminDigits: 2, matchingdouble/float/Half. §5.12.2 constrains the syntax and requires the roundtrip to recover the quantum; it says nothing about padding. Permitted — recorded, not failed. - 446 are Infinity/NaN-payload spellings, a .NET-vocabulary divergence rather than a 5.12 question; 339 are engineering notation, which .NET has no format for; 94 are my own port’s spellings and not GDAS ops at all.
Zero vectors fail to recover the quantum. That is the clause 5.12 requirement, and it is met.
5.4 fusedMultiplyAdd — STANDS FIXED
Re-verified on this build: single rounding across the sweep’s FMA vectors, bit-exact.
MultiplyAddEstimate still exists alongside; its fate remains the API-hygiene question
noted last edition (section 8.1).
5.5 totalOrder (clause 5.10) — STANDS FIXED
The 26376.106 edition found the comparer’s generic path returning 0 for cohort members,
with a TODO in the shipped source; 26380.103 closed it via runtime PRs #131087 and
#131205, which gave TotalOrderIeee754Comparer<T> real decimal semantics. Re-run
here against this build, the full total-order corpus still passes through the comparer —
3-way comparetotal/comparetotmag and the boolean totalOrder predicates — with
cohort members ordered by exponent in both signs, −0 before +0, NaN placement at both
extremes, payload-ordered NaNs. There is still no TotalOrder method on Decimal128
itself, and CompareTo remains (correctly, for .NET) a value order — the comparer is
the standard vehicle, and it remains a conformant one.
5.6 Clause 9.2 recommended functions — UNCHANGED
Same verdict and same counts as last edition: Exp/Pow/RootN faithfully rounded
(sporadic last-place misses, counted in the sweep), exact results returned as the
full-precision cohort member where GDAS prescribes the ideal exponent. Correctly
rounded remains the bar that matters for a 34-digit format; the counts did not move.
5.7 Log/Log2/Log10 near 1 — UNCHANGED
Log(1 + 1e-28) still comes back with about 10 correct digits; LogP1 on the same
engine is still flawless, so the LogP1(x − 1) workaround stands. The binary working
format conversion diagnosed last edition is unchanged. For the log-return arithmetic
this type targets, this remains the accuracy finding to watch.
5.8 RootN (own-bug disclosure) — port fix in progress
Last edition disclosed that Microsoft’s RootN with negative n is correctly rounded
where my implementations are 1–2 ulp off. That bug remains on my punchlist (fix and
oracle regeneration in progress); nothing in this edition’s tables depends on RootN.
6. Permitted & Intentional Divergences — unchanged
Non-canonical propagation through the clause 5.5.2 re-encoding operations (permitted), the min/max cohort-member choice (158 cases, permitted per the clause 9.6 NOTE), and the negative canonical NaN (implementation-defined) all reproduce exactly as published.
7. Performance
The previous edition’s performance story was a 4× divide speedup, attributed and dissected across two sections. This edition has no story, and two builds into a code freeze that is itself the finding — the optimization phase that edition documented appears to be over.
The evidence is also better than a single machine’s A/B. Against 26380.103 — the RC 1
daily two days earlier, measured but not published as an edition — holding the engine
fixed at 64e891a and changing only the runtime, the Decimal128 column over all 52
comparable cells moves by:
| machine | geomean | median | per-cell range |
|---|---|---|---|
M3 Pro (arm64), Rcsbid6 → Rcsbid9 |
1.009 | 1.011 | 0.88–1.07 |
i9-9880H (x86-64), xRcsbid6 → xRcsbid9 |
1.011 | 1.005 | 0.96–1.30 |
Two machines, two microarchitectures, two independent sessions, agreeing to two decimal places that nothing happened. The i9 additionally measured the intervening daily and found the same thing (geomean 1.004), so the interval is covered build by build rather than only at its endpoints.
A caution on how much that buys: 26380.103 and 26401.101 are two calendar days apart. Three dailies inside one code-freeze week showing no movement is consistent with a freeze; it is not, by itself, proof that nothing will move before GA.
The per-cell range is wider than the geomean on both machines and is not evidence of anything: the arm64 outliers are P-gen add/sub sign-split cells, a known bimodality on this M3 that earlier stabilization runs were commissioned to characterise, and the i9’s 1.30 tail is that machine’s documented day-scale drift. The aggregate is the signal; no individual cell in that spread is worth a sentence.
What did not change matters more than what did. The strip-loop signature diagnosed in the previous edition is intact: exact-quotient and power-of-ten divides — the bands with the least mathematical work in them — still cost 90–95 ns on arm64, about 1.7× the width-generic CD/WD bands and ~8× Intel’s libbid, because the trailing-zero strip still removes one digit per iteration. Far-quantum add/sub still costs 710–730 ns against libbid’s 9–10. Extra-wide multiply still fails both of the preview-7 window’s fast-path guards and runs ~720 ns, against 31 ns one band down. FMA still runs the wide machinery end-to-end at 1.95–2.33 µs. The conclusion from last edition stands verbatim: within the one-limb design constraint, the tuning that was available has been taken; the next multiple is an algorithm change, and RC is not where algorithm changes land. Unless something unexpected happens post-GA, this is the performance profile .NET 11 ships.
A reading note for the tables: the port column is decimal128-csharp-bid (section 3) — same-run, same-SDK as the Decimal128 column. Where a number is needed as an external yardstick, use the libbid column; the Cornea et al. paper cited last edition remains the standing description of what a mature software implementation of this format does.
7.1 P-fin (financial profile)
One observation worth a sentence: on this profile’s compact traffic — the add/sub MIX and compact multiply that dominate real financial workloads — Decimal128 now runs ahead of Intel’s libbid on both Apple silicon and Intel i9. The gap to close is concentrated in divide on both machines, and on the i9 in wide multiply as well.
Realistic financial mix (P-fin) — M3 Pro (arm64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| add | MIX | 8.84 | 10.72 | 3.23 |
| sub | MIX | 9.16 | 11.80 | 3.42 |
| mul | CP | 8.79 | 23.57 | 2.01 |
| mul | WP | 28.18 | 34.52 | 19.49 |
| div | CD | 59.81 | 35.07 | 24.49 |
| div | WD | 49.80 | 40.37 | 37.88 |
| div | ET | 99.99 | 6.09 | 6.38 |
| div | PT | 103.30 | 6.09 | 5.28 |
Realistic financial mix (P-fin) — i9-9880H (x86_64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| add | MIX | 24.78 | 30.43 | 11.13 |
| sub | MIX | 28.06 | 31.78 | 14.88 |
| mul | CP | 35.64 | 50.57 | 6.84 |
| mul | WP | 105.33 | 66.15 | 45.50 |
| div | CD | 211.87 | 83.77 | 96.45 |
| div | WD | 169.67 | 87.09 | 117.52 |
| div | ET | 404.16 | 21.12 | 23.59 |
| div | PT | 441.59 | 21.66 | 21.53 |
7.2 Add (P-gen, sign-split)
Add (P-gen, sign-split) — M3 Pro (arm64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| add | SQss | 11.10 | 7.96 | 1.39 |
| add | SQos | 12.79 | 8.69 | 3.25 |
| add | NQss | 13.29 | 9.36 | 6.44 |
| add | NQos | 13.91 | 9.78 | 7.00 |
| add | MQss | 14.95 | 9.75 | 8.46 |
| add | MQos | 14.06 | 9.71 | 12.31 |
| add | OQss | 100.28 | 13.66 | 17.68 |
| add | OQos | 98.84 | 15.31 | 27.62 |
| add | FQss | 709.59 | 9.32 | 9.99 |
| add | FQos | 729.33 | 10.38 | 13.10 |
Add (P-gen, sign-split) — i9-9880H (x86_64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| add | SQss | 48.25 | 29.79 | 5.20 |
| add | SQos | 51.30 | 27.73 | 14.30 |
| add | NQss | 55.33 | 31.59 | 16.11 |
| add | NQos | 59.06 | 30.34 | 23.12 |
| add | MQss | 57.19 | 29.68 | 21.57 |
| add | MQos | 61.89 | 29.14 | 42.76 |
| add | OQss | 293.92 | 46.91 | 55.32 |
| add | OQos | 294.26 | 46.97 | 74.95 |
| add | FQss | 1891.38 | 31.96 | 33.76 |
| add | FQos | 1916.48 | 31.97 | 43.44 |
7.3 Subtract (P-gen, sign-split)
Subtract (P-gen, sign-split) — M3 Pro (arm64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| sub | SQss | 12.65 | 9.68 | 2.15 |
| sub | SQos | 11.24 | 10.04 | 1.63 |
| sub | NQss | 13.82 | 10.87 | 7.65 |
| sub | NQos | 13.69 | 11.42 | 6.24 |
| sub | MQss | 14.31 | 9.94 | 12.95 |
| sub | MQos | 15.01 | 9.78 | 8.39 |
| sub | OQss | 98.99 | 15.79 | 27.04 |
| sub | OQos | 97.19 | 14.84 | 17.82 |
| sub | FQss | 724.72 | 9.39 | 12.88 |
| sub | FQos | 712.80 | 9.48 | 9.50 |
Subtract (P-gen, sign-split) — i9-9880H (x86_64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| sub | SQss | 69.24 | 32.57 | 11.39 |
| sub | SQos | 50.70 | 32.72 | 5.16 |
| sub | NQss | 60.44 | 34.44 | 20.99 |
| sub | NQos | 57.24 | 35.18 | 15.98 |
| sub | MQss | 59.79 | 34.00 | 42.60 |
| sub | MQos | 57.14 | 34.35 | 23.88 |
| sub | OQss | 291.93 | 51.38 | 75.17 |
| sub | OQos | 292.36 | 52.36 | 54.68 |
| sub | FQss | 1888.12 | 37.10 | 42.85 |
| sub | FQos | 1878.31 | 36.63 | 30.10 |
7.4 Multiply (P-gen)
Multiply (P-gen) — M3 Pro (arm64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| mul | CP | 8.93 | 23.13 | 2.77 |
| mul | WP | 30.80 | 35.06 | 19.60 |
| mul | XP | 720.83 | 45.24 | 49.79 |
Multiply (P-gen) — i9-9880H (x86_64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| mul | CP | 35.90 | 51.55 | 9.97 |
| mul | WP | 110.49 | 73.72 | 49.40 |
| mul | XP | 1814.38 | 105.10 | 96.34 |
7.5 Divide (P-gen)
Divide (P-gen) — M3 Pro (arm64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| div | CD | 55.84 | 36.77 | 25.83 |
| div | WD | 51.69 | 37.54 | 39.96 |
| div | XD | 107.36 | 38.97 | 38.60 |
| div | ET | 95.38 | 11.68 | 7.21 |
| div | PT | 89.98 | 11.45 | 5.71 |
Divide (P-gen) — i9-9880H (x86_64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| div | CD | 220.19 | 87.11 | 97.94 |
| div | WD | 201.85 | 88.01 | 122.10 |
| div | XD | 334.09 | 90.00 | 122.05 |
| div | ET | 385.10 | 32.13 | 32.62 |
| div | PT | 382.08 | 32.81 | 15.88 |
7.6 FMA
Correct, still the most expensive operation in the suite, and still the widest gap to the cohort — roughly 24× libbid on arm64 and 29× on x86-64. The FF-costs-more-than-FN inversion diagnosed in the previous edition (the per-digit drop loop, at full amplitude) reproduces exactly. FN sits at 1.95 µs against 26380.103’s 1.89 — a drift of the size these cells routinely show, and in the opposite direction from the one measured across the preview-7 window, so there is nothing here to attribute.
FMA — M3 Pro (arm64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| fma | FN | 1952.18 | 82.83 | 103.25 |
| fma | FF | 2331.39 | 58.46 | 74.68 |
FMA — i9-9880H (x86_64):
| op | cat | Decimal128 (.NET 11) | libbid C | decimal128-csharp-bid |
|---|---|---|---|---|
| fma | FN | 5072.81 | 176.07 | 210.49 |
| fma | FF | 6288.99 | 135.29 | 167.72 |
8. Recommendations
8.1 Now — before GA
- API hygiene: decide
MultiplyAddEstimate’s fate now thatFusedMultiplyAddexists (unchanged from last edition; GA locks surfaces in). It is a judgement call about surface area rather than a conformance defect, and it is the only thing left that an RC can still absorb.
Everything the series raised that could be fixed inside this release has been:
fusedMultiplyAdd (26376.106), totalOrder (26380.103), quantum-preserving ToString
(here). That is a good record, and it should not be read as the ledger being clear —
the two items below were never candidates for this window.
8.2 After GA — required, and now unavoidably deferred
These are not enhancements. Rounding attributes and status flags are required by IEEE 754-2019, and shipping without them is why section 9 does not call this type conformant. They appear here rather than in 8.1 for a practical reason, not a principled one: both are API surface, both would need design review, and neither survives a code freeze. GA will lock the surface they have to extend.
- Rounding-direction attributes (clause 4.3, section 5.1): four of the five are missing. Retrofitting them after GA is additive but harder than doing it now, because every arithmetic entry point already has a signature that assumes a single attribute.
- Status flags (clause 7, section 5.2): the whole surface is absent. Still the cheapest thing on this list to add compatibly, and still the one that would let callers — and reviewers — verify inexactness and underflow independently of the returned value.
- Near-1 logarithm accuracy (section 5.7): unlike the two above this one is
optional — §9.2 recommends correct rounding without requiring it — and the
LogP1workaround carries users meanwhile.
9. Conclusion
This series opened by arguing that Decimal128 was numerically sound and specificationally
incomplete. Three editions later, both halves of that sentence still hold — but the second
half is now much narrower, and what is left of it is not going to move before November.
The narrowing is real and worth saying plainly. fusedMultiplyAdd is genuinely fused,
totalOrder has real cohort ordering, and string conversion preserves the quantum. Every
one of those was raised here, argued for on a pre-GA timetable, and fixed on that
timetable. That is a responsive team, and three editions ago I would not have predicted
the clean sweep.
What remains is not a rounding error in that record. Rounding-direction attributes (clause 4.3) and status flags (clause 7) are both required by IEEE 754-2019, and both are entirely absent — four of five rounding attributes missing, no flag surface at all. So the honest verdict on the type that ships in November is not “conformant with minor gaps”; it is numerically correct within the subset it implements, and not a conforming IEEE 754 implementation. The distinction matters for anyone choosing this type on the strength of the standard’s name: what you get is correctly-rounded tiesToEven decimal arithmetic with an IEEE 754 encoding and an IEEE 754 comparison model, not IEEE 754.
Nor is that a criticism of what the team did with the time. Both gaps are API surface; API surface is what an RC exists to stop changing; and the window in which they could have landed closed before this series began. Deferring them was the only available choice by the time anyone was looking. It does mean the deferral is now permanent for .NET 11: they become additive work against a surface GA will have locked, which is strictly harder than it would have been in preview.
Performance is stationary — 1.009 and 1.011 build-over-build on two machines — and the wide-path and divide economics in section 7 are, I expect, what ships. Those follow from the one-limb representation and the frozen BID in-memory choice, and were never something an RC was going to change either.
So: a good release, a genuinely improved type, and a conformance claim that should be made carefully. I wish the team a quiet freeze — and I would like to spend .NET 12 arguing about rounding attributes.