System.Numerics.Decimal128 — .NET 11 RC 1, build 26411.119

IEEE 754-2019 compliant decimal128 high-performance software solution created by Miguel.

Fourth in a series — one report per release candidate through GA. Previous editions: preview 7, build 26366.102, preview 7, build 26376.106, RC 1, build 26401.101 (all frozen as published).

Author stance: independent, self-appointed decimal128 specialist. Not affiliated with the .NET team. Findings are anchored to IEEE 754-2019 and GDAS (Cowlishaw) and cross-checked with an independent implementation that uses BID implementation (decimal128-csharp-bid) plus Intel libbid.


1. Framing

This edition reviews SDK 11.0.100-rc.1.26411.119 (daily build 26411.119, 2026-08-11), ten dailies on from the build the previous edition covered.

It is a short edition, because nothing changed. Every finding was retested and every one holds its prior verdict. The conformance sweep reproduces vector for vector. Both machines were re-measured and neither found movement. If you have read the 26401.101 edition, you already know everything this type is going to do in November; this edition exists to say so with evidence rather than by assumption, and to put a fourth data point under the claim that the optimization phase is over.

A stationary reading is not a non-result at this stage of a release. The interesting question two months from GA is no longer what improved but whether anything still moves, and that question can only be answered by measuring builds that turn out to be identical and reporting them as such.

2. Executive Summary

  • No finding changed status. The ledger in section 5 carries every verdict forward unchanged: rounding attributes and status flags still absent, ToString still fixed, fusedMultiplyAdd and totalOrder still standing.
  • The conformance sweep reproduces exactly: 54,820 vectors run, 54,820 pass, 0 fail, 0 undecided. Not merely the same headline — the same counts, band by band and divergence by divergence, as the previous build (section 5).
  • Performance is stationary for a fourth consecutive build. On arm64 the Decimal128 column moves by a geomean of 0.988 across all 52 comparable cells, with not one cell moving as much as 10%. On x86-64 both columns fell together by ~8%, which is that machine’s documented day-scale drift rather than the runtime; the drift-immune view puts the change at 0.995 (section 7).
  • Every pathology ships as diagnosed. Far-quantum add/sub ~700 ns, extra-wide multiply ~710 ns, exact and power-of-ten divide ~89 ns, FMA 1.9–2.3 µs — all on arm64, all reproducing the previous edition’s numbers within noise (section 7).
  • The conclusion is unchanged and now effectively fixed for .NET 11. Two required clauses remain unimplemented — rounding-direction attributes (clause 4.3) and status flags (clause 7) — and both are API surface that a release candidate will not take. What ships in November is numerically correct within the subset it implements, and not a conforming IEEE 754 implementation (sections 5.1, 5.2, 9).

3. Scope & Methodology

  • Version under test: .NET 11 RC 1 line, SDK 11.0.100-rc.1.26411.119. For anyone reproducing this: the commit stamped in that SDK is a dotnet/dotnet VMR commit (7cdb2174, 2026-08-11), not a dotnet/runtime one — the corresponding runtime source is 067d74c2, resolved through the VMR’s source manifest. The two are easy to confuse and only one of them will check out.
  • Harness — unchanged this edition. The previous edition made one deliberate change (driving the to-scientific-string family for real instead of excluding it); that change is carried forward as-is and nothing else was touched. Total-order operations still run through TotalOrderIeee754Comparer, and every bit-level verdict still flows through Microsoft’s own clause 5.5.2 DecodeBinary/EncodeBinary.
  • Comparison cohort — unchanged: System.Numerics.Decimal128 (system under test), Intel libbid, and decimal128-csharp-bid — my own C# implementation, BID128-native in memory. All three compute directly on the IEEE 754 BID128 encoding. The Decimal128 (.NET 11) column is comparable build-over-build; the port column is a methodology-disclosed comparison, not the subject of this review.
  • Benchmark pairing: arm64 run Rcsbid10 and x86-64 run xRcsbid10, both on this SDK against an unchanged engine, each measuring the Decimal128 and port columns in the same session with one process per cell. This is the third consecutive edition in which both machines sit on the same build.
  • The libbid caveat is narrower this edition, and still present. Last edition had to warn that x86-64 libbid was measured in a different session from the system under test, good to roughly ±10% on a machine that drifts at day scale. The i9 has since re-measured all nine ports in a single session, and libbid from that sweep agrees with the older same-session arm to geomean 1.003 — so the tables below now take libbid from it on both arches. Those rows are still a week apart from xRcsbid10, so treat x86-64 libbid ratios as sound to a few percent rather than to the digit. On arm64 the M3’s day-scale repeatability makes this a non-issue.
  • Reproducibility: decimal128-csharp-bid is the engine under github.com/migueldecimal128/decimal128-csharp-bid, commit 64e891a — held fixed across every build compared in this series, which is what makes build-over-build movement attributable to the runtime.

4. What They Got Right

Unchanged from the previous edition, and re-verified rather than carried on paper: numerically correct required arithmetic, canonical BID encoding, honest subnormals, truly fused FusedMultiplyAdd, correctly-rounded Sqrt and RootN, real cohort ordering in TotalOrderIeee754Comparer, and a ToString that preserves the quantum. Three findings this series raised were fixed inside the release window; all three still hold on this build.

5. IEEE 754 Compliance — prior findings retested

# Finding (clause) 26366.102 26376.106 26380.103 26401.101 26411.119
5.1 Rounding attributes / double-rounding (4.3, 5.1) ❌ ❌ ❌ ❌ UNCHANGED
5.2 Status flags absent (7) ❌ ❌ ❌ ❌ UNCHANGED
5.3 Quantum-preserving ToString / string conversion (5.12) ❌ ❌ ❌ ✅ FIXED STANDS
5.4 fusedMultiplyAdd (5.4.1) ❌ ✅ FIXED ✅ STANDS STANDS STANDS
5.5 totalOrder (5.10) ❌ ⚠️ partial ✅ FIXED STANDS STANDS
5.6 Clause 9.2 ops faithful, not correctly rounded — ⚠️ ⚠️ ⚠️ UNCHANGED
5.7 Log/Log10 catastrophic loss near 1 — ❌ ❌ ❌ UNCHANGED

Every row was re-run against this build. The sweep runs 54,820 vectors with 0 failures and 0 undecided, with 17,731 skips — the same totals as the previous build, and the same distribution behind them: 8,360 vectors skipped because they specify a rounding direction this type cannot express, 4,342 because an invalid integer conversion has no observable flag channel to check.

That exactness is the point of this section. Two builds producing the same pass count could still differ underneath; these do not. Every permitted-divergence bucket reproduces to the vector — 158 min/max cohort choices, 193 exponent-padding spellings, 95 clause-9.2 quantum divergences, 14 clause-9.2 value divergences, 4 negative canonical NaNs. Nothing moved, and nothing traded places.

5.1 Rounding — UNCHANGED

Arithmetic rounds tiesToEven only; there is no rounding-attribute surface; and composing a directed result out of already-rounded arithmetic double-rounds, so the workaround does not recover the missing behavior. Clause 4.3 requires all five attributes for a decimal format; supplying one is not a subset of conformance but non-conformance with respect to that clause. This is a long-horizon engagement item and I do not expect it to move in .NET 11 — but it is required, and it does not stop being required because the window closed.

5.2 Flags — UNCHANGED

No status-flag surface, so the harness still runs value-only. Clause 7 requires status flags for the five exceptions — invalid, divideByZero, overflow, underflow, inexact — and their absence is the second respect in which this type does not conform. A caller cannot detect that a result was rounded; a reviewer cannot verify inexactness independently of the returned value.

5.3 String conversion (clause 5.12) — STANDS FIXED

PR #131422 re-verified on this build. MaxValue.ToString() is 41 characters, 1E+7 round-trips, 0E+2 survives, (-0).ToString("F2") is -0.00, and zero vectors fail to recover the quantum. The 193 vectors that recover the quantum but pad the exponent to 1E+07 are permitted and remain recorded rather than failed.

5.4 fusedMultiplyAdd (5.4.1), 5.5 totalOrder (5.10) — STAND FIXED

Both re-verified: single rounding across the FMA vectors, bit-exact; and the full total-order corpus still passes through TotalOrderIeee754Comparer with cohort members ordered by exponent in both signs, −0 before +0, payload-ordered NaNs. MultiplyAddEstimate still exists alongside FusedMultiplyAdd, still an API-hygiene question rather than a defect (section 8.1).

Same verdicts, same counts, same worked examples as the previous edition. Exp/Pow/ RootN are faithfully rounded with sporadic last-place misses; exact results come back as the full-precision cohort member where GDAS prescribes the ideal exponent; Log(1 + 1e-28) still returns about 10 correct digits while LogP1 on the same engine is flawless, so the LogP1(x − 1) workaround still stands.

5.8 RootN (own-bug disclosure) — unchanged

Microsoft’s RootN with negative n is correctly rounded where my own implementations are 1–2 ulp off. That bug is still on my punchlist and still not fixed. Nothing in this edition’s tables depends on RootN.

6. Permitted & Intentional Divergences — unchanged

Non-canonical propagation through the clause 5.5.2 re-encoding operations (permitted), the min/max cohort-member choice (158 cases, permitted per the clause 9.6 NOTE), and the negative canonical NaN (implementation-defined) all reproduce exactly, at identical counts.

7. Performance

The previous edition reported that performance had stopped moving and cautioned that three dailies inside one freeze week was consistent with a freeze without proving one. This edition extends the interval to ten dailies and finds the same thing.

On arm64, holding the engine at 64e891a and changing only the runtime:

machine geomean median per-cell range cells beyond ±10%
M3 Pro (arm64), Rcsbid9 → Rcsbid10 0.988 0.993 0.91–1.09 0 of 52

Not one cell of fifty-two moved by a tenth. That is a stronger statement than the aggregate: a geomean near 1 can hide offsetting movement, and here there is none to hide.

The x86-64 machine needs its own paragraph, because its raw numbers look like an 8% across-the-board improvement and are not one. Both columns fell together — Decimal128 by a geomean of 0.916, my port by 0.912, near-identical distributions — and software does not speed up two unrelated implementations by the same amount on the same day. That box is documented as sitting in a slower state at day scale, and the eleven days between these two runs are enough for it to have moved. The drift-immune view, comparing each build’s Decimal128-to-port ratio against the previous build’s, removes the common factor:

machine ratio-of-ratios geomean median
i9-9880H (x86-64), xRcsbid9 → xRcsbid10 0.995 0.992

Which agrees with arm64: nothing happened. The arm64 twin is what makes this readable at all — the M3 does not drift at day scale, so when the two machines disagree about absolute movement and agree about the ratio, the disagreement is the machine.

One candidate finding was raised by that x86 data and then killed, which is worth recording because it is exactly the shape of thing a single-machine review would have published. The drift-immune view threw one standout: sub SQss on the general profile at 1.45, driven by the Decimal128 column falling from 69.24 ns to 48.33 ns — its largest single move, several standard deviations out, and superficially a real fast-path win. The arm64 twin reads 1.002 on that same cell (12.65 ns → 12.62 ns). It is an x86-local measurement artifact, and it is not promoted.

What did not change remains the substance. The strip-loop signature is intact: exact-quotient and power-of-ten divides — the bands with the least mathematical work in them — still cost about 89 ns on arm64, roughly 1.7× the width-generic CD/WD bands and ~7.6× Intel’s libbid, because the trailing-zero strip still removes one digit per iteration. Far-quantum add/sub still costs ~700–710 ns against libbid’s 9–10. Extra-wide multiply still fails both fast-path guards and runs ~710 ns, against 31 ns one band down. FMA still runs the wide machinery end-to-end at 1.90–2.27 µs, some 23× libbid on the near band and 39× on the far one. The conclusion from the last two editions stands verbatim: within the one-limb design constraint, the tuning that was available has been taken; the next multiple is an algorithm change, and RC is not where algorithm changes land.

A reading note for the tables: the port column is decimal128-csharp-bid (section 3) — same-run, same-SDK as the Decimal128 column. Where a number is needed as an external yardstick, use the libbid column.

7.1 P-fin (financial profile)

The observation from the previous edition reproduces unchanged: on this profile’s compact traffic — the add/sub MIX and compact multiply that dominate real financial workloads — Decimal128 runs ahead of Intel’s libbid on both machines, and on arm64 that extends to wide multiply. The gap to close is concentrated in divide on both machines, and on the i9 in wide multiply.

Realistic financial mix (P-fin) — M3 Pro (arm64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
add MIX 8.92 10.72 3.20
sub MIX 9.22 11.80 3.57
mul CP 8.80 23.57 2.09
mul WP 30.68 34.52 19.79
div CD 55.28 35.07 24.48
div WD 46.80 40.37 37.90
div ET 102.82 6.09 6.48
div PT 100.87 6.09 5.21

Realistic financial mix (P-fin) — i9-9880H (x86_64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
add MIX 22.70 30.56 11.15
sub MIX 25.21 31.87 12.99
mul CP 42.83 50.91 6.66
mul WP 98.05 66.63 41.78
div CD 197.07 82.11 85.37
div WD 157.40 86.90 104.34
div ET 376.89 21.25 21.81
div PT 399.65 21.42 13.88

7.2 Add (P-gen, sign-split)

Add (P-gen, sign-split) — M3 Pro (arm64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
add SQss 11.24 7.96 1.38
add SQos 12.38 8.69 3.20
add NQss 13.57 9.36 6.45
add NQos 14.07 9.78 7.15
add MQss 15.35 9.75 8.45
add MQos 14.59 9.71 12.34
add OQss 94.93 13.66 17.94
add OQos 95.81 15.31 28.15
add FQss 701.57 9.32 10.50
add FQos 710.21 10.38 13.04

Add (P-gen, sign-split) — i9-9880H (x86_64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
add SQss 44.10 30.12 4.29
add SQos 48.82 27.81 13.80
add NQss 62.21 31.74 14.63
add NQos 54.84 30.17 20.18
add MQss 51.91 29.55 19.81
add MQos 55.38 28.78 38.13
add OQss 266.41 46.76 55.41
add OQos 263.65 47.44 67.81
add FQss 1719.40 33.02 29.58
add FQos 1717.41 32.54 41.59

7.3 Subtract (P-gen, sign-split)

Subtract (P-gen, sign-split) — M3 Pro (arm64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
sub SQss 12.62 9.68 2.15
sub SQos 11.20 10.04 1.40
sub NQss 14.11 10.87 7.53
sub NQos 13.53 11.42 6.02
sub MQss 14.30 9.94 13.02
sub MQos 15.24 9.78 8.12
sub OQss 98.45 15.79 27.84
sub OQos 93.07 14.84 17.66
sub FQss 743.83 9.39 13.02
sub FQos 728.07 9.48 9.74

Subtract (P-gen, sign-split) — i9-9880H (x86_64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
sub SQss 48.33 33.73 11.56
sub SQos 43.04 32.99 4.89
sub NQss 53.70 34.85 20.51
sub NQos 50.66 35.61 13.61
sub MQss 53.83 34.98 37.81
sub MQos 52.44 34.66 20.72
sub OQss 265.48 51.72 65.73
sub OQos 263.25 50.81 47.67
sub FQss 1743.93 37.25 45.62
sub FQos 1717.77 36.53 28.95

7.4 Multiply (P-gen)

Multiply (P-gen) — M3 Pro (arm64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
mul CP 8.91 23.13 2.73
mul WP 30.89 35.06 19.51
mul XP 710.01 45.24 48.17

Multiply (P-gen) — i9-9880H (x86_64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
mul CP 34.06 50.55 9.36
mul WP 100.71 73.62 44.03
mul XP 1664.78 104.96 86.52

7.5 Divide (P-gen)

Divide (P-gen) — M3 Pro (arm64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
div CD 52.41 36.77 25.91
div WD 51.91 37.54 40.04
div XD 108.65 38.97 38.68
div ET 89.06 11.68 7.31
div PT 89.23 11.45 5.80

Divide (P-gen) — i9-9880H (x86_64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
div CD 198.44 86.65 88.32
div WD 170.00 88.12 119.63
div XD 293.74 88.89 109.77
div ET 348.53 31.98 32.97
div PT 342.14 32.28 14.45

7.6 FMA

Correct, still the most expensive operation in the suite, and still the widest gap to the cohort. The FF-costs-more-than-FN inversion diagnosed two editions ago — the per-digit drop loop at full amplitude — reproduces exactly.

FMA — M3 Pro (arm64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
fma FN 1896.17 82.83 105.53
fma FF 2266.88 58.46 84.02

FMA — i9-9880H (x86_64):

op cat Decimal128 (.NET 11) libbid C decimal128-csharp-bid
fma FN 4634.12 176.51 190.68
fma FF 5763.45 135.93 165.51

8. Recommendations

8.1 Now — before GA

  1. API hygiene: decide MultiplyAddEstimate’s fate now that FusedMultiplyAdd exists. Unchanged for three editions, and still the only item on this list — a judgement call about surface area rather than a conformance defect, and the only thing left that a release candidate can still absorb.

8.2 After GA — required, and now unavoidably deferred

Unchanged. Rounding-direction attributes (clause 4.3) and status flags (clause 7) are required by IEEE 754-2019 and entirely absent; both are API surface, both would need design review, and neither survives a code freeze. Near-1 logarithm accuracy (section 5.7) is the third item and is genuinely optional — clause 9.2 recommends correct rounding without requiring it — with the LogP1 workaround carrying users meanwhile.

9. Conclusion

There is little to add to the previous edition, and that is this edition’s finding.

Ten dailies produced no change in any finding, no change in the sweep, and no change in performance that survives contact with a second machine. The type is doing what a type does inside a code freeze, which is nothing, and it is doing it consistently enough that I would now be surprised by movement before GA.

So the verdict carries over intact, and I think it can be stated as settled rather than provisional: what ships in November is numerically correct within the subset it implements, and not a conforming IEEE 754 implementation — four of five rounding attributes missing (clause 4.3), no status-flag surface at all (clause 7). Everything this series raised that could be fixed inside the release window was fixed, and fixed well; what remains was out of reach before the series started.

The next edition worth writing is the one that covers a build where something moves. On this evidence that is unlikely to be an RC 1 build, and the interesting conversation — rounding attributes, flags, and what a conformance claim should say — belongs to .NET 12.