diff --git a/docs/BENCHMARKING.md b/docs/BENCHMARKING.md index 9bb5605..d5e27aa 100644 --- a/docs/BENCHMARKING.md +++ b/docs/BENCHMARKING.md @@ -57,7 +57,7 @@ Rule of thumb: The three canonical workflows compose around one artifact schema, metric set, and renderer: -| Recipe | Measure | Retain CSV/JSON | Promote release docs | +| Recipe | Measure | Retain complete-run JSON | Promote release docs | |--------|---------|-----------------|----------------------| | `performance-local` | Yes | Yes | No | | `performance-doc` | No | Consumes retained inputs | Yes | @@ -87,10 +87,10 @@ rows so validation and adapter costs remain visible. Every row uses `iter` with borrowed, preconstructed inputs; only `construct_then_angle` includes construction. This focused signal is outside the release-report schema. -Newly rendered reports use one table per selected suite. Dimension and -adversarial-input group appear in a `Case` column instead of creating a separate -table for every group. The `vs_linalg` table is the one wider variant because it -adds nalgebra and faer context columns where matching peer measurements exist. +Complete-run reports use a table for each statistic, with columns for the +selected series. Benchmark identifiers retain the suite, dimension, and input +case. Saved-baseline comparisons and historical reports retain their separate +layouts. **`vs_linalg`** (`benches/vs_linalg.rs`) compares `la-stack` against `nalgebra` and `faer` across D=2-64 for LU, solve, determinant, dot, norm, and @@ -166,8 +166,9 @@ tag. Both revisions must pass the same benchmark-input tests used by correctness gate over the deterministic fixtures and operations, not validation of each timed Criterion sample. It writes `target/bench-reports/performance.md` plus retained `performance.run.json` and -`performance.evidence.json` complete-run inputs. The report and sidecar embed -both commits, CPU, operating system, Rust toolchain, lockfile and harness +`performance.evidence.json` complete-run inputs. The Markdown report includes +the run identity and each series' phase, release, and revision. The evidence +JSON retains the CPU, operating system, Rust toolchain, lockfile and harness digests, Criterion selection/commands, and both correctness-gate results. The report reader rejects malformed or mismatched provenance and incomplete selected-suite coverage. @@ -315,7 +316,7 @@ See `uv run --locked criterion-dim-plot --help` for plotting options. ### Create The Release Performance Report -Release PRs promote one curated release-to-release comparison into committed +Release PRs promote one complete release-to-release comparison into committed docs: ```bash @@ -342,18 +343,31 @@ dependency versions, host identity, and independent gate results. Promotion publishes immutable `run.json`, `evidence.json`, and `report.md` under `docs/performance-runs/runs//`, plus the shared index, latest pointer, history index, and `docs/performance.md` in one transaction. -Repeated runs for one release pair coexist. The previous curated report is -retained under `docs/archive/performance/` when its archive path is new. +Repeated runs for one release pair coexist. `tooling/performance-report.toml` +selects the current report path and title. The shared renderer produces mean +and median tables with named series and recorded marginal intervals; unavailable +cases appear as an em dash. Each series identifies its measurement phase, +release, and revision. Full environment, harness, command, and gate provenance +remains in that run's `evidence.json`, beside `run.json` and `report.md`. +New reports no longer use the custom percentage-change, ratio, or interval-overlap +columns. The README dimension plot and its table retain their existing format. Serialization, validation, rendering, and promotion failures preserve existing published reports and selection. Review and commit the whole new run and index. Legacy snapshots under `docs/performance/` retain their original bytes, hashes, framing, and reader; see the [legacy index](performance/README.md). +The last published legacy report is preserved in the +[Markdown archive](archive/performance/README.md). The checked-in current report +keeps its historical format until a complete run is promoted; old summaries do +not contain the raw samples required to reconstruct complete-run evidence. +Existing history and optimization studies are retained. No automatic pruning is +performed; removing indexed runs requires a coordinated change to the shared +index and latest pointer, while preserving evidence referenced by active reports. The Markdown check and format recipes exclude the generated -`docs/performance-runs/` archive because shared validation requires exact report -bytes. Keep both recipe exclusions aligned with `archive` in -`tooling/performance-report.toml` when relocating it. The curated -`docs/performance.md` remains subject to active Markdown checks. +`docs/performance-runs/` archive and `docs/performance.md` because these outputs +belong to the shared renderer. Shared validation requires exact retained report +bytes. Keep both recipe exclusions aligned with `archive` and `current` in +`tooling/performance-report.toml` when relocating them. To reproduce and promote the report without running Cargo or creating Git worktrees, use: @@ -366,7 +380,7 @@ This command fails closed on a missing, partial, malformed, mismatched, or unsupported artifact pair. It consumes the default complete-run payload/envelope or follows the validated `docs/performance-runs/latest.json` when all scratch inputs are absent. A repository with only legacy evidence continues through -the historical reader. Report replay runs offline and updates the curated +the historical reader. Report replay runs offline and updates the current document and shared history transactionally. Promotion requires distinct current and baseline package versions, so a same-version local comparison is retained and reproducible but cannot become release documentation. Use promotion @@ -406,10 +420,10 @@ shared-harness workflow before attributing a difference solely to library code. | `target/bench-reports/performance.evidence.json` | No | `performance-local`, `performance-release` | Shared evidence envelope with source/harness and phase provenance. | | `target/bench-reports/performance-non-exact.*` | No | `performance-local-non-exact` | Narrowed non-exact report and retained peer-context comparison inputs. | | `target/bench-reports/github-assets-performance.md` | No | `performance-github-assets` | Local report from published release artifacts. | -| `docs/performance.md` | Yes | `performance-release`, `performance-doc` | Latest curated release-to-release comparison. | +| `docs/performance.md` | Yes | `performance-release`, `performance-doc` | Current shared full report; last legacy report until complete-run promotion. | | `docs/performance-runs/` | Yes | `performance-release`, `performance-doc` | Immutable complete runs, reports, index, and validated latest pointer. | | `docs/performance/` | Yes | Legacy reader | Unchanged historical CSV/JSON evidence and hashes. | -| `docs/archive/performance/` | Yes | `performance-release`, `performance-doc` | Older curated release-to-release comparisons. | +| `docs/archive/performance/` | Yes | Legacy reader | Preserved reports in the historical format. | | `docs/archive/performance/studies/` | Yes | Maintainer investigations | Completed optimization studies and decisions. | | `docs/assets/bench/` | Yes | `performance-readme` | README benchmark CSV/SVG assets and JSON provenance. | | GitHub Release | Remote | `.github/workflows/release-benchmarks.yml` | Criterion baseline archive. | diff --git a/docs/RELEASING.md b/docs/RELEASING.md index b459226..eab0c56 100644 --- a/docs/RELEASING.md +++ b/docs/RELEASING.md @@ -119,8 +119,8 @@ just performance-release ``` The no-argument form compares the current package version with the previous -stable published release. Review `docs/performance.md`, any archived comparison -under `docs/archive/performance/`, and the immutable complete-run evidence, +stable published release. Review the shared report at `docs/performance.md` +and the immutable complete-run evidence, reports, index, and latest pointer under `docs/performance-runs/`. Include the entire new run and index changes in the release PR. Shared complete evidence retains both phases, mean and median estimates, 100 raw samples per case, and diff --git a/docs/archive/performance/README.md b/docs/archive/performance/README.md index 6456dbc..9d0bf4d 100644 --- a/docs/archive/performance/README.md +++ b/docs/archive/performance/README.md @@ -1,12 +1,15 @@ # Archived Performance Reports -Older release-to-release benchmark comparisons are archived here. -`docs/performance.md` contains the latest curated comparison. +Legacy release-to-release benchmark comparisons are archived here, including +the last published report preserved before adopting the shared report format. +`docs/performance.md` remains the current report. New complete runs use the +shared history described in [Benchmarking](../../BENCHMARKING.md). - [v0.4.1-vs-v0.4.0](v0.4.1-vs-v0.4.0.md) - [v0.4.2-vs-v0.4.1](v0.4.2-vs-v0.4.1.md) - [v0.4.3-vs-v0.4.2](v0.4.3-vs-v0.4.2.md) - [v0.4.4-vs-v0.4.3](v0.4.4-vs-v0.4.3.md) - [v0.4.5-vs-v0.4.4](v0.4.5-vs-v0.4.4.md) +- [v0.4.6-vs-v0.4.5](v0.4.6-vs-v0.4.5.md) Completed optimization investigations are in [archived studies](studies/README.md). diff --git a/docs/archive/performance/v0.4.6-vs-v0.4.5.md b/docs/archive/performance/v0.4.6-vs-v0.4.5.md new file mode 100644 index 0000000..e49e34f --- /dev/null +++ b/docs/archive/performance/v0.4.6-vs-v0.4.5.md @@ -0,0 +1,334 @@ +# Benchmark Performance + +**la-stack** v0.4.6 · `d36a9e9` (HEAD) +**Source revision timestamp**: 2026-09-08 01:59:08 UTC (deterministic report metadata; not the benchmark measurement time) +**Benchmark measurement timestamp**: not recorded by Criterion; use the provenance below to identify the measured revisions and environment. +**Statistic**: median +**Suite**: all +**Scope**: release-signal + +## Benchmark Results + +Comparison against baseline **v0.4.5**: + +Negative point-estimate change means the current point estimate is smaller; a baseline/current point-estimate ratio above 1.00 has the same meaning. +The CI-relation column reports only whether the two marginal Criterion intervals overlap. These are not paired confidence intervals +for the change, so the report makes no statistical-significance or performance-improvement claim from interval separation. + +### Reproducibility Provenance + +**Measurement environment**: recorded for both samples under one shared current harness. + +- CPU: `Apple M4 Max (arm64)` +- OS: `Darwin 25.6.0 arm64` +- rustc: `rustc 1.98.1 (48a229cea 2026-09-01)` +- Current commit: `d36a9e99ec223c4cbbf8a9191818f186159c27f7` +- Current Git clean: `false` +- Current source-state SHA-256: `b8fcf706fa6e287ca6682a1ecef5a69576413b9c88040184a9b1ecb8f6e92336` +- Baseline commit: `78ba1c25beaed7321e175113e0876aac6d763a59` +- Baseline Git clean: `false` +- Baseline source-state SHA-256: `7d72086128b12e2a89d5ca78f0e8cc0cb393a44f1062d027a069722ff9dae007` +- Cargo.lock SHA-256: `37c6f3339dc4eedcc5587e44133a7d0299fef2bd9da52fc0464e2ceeb0138edb` +- Benchmark harness SHA-256: `27acd15e854e46899ff1c1f8e1b399ff240b0b9732e56a36cf8edb39c4a58db3` + +**Publication and validation environment**: + +- Publication CPU: `Apple M4 Max (arm64)` +- Publication OS: `Darwin 25.6.0 arm64` +- Publication rustc: `rustc 1.98.1 (48a229cea 2026-09-01)` +- Publication commit: `d36a9e99ec223c4cbbf8a9191818f186159c27f7` +- Publication Git clean: `false` +- Publication source-state SHA-256: `b8fcf706fa6e287ca6682a1ecef5a69576413b9c88040184a9b1ecb8f6e92336` +- Publication Cargo.lock SHA-256: `37c6f3339dc4eedcc5587e44133a7d0299fef2bd9da52fc0464e2ceeb0138edb` +- Publication harness SHA-256: `27acd15e854e46899ff1c1f8e1b399ff240b0b9732e56a36cf8edb39c4a58db3` +- Criterion suite/scope: `all` / `release-signal` +- Criterion statistic/sample: `median` / `new` +- Criterion dependency version: `0.8.2` +- Baseline command: `just bench-save-baseline v0.4.5` +- Current command: `just bench-latest` +- Correctness gate: `just test-bench-inputs` passed against both the current and baseline revisions using the shared current fixture harness. +- Validated current revision: `d36a9e99ec223c4cbbf8a9191818f186159c27f7` (Git clean: `false`; + source-state SHA-256: `b8fcf706fa6e287ca6682a1ecef5a69576413b9c88040184a9b1ecb8f6e92336`) +- Validated baseline revision: `78ba1c25beaed7321e175113e0876aac6d763a59` (Git clean: `false`; + source-state SHA-256: `7d72086128b12e2a89d5ca78f0e8cc0cb393a44f1062d027a069722ff9dae007`) +- Baseline API compatibility: `la_stack_pre_rational_input_api` selects only source-compatible benchmark calls; + one-sided rows outside the baseline's correctness domain are identified by the retained CSV coverage status and note. + +## Exact arithmetic + +| Case | Benchmark | v0.4.5 (point + CI) | Latest (point + CI) | Point-estimate change | CI relation | Point-estimate ratio | +|:-----|:----------|-------:|-------:|-------:|:-----------|--------:| +| D=2 | det | 0.4 ns [0.4 ns, 0.4 ns] | 0.4 ns [0.4 ns, 0.4 ns] | -0.2% | marginal CIs overlap | 1.00x | +| D=2 | det_direct | 0.4 ns [0.4 ns, 0.4 ns] | 0.4 ns [0.4 ns, 0.4 ns] | -0.3% | marginal CIs overlap | 1.00x | +| D=2 | det_direct_with_errbound | 1.7 ns [1.7 ns, 1.7 ns] | 1.7 ns [1.7 ns, 1.7 ns] | -0.0% | marginal CIs overlap | 1.00x | +| D=2 | det_errbound | 1.7 ns [1.6 ns, 1.7 ns] | 1.6 ns [1.6 ns, 1.6 ns] | -2.5% | faster point estimate; marginal CIs separated | 1.03x | +| D=2 | det_exact | 110.5 ns [110.4 ns, 110.7 ns] | 100.4 ns [100.3 ns, 100.5 ns] | -9.2% | faster point estimate; marginal CIs separated | 1.10x | +| D=2 | det_exact_f64_result | 71.3 ns [71.2 ns, 71.5 ns] | 73.1 ns [72.9 ns, 73.2 ns] | +2.5% | slower point estimate; marginal CIs separated | 0.98x | +| D=2 | det_exact_rounded_f64 | 72.7 ns [72.6 ns, 72.7 ns] | 73.3 ns [73.2 ns, 73.4 ns] | +0.8% | slower point estimate; marginal CIs separated | 0.99x | +| D=2 | det_sign_exact | 2.7 ns [2.7 ns, 2.7 ns] | 3.0 ns [3.0 ns, 3.0 ns] | +9.1% | slower point estimate; marginal CIs separated | 0.92x | +| D=2 | solve_exact | 7.06 µs [7.05 µs, 7.07 µs] | 7.10 µs [7.09 µs, 7.12 µs] | +0.7% | slower point estimate; marginal CIs separated | 0.99x | +| D=2 | solve_exact_f64_result | 8.28 µs [8.26 µs, 8.28 µs] | 7.33 µs [7.32 µs, 7.34 µs] | -11.5% | faster point estimate; marginal CIs separated | 1.13x | +| D=2 | solve_exact_rounded_f64 | 7.48 µs [7.47 µs, 7.49 µs] | 7.55 µs [7.54 µs, 7.56 µs] | +0.8% | slower point estimate; marginal CIs separated | 0.99x | +| D=3 | det | 0.8 ns [0.8 ns, 0.8 ns] | 0.8 ns [0.8 ns, 0.8 ns] | +0.1% | marginal CIs overlap | 1.00x | +| D=3 | det_direct | 0.8 ns [0.8 ns, 0.8 ns] | 0.8 ns [0.8 ns, 0.8 ns] | +2.9% | slower point estimate; marginal CIs separated | 0.97x | +| D=3 | det_direct_with_errbound | 3.4 ns [3.3 ns, 3.4 ns] | 3.3 ns [3.3 ns, 3.3 ns] | -0.5% | faster point estimate; marginal CIs separated | 1.01x | +| D=3 | det_errbound | 3.3 ns [3.3 ns, 3.3 ns] | 3.3 ns [3.3 ns, 3.3 ns] | +0.1% | slower point estimate; marginal CIs separated | 1.00x | +| D=3 | det_exact | 370.6 ns [369.9 ns, 371.5 ns] | 338.9 ns [338.5 ns, 339.4 ns] | -8.6% | faster point estimate; marginal CIs separated | 1.09x | +| D=3 | det_exact_f64_result | 338.5 ns [338.0 ns, 339.2 ns] | 313.9 ns [313.2 ns, 314.4 ns] | -7.3% | faster point estimate; marginal CIs separated | 1.08x | +| D=3 | det_exact_rounded_f64 | 340.1 ns [338.8 ns, 341.2 ns] | 317.2 ns [316.0 ns, 318.2 ns] | -6.7% | faster point estimate; marginal CIs separated | 1.07x | +| D=3 | det_sign_exact | 4.7 ns [4.7 ns, 4.7 ns] | 4.7 ns [4.7 ns, 4.7 ns] | -0.0% | marginal CIs overlap | 1.00x | +| D=3 | solve_exact | 30.84 µs [30.78 µs, 30.89 µs] | 30.60 µs [30.51 µs, 30.67 µs] | -0.8% | faster point estimate; marginal CIs separated | 1.01x | +| D=3 | solve_exact_f64_result | 32.80 µs [32.76 µs, 32.87 µs] | 30.92 µs [30.86 µs, 30.96 µs] | -5.7% | faster point estimate; marginal CIs separated | 1.06x | +| D=3 | solve_exact_rounded_f64 | 31.46 µs [31.35 µs, 31.60 µs] | 31.60 µs [31.55 µs, 31.72 µs] | +0.5% | marginal CIs overlap | 1.00x | +| D=4 | det | 3.2 ns [3.2 ns, 3.2 ns] | 3.2 ns [3.2 ns, 3.2 ns] | -0.1% | marginal CIs overlap | 1.00x | +| D=4 | det_direct | 2.4 ns [2.4 ns, 2.4 ns] | 2.4 ns [2.4 ns, 2.4 ns] | -0.0% | marginal CIs overlap | 1.00x | +| D=4 | det_direct_with_errbound | 6.8 ns [6.8 ns, 6.8 ns] | 6.8 ns [6.8 ns, 6.8 ns] | +0.2% | slower point estimate; marginal CIs separated | 1.00x | +| D=4 | det_errbound | 6.7 ns [6.7 ns, 6.7 ns] | 6.7 ns [6.7 ns, 6.7 ns] | +0.1% | slower point estimate; marginal CIs separated | 1.00x | +| D=4 | det_exact | 1.22 µs [1.21 µs, 1.22 µs] | 1.01 µs [1.01 µs, 1.01 µs] | -17.0% | faster point estimate; marginal CIs separated | 1.20x | +| D=4 | det_exact_f64_result | 1.17 µs [1.16 µs, 1.17 µs] | 1.06 µs [1.05 µs, 1.06 µs] | -9.6% | faster point estimate; marginal CIs separated | 1.11x | +| D=4 | det_exact_rounded_f64 | 1.17 µs [1.17 µs, 1.17 µs] | 1.01 µs [1.01 µs, 1.01 µs] | -13.8% | faster point estimate; marginal CIs separated | 1.16x | +| D=4 | det_sign_exact | 7.7 ns [7.7 ns, 7.7 ns] | 8.0 ns [8.0 ns, 8.0 ns] | +3.6% | slower point estimate; marginal CIs separated | 0.97x | +| D=4 | solve_exact | 80.39 µs [80.27 µs, 80.49 µs] | 79.18 µs [79.08 µs, 79.34 µs] | -1.5% | faster point estimate; marginal CIs separated | 1.02x | +| D=4 | solve_exact_f64_result | 83.39 µs [83.24 µs, 83.63 µs] | 79.34 µs [79.13 µs, 79.50 µs] | -4.9% | faster point estimate; marginal CIs separated | 1.05x | +| D=4 | solve_exact_rounded_f64 | 81.02 µs [80.74 µs, 81.52 µs] | 79.86 µs [79.66 µs, 80.05 µs] | -1.4% | faster point estimate; marginal CIs separated | 1.01x | +| D=5 | det | 24.7 ns [24.7 ns, 24.8 ns] | 25.0 ns [25.0 ns, 25.1 ns] | +1.4% | slower point estimate; marginal CIs separated | 0.99x | +| D=5 | det_exact | 3.54 µs [3.52 µs, 3.54 µs] | 3.74 µs [3.71 µs, 3.79 µs] | +5.7% | slower point estimate; marginal CIs separated | 0.95x | +| D=5 | det_exact_f64_result | 3.49 µs [3.48 µs, 3.50 µs] | 3.54 µs [3.54 µs, 3.55 µs] | +1.7% | slower point estimate; marginal CIs separated | 0.98x | +| D=5 | det_exact_rounded_f64 | 3.47 µs [3.46 µs, 3.48 µs] | 3.61 µs [3.60 µs, 3.61 µs] | +3.8% | slower point estimate; marginal CIs separated | 0.96x | +| D=5 | det_sign_exact | 3.58 µs [3.57 µs, 3.59 µs] | 3.70 µs [3.70 µs, 3.71 µs] | +3.5% | slower point estimate; marginal CIs separated | 0.97x | +| D=5 | solve_exact | 159.88 µs [159.63 µs, 160.24 µs] | 157.60 µs [157.28 µs, 157.89 µs] | -1.4% | faster point estimate; marginal CIs separated | 1.01x | +| D=5 | solve_exact_f64_result | 163.37 µs [162.91 µs, 163.71 µs] | 157.72 µs [157.38 µs, 157.93 µs] | -3.5% | faster point estimate; marginal CIs separated | 1.04x | +| D=5 | solve_exact_rounded_f64 | 161.50 µs [161.20 µs, 161.73 µs] | 164.26 µs [163.54 µs, 164.87 µs] | +1.7% | slower point estimate; marginal CIs separated | 0.98x | +| Hilbert 4x4 | det_exact | 1.34 µs [1.33 µs, 1.35 µs] | 1.04 µs [1.04 µs, 1.05 µs] | -22.3% | faster point estimate; marginal CIs separated | 1.29x | +| Hilbert 4x4 | det_sign_exact | 7.7 ns [7.7 ns, 7.7 ns] | 8.0 ns [8.0 ns, 8.0 ns] | +3.4% | slower point estimate; marginal CIs separated | 0.97x | +| Hilbert 4x4 | solve_exact | 60.21 µs [60.13 µs, 60.29 µs] | 59.56 µs [59.47 µs, 59.63 µs] | -1.1% | faster point estimate; marginal CIs separated | 1.01x | +| Hilbert 4x4 | solve_exact_f64_result | 62.75 µs [62.65 µs, 62.89 µs] | 59.74 µs [59.59 µs, 59.79 µs] | -4.8% | faster point estimate; marginal CIs separated | 1.05x | +| Hilbert 4x4 | solve_exact_rounded_f64 | 61.09 µs [60.95 µs, 61.18 µs] | 60.47 µs [60.26 µs, 60.62 µs] | -1.0% | faster point estimate; marginal CIs separated | 1.01x | +| Hilbert 5x5 | det_exact | 3.68 µs [3.66 µs, 3.69 µs] | 3.75 µs [3.74 µs, 3.76 µs] | +1.9% | slower point estimate; marginal CIs separated | 0.98x | +| Hilbert 5x5 | det_sign_exact | 3.87 µs [3.87 µs, 3.88 µs] | 3.92 µs [3.92 µs, 3.93 µs] | +1.1% | slower point estimate; marginal CIs separated | 0.99x | +| Hilbert 5x5 | solve_exact | 124.69 µs [124.44 µs, 124.97 µs] | 122.82 µs [122.53 µs, 122.95 µs] | -1.5% | faster point estimate; marginal CIs separated | 1.02x | +| Hilbert 5x5 | solve_exact_f64_result | 127.73 µs [127.44 µs, 128.19 µs] | 122.06 µs [121.91 µs, 122.40 µs] | -4.4% | faster point estimate; marginal CIs separated | 1.05x | +| Hilbert 5x5 | solve_exact_rounded_f64 | 126.57 µs [126.11 µs, 126.93 µs] | 124.52 µs [124.21 µs, 124.64 µs] | -1.6% | faster point estimate; marginal CIs separated | 1.02x | +| Large entries 3x3 | det_exact | 338.5 ns [337.7 ns, 339.5 ns] | 341.6 ns [340.6 ns, 343.9 ns] | +0.9% | slower point estimate; marginal CIs separated | 0.99x | +| Large entries 3x3 | det_sign_exact | 320.8 ns [319.6 ns, 323.0 ns] | 334.9 ns [334.1 ns, 337.0 ns] | +4.4% | slower point estimate; marginal CIs separated | 0.96x | +| Large entries 3x3 | solve_exact | 93.99 µs [93.81 µs, 94.26 µs] | 93.14 µs [92.93 µs, 93.30 µs] | -0.9% | faster point estimate; marginal CIs separated | 1.01x | +| Large entries 3x3 | solve_exact_f64_result | 94.71 µs [94.65 µs, 94.80 µs] | 93.37 µs [93.24 µs, 93.46 µs] | -1.4% | faster point estimate; marginal CIs separated | 1.01x | +| Large entries 3x3 | solve_exact_rounded_f64 | 94.49 µs [94.43 µs, 94.61 µs] | 93.32 µs [93.20 µs, 93.38 µs] | -1.2% | faster point estimate; marginal CIs separated | 1.01x | +| Near-singular 3x3 | det_exact | 342.0 ns [341.2 ns, 342.8 ns] | 334.5 ns [333.8 ns, 336.3 ns] | -2.2% | faster point estimate; marginal CIs separated | 1.02x | +| Near-singular 3x3 | det_sign_exact | 398.3 ns [394.8 ns, 401.2 ns] | 360.9 ns [360.5 ns, 361.6 ns] | -9.4% | faster point estimate; marginal CIs separated | 1.10x | +| Near-singular 3x3 | solve_exact | 2.60 µs [2.60 µs, 2.61 µs] | 2.81 µs [2.80 µs, 2.81 µs] | +7.9% | slower point estimate; marginal CIs separated | 0.93x | +| Near-singular 3x3 | solve_exact_f64_result | 2.70 µs [2.70 µs, 2.71 µs] | 2.81 µs [2.80 µs, 2.81 µs] | +3.8% | slower point estimate; marginal CIs separated | 0.96x | +| Near-singular 3x3 | solve_exact_rounded_f64 | 2.60 µs [2.60 µs, 2.61 µs] | 2.78 µs [2.77 µs, 2.80 µs] | +6.9% | slower point estimate; marginal CIs separated | 0.94x | +| Random corpus D=2 | det_exact | 3.12 µs [3.12 µs, 3.12 µs] | 3.16 µs [3.14 µs, 3.18 µs] | +1.2% | slower point estimate; marginal CIs separated | 0.99x | +| Random corpus D=2 | det_sign_exact | 136.4 ns [135.6 ns, 137.0 ns] | 149.9 ns [148.9 ns, 150.9 ns] | +9.9% | slower point estimate; marginal CIs separated | 0.91x | +| Random corpus D=2 | solve_exact | 67.47 µs [67.33 µs, 67.59 µs] | 78.46 µs [78.37 µs, 78.65 µs] | +16.3% | slower point estimate; marginal CIs separated | 0.86x | +| Random corpus D=2 | solve_exact_f64_result | 78.30 µs [77.84 µs, 78.74 µs] | 79.81 µs [79.67 µs, 79.96 µs] | +1.9% | slower point estimate; marginal CIs separated | 0.98x | +| Random corpus D=2 | solve_exact_rounded_f64 | 68.56 µs [68.41 µs, 68.68 µs] | 82.23 µs [82.02 µs, 82.37 µs] | +19.9% | slower point estimate; marginal CIs separated | 0.83x | +| Random corpus D=3 | det_exact | 8.35 µs [8.08 µs, 8.64 µs] | 8.47 µs [8.44 µs, 8.49 µs] | +1.5% | marginal CIs overlap | 0.99x | +| Random corpus D=3 | det_sign_exact | 225.1 ns [225.0 ns, 225.2 ns] | 227.9 ns [227.4 ns, 228.3 ns] | +1.2% | slower point estimate; marginal CIs separated | 0.99x | +| Random corpus D=3 | solve_exact | 214.24 µs [213.80 µs, 214.55 µs] | 239.62 µs [239.14 µs, 240.19 µs] | +11.9% | slower point estimate; marginal CIs separated | 0.89x | +| Random corpus D=3 | solve_exact_f64_result | 229.92 µs [229.52 µs, 230.51 µs] | 240.54 µs [240.00 µs, 241.28 µs] | +4.6% | slower point estimate; marginal CIs separated | 0.96x | +| Random corpus D=3 | solve_exact_rounded_f64 | 217.63 µs [217.29 µs, 217.94 µs] | 242.67 µs [242.06 µs, 243.32 µs] | +11.5% | slower point estimate; marginal CIs separated | 0.90x | +| Random corpus D=4 | det_exact | 25.26 µs [25.21 µs, 25.30 µs] | 21.20 µs [21.16 µs, 21.27 µs] | -16.1% | faster point estimate; marginal CIs separated | 1.19x | +| Random corpus D=4 | det_sign_exact | 424.5 ns [423.1 ns, 425.6 ns] | 454.6 ns [452.7 ns, 456.0 ns] | +7.1% | slower point estimate; marginal CIs separated | 0.93x | +| Random corpus D=4 | solve_exact | 489.45 µs [488.84 µs, 490.29 µs] | 535.37 µs [533.84 µs, 537.11 µs] | +9.4% | slower point estimate; marginal CIs separated | 0.91x | +| Random corpus D=4 | solve_exact_f64_result | 507.75 µs [507.05 µs, 508.51 µs] | 534.53 µs [533.89 µs, 535.33 µs] | +5.3% | slower point estimate; marginal CIs separated | 0.95x | +| Random corpus D=4 | solve_exact_rounded_f64 | 491.51 µs [490.45 µs, 491.98 µs] | 535.39 µs [535.09 µs, 535.96 µs] | +8.9% | slower point estimate; marginal CIs separated | 0.92x | +| Random corpus D=5 | det_exact | 47.41 µs [47.34 µs, 47.50 µs] | 55.99 µs [55.88 µs, 56.13 µs] | +18.1% | slower point estimate; marginal CIs separated | 0.85x | +| Random corpus D=5 | det_sign_exact | 49.54 µs [48.76 µs, 50.88 µs] | 61.93 µs [61.88 µs, 62.04 µs] | +25.0% | slower point estimate; marginal CIs separated | 0.80x | +| Random corpus D=5 | solve_exact | 965.82 µs [964.45 µs, 966.73 µs] | 1.04 ms [1.04 ms, 1.04 ms] | +8.1% | slower point estimate; marginal CIs separated | 0.92x | +| Random corpus D=5 | solve_exact_f64_result | 987.58 µs [986.29 µs, 989.22 µs] | 1.04 ms [1.04 ms, 1.04 ms] | +5.5% | slower point estimate; marginal CIs separated | 0.95x | +| Random corpus D=5 | solve_exact_rounded_f64 | 966.85 µs [965.47 µs, 968.31 µs] | 1.04 ms [1.04 ms, 1.05 ms] | +8.1% | slower point estimate; marginal CIs separated | 0.93x | + +## vs_linalg + +| Case | Benchmark | v0.4.5 (point + CI) | Latest (point + CI) | Point-estimate change | CI relation | Point-estimate ratio | v0.4.5 nalgebra | v0.4.5 faer | +|:-----|:----------|-------:|-------:|-------:|:-----------|--------:|-------:|-------:| +| D=16 | la_stack_det | 433.3 ns [431.0 ns, 435.4 ns] | 463.8 ns [460.4 ns, 472.3 ns] | +7.0% | slower point estimate; marginal CIs separated | 0.93x | — | — | +| D=16 | la_stack_det_from_ldlt | 2.9 ns [2.9 ns, 2.9 ns] | 2.9 ns [2.9 ns, 2.9 ns] | -0.1% | marginal CIs overlap | 1.00x | 1.7 ns [1.7 ns, 1.7 ns] | 4.3 ns [4.3 ns, 4.4 ns] | +| D=16 | la_stack_det_from_lu | 3.3 ns [3.3 ns, 3.3 ns] | 3.3 ns [3.3 ns, 3.3 ns] | +0.2% | marginal CIs overlap | 1.00x | 1.8 ns [1.8 ns, 1.8 ns] | 4.8 ns [4.7 ns, 4.8 ns] | +| D=16 | la_stack_det_via_lu | 403.1 ns [402.1 ns, 405.0 ns] | 405.2 ns [403.8 ns, 406.9 ns] | +0.5% | marginal CIs overlap | 0.99x | 449.5 ns [448.4 ns, 450.3 ns] | 671.6 ns [665.6 ns, 677.2 ns] | +| D=16 | la_stack_dot | 2.3 ns [2.3 ns, 2.3 ns] | 2.3 ns [2.3 ns, 2.3 ns] | +0.5% | slower point estimate; marginal CIs separated | 1.00x | 1.9 ns [1.9 ns, 1.9 ns] | 3.3 ns [3.3 ns, 3.3 ns] | +| D=16 | la_stack_inf_norm | 32.2 ns [32.2 ns, 32.3 ns] | 32.2 ns [32.1 ns, 32.2 ns] | -0.3% | marginal CIs overlap | 1.00x | 31.4 ns [31.3 ns, 31.4 ns] | 32.7 ns [32.5 ns, 32.7 ns] | +| D=16 | la_stack_ldlt | 383.4 ns [382.4 ns, 383.9 ns] | 403.8 ns [403.4 ns, 404.5 ns] | +5.3% | slower point estimate; marginal CIs separated | 0.95x | 397.0 ns [396.6 ns, 397.6 ns] | 428.2 ns [426.1 ns, 430.8 ns] | +| D=16 | la_stack_ldlt_solve | 425.4 ns [424.9 ns, 425.8 ns] | 447.6 ns [446.3 ns, 448.1 ns] | +5.2% | slower point estimate; marginal CIs separated | 0.95x | 607.4 ns [606.6 ns, 608.8 ns] | 626.1 ns [624.9 ns, 627.1 ns] | +| D=16 | la_stack_lu | 393.8 ns [392.8 ns, 395.0 ns] | 395.6 ns [394.4 ns, 397.1 ns] | +0.5% | marginal CIs overlap | 1.00x | 453.8 ns [452.9 ns, 454.1 ns] | 641.8 ns [639.4 ns, 645.7 ns] | +| D=16 | la_stack_lu_solve | 640.8 ns [639.4 ns, 641.5 ns] | 643.5 ns [641.5 ns, 644.9 ns] | +0.4% | slower point estimate; marginal CIs separated | 1.00x | 569.4 ns [568.9 ns, 569.9 ns] | 903.1 ns [900.3 ns, 905.3 ns] | +| D=16 | la_stack_norm2_sq | 2.1 ns [2.1 ns, 2.1 ns] | 2.1 ns [2.1 ns, 2.1 ns] | +0.0% | marginal CIs overlap | 1.00x | 1.5 ns [1.5 ns, 1.5 ns] | 4.1 ns [4.1 ns, 4.1 ns] | +| D=16 | la_stack_solve_from_ldlt | 27.7 ns [27.7 ns, 27.8 ns] | 27.7 ns [27.7 ns, 27.7 ns] | -0.1% | marginal CIs overlap | 1.00x | 116.1 ns [115.7 ns, 116.2 ns] | 181.7 ns [181.4 ns, 182.2 ns] | +| D=16 | la_stack_solve_from_lu | 188.0 ns [187.6 ns, 188.6 ns] | 188.8 ns [188.3 ns, 189.4 ns] | +0.4% | marginal CIs overlap | 1.00x | 92.2 ns [92.1 ns, 92.3 ns] | 258.8 ns [255.7 ns, 262.5 ns] | +| D=2 | la_stack_det | 0.6 ns [0.6 ns, 0.6 ns] | 0.6 ns [0.6 ns, 0.6 ns] | +0.0% | marginal CIs overlap | 1.00x | — | — | +| D=2 | la_stack_det_from_ldlt | 0.5 ns [0.5 ns, 0.5 ns] | 0.5 ns [0.5 ns, 0.5 ns] | +0.0% | marginal CIs overlap | 1.00x | 0.4 ns [0.4 ns, 0.4 ns] | 0.5 ns [0.5 ns, 0.6 ns] | +| D=2 | la_stack_det_from_lu | 0.5 ns [0.5 ns, 0.5 ns] | 0.5 ns [0.5 ns, 0.5 ns] | -0.1% | marginal CIs overlap | 1.00x | 0.4 ns [0.4 ns, 0.4 ns] | 0.7 ns [0.7 ns, 0.7 ns] | +| D=2 | la_stack_det_via_lu | 1.8 ns [1.8 ns, 1.8 ns] | 1.8 ns [1.8 ns, 1.8 ns] | +0.3% | slower point estimate; marginal CIs separated | 1.00x | 0.8 ns [0.8 ns, 0.8 ns] | 131.7 ns [131.1 ns, 132.0 ns] | +| D=2 | la_stack_dot | 0.6 ns [0.6 ns, 0.6 ns] | 0.6 ns [0.6 ns, 0.6 ns] | -0.2% | marginal CIs overlap | 1.00x | 0.6 ns [0.6 ns, 0.6 ns] | 2.5 ns [2.5 ns, 2.7 ns] | +| D=2 | la_stack_inf_norm | 0.6 ns [0.6 ns, 0.6 ns] | 0.6 ns [0.6 ns, 0.6 ns] | -0.0% | marginal CIs overlap | 1.00x | 0.5 ns [0.5 ns, 0.5 ns] | 0.7 ns [0.7 ns, 0.7 ns] | +| D=2 | la_stack_ldlt | 6.5 ns [6.5 ns, 6.5 ns] | 6.6 ns [6.6 ns, 6.6 ns] | +2.6% | slower point estimate; marginal CIs separated | 0.97x | 1.7 ns [1.7 ns, 1.7 ns] | 98.0 ns [97.6 ns, 98.4 ns] | +| D=2 | la_stack_ldlt_solve | 9.3 ns [9.3 ns, 9.3 ns] | 9.4 ns [9.4 ns, 9.5 ns] | +1.2% | slower point estimate; marginal CIs separated | 0.99x | 2.5 ns [2.5 ns, 2.5 ns] | 145.9 ns [143.5 ns, 147.3 ns] | +| D=2 | la_stack_lu | 1.6 ns [1.6 ns, 1.6 ns] | 1.6 ns [1.6 ns, 1.6 ns] | -0.3% | faster point estimate; marginal CIs separated | 1.00x | 1.6 ns [1.6 ns, 1.6 ns] | 120.9 ns [120.0 ns, 121.7 ns] | +| D=2 | la_stack_lu_solve | 2.0 ns [2.0 ns, 2.0 ns] | 2.0 ns [2.0 ns, 2.0 ns] | -0.5% | marginal CIs overlap | 1.01x | 4.5 ns [4.5 ns, 4.5 ns] | 182.5 ns [182.1 ns, 183.0 ns] | +| D=2 | la_stack_norm2_sq | 0.4 ns [0.4 ns, 0.4 ns] | 0.4 ns [0.4 ns, 0.4 ns] | -0.0% | marginal CIs overlap | 1.00x | 0.4 ns [0.4 ns, 0.4 ns] | 4.1 ns [4.1 ns, 4.1 ns] | +| D=2 | la_stack_solve_from_ldlt | 1.2 ns [1.2 ns, 1.2 ns] | 1.2 ns [1.2 ns, 1.2 ns] | +1.6% | slower point estimate; marginal CIs separated | 0.98x | 1.2 ns [1.2 ns, 1.2 ns] | 44.0 ns [43.7 ns, 44.2 ns] | +| D=2 | la_stack_solve_from_lu | 1.2 ns [1.2 ns, 1.2 ns] | 1.2 ns [1.2 ns, 1.2 ns] | -0.0% | marginal CIs overlap | 1.00x | 2.9 ns [2.9 ns, 2.9 ns] | 54.8 ns [54.6 ns, 54.9 ns] | +| D=3 | la_stack_det | 1.3 ns [1.3 ns, 1.3 ns] | 1.3 ns [1.3 ns, 1.3 ns] | -0.2% | marginal CIs overlap | 1.00x | — | — | +| D=3 | la_stack_det_from_ldlt | 0.6 ns [0.6 ns, 0.6 ns] | 0.6 ns [0.6 ns, 0.6 ns] | +0.0% | marginal CIs overlap | 1.00x | 0.5 ns [0.5 ns, 0.5 ns] | 0.7 ns [0.7 ns, 0.7 ns] | +| D=3 | la_stack_det_from_lu | 0.7 ns [0.7 ns, 0.7 ns] | 0.7 ns [0.7 ns, 0.7 ns] | +0.0% | marginal CIs overlap | 1.00x | 0.5 ns [0.5 ns, 0.5 ns] | 1.0 ns [1.0 ns, 1.0 ns] | +| D=3 | la_stack_det_via_lu | 8.7 ns [8.7 ns, 8.7 ns] | 8.7 ns [8.7 ns, 8.8 ns] | +0.4% | marginal CIs overlap | 1.00x | 17.8 ns [17.7 ns, 17.9 ns] | 162.9 ns [161.8 ns, 164.2 ns] | +| D=3 | la_stack_dot | 0.7 ns [0.7 ns, 0.7 ns] | 0.7 ns [0.7 ns, 0.7 ns] | +0.1% | marginal CIs overlap | 1.00x | 0.7 ns [0.7 ns, 0.7 ns] | 3.0 ns [3.0 ns, 3.0 ns] | +| D=3 | la_stack_inf_norm | 1.3 ns [1.3 ns, 1.3 ns] | 1.3 ns [1.3 ns, 1.3 ns] | +0.0% | marginal CIs overlap | 1.00x | 1.1 ns [1.1 ns, 1.1 ns] | 1.3 ns [1.2 ns, 1.3 ns] | +| D=3 | la_stack_ldlt | 13.4 ns [13.3 ns, 13.4 ns] | 13.6 ns [13.5 ns, 13.7 ns] | +1.6% | slower point estimate; marginal CIs separated | 0.98x | 3.9 ns [3.8 ns, 3.9 ns] | 108.2 ns [107.8 ns, 109.1 ns] | +| D=3 | la_stack_ldlt_solve | 12.5 ns [12.5 ns, 12.6 ns] | 18.3 ns [17.8 ns, 20.3 ns] | +45.7% | slower point estimate; marginal CIs separated | 0.69x | 5.6 ns [5.6 ns, 5.6 ns] | 155.4 ns [154.6 ns, 155.9 ns] | +| D=3 | la_stack_lu | 8.3 ns [8.3 ns, 8.3 ns] | 17.7 ns [16.7 ns, 18.4 ns] | +114.1% | slower point estimate; marginal CIs separated | 0.47x | 14.8 ns [14.7 ns, 14.8 ns] | 156.6 ns [155.4 ns, 158.0 ns] | +| D=3 | la_stack_lu_solve | 10.1 ns [10.0 ns, 10.2 ns] | 10.0 ns [9.9 ns, 10.0 ns] | -1.1% | faster point estimate; marginal CIs separated | 1.01x | 22.8 ns [22.7 ns, 23.0 ns] | 217.6 ns [216.2 ns, 219.0 ns] | +| D=3 | la_stack_norm2_sq | 0.4 ns [0.4 ns, 0.4 ns] | 0.4 ns [0.4 ns, 0.4 ns] | +0.1% | marginal CIs overlap | 1.00x | 0.4 ns [0.4 ns, 0.4 ns] | 4.1 ns [4.1 ns, 4.1 ns] | +| D=3 | la_stack_solve_from_ldlt | 1.8 ns [1.8 ns, 1.8 ns] | 1.8 ns [1.8 ns, 1.8 ns] | +0.0% | marginal CIs overlap | 1.00x | 2.7 ns [2.7 ns, 2.8 ns] | 45.8 ns [45.4 ns, 46.5 ns] | +| D=3 | la_stack_solve_from_lu | 2.1 ns [2.1 ns, 2.1 ns] | 2.1 ns [2.1 ns, 2.1 ns] | +0.1% | marginal CIs overlap | 1.00x | 4.5 ns [4.5 ns, 4.5 ns] | 56.7 ns [56.5 ns, 57.0 ns] | +| D=32 | la_stack_det | 2.18 µs [2.17 µs, 2.18 µs] | 2.17 µs [2.16 µs, 2.18 µs] | -0.5% | marginal CIs overlap | 1.01x | — | — | +| D=32 | la_stack_det_from_ldlt | 7.7 ns [7.7 ns, 7.7 ns] | 7.8 ns [7.7 ns, 7.8 ns] | +0.6% | slower point estimate; marginal CIs separated | 0.99x | 3.0 ns [3.0 ns, 3.0 ns] | 8.3 ns [8.2 ns, 8.3 ns] | +| D=32 | la_stack_det_from_lu | 9.1 ns [9.0 ns, 9.1 ns] | 9.0 ns [9.0 ns, 9.0 ns] | -0.1% | marginal CIs overlap | 1.00x | 3.1 ns [3.1 ns, 3.1 ns] | 8.6 ns [8.6 ns, 8.7 ns] | +| D=32 | la_stack_det_via_lu | 2.05 µs [2.04 µs, 2.05 µs] | 2.05 µs [2.04 µs, 2.06 µs] | +0.3% | marginal CIs overlap | 1.00x | 2.19 µs [2.19 µs, 2.20 µs] | 2.31 µs [2.30 µs, 2.32 µs] | +| D=32 | la_stack_dot | 4.0 ns [4.0 ns, 4.0 ns] | 4.0 ns [4.0 ns, 4.0 ns] | -0.1% | marginal CIs overlap | 1.00x | 4.7 ns [4.7 ns, 4.7 ns] | 4.9 ns [4.9 ns, 4.9 ns] | +| D=32 | la_stack_inf_norm | 127.0 ns [126.9 ns, 127.2 ns] | 127.6 ns [127.4 ns, 127.7 ns] | +0.5% | slower point estimate; marginal CIs separated | 0.99x | 158.5 ns [154.9 ns, 159.6 ns] | 163.6 ns [163.4 ns, 164.2 ns] | +| D=32 | la_stack_ldlt | 2.44 µs [2.43 µs, 2.44 µs] | 2.45 µs [2.45 µs, 2.45 µs] | +0.5% | slower point estimate; marginal CIs separated | 0.99x | 2.06 µs [2.05 µs, 2.06 µs] | 1.45 µs [1.45 µs, 1.46 µs] | +| D=32 | la_stack_ldlt_solve | 2.78 µs [2.77 µs, 2.79 µs] | 2.77 µs [2.76 µs, 2.79 µs] | -0.4% | marginal CIs overlap | 1.00x | 2.69 µs [2.69 µs, 2.69 µs] | 1.94 µs [1.94 µs, 1.94 µs] | +| D=32 | la_stack_lu | 2.11 µs [2.10 µs, 2.12 µs] | 2.12 µs [2.11 µs, 2.13 µs] | +0.4% | marginal CIs overlap | 1.00x | 2.14 µs [2.13 µs, 2.14 µs] | 2.26 µs [2.26 µs, 2.27 µs] | +| D=32 | la_stack_lu_solve | 2.78 µs [2.77 µs, 2.80 µs] | 2.76 µs [2.75 µs, 2.76 µs] | -1.0% | faster point estimate; marginal CIs separated | 1.01x | 2.71 µs [2.70 µs, 2.72 µs] | 2.92 µs [2.91 µs, 2.93 µs] | +| D=32 | la_stack_norm2_sq | 4.0 ns [4.0 ns, 4.0 ns] | 4.0 ns [4.0 ns, 4.0 ns] | +0.0% | marginal CIs overlap | 1.00x | 3.8 ns [3.8 ns, 3.8 ns] | 4.2 ns [4.2 ns, 4.2 ns] | +| D=32 | la_stack_solve_from_ldlt | 307.5 ns [307.0 ns, 307.8 ns] | 310.4 ns [309.5 ns, 311.1 ns] | +0.9% | slower point estimate; marginal CIs separated | 0.99x | 533.2 ns [531.2 ns, 534.3 ns] | 466.9 ns [466.3 ns, 467.6 ns] | +| D=32 | la_stack_solve_from_lu | 661.0 ns [660.0 ns, 662.2 ns] | 652.5 ns [650.7 ns, 653.7 ns] | -1.3% | faster point estimate; marginal CIs separated | 1.01x | 327.4 ns [327.0 ns, 327.8 ns] | 631.0 ns [629.5 ns, 632.5 ns] | +| D=4 | la_stack_det | 2.6 ns [2.6 ns, 2.6 ns] | 2.6 ns [2.6 ns, 2.6 ns] | +0.1% | slower point estimate; marginal CIs separated | 1.00x | — | — | +| D=4 | la_stack_det_from_ldlt | 0.8 ns [0.8 ns, 0.8 ns] | 0.8 ns [0.8 ns, 0.8 ns] | +0.0% | marginal CIs overlap | 1.00x | 0.5 ns [0.5 ns, 0.5 ns] | 1.0 ns [1.0 ns, 1.0 ns] | +| D=4 | la_stack_det_from_lu | 0.9 ns [0.9 ns, 0.9 ns] | 0.9 ns [0.9 ns, 0.9 ns] | +0.0% | marginal CIs overlap | 1.00x | 0.6 ns [0.6 ns, 0.6 ns] | 1.2 ns [1.2 ns, 1.2 ns] | +| D=4 | la_stack_det_via_lu | 14.4 ns [14.4 ns, 14.5 ns] | 14.5 ns [14.4 ns, 14.5 ns] | +0.1% | marginal CIs overlap | 1.00x | 30.0 ns [30.0 ns, 30.0 ns] | 179.5 ns [179.0 ns, 180.9 ns] | +| D=4 | la_stack_dot | 0.7 ns [0.7 ns, 0.7 ns] | 0.7 ns [0.7 ns, 0.7 ns] | +0.0% | marginal CIs overlap | 1.00x | 0.6 ns [0.6 ns, 0.6 ns] | 2.6 ns [2.6 ns, 2.7 ns] | +| D=4 | la_stack_inf_norm | 2.2 ns [2.2 ns, 2.2 ns] | 2.2 ns [2.2 ns, 2.2 ns] | +0.1% | marginal CIs overlap | 1.00x | 2.0 ns [2.0 ns, 2.0 ns] | 2.0 ns [2.0 ns, 2.0 ns] | +| D=4 | la_stack_ldlt | 21.1 ns [21.1 ns, 21.2 ns] | 21.3 ns [21.3 ns, 21.4 ns] | +0.8% | slower point estimate; marginal CIs separated | 0.99x | 7.5 ns [7.4 ns, 7.5 ns] | 126.4 ns [126.1 ns, 126.9 ns] | +| D=4 | la_stack_ldlt_solve | 22.9 ns [22.9 ns, 23.0 ns] | 38.6 ns [38.0 ns, 39.7 ns] | +68.2% | slower point estimate; marginal CIs separated | 0.59x | 10.4 ns [10.4 ns, 10.4 ns] | 172.6 ns [172.4 ns, 173.2 ns] | +| D=4 | la_stack_lu | 14.0 ns [13.9 ns, 14.0 ns] | 13.9 ns [13.8 ns, 13.9 ns] | -0.9% | faster point estimate; marginal CIs separated | 1.01x | 29.1 ns [29.1 ns, 29.2 ns] | 175.5 ns [174.3 ns, 176.3 ns] | +| D=4 | la_stack_lu_solve | 21.9 ns [21.9 ns, 22.0 ns] | 22.1 ns [22.1 ns, 22.1 ns] | +0.8% | slower point estimate; marginal CIs separated | 0.99x | 51.7 ns [51.5 ns, 51.9 ns] | 241.5 ns [239.3 ns, 243.2 ns] | +| D=4 | la_stack_norm2_sq | 0.5 ns [0.5 ns, 0.5 ns] | 0.5 ns [0.5 ns, 0.5 ns] | +0.1% | marginal CIs overlap | 1.00x | 0.5 ns [0.5 ns, 0.5 ns] | 4.1 ns [4.1 ns, 4.1 ns] | +| D=4 | la_stack_solve_from_ldlt | 2.5 ns [2.5 ns, 2.5 ns] | 2.5 ns [2.5 ns, 2.5 ns] | -0.0% | marginal CIs overlap | 1.00x | 5.3 ns [5.2 ns, 5.3 ns] | 44.8 ns [44.6 ns, 44.9 ns] | +| D=4 | la_stack_solve_from_lu | 4.0 ns [4.0 ns, 4.0 ns] | 3.9 ns [3.9 ns, 3.9 ns] | -1.8% | faster point estimate; marginal CIs separated | 1.02x | 5.7 ns [5.6 ns, 5.7 ns] | 59.9 ns [59.5 ns, 60.2 ns] | +| D=5 | la_stack_det | 39.4 ns [39.3 ns, 39.6 ns] | 40.6 ns [40.0 ns, 58.5 ns] | +2.9% | slower point estimate; marginal CIs separated | 0.97x | — | — | +| D=5 | la_stack_det_from_ldlt | 0.9 ns [0.9 ns, 0.9 ns] | 0.9 ns [0.9 ns, 0.9 ns] | +0.1% | marginal CIs overlap | 1.00x | 0.6 ns [0.6 ns, 0.6 ns] | 1.2 ns [1.2 ns, 1.2 ns] | +| D=5 | la_stack_det_from_lu | 1.1 ns [1.1 ns, 1.1 ns] | 1.1 ns [1.1 ns, 1.1 ns] | +0.1% | marginal CIs overlap | 1.00x | 0.7 ns [0.7 ns, 0.7 ns] | 1.5 ns [1.5 ns, 1.5 ns] | +| D=5 | la_stack_det_via_lu | 32.1 ns [32.0 ns, 32.3 ns] | 32.7 ns [32.6 ns, 32.7 ns] | +1.7% | slower point estimate; marginal CIs separated | 0.98x | 54.8 ns [54.7 ns, 54.9 ns] | 223.3 ns [221.9 ns, 224.9 ns] | +| D=5 | la_stack_dot | 0.8 ns [0.8 ns, 0.8 ns] | 0.8 ns [0.8 ns, 0.8 ns] | +5.2% | slower point estimate; marginal CIs separated | 0.95x | 0.7 ns [0.7 ns, 0.8 ns] | 2.8 ns [2.8 ns, 2.8 ns] | +| D=5 | la_stack_inf_norm | 3.4 ns [3.4 ns, 3.4 ns] | 3.5 ns [3.5 ns, 3.5 ns] | +1.9% | slower point estimate; marginal CIs separated | 0.98x | 3.2 ns [3.2 ns, 3.2 ns] | 3.2 ns [3.2 ns, 3.2 ns] | +| D=5 | la_stack_ldlt | 40.5 ns [40.4 ns, 40.6 ns] | 41.9 ns [41.6 ns, 42.3 ns] | +3.6% | slower point estimate; marginal CIs separated | 0.97x | 26.1 ns [25.1 ns, 26.5 ns] | 144.0 ns [143.6 ns, 144.6 ns] | +| D=5 | la_stack_ldlt_solve | 52.3 ns [51.9 ns, 52.5 ns] | 45.4 ns [45.3 ns, 45.5 ns] | -13.3% | faster point estimate; marginal CIs separated | 1.15x | 40.4 ns [40.3 ns, 40.4 ns] | 234.3 ns [231.2 ns, 246.0 ns] | +| D=5 | la_stack_lu | 32.0 ns [31.9 ns, 32.1 ns] | 32.3 ns [32.3 ns, 32.4 ns] | +1.0% | slower point estimate; marginal CIs separated | 0.99x | 55.1 ns [54.9 ns, 55.2 ns] | 217.7 ns [214.6 ns, 220.0 ns] | +| D=5 | la_stack_lu_solve | 44.3 ns [43.9 ns, 44.8 ns] | 46.1 ns [46.0 ns, 46.1 ns] | +4.0% | slower point estimate; marginal CIs separated | 0.96x | 69.0 ns [68.9 ns, 69.2 ns] | 326.8 ns [324.0 ns, 361.6 ns] | +| D=5 | la_stack_norm2_sq | 0.5 ns [0.5 ns, 0.5 ns] | 0.5 ns [0.5 ns, 0.5 ns] | +0.1% | slower point estimate; marginal CIs separated | 1.00x | 0.6 ns [0.6 ns, 0.6 ns] | 4.1 ns [4.1 ns, 4.1 ns] | +| D=5 | la_stack_solve_from_ldlt | 3.9 ns [3.9 ns, 3.9 ns] | 3.9 ns [3.9 ns, 3.9 ns] | +0.9% | slower point estimate; marginal CIs separated | 0.99x | 8.9 ns [8.8 ns, 8.9 ns] | 63.4 ns [63.3 ns, 63.6 ns] | +| D=5 | la_stack_solve_from_lu | 6.0 ns [6.0 ns, 6.0 ns] | 6.0 ns [6.0 ns, 6.0 ns] | -0.7% | faster point estimate; marginal CIs separated | 1.01x | 8.8 ns [8.8 ns, 8.8 ns] | 93.6 ns [91.6 ns, 95.0 ns] | +| D=64 | la_stack_det | 15.21 µs [15.17 µs, 15.25 µs] | 15.11 µs [15.05 µs, 15.16 µs] | -0.7% | faster point estimate; marginal CIs separated | 1.01x | — | — | +| D=64 | la_stack_det_from_ldlt | 22.6 ns [22.5 ns, 22.6 ns] | 22.8 ns [22.8 ns, 22.8 ns] | +1.0% | slower point estimate; marginal CIs separated | 0.99x | 8.2 ns [8.2 ns, 8.3 ns] | 20.2 ns [20.2 ns, 20.3 ns] | +| D=64 | la_stack_det_from_lu | 23.0 ns [22.9 ns, 23.0 ns] | 22.9 ns [22.9 ns, 23.0 ns] | -0.2% | marginal CIs overlap | 1.00x | 8.5 ns [8.5 ns, 8.6 ns] | 21.1 ns [21.1 ns, 21.2 ns] | +| D=64 | la_stack_det_via_lu | 15.21 µs [15.17 µs, 15.23 µs] | 15.03 µs [14.99 µs, 15.06 µs] | -1.1% | faster point estimate; marginal CIs separated | 1.01x | 13.41 µs [13.40 µs, 13.43 µs] | 10.55 µs [10.55 µs, 10.56 µs] | +| D=64 | la_stack_dot | 10.7 ns [10.6 ns, 10.7 ns] | 10.8 ns [10.8 ns, 10.8 ns] | +1.4% | slower point estimate; marginal CIs separated | 0.99x | 8.9 ns [8.9 ns, 8.9 ns] | 8.3 ns [8.3 ns, 8.3 ns] | +| D=64 | la_stack_inf_norm | 612.7 ns [612.2 ns, 614.1 ns] | 611.1 ns [610.2 ns, 613.2 ns] | -0.3% | marginal CIs overlap | 1.00x | 1.09 µs [1.08 µs, 1.09 µs] | 1.54 µs [1.54 µs, 1.54 µs] | +| D=64 | la_stack_ldlt | 19.64 µs [19.60 µs, 19.65 µs] | 20.87 µs [20.81 µs, 20.91 µs] | +6.3% | slower point estimate; marginal CIs separated | 0.94x | 11.48 µs [11.46 µs, 11.50 µs] | 8.81 µs [8.79 µs, 8.82 µs] | +| D=64 | la_stack_ldlt_solve | 21.85 µs [21.79 µs, 21.90 µs] | 23.16 µs [23.08 µs, 23.26 µs] | +6.0% | slower point estimate; marginal CIs separated | 0.94x | 12.12 µs [12.09 µs, 12.14 µs] | 10.15 µs [10.13 µs, 10.16 µs] | +| D=64 | la_stack_lu | 15.02 µs [14.99 µs, 15.03 µs] | 14.42 µs [14.40 µs, 14.44 µs] | -4.0% | faster point estimate; marginal CIs separated | 1.04x | 12.93 µs [12.91 µs, 12.94 µs] | 10.41 µs [10.40 µs, 10.43 µs] | +| D=64 | la_stack_lu_solve | 17.22 µs [17.16 µs, 17.44 µs] | 17.52 µs [17.36 µs, 17.56 µs] | +1.7% | marginal CIs overlap | 0.98x | 14.31 µs [14.30 µs, 14.35 µs] | 12.11 µs [12.09 µs, 12.12 µs] | +| D=64 | la_stack_norm2_sq | 10.6 ns [10.6 ns, 10.7 ns] | 10.8 ns [10.7 ns, 10.8 ns] | +1.3% | slower point estimate; marginal CIs separated | 0.99x | 7.5 ns [7.5 ns, 7.5 ns] | 6.2 ns [6.2 ns, 6.2 ns] | +| D=64 | la_stack_solve_from_ldlt | 1.46 µs [1.45 µs, 1.46 µs] | 1.08 µs [1.07 µs, 1.08 µs] | -26.1% | faster point estimate; marginal CIs separated | 1.35x | 1.25 µs [1.24 µs, 1.25 µs] | 1.26 µs [1.26 µs, 1.27 µs] | +| D=64 | la_stack_solve_from_lu | 2.52 µs [2.51 µs, 2.52 µs] | 2.47 µs [2.47 µs, 2.48 µs] | -1.7% | faster point estimate; marginal CIs separated | 1.02x | 793.3 ns [792.9 ns, 793.9 ns] | 1.67 µs [1.67 µs, 1.68 µs] | +| D=8 | la_stack_det | 90.7 ns [90.5 ns, 91.0 ns] | 90.5 ns [90.3 ns, 90.6 ns] | -0.3% | marginal CIs overlap | 1.00x | — | — | +| D=8 | la_stack_det_from_ldlt | 1.3 ns [1.3 ns, 1.3 ns] | 1.3 ns [1.3 ns, 1.3 ns] | -0.0% | marginal CIs overlap | 1.00x | 0.9 ns [0.9 ns, 0.9 ns] | 1.9 ns [1.9 ns, 1.9 ns] | +| D=8 | la_stack_det_from_ldlt_balanced_range | 8.9 ns [8.9 ns, 8.9 ns] | 8.9 ns [8.9 ns, 9.0 ns] | +0.4% | slower point estimate; marginal CIs separated | 1.00x | — | — | +| D=8 | la_stack_det_from_lu | 1.5 ns [1.5 ns, 1.5 ns] | 1.5 ns [1.5 ns, 1.5 ns] | +0.3% | slower point estimate; marginal CIs separated | 1.00x | 1.0 ns [1.0 ns, 1.0 ns] | 2.2 ns [2.2 ns, 2.2 ns] | +| D=8 | la_stack_det_from_lu_balanced_range | 9.7 ns [9.7 ns, 9.8 ns] | 9.1 ns [9.0 ns, 9.2 ns] | -6.4% | faster point estimate; marginal CIs separated | 1.07x | — | — | +| D=8 | la_stack_det_via_lu | 84.1 ns [83.9 ns, 84.4 ns] | 84.6 ns [84.3 ns, 85.0 ns] | +0.6% | marginal CIs overlap | 0.99x | 144.2 ns [143.9 ns, 144.6 ns] | 294.0 ns [292.1 ns, 294.7 ns] | +| D=8 | la_stack_dot | 0.9 ns [0.9 ns, 0.9 ns] | 1.0 ns [1.0 ns, 1.0 ns] | +3.2% | slower point estimate; marginal CIs separated | 0.97x | 1.1 ns [1.1 ns, 1.1 ns] | 2.6 ns [2.6 ns, 2.7 ns] | +| D=8 | la_stack_inf_norm | 8.4 ns [8.4 ns, 8.4 ns] | 8.4 ns [8.4 ns, 8.4 ns] | -0.1% | marginal CIs overlap | 1.00x | 8.0 ns [8.0 ns, 8.0 ns] | 8.0 ns [8.0 ns, 8.1 ns] | +| D=8 | la_stack_ldlt | 90.7 ns [89.7 ns, 92.3 ns] | 90.5 ns [88.6 ns, 91.4 ns] | -0.2% | marginal CIs overlap | 1.00x | 97.7 ns [97.3 ns, 98.1 ns] | 215.2 ns [214.5 ns, 216.1 ns] | +| D=8 | la_stack_ldlt_ill_conditioned | 91.2 ns [89.9 ns, 92.6 ns] | 90.3 ns [88.8 ns, 91.4 ns] | -1.0% | marginal CIs overlap | 1.01x | — | — | +| D=8 | la_stack_ldlt_solve | 100.9 ns [100.6 ns, 101.1 ns] | 99.3 ns [98.9 ns, 99.5 ns] | -1.6% | faster point estimate; marginal CIs separated | 1.02x | 142.0 ns [134.5 ns, 144.4 ns] | 296.7 ns [291.9 ns, 300.7 ns] | +| D=8 | la_stack_lu | 82.2 ns [82.0 ns, 82.4 ns] | 82.9 ns [82.6 ns, 83.3 ns] | +0.9% | slower point estimate; marginal CIs separated | 0.99x | 144.0 ns [143.7 ns, 144.5 ns] | 275.8 ns [273.4 ns, 277.3 ns] | +| D=8 | la_stack_lu_ill_conditioned | 81.9 ns [81.5 ns, 82.1 ns] | 82.3 ns [82.2 ns, 82.6 ns] | +0.6% | slower point estimate; marginal CIs separated | 0.99x | — | — | +| D=8 | la_stack_lu_pivoting | 90.8 ns [90.6 ns, 90.9 ns] | 90.4 ns [90.2 ns, 90.7 ns] | -0.4% | marginal CIs overlap | 1.00x | — | — | +| D=8 | la_stack_lu_solve | 137.4 ns [135.7 ns, 143.1 ns] | 128.4 ns [128.2 ns, 128.7 ns] | -6.5% | faster point estimate; marginal CIs separated | 1.07x | 165.9 ns [164.8 ns, 168.9 ns] | 401.4 ns [396.0 ns, 407.1 ns] | +| D=8 | la_stack_norm2_sq | 0.7 ns [0.7 ns, 0.7 ns] | 0.7 ns [0.7 ns, 0.7 ns] | +0.4% | marginal CIs overlap | 1.00x | 0.7 ns [0.7 ns, 0.7 ns] | 4.2 ns [4.2 ns, 4.2 ns] | +| D=8 | la_stack_solve_from_ldlt | 8.1 ns [8.1 ns, 8.1 ns] | 8.0 ns [8.0 ns, 8.1 ns] | -0.4% | faster point estimate; marginal CIs separated | 1.00x | 21.0 ns [21.0 ns, 21.0 ns] | 71.1 ns [70.8 ns, 71.1 ns] | +| D=8 | la_stack_solve_from_lu | 13.3 ns [13.3 ns, 13.3 ns] | 13.5 ns [13.5 ns, 13.5 ns] | +1.2% | slower point estimate; marginal CIs separated | 0.99x | 16.0 ns [15.9 ns, 16.0 ns] | 99.5 ns [99.0 ns, 99.6 ns] | + +## Coverage Notes + +One-sided rows retain the available measurement but are excluded from point-estimate change and ratio calculations. + +| Benchmark | Coverage | v0.4.5 (point + CI) | Latest (point + CI) | Note | +|:----------|:---------|-----------------------------:|--------------------:|:-----| +| rational_input_d2/det_big_rational_gaussian | current-only | — | 1.18 µs [1.18 µs, 1.18 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d2/det_row_cleared_bareiss | current-only | — | 191.3 ns [190.9 ns, 192.0 ns] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d2/det_sign_row_cleared_bareiss | current-only | — | 82.3 ns [82.2 ns, 82.4 ns] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d2/solve_big_rational_gaussian | current-only | — | 2.22 µs [2.22 µs, 2.22 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d2/solve_row_cleared_bareiss | current-only | — | 967.5 ns [966.5 ns, 968.2 ns] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d3/det_big_rational_gaussian | current-only | — | 4.89 µs [4.88 µs, 4.90 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d3/det_row_cleared_bareiss | current-only | — | 578.6 ns [578.0 ns, 579.1 ns] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d3/det_sign_row_cleared_bareiss | current-only | — | 368.6 ns [367.8 ns, 369.6 ns] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d3/solve_big_rational_gaussian | current-only | — | 8.44 µs [8.43 µs, 8.46 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d3/solve_row_cleared_bareiss | current-only | — | 2.50 µs [2.49 µs, 2.51 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d4/det_big_rational_gaussian | current-only | — | 14.16 µs [14.14 µs, 14.19 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d4/det_row_cleared_bareiss | current-only | — | 1.17 µs [1.16 µs, 1.17 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d4/det_sign_row_cleared_bareiss | current-only | — | 769.1 ns [767.2 ns, 770.7 ns] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d4/solve_big_rational_gaussian | current-only | — | 24.94 µs [24.90 µs, 24.98 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d4/solve_row_cleared_bareiss | current-only | — | 5.51 µs [5.51 µs, 5.52 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d5/det_big_rational_gaussian | current-only | — | 24.97 µs [24.55 µs, 25.75 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d5/det_row_cleared_bareiss | current-only | — | 2.15 µs [2.15 µs, 2.16 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d5/det_sign_row_cleared_bareiss | current-only | — | 1.58 µs [1.58 µs, 1.59 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d5/solve_big_rational_gaussian | current-only | — | 41.11 µs [41.07 µs, 41.14 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d5/solve_row_cleared_bareiss | current-only | — | 9.67 µs [9.64 µs, 9.72 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d6/det_big_rational_gaussian | current-only | — | 57.91 µs [57.84 µs, 58.02 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d6/det_row_cleared_bareiss | current-only | — | 3.75 µs [3.75 µs, 3.76 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d6/det_sign_row_cleared_bareiss | current-only | — | 3.10 µs [3.09 µs, 3.10 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d6/solve_big_rational_gaussian | current-only | — | 92.50 µs [92.35 µs, 92.60 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d6/solve_row_cleared_bareiss | current-only | — | 18.03 µs [18.01 µs, 18.05 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d7/det_big_rational_gaussian | current-only | — | 112.35 µs [112.25 µs, 112.50 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d7/det_row_cleared_bareiss | current-only | — | 6.15 µs [6.14 µs, 6.16 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d7/det_sign_row_cleared_bareiss | current-only | — | 5.30 µs [5.29 µs, 5.31 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d7/solve_big_rational_gaussian | current-only | — | 165.07 µs [164.88 µs, 165.48 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d7/solve_row_cleared_bareiss | current-only | — | 29.46 µs [29.44 µs, 29.50 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d8/det_big_rational_gaussian | current-only | — | 194.98 µs [193.24 µs, 197.90 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d8/det_row_cleared_bareiss | current-only | — | 9.31 µs [9.29 µs, 9.33 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d8/det_sign_row_cleared_bareiss | current-only | — | 8.47 µs [8.46 µs, 8.48 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d8/solve_big_rational_gaussian | current-only | — | 272.40 µs [272.07 µs, 272.65 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | +| rational_input_d8/solve_row_cleared_bareiss | current-only | — | 43.37 µs [43.31 µs, 43.43 µs] | Baseline v0.4.5 has no correctness-compatible benchmark row under la_stack_pre_rational_input_api. | + +## How to Update + +Local performance reports are generated in isolated temporary worktrees: + +```bash +# Local development: compare the current tree with the latest release +just performance-local + +# Release PR: update docs/performance.md and archive the previous report +just performance-release + +# Build release docs from retained CSV/JSON inputs (no benchmarks) +just performance-doc + +# GitHub Actions release assets +just performance-github-assets + +# Explicit repair +just performance-release +``` + +`just performance-local` writes `performance.md` plus retained `performance.csv` and +`performance.provenance.json` comparison inputs under `target/bench-reports/` without promoting documentation. +It applies staged and unstaged tracked changes; untracked files are excluded. +`just performance-github-assets` writes `target/bench-reports/github-assets-performance.md`. +`just performance-release` also preserves every recorded local benchmark summary and the report inputs under +`docs/performance/-vs-//`, then promotes distinct-release documentation. +`just performance-doc` consumes the retained pair from either workflow without benchmarking and promotes it when the package versions differ. +After `just clean`, `performance-doc` and `performance-readme` use the latest complete snapshot in `docs/performance/`. +For a distinct pair, `performance-local` followed by `performance-doc` is equivalent to the atomic `performance-release` workflow. + +See the [local summary index](https://github.com/acgetchell/la-stack/tree/main/docs/performance) +for the saved-data schema, provenance, and historical availability. +Older curated release-to-release reports are archived in `docs/archive/performance/`. + +See `docs/BENCHMARKING.md` for the full comparison workflow. diff --git a/docs/code_organization.md b/docs/code_organization.md index 2703b17..dd916ec 100644 --- a/docs/code_organization.md +++ b/docs/code_organization.md @@ -154,8 +154,9 @@ retain scientific schemas and consumer policy checks. Performance consumers use shared Criterion parsing and estimate/comparison validation, digest verification, archive extraction, byte-preserving document -sections, and multi-file transactions. Local rendering produces complete candidate -outputs before publication. Historical artifact schemas and fingerprint framing, +sections, and multi-file transactions. Shared rendering produces complete-run +reports; local rendering remains for historical reports and dimension plots. +Historical artifact schemas and fingerprint framing, benchmark selection and eligibility remain in the consumer. Published v0.1.8 owns common-harness orchestration, completeness, worktrees, complete run retention, validated latest selection, and transactional report publication. @@ -176,9 +177,9 @@ gates for both local measurements and release inventories. `tooling/performance.toml` uses the shared measurement schema for source/harness inputs, probes, timeouts, and compatibility fields. `scripts/performance_runs.py` declares scientific coverage, -release compatibility, and named reference phases, and renders the local tables. +release compatibility, and named reference phases used by shared reports and local plots. `scripts/archive_performance.py` composes the shared APIs for consumer commands. -`tooling/performance-report.toml` owns the new retention paths. +`tooling/performance-report.toml` owns the canonical report path, title, and shared history path. `scripts/benchmark_summaries.py` only reads historical CSV/JSON snapshots; `scripts/criterion_measurements.py` and the old generic measurement/retention engines are removed. `scripts/tests/test_performance_workflow.py` verifies diff --git a/docs/performance/README.md b/docs/performance/README.md index 4cd2824..398f9bc 100644 --- a/docs/performance/README.md +++ b/docs/performance/README.md @@ -1,8 +1,9 @@ -# Local Benchmark Summaries +# Legacy Local Benchmark Summaries -This directory preserves complete local measurement summaries across releases -and `just clean`. The generated [performance report](../performance.md) is a -selected view of those measurements; the README chart uses its LU-solve subset. +This directory preserves historical local measurement summaries across +`just clean`. The last published report and README chart used these inputs. +New measurements use shared complete-run evidence under `docs/performance-runs/`; +the [Benchmarking guide](../BENCHMARKING.md) owns the current workflow. ## Contents @@ -13,10 +14,10 @@ selected view of those measurements; the README chart uses its LU-solve subset. ## Saved runs -`just performance-release` saves each successful local comparison under -`-vs-//`. The digest identifies the complete -contents, so another run or machine creates a separate snapshot. Identical -promotion is idempotent. Earlier snapshots remain available. +Historical comparisons were saved under +`-vs-//`. The digest identifies the snapshot's +contents. These files retain their original schemas and hashes; new runs do not +write this format. Each snapshot contains: @@ -34,14 +35,13 @@ Markdown. Each recorded case must have valid metadata, estimates, and 100 raw samples before publication. Baseline-only peers remain explicitly baseline measurements; missing current measurements are never filled from an older run. -`latest.json` points to the most recently promoted complete snapshot. It is -updated with the report in the same rollback-protected publication operation. -Review and commit the snapshot and pointer along with the release report. +`latest.json` identifies the last saved legacy snapshot. New complete-run +publication maintains its own shared index and latest pointer instead. ## Regenerating reports -After a successful release promotion, these commands can use the committed -snapshot when the scratch report inputs are absent: +While no shared complete run has been retained, these commands can use the +committed legacy snapshot when scratch report inputs are absent: ```bash just performance-doc @@ -56,10 +56,9 @@ release version. Committing only the saved artifacts does not invalidate a snapshot: its original measurement commit remains in provenance. Legacy inputs without a benchmark-contract digest also require the original commit. -`performance-local` initially keeps complete summaries beside its selected -inputs under `target/bench-reports/`. Promote a distinct-release comparison with -`performance-doc` before cleaning if it should become committed release history. -Same-version local experiments remain scratch output. +Once a shared complete run is promoted, the commands select that history. +Legacy summaries cannot be converted losslessly into complete-run evidence: +they retain aggregate statistics and sample counts, but not the raw samples. ## Comparing measurements diff --git a/justfile b/justfile index 4dcc328..f76755d 100644 --- a/justfile +++ b/justfile @@ -329,7 +329,7 @@ markdown-check: tools-check _ensure-uv files=() while IFS= read -r -d '' file; do case "$file" in - CHANGELOG.md|docs/archive/*|docs/archives/changelog/*|docs/performance-runs/*) continue ;; + CHANGELOG.md|docs/archive/*|docs/archives/changelog/*|docs/performance.md|docs/performance-runs/*) continue ;; esac if [ -f "$file" ]; then files+=("$file") @@ -353,7 +353,7 @@ markdown-fix: tools-check files=() while IFS= read -r -d '' file; do case "$file" in - CHANGELOG.md|docs/archive/*|docs/archives/changelog/*|docs/performance-runs/*) continue ;; + CHANGELOG.md|docs/archive/*|docs/archives/changelog/*|docs/performance.md|docs/performance-runs/*) continue ;; esac if [ -f "$file" ]; then files+=("$file") diff --git a/scripts/README.md b/scripts/README.md index 4499fa2..8c30e34 100644 --- a/scripts/README.md +++ b/scripts/README.md @@ -68,10 +68,10 @@ are excluded; historical Markdown reports remain available in the archive. # Local development: compare the current tree with the latest release just performance-local -# Release PR: update docs/performance.md and archive the previous report +# Release PR: publish docs/performance.md and retain the complete run just performance-release -# Build release docs from retained CSV/JSON inputs +# Build release docs from retained complete-run evidence (or historical inputs) just performance-doc # GitHub Actions release assets, without local cargo benchmark runs @@ -87,8 +87,8 @@ its commit/ref and source-state provenance still distinguish the revisions. The local release workflows run the independent benchmark-input correctness gate and then measure both library revisions with one hashed current benchmark -harness. Reports record source-state, environment, toolchain, dependency, -Criterion, harness, and validation provenance and fail on incomplete selected +harness. Retained evidence records source-state, environment, toolchain, dependency, +Criterion, harness, and validation provenance and fails on incomplete selected coverage. `performance-local` writes `performance.md`, `performance.run.json`, and `performance.evidence.json` under `target/bench-reports/`. The shared complete-run payload preserves every semantic Criterion case, 100 raw samples, @@ -96,8 +96,11 @@ mean and median estimates, and 95% intervals from both phases. `performance-release` requires distinct releases and publishes immutable `run.json`, `evidence.json`, and `report.md` files under `docs/performance-runs/runs//`, with a validated index and latest -pointer. Its scientific tables are published to `docs/performance.md` in the -same transaction. Repeated runs for one pair coexist; publication failures +pointer. The shared full report is published to `docs/performance.md` in the +same transaction, using the path and title in `tooling/performance-report.toml`. +Its mean and median tables show named series, marginal intervals, and unavailable +measurements, with each series attributed to its original phase. +Full provenance remains in the evidence envelope. Repeated runs for one pair coexist; publication failures preserve the previous reports and selection. `performance-doc` replays this evidence without Cargo. Same-version local comparisons require distinct source identities and cannot be promoted as release reports. @@ -352,9 +355,10 @@ and recipe forwarding. `tests/test_cargo_update_integration.py` executes native Cargo upgrades against a disposable local registry; Python updates have a matching real-uv fixture in the toolchain tests. Common parser, transaction, Markdown, and fixture regressions belong to the shared package. Scientific -eligibility, Cargo commands/features, release adapters, selected rows, and report -layouts remain consumer-owned. The shared package owns common-harness phases, -completeness, worktrees, run identities, retention, and publication. +eligibility, Cargo commands/features, release adapters, selected rows, historical +report layouts, and dimension plots remain consumer-owned. The shared package +owns complete-run reports, common-harness phases, completeness, worktrees, run +identities, retention, and publication. Notebook tooling remains outside scope. The same pinned release owns opt-in CodeRabbit review orchestration through @@ -392,7 +396,7 @@ preview an annotation without creating a tag. | `benchmark_contract.py` | Select the Cargo checkout and hash the historical benchmark inventory contract | | `benchmark_summaries.py` | Read historical complete CSV/JSON summaries without changing framing | | `criterion_dim_plot.py` | Plot Criterion benchmark results (CSV + SVG + README table) | -| `performance_runs.py` | Declare scientific run policy and adapt retained series to local tables | +| `performance_runs.py` | Declare scientific run policy and adapt retained series for dimension plots | | `performance_artifacts.py` | Validate and publish schema-versioned performance-comparison CSV/JSON inputs | | `release_baseline.py` | Inventory full Criterion suites and validate complete raw release baselines before packaging | @@ -424,6 +428,10 @@ summary retention writer, and their duplicate generic tests are removed. policies, failure preservation, historical golden bytes, and offline replay. The shared multi-series plotting extension is deferred, so the existing local CSV/SVG/README adapter preserves figure format and baseline-phase peer labels. +Complete-run Markdown uses the shared renderer directly. Upstream follow-ups +cover [coordinate plots](https://github.com/acgetchell/research-repo-tools/issues/95), +[provenance summaries](https://github.com/acgetchell/research-repo-tools/issues/96), and +[complete-run document publication](https://github.com/acgetchell/research-repo-tools/issues/97). Performance scripts also use the published shared Criterion parser and estimate validation, comparison arithmetic, exact-byte digest verification, safe archive diff --git a/scripts/archive_performance.py b/scripts/archive_performance.py index 4a0ecb5..23e7533 100644 --- a/scripts/archive_performance.py +++ b/scripts/archive_performance.py @@ -14,7 +14,7 @@ from research_repo_tools.archives import extract_archive from research_repo_tools.common_measurement import measure_prepared_pair -from research_repo_tools.complete_runs import run_identity +from research_repo_tools.complete_runs import render_run, run_identity from research_repo_tools.evidence import Evidence, serialize_evidence from research_repo_tools.files import replace_many from research_repo_tools.measurement import resolve_revision @@ -34,7 +34,6 @@ REPORT_CONFIG, archive_directory, measurement_plan, - render_scientific_report, retained_run, scratch_paths, validate_scientific_run, @@ -104,10 +103,11 @@ def measure(root: Path, pair: ReleasePair, suite: str, scope: str) -> Evidence: def save_scratch(root: Path, stem: Path, evidence: Evidence) -> None: - """Save new complete evidence and scientific prose without touching legacy files.""" + """Save complete evidence and the shared report without touching legacy files.""" + validate_scientific_run(evidence) payload, manifest = scratch_paths(stem) data, envelope = serialize_evidence(evidence) - report = render_scientific_report(evidence) + report = render_run(evidence) plan = plan_outputs( root, {path.relative_to(root).as_posix(): value for path, value in ((payload, data), (manifest, envelope), (stem.with_suffix(".md"), report))}, @@ -117,7 +117,7 @@ def save_scratch(root: Path, stem: Path, evidence: Evidence) -> None: def promote(root: Path, stem: Path) -> None: - """Publish shared history, latest selection and scientific prose in one plan.""" + """Validate consumer policy and publish the shared report and history plan.""" evidence = retained_run(root, stem) validate_scientific_run(evidence, release=True) payload, manifest = scratch_paths(stem) @@ -127,30 +127,9 @@ def promote(root: Path, stem: Path) -> None: history = load_run_report_plan(root, REPORT_CONFIG, **kwargs) outputs = dict(history.outputs) if json.loads(outputs[f"{archive_directory(root)}/latest.json"])["run"]["id"] != run_identity(evidence): - msg = "selected evidence changed while composing scientific report" - raise ValueError(msg) - outputs["docs/performance.md"] = render_scientific_report(evidence) - # Same-pair reruns are retained in shared history. Archive the previous - # curated pair when advancing releases, validating even existing archives. - previous = root / "docs/performance.md" - previous_bytes = previous.read_bytes() if previous.exists() else None - if previous_bytes is not None and previous_bytes != outputs["docs/performance.md"]: - identity = parse_report_id(previous_bytes.decode("utf-8")) - if identity != parse_report_id(outputs["docs/performance.md"].decode("utf-8")): - name = f"docs/archive/performance/{identity.current_tag}-vs-{identity.baseline_tag}.md" - outputs[name] = previous_bytes - originals = dict(history.originals) - inputs = {name: value for name, value in originals.items() if name not in outputs and value is not None} - immutable = tuple(name for name in outputs if "/runs/" in name or name.startswith("docs/archive/performance/")) - combined = plan_outputs(root, outputs, inputs=inputs, immutable=immutable) - combined = replace(combined, git_checks=history.git_checks, glob_checks=history.glob_checks) - if dict(combined.originals)["docs/performance.md"] != previous_bytes: - msg = "curated report changed during publication planning" - raise ValueError(msg) - if any(dict(combined.originals).get(name) != value for name, value in history.originals): - msg = "publication inputs changed while composing scientific report" + msg = "selected evidence changed while planning report publication" raise ValueError(msg) - publish_publication(combined) + publish_publication(history) def parse_report_id(text: str) -> ReportId: diff --git a/scripts/performance_runs.py b/scripts/performance_runs.py index 53eeb71..f95f53a 100644 --- a/scripts/performance_runs.py +++ b/scripts/performance_runs.py @@ -1,4 +1,4 @@ -"""la-stack scientific policy and rendering over shared complete-run contracts.""" +"""la-stack scientific policy and series selection over shared complete runs.""" import json import os @@ -8,7 +8,7 @@ from typing import TYPE_CHECKING from research_repo_tools.common_measurement import CommonHarnessPlan, MeasurementPhase -from research_repo_tools.complete_runs import CompletePolicy, CompleteRun, RunSeries, run_identity, validate_run_evidence +from research_repo_tools.complete_runs import CompletePolicy, CompleteRun, RunSeries, validate_run_evidence from research_repo_tools.evidence import Evidence, fingerprint_files, load_evidence from research_repo_tools.measurement import MeasurementConfig, load_measurement from research_repo_tools.process import run_command @@ -272,75 +272,3 @@ def estimates_by_series(run: CompleteRun, statistic: str = "median") -> dict[str raise ValueError(f"unsupported statistic: {statistic}") phases = {phase: {case.full_id: case for case in sample.cases} for phase, sample in run.phases} return {series.name: {label: phases[series.phase][full_id].estimate(statistic) for label, full_id in series.rows} for series in run.series} - - -def render_scientific_report(evidence: Evidence) -> bytes: - """Retain la-stack's comparison tables and cautious statistical interpretation.""" - run = validate_scientific_run(evidence) - sources = dict(evidence.sources) - current, baseline = (dict(sources[phase].context)["release"] for phase in ("current", "baseline")) - values = estimates_by_series(run) - comparisons = [] - current_only = [] - for name, estimate in values["la-stack current"].items(): - group, bench = name.split("/", 1) - prior = values["la-stack baseline"].get(name) - if prior is None: - current_only.append(f"- `{name}`: current {report.format_time(estimate.point)}; baseline API unavailable.") - continue - - comparisons.append( - report.Comparison( - suite="exact" if group.startswith(("exact_", *UNAVAILABLE_PREFIXES)) else "vs_linalg", - group=group, - bench=bench, - baseline_bench=bench, - baseline=prior, - current=estimate, - assessment=report.assess_change(prior, estimate), - baseline_nalgebra=values.get("nalgebra", {}).get(name), - baseline_faer=values.get("faer", {}).get(name), - ) - ) - lines = [ - "# Benchmark Performance", - "", - f"**la-stack** {current} · `{sources['current'].revision}`", - "", - f"Comparison against baseline **{baseline}**:", - "", - "**Statistic**: median; complete mean and median measurements are retained.", - "", - "Negative point-estimate change means a smaller current estimate; a baseline/current ratio above 1 has the same meaning.", - "Marginal Criterion interval separation is not a paired confidence interval or a statistical-significance claim.", - "", - "nalgebra/faer reference measurements originate in the baseline phase under the captured current harness; they were not rerun in the current phase.", - "", - f"**Complete run**: `{run_identity(evidence)}`", - "", - ] - for phase, source in evidence.sources: - context = dict(source.context) - lines.extend( - [ - f"- {phase}:", - f" - Revision: `{source.revision}`.", - f" - Source: `{source.source_sha256}`.", - f" - Harness: `{source.harness_sha256}`.", - f" - Gate status: {context['gate-status']}.", - f" - Gate: `{context['gate-command']}`.", - f" - Command: `{context['command']}`.", - ] - ) - lines.extend( - [ - "", - report.comparison_tables(comparisons, baseline), - "", - *current_only, - "", - "Regenerate from retained evidence with `just performance-doc`; update README with `just performance-readme`.", - "", - ] - ) - return "\n".join(lines).encode() diff --git a/scripts/tests/test_performance_workflow.py b/scripts/tests/test_performance_workflow.py index 4627a83..a20d6ae 100644 --- a/scripts/tests/test_performance_workflow.py +++ b/scripts/tests/test_performance_workflow.py @@ -14,7 +14,7 @@ import pytest from research_repo_tools.common_measurement import CommonHarnessPlan, measure_prepared_pair -from research_repo_tools.complete_runs import RunSeries, run_identity, serialize_run +from research_repo_tools.complete_runs import RunSeries, render_run, run_identity, serialize_run from research_repo_tools.just_inspect import dry_run from research_repo_tools.process import run_command from research_repo_tools.release_pairs import ReleasePair @@ -132,6 +132,16 @@ def complete_run(tmp_path: Path, measured_run: tuple[Path, Evidence]) -> tuple[P return destination, evidence +@pytest.fixture +def legacy_artifacts(tmp_path: Path) -> ArtifactPaths: + """Keep historical promotion coverage on the last published CSV schema.""" + history = ROOT / "docs/performance" + run = json.loads((history / "latest.json").read_bytes())["run"] + for name in ("performance.csv", "performance.provenance.json"): + shutil.copyfile(history / run / name, tmp_path / name) + return ArtifactPaths(tmp_path / "performance.csv", tmp_path / "performance.provenance.json") + + def test_common_harness_and_reference_phase(tmp_path: Path) -> None: baseline, current, plan = prepared(tmp_path) (baseline / "benches/exact.rs").write_text("obsolete harness\n", encoding="utf-8") @@ -143,10 +153,6 @@ def test_common_harness_and_reference_phase(tmp_path: Path) -> None: assert (baseline / "benches/exact.rs").read_bytes() == (current / "benches/exact.rs").read_bytes() assert {series.phase for series in run.series if series.name in {"nalgebra", "faer"}} == {"baseline"} assert all(sample.policy.statistics == ("mean", "median") for _, sample in run.phases) - report = policy.render_scientific_report(evidence) - assert b"not rerun in the current phase" in report - assert b"baseline API unavailable" in report - assert b"paired confidence interval" in report @pytest.mark.parametrize("fault", ["gate", "stale", "missing-row", "missing-statistic", "missing-sample"]) @@ -159,14 +165,31 @@ def test_declared_gates_and_completeness_fail_closed(tmp_path: Path, fault: str) assert not (baseline / "target/criterion").exists() -def test_shared_retention_repeated_runs_and_offline_report(tmp_path: Path) -> None: - root, first = measured(tmp_path) +@pytest.mark.parametrize("report_path", ["docs/performance.md", "docs/reports/current.md"]) +def test_shared_retention_repeated_runs_and_offline_report(complete_run: tuple[Path, Evidence], report_path: str) -> None: + root, first = complete_run config = root / policy.REPORT_CONFIG - config.write_text(config.read_text(encoding="utf-8").replace("docs/performance-runs", "docs/performance-history"), encoding="utf-8") + config.write_text( + config.read_text(encoding="utf-8") + .replace("docs/performance-runs", "docs/performance-history") + .replace("docs/performance.md", report_path) + .replace("la-stack complete benchmark measurements", "Configured performance report"), + encoding="utf-8", + newline="\n", + ) stem = root / "target/bench-reports/performance" workflow.save_scratch(root, stem, first) + assert stem.with_suffix(".md").read_bytes() == render_run(first) workflow.promote(root, stem) - first_report = (root / "docs/performance.md").read_bytes() + first_report = (root / report_path).read_bytes() + assert first_report == render_run(first, title="Configured performance report") + assert b"## Mean (ns)" in first_report + assert b"## Median (ns)" in first_report + assert b"nalgebra: measured in **baseline** phase" in first_report + assert b"faer: measured in **baseline** phase" in first_report + assert not (root / policy.archive_directory(root) / "current.md").exists() + if report_path != "docs/performance.md": + assert not (root / "docs/performance.md").exists() retained = root / policy.archive_directory(root) / "runs" / run_identity(first) original = {path: path.read_bytes() for path in retained.iterdir()} # A second valid measurement with the same release pair has new provenance. @@ -177,25 +200,22 @@ def test_shared_retention_repeated_runs_and_offline_report(tmp_path: Path) -> No assert len(list((root / policy.archive_directory(root) / "runs").iterdir())) == 2 assert {path: path.read_bytes() for path in original} == original assert (retained / "run.json").is_file() - assert first_report != (root / "docs/performance.md").read_bytes() - # Same-pair reruns live in complete-run history, leaving the canonical - # release-pair archive available for a later release transition. - assert not (root / "docs/archive/performance/v0.4.6-vs-v0.4.5.md").exists() + assert first_report != (root / report_path).read_bytes() + assert not (root / "docs/archive/performance").exists() shutil.rmtree(root / "target") assert run_identity(policy.retained_run(root, stem)) == run_identity(second) - before = (root / "docs/performance.md").read_bytes() + before = (root / report_path).read_bytes() workflow.promote(root, stem) - assert (root / "docs/performance.md").read_bytes() == before + assert (root / report_path).read_bytes() == before (root / policy.archive_directory(root) / "latest.json").write_text("{}", encoding="utf-8") with pytest.raises(ValueError, match=r"latest|pointer|fields|schema"): policy.retained_run(root, stem) @pytest.mark.parametrize("destination", ["missing", "identical", "different", "directory", "symlink", "relative-symlink"]) -def test_promotion_preserves_prior_report_on_archive_collision(complete_run: tuple[Path, Evidence], destination: str) -> None: - root, evidence = complete_run +def test_legacy_promotion_preserves_prior_report_on_archive_collision(tmp_path: Path, legacy_artifacts: ArtifactPaths, destination: str) -> None: + root = tmp_path stem = root / "target/bench-reports/performance" - workflow.save_scratch(root, stem, evidence) current = root / "docs/performance.md" current.parent.mkdir(parents=True, exist_ok=True) previous = b"**la-stack** v0.4.5\n\nComparison against baseline **v0.4.4**:\n\nPrior curated report.\n" @@ -213,16 +233,15 @@ def test_promotion_preserves_prior_report_on_archive_collision(complete_run: tup original_link_target = archived.readlink() if archived.is_symlink() else None before = {path: path.read_bytes() for path in root.rglob("*") if path.is_file() and ".git" not in path.parts} if destination in {"missing", "identical"}: - workflow.promote(root, stem) + workflow.render_and_promote_artifacts(artifacts=legacy_artifacts, output=stem.with_suffix(".md"), current=current, archive_dir=archived.parent) assert archived.read_bytes() == previous - assert current.read_bytes() == policy.render_scientific_report(evidence) - assert run_identity(policy.retained_run(root, stem)) == run_identity(evidence) + assert current.read_text(encoding="utf-8") == render_release_artifacts(legacy_artifacts) else: errors = {"different": "immutable output differs", "directory": "regular file", "symlink": "symlink", "relative-symlink": "symlink"} with pytest.raises(ValueError, match=errors[destination]): - workflow.promote(root, stem) + workflow.render_and_promote_artifacts(artifacts=legacy_artifacts, output=stem.with_suffix(".md"), current=current, archive_dir=archived.parent) assert {path: path.read_bytes() for path in root.rglob("*") if path.is_file() and ".git" not in path.parts} == before - assert not (root / policy.archive_directory(root)).exists() + assert not (root / "docs/performance-runs").exists() if destination == "directory": assert archived.is_dir() elif destination in {"symlink", "relative-symlink"}: @@ -232,21 +251,22 @@ def test_promotion_preserves_prior_report_on_archive_collision(complete_run: tup @pytest.mark.parametrize("duplicate", ["**la-stack** v0.4.4", "Comparison against baseline **v0.4.3**:"]) -def test_duplicate_report_identity_prevents_publication(complete_run: tuple[Path, Evidence], duplicate: str) -> None: - root, evidence = complete_run +def test_duplicate_legacy_report_identity_prevents_publication(tmp_path: Path, legacy_artifacts: ArtifactPaths, duplicate: str) -> None: + root = tmp_path stem = root / "target/bench-reports/performance" - workflow.save_scratch(root, stem, evidence) current = root / "docs/performance.md" current.parent.mkdir(parents=True, exist_ok=True) text = f"**la-stack** v0.4.5\n\nComparison against baseline **v0.4.4**:\n\n{duplicate}\n" - current.write_text(text, encoding="utf-8") + current.write_text(text, encoding="utf-8", newline="\n") before = {path: path.read_bytes() for path in root.rglob("*") if path.is_file() and ".git" not in path.parts} with pytest.raises(ValueError, match="exactly one"): workflow.parse_report_id(text) with pytest.raises(ValueError, match="exactly one"): - workflow.promote(root, stem) + workflow.render_and_promote_artifacts( + artifacts=legacy_artifacts, output=stem.with_suffix(".md"), current=current, archive_dir=root / "docs/archive/performance" + ) assert {path: path.read_bytes() for path in root.rglob("*") if path.is_file() and ".git" not in path.parts} == before - assert not (root / policy.archive_directory(root)).exists() + assert not (root / "docs/performance-runs").exists() def test_partial_scratch_and_stale_source_cannot_select_old_history(tmp_path: Path) -> None: @@ -358,6 +378,7 @@ def test_complete_report_passes_markdown_recipes_without_rewriting_history(compl shutil.copyfile(ROOT / name, root / name) history = root / policy.archive_directory(root) originals = {path: path.read_bytes() for path in history.rglob("*") if path.is_file()} + originals[root / "docs/performance.md"] = (root / "docs/performance.md").read_bytes() for recipe in ("markdown-check", "markdown-fix", "markdown-check"): result = run_command( "just", @@ -443,6 +464,7 @@ def test_native_inventory_and_phase_commands_use_the_measured_checkout(tmp_path: f' *"--bench vs_linalg"*) printf "%s: benchmark\\n" {shlex.join(baseline_ids)} ;;\n' "esac\n", encoding="utf-8", + newline="\n", ) cargo.chmod(0o755) record = tmp_path / "commands.txt" @@ -501,7 +523,7 @@ def test_pre_just_run_remains_readable_without_its_python_driver(tmp_path: Path) sources=tuple(sources), payload=serialize_run(replace(run, compatible=tuple(field for field in run.compatible if field != "context.tool.just"))), ) - assert b"not rerun in the current phase" in policy.render_scientific_report(legacy) + policy.validate_scientific_run(legacy) bad = replace(sources[0][1], context=tuple((key, "[]" if key == "baseline-cargo-commands" else value) for key, value in sources[0][1].context)) with pytest.raises(ValueError, match="consumer policy"): policy.validate_scientific_run(replace(legacy, sources=(("baseline", bad), sources[1]))) diff --git a/tooling/performance-report.toml b/tooling/performance-report.toml index a2ca6e8..9a41123 100644 --- a/tooling/performance-report.toml +++ b/tooling/performance-report.toml @@ -1,5 +1,5 @@ schema = 2 -current = "docs/performance-runs/current.md" -# Keep the generated-archive exclusions in both Markdown recipes aligned. +current = "docs/performance.md" +# Keep current/archive exclusions in both Markdown recipes aligned. archive = "docs/performance-runs" title = "la-stack complete benchmark measurements"