Skip to content

Upgrades, bugfixes and improvements - #76

Open
rnbrady wants to merge 47 commits into
bitauth:masterfrom
ParyonUSD:master
Open

rnbrady wants to merge 47 commits into
bitauth:masterfrom
ParyonUSD:master

Conversation

@rnbrady

@rnbrady rnbrady commented May 10, 2026 •

Copy link
Copy Markdown

Closes #71
Closes #72
Closes #73
Closes #74
Closes #75
Closes #55
Closes #77
Closes #78
Closes #79

Introduces new columns and indexes for performance.

Leaving these here for others who may find them helpful, courtesy of ParyonUSD.

Code by various Codex (5.5 - 6.1) and Claude (Fable 5, Opus 5.5) models. See commit messages for attribution.

rnbrady added 27 commits May 7, 2026 13:22
rnbrady and others added 17 commits June 12, 2026 15:11
The default RollingUpdate strategy runs old and new agent pods
concurrently until the new pod passes readiness (minutes of node
init). The agent is a singleton writer — inserts are idempotent but
reorg handling assumes exclusive access to node_block, so an
overlapping agent can transiently resurrect stale-chain acceptance
rows. Recreate guarantees at most one agent at the cost of a brief
indexing gap per rollout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
3.5 years of CVE and performance fixes over v2.16.1, staying on the v2
line. First startup against an existing database performs a one-way
metadata catalog upgrade (47 -> 48); v2.16.1 will not start against
the upgraded catalog, so snapshot hdb_catalog before deploying.
Rehearsed against a schema-only copy of the production database
(including its catalog-47 hdb_catalog): migration completes, metadata
is consistent, and existing query shapes execute unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Restore hashes with encode(hash, 'hex') in Postgres instead of
materializing ~1.5M Buffers and hex-encoding each on the agent event
loop, which dominated startup on large databases. Log the restore
count and duration on completion. Replace the misleading "database
configuration or connectivity problem" warning — which fired after
just 6 seconds of a legitimately slow restore — with an accurate
"still restoring" message gated to 15s. e2e: verify the SQL-side
encoding matches the previous client-side conversion exactly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The "misconfigured or unresponsive" warning fired for nodes whose P2P
connection was ready but whose chain state was still being restored
from the database — legitimate work proportional to chain length
(~9s for a 960k-block chain sharing the pool with the block hash
restore). Split the warning: nodes without a ready P2P connection keep
the misconfigured/unresponsive text; connected nodes get an accurate
"still restoring chain state" message. Log each node's registration +
chain restore duration on completion, and let the block hash restore
warning use the same 5-second base patience as node warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Node 18 has been EOL since April 2025. The e2e suite already runs
under Node 24 locally (89 tests passing), and the agent boots cleanly
in the node:24-alpine image. engines raised to >=22.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Same-major driver bump carrying 3.5 years of fixes; e2e suite (89
tests, real sync against Postgres 18) passes unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
/health-check answers ~1s after start, but both probes waited 30s
before the first check, so pods showed NotReady for 30+ seconds.
Readiness now 5s delay / 10s period / 5s timeout; liveness stays
coarse so event-loop stalls during catch-up don't get pods killed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(sql): maintain per-node output membership

* fix(sql): inline output membership roots

* fix(sql): make readiness guard migration idempotent

* fix(sql): allow membership validation reruns
…actly

search_output compares each input, unchanged, with the first 25 bytes of
every locking bytecode, so inputs longer than 25 bytes (all P2SH32 outputs)
never match, and 25-byte inputs also match longer scripts sharing that
prefix. Add a failing e2e test over synthetic outputs (inserted in a
rolled-back transaction).

Refs bitauth#78

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Truncate each input to 25 bytes for the indexed comparison against
output_search_index, then recheck full equality, so P2SH32 and other scripts
longer than 25 bytes are found and 25-byte inputs no longer match longer
scripts sharing that prefix. Rewrite as an inlinable SQL function so
Hasura's limit and filters apply inside the query.

Fixes bitauth#78

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
search_output_prefix matches with LIKE, so bytes 0x5c (escape), 0x25 (%)
and 0x5f (_) in the decoded prefix act as pattern syntax: a prefix
containing 0x5c returns nothing and 0x25/0x5f over-match. Prefixes longer
than 25 bytes also never match. Add a failing e2e test.

Refs bitauth#79

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replace LIKE with a byte range on the 25-byte expression indexed by
output_search_index plus an exact prefix recheck, so bytes 0x5c, 0x25 and
0x5f are matched literally and prefixes longer than 25 bytes work. The
index is still used.

Fixes bitauth#79

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
substring(locking_bytecode, 0, 26) returns the first 25 bytes because
PostgreSQL positions start at 1, but reads as 26 bytes. Require the managed
prefix indexes to use substring(locking_bytecode from 1 for 25) and both
search functions to be served by them (with sequential scans disabled, a
mismatched expression shows up as a Seq Scan). Also expect the renamed
managed indexes: output_acceptance_index, unspent_output_index,
unspent_output_search_index and unspent_output_category_index (now covering
all unspent token outputs).

Refs bitauth#78

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Define output_search_index and unspent_output_search_index on
substring(locking_bytecode from 1 for 25) and switch search_output and
search_output_prefix to the same expression (migration 1791100002000).
Rename the managed membership indexes (output_acceptance_index,
unspent_output_index, unspent_output_search_index) and generalise
unspent_output_category_index to all unspent token outputs.

PostgreSQL only uses an expression index for the identical expression, so
existing deployments must drop/rename the old indexes, upgrade the agent
(which builds the new ones), and only then upgrade Hasura. See README
"Upgrading to the 1-based locking bytecode prefix indexes".

Fixes bitauth#78

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
refresh() was pinned to force_generic_plan, which estimates
`= ANY(transaction_internal_ids)` at 10 elements. Header catch-up inserts
node_block rows for thousands of blocks at once (~12k transactions,
~120k affected outputs in the e2e chain); the generic plan estimated
spent_state at 1 row and nested-looped over it per output (3.6e9 rows
removed by join filter, ~138s), while the next catch-up batch waited on
the membership advisory lock. This made "[e2e] catches up a new node via
headers" time out. With force_custom_plan the same refresh takes ~2s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment