Relationship to existing work
#1750 reports that get_code_snippet and search_code slice the current file with indexed line coordinates (symptom 1b below, which still reproduces at ec4fbd3a). Open PRs #1788 and #2216 address it by checking file metadata before using stored coordinates. #1461 proposes real content hashing; its first item appears done, because file_hashes.sha256 holds content hashes at ec4fbd3a.
This issue is broader. Fixing stale snippets alone does not make results trustworthy. Both PRs decide staleness from metadata: #1788 reuses the mtime and size comparison behind check_index_coverage. Symptom 1c is a content change that passes that comparison, so a metadata-gated snippet would still return wrong text for it. Separately, symptoms 1a and 1d show that a client cannot tell which index publication produced a result, or trust freshness after a re-index. We found no existing issue for 1a or 1d.
Symptoms
1a. metadata.generation does not change on incremental re-index.
generation after full index: 2026-10-10T22:58:22Z
generation after edit 1 + index: 2026-10-10T22:58:22Z
generation after edit 2 + index: 2026-10-10T22:58:22Z
recorded_at advances, but only at one-second resolution, so it cannot distinguish publications either.
1b. get_code_snippet applies indexed line numbers to current disk bytes (#1750). After inserting three lines at the top of the file, without re-indexing, the snippet for balance returns lines from post_entry:
indexed snippet of balance: (6, 7, 'def balance(ledger):\n return sum(ledger)\n')
after inserting 3 lines at top, no re-index; freshness=metadata_changed
snippet of balance now: (6, 7, ' return balance(ledger)\n\n')
1c. Freshness misses a same-size content change that keeps the original mtime.
content changed: 1 1 src/ledger.py (same size, original mtime)
freshness: metadata_match
Freshness compares only mtime and size (coverage_path_freshness in src/mcp/mcp.c), although file_hashes already stores the indexed SHA-256. mtime is not a reliable change signal: tools that preserve timestamps (cp -p, rsync -t, archive extraction) or coarse-resolution filesystems can produce this case without anyone intending to.
1d. Identical bytes with a new mtime stay metadata_changed after re-indexing.
git status after restore: ''
freshness after re-index: metadata_changed
freshness after second re-index: metadata_changed
Unchanged content appears to skip the file_hashes mtime and size refresh, so re-indexing cannot clear the warning.
Correctness objective
A client combining query_graph, search_graph, get_code_snippet, search_code, and check_index_coverage should be able to establish, from tool responses alone, that:
- every result came from one identified index publication, and the identity changes whenever the published graph changes;
- source text returned for a symbol is the text that was indexed, or the response says it is not; and
- freshness reflects file content, so it neither passes changed content (1c) nor stays stale after unchanged content is re-indexed (1d).
Possible incremental steps
These are suggestions; naming and shape are yours.
- Return a publication identity on tool responses that changes on every incremental re-index.
db_uid and mutation_gen already exist in store_meta.
- Optionally, a digest over the indexed content hashes, indexing configuration, and build, so two publications with the same content can be recognized as equivalent.
- Serve snippets from indexed content, or verify the file's current SHA-256 against
file_hashes.sha256 before slicing, and return the indexed hash. This closes 1b without leaving 1c open.
- Compare content hashes when mtime or size is ambiguous, or expose the indexed
sha256 so clients can.
- Refresh the mtime and size record when unchanged content is re-indexed.
Environment: built from main at ec4fbd3a (2026-10-10), Linux x86-64, CLI mode, fresh cache. Upstream main at 76a24859 adds no changes that appear to touch these code paths.
Reproduction
CBM="codebase-memory-mcp cli" ./repro.sh 1a 1b 1c 1d (needs git and python3; each case uses a fresh throwaway repository). Each check prints REPRODUCED or NOT REPRODUCED. Exit status is 0 when all reproduce, 1 when any does not, and 2 on a command failure.
repro.sh
#!/bin/bash
# Standalone reproductions for codebase-memory-mcp issue reports.
# Usage: CBM="codebase-memory-mcp cli" ./repro.sh [1a|1b|1c|1d|2|3|4|5 ...]
# Needs: git, python3. Each case uses a fresh throwaway repository.
#
# Every case asserts the defective outcome it reports:
# REPRODUCED the defect is present (expected on the reported build)
# NOT REPRODUCED the observed value differs (fixed, or environment differs)
# Exit status: 0 all reproduced, 1 any not reproduced, 2 a command failed.
set -eEuo pipefail
shopt -s inherit_errexit
CBM=${CBM:-"codebase-memory-mcp cli"}
WORK=${WORK:-$(mktemp -d)}
REPO="" PROJ="" MISSED=0
# A failure anywhere, including inside $(...), leaves a marker so the case
# reports a command error instead of a misleading NOT REPRODUCED.
die() { echo "ERROR: $*" >&2; : > "$WORK/failed"; exit 2; }
trap 'die "command failed at line $LINENO: $BASH_COMMAND"' ERR
# Run one CLI tool; a nonzero exit or non-JSON output stops the script.
# index_repository always answers in JSON and takes no --format flag.
cbm() {
local out format=(--format json)
[ "$1" = index_repository ] && format=()
out=$($CBM "$@" "${format[@]}" 2>"$WORK/stderr") || { cat "$WORK/stderr" >&2; die "$CBM $1 exited nonzero"; }
python3 -c 'import json,sys; json.loads(sys.argv[1])' "$out" 2>/dev/null || die "$CBM $1 returned non-JSON: $out"
printf '%s' "$out"
}
py() { python3 -c "import json,sys; d=json.loads(sys.argv[1]); print($2)" "$1"; }
redact() { sed -e "s|$REPO|<repo>|g" -e "s|$PROJ|<project>|g"; }
expect() { # expect <claim> <actual> <value-that-shows-the-defect>
if [ "$2" = "$3" ]; then echo "REPRODUCED: $1"
else echo "NOT REPRODUCED: $1"; echo " expected: $3"; echo " actual: $2"; MISSED=1; fi
}
new_repo() {
REPO=$(mktemp -d "$WORK/repo.XXXX")
cd "$REPO"
git init -q && git config user.email t@t && git config user.name t
mkdir -p src/util docs
printf 'def post_entry(ledger, amount):\n ledger.append(amount)\n return balance(ledger)\n\n\ndef balance(ledger):\n return sum(ledger)\n' > src/ledger.py
printf 'def clamp(x, lo, hi):\n return max(lo, min(hi, x))\n' > src/util/rounding.py
}
commit() { git add -A && git commit -qm "$1"; }
index() { local out; out=$(cbm index_repository --repo-path "$REPO"); PROJ=$(py "$out" 'd["project"]'); }
coverage() { cbm check_index_coverage --project "$PROJ" --paths "[\"$1\"]"; }
fresh() { local out; out=$(coverage "$1"); py "$out" 'd["paths"][0]["freshness"]'; }
generation() { local out; out=$(coverage src/ledger.py); py "$out" 'd["metadata"]["generation"]'; }
cypher() { local out; out=$(cbm query_graph --project "$PROJ" --query "$1"); py "$out" 'd["rows"]' | redact; }
snippet() { local out; out=$(cbm get_code_snippet --project "$PROJ" --qualified-name "$1" --source-mode full)
py "$out" 'repr((d["start_line"], d["end_line"], d["source"]))'; }
doc_links() { local out; out=$(cbm index_status --project "$PROJ" --diagnostics full); py "$out" 'repr((d["doc_links"]["mentions"], d["doc_links"]["unresolved"]))'; }
case_1a() { # exposed generation does not change on incremental re-index
new_repo; commit init; index
local first; first=$(generation); echo "generation after full index: $first"
for i in 1 2; do
sleep 1.1; echo "# edit $i" >> src/ledger.py; index
local now; now=$(generation); echo "generation after edit $i + index: $now"
expect "generation unchanged after incremental re-index $i" "$now" "$first"
done
}
case_1b() { # snippet mixes indexed line range with current disk bytes
new_repo; commit init; index
local before after; before=$(snippet balance); echo "indexed snippet of balance: $before"
printf '# a\n# b\n# c\n' | cat - src/ledger.py > "$WORK/shifted" && mv "$WORK/shifted" src/ledger.py
local state; state=$(fresh src/ledger.py); echo "after inserting 3 lines at top, no re-index; freshness=$state"
after=$(snippet balance); echo "snippet of balance now: $after"
expect "snippet for balance returns another function's lines" "$after" "(6, 7, ' return balance(ledger)\\n\\n')"
}
case_1c() { # same-size edit with restored mtime passes freshness
new_repo; commit init; index
touch -r src/ledger.py "$WORK/mtime-ref"
sed -i 's/return sum(ledger)/return max(ledger)/' src/ledger.py
touch -r "$WORK/mtime-ref" src/ledger.py
echo "content changed: $(git diff --numstat -- src/ledger.py) (same size, original mtime)"
expect "freshness reports a content change as current" "$(fresh src/ledger.py)" "metadata_match"
}
case_1d() { # identical bytes with new mtime stay metadata_changed after re-index
new_repo; commit init; index
echo "# tmp" >> src/ledger.py; git checkout -q -- src/ledger.py # identical bytes, new mtime
echo "git status after restore: '$(git status --short)'"
index; expect "freshness stale after re-index of identical bytes" "$(fresh src/ledger.py)" "metadata_changed"
index; expect "freshness stale after second re-index" "$(fresh src/ledger.py)" "metadata_changed"
}
case_2() { # repeated headings collapse; references misattributed
new_repo
printf '# Service A\n\n## Implementation\n\nUses `src/ledger.py::balance`.\n\n# Service B\n\n## Implementation\n\nUses `src/ledger.py::post_entry`.\n' > docs/repeated.md
commit init; index
local sections sources
sections=$(cypher 'MATCH (s:Section) WHERE s.file_path = "docs/repeated.md" AND s.name = "Implementation" RETURN s.qualified_name, s.start_line')
sources=$(cypher 'MATCH (s)-[:MENTIONS]->(c) WHERE s.file_path = "docs/repeated.md" RETURN s.label, c.name ORDER BY c.name')
echo "Implementation sections: $sections"; echo "MENTIONS sources: $sources"
expect "two Implementation headings yield one Section" "$sections" "[['<project>.docs.repeated.Implementation', '9']]"
expect "references under them are attributed to the File" "$sources" "[['File', 'balance'], ['File', 'post_entry']]"
}
case_3() { # missing member silently becomes a file edge
new_repo
printf '# Notes\n\nSee `src/util/rounding.py::no_such_function`.\n' > docs/notes.md
commit init; index
local edges; edges=$(cypher 'MATCH (s)-[r:MENTIONS]->(c) WHERE s.file_path = "docs/notes.md" RETURN c.label, c.name, r.tier')
echo "MENTIONS: $edges"; echo "doc_links (mentions, unresolved): $(doc_links)"
expect "missing member resolves as an exact File edge" "$edges" "[['File', 'rounding.py', 'exact']]"
expect "no unresolved row records the missing member" "$(doc_links)" "(1, {})"
}
case_4() { # references into a moved directory vanish without unresolved rows
new_repo
printf '# Notes\n\nSee `src/util/rounding.py::clamp`.\n' > docs/notes.md
commit init; index
echo "before move: $(cypher 'MATCH (s)-[:MENTIONS]->(c) RETURN c.name')"
git mv src/util src/helpers; index
local edges; edges=$(cypher 'MATCH (s)-[:MENTIONS]->(c) RETURN c.name')
echo "after git mv src/util src/helpers and re-index: MENTIONS=$edges doc_links=$(doc_links)"
expect "reference into moved directory disappears" "$edges" "[]"
expect "no unresolved row records it" "$(doc_links)" "(0, {})"
}
case_5() { # enhancement context: a doc-only edit yields no detect_changes seeds
new_repo
printf '# Notes\n\n## Posting\n\nPosting uses `src/ledger.py::post_entry`.\n' > docs/notes.md
commit init; index
echo "MENTIONS from the section: $(cypher 'MATCH (s:Section)-[:MENTIONS]->(c) RETURN s.name, c.name')"
sed -i 's/Posting uses/Posting always uses/' docs/notes.md; index
local out; out=$(cbm detect_changes --project "$PROJ" --edge-types '["CALLS","MENTIONS"]')
echo "doc-only edit: $(py "$out" 'repr((d["changed_files"], d["seed_symbols"]))')"
expect "doc-only edit produces zero seeds" "$(py "$out" 'd["seed_symbols"]')" "0"
}
trap - ERR
for c in "${@:-1a 1b 1c 1d 2 3 4 5}"; do for one in $c; do
echo "=== case $one"
rm -f "$WORK/failed"
set +e
( set -e; trap 'die "command failed at line $LINENO: $BASH_COMMAND"' ERR; "case_$one"; exit "$MISSED" )
status=$?
set -e
if [ -e "$WORK/failed" ] || [ "$status" -gt 1 ]; then echo "case $one: COMMAND FAILED" >&2; exit 2; fi
[ "$status" -eq 1 ] && MISSED=1
done; done
[ "$MISSED" -eq 0 ] && echo "=== all cases reproduced" || echo "=== some cases did not reproduce"
exit "$MISSED"
Relationship to existing work
#1750 reports that
get_code_snippetandsearch_codeslice the current file with indexed line coordinates (symptom 1b below, which still reproduces atec4fbd3a). Open PRs #1788 and #2216 address it by checking file metadata before using stored coordinates. #1461 proposes real content hashing; its first item appears done, becausefile_hashes.sha256holds content hashes atec4fbd3a.This issue is broader. Fixing stale snippets alone does not make results trustworthy. Both PRs decide staleness from metadata: #1788 reuses the mtime and size comparison behind
check_index_coverage. Symptom 1c is a content change that passes that comparison, so a metadata-gated snippet would still return wrong text for it. Separately, symptoms 1a and 1d show that a client cannot tell which index publication produced a result, or trust freshness after a re-index. We found no existing issue for 1a or 1d.Symptoms
1a.
metadata.generationdoes not change on incremental re-index.recorded_atadvances, but only at one-second resolution, so it cannot distinguish publications either.1b.
get_code_snippetapplies indexed line numbers to current disk bytes (#1750). After inserting three lines at the top of the file, without re-indexing, the snippet forbalancereturns lines frompost_entry:1c. Freshness misses a same-size content change that keeps the original mtime.
Freshness compares only mtime and size (
coverage_path_freshnessinsrc/mcp/mcp.c), althoughfile_hashesalready stores the indexed SHA-256. mtime is not a reliable change signal: tools that preserve timestamps (cp -p,rsync -t, archive extraction) or coarse-resolution filesystems can produce this case without anyone intending to.1d. Identical bytes with a new mtime stay
metadata_changedafter re-indexing.Unchanged content appears to skip the
file_hashesmtime and size refresh, so re-indexing cannot clear the warning.Correctness objective
A client combining
query_graph,search_graph,get_code_snippet,search_code, andcheck_index_coverageshould be able to establish, from tool responses alone, that:Possible incremental steps
These are suggestions; naming and shape are yours.
db_uidandmutation_genalready exist instore_meta.file_hashes.sha256before slicing, and return the indexed hash. This closes 1b without leaving 1c open.sha256so clients can.Environment: built from
mainatec4fbd3a(2026-10-10), Linux x86-64, CLI mode, fresh cache. Upstreammainat76a24859adds no changes that appear to touch these code paths.Reproduction
CBM="codebase-memory-mcp cli" ./repro.sh 1a 1b 1c 1d(needs git and python3; each case uses a fresh throwaway repository). Each check printsREPRODUCEDorNOT REPRODUCED. Exit status is 0 when all reproduce, 1 when any does not, and 2 on a command failure.repro.sh