Skip to content

Windows: index_repository fails with pipeline.stage lock_failed errno=22 on w64devkit #2584

Description

@tmonestudio

Summary

On Windows 11, a clean production build of origin/main fails index_repository before source discovery. The supervised worker exits cleanly, but the parent returns the generic Pipeline failed envelope. The worker log contains:

pipeline.stage action=lock_failed errno=22 path=...db.stage.<suffix>

The same checkout, arguments, repository path, and fresh cache/runtime succeed with the previously served binary built before the staging-lock change.

Reproduction

Environment:

  • Windows 11
  • w64devkit 2.8.0 / MinGW GCC 16.1.0
  • zlib 1.3.2
  • commit: 95c91b8d8d0fc11a01b9109b43d0eea94393f538 (current origin/main)
  • clean production binary SHA256: A113106F4DAB34D802241631970BF6D1ADB3765145E116F0696375BC75F9FD6D

Use fresh, owner-private cache and runtime directories and any local repository checkout:

$env:CBM_CACHE_DIR = 'C:\tmp\cbm-cache'
$env:CBM_RUNTIME_DIR = 'C:\tmp\cbm-runtime'
.\codebase-memory-mcp.exe cli --progress --verbose --json index_repository `
  --repo-path 'C:\src\codebase-memory-mcp' `
  --mode fast --name 'origin-main-canary' --persistence false

Expected: status=indexed.

Actual: exit code 1, isError=true, generic Pipeline failed; the worker log reports pipeline.stage action=lock_failed errno=22. The staging path code retries all generated suffixes and then gives up.

Control

The older served binary, built before the staging-lock change, indexes the exact same checkout with the exact same arguments and fresh cache/runtime:

  • SHA256: 3F935E7FEBB8AB2C4D1A9589CDD479F50B7B501A6505507CF5B5D5C09531B351
  • exit code 0
  • status=indexed
  • 22,387 nodes and 125,496 edges
  • elapsed time: approximately 8.3 seconds

This makes the failure specific to the newer runtime rather than the CLI invocation, repository contents, or cache reuse.

Source correlation

create_staging_path() in src/pipeline/pipeline.c now acquires the sidecar lock before creating the stage file. On Windows, cbm_lockfile_open() in src/foundation/compat_fs.c calls:

_wsopen(wpath,
        _O_RDWR | _O_BINARY | _O_NOINHERIT | _O_CREAT,
        _SH_DENYRW,
        _S_IREAD | _S_IWRITE);

With the supported w64devkit toolchain, this call returns -1 with errno=22 for each generated staging sidecar. Consequently, all attempts fail before the stage file becomes visible, and the worker returns the generic pipeline error.

Relevant history:

  • 0c4831a3 (fix(pipeline): own every staging file through an exclusive sidecar lock) introduced the Windows sidecar-lock primitive.
  • 91e31211 (fix(pipeline): take the stage lock before the stage file is visible) changed the ordering to lock-first. Its commit message mentions Windows red runs with errno=13; this environment reports errno=22.

The graph/index was refreshed against the current checkout before this report; the failure reproduces through the actual MCP index_repository surface as well as the CLI.

Request

Could you confirm the intended Windows lock primitive and flags for the supported w64devkit toolchain, and add a focused Windows regression test for cbm_lockfile_open() plus the index_repository staging path? It would also help if the user-facing error exposed the stage-lock errno instead of only returning Pipeline failed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    windowsWindows-specific issues

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions