fix(mssql,azuresql): keep OpenMetadata's own queries out of usage and lineage - #31163
Conversation
… lineage SQL Server reported the connector's own reflection SQL back as user queries. Two independent defects had to be fixed together; either one alone leaves the behaviour unchanged. sys.query_store_query_text stores one row per statement, and a leading comment belongs to the batch rather than to any statement, so the header was discarded before it could ever be matched. Sending it after the first token keeps it inside the statement. This is the same reason Vertica already has an override, though there the history table drops the comment outright. The plan-cache path did keep the header, but sp_executesql puts the parameter declarations in front of it, so the position-0 anchor in the four usage, lineage and stored-procedure filters missed every parameterised statement. The patterns now match the header anywhere in the text. Azure SQL reuses MssqlUsageSource and so shares those queries, but its schema never declared supportsQueryComment, which left create_generic_db_connection skipping header injection for it entirely. Declaring the field turns tagging on. Verified against SQL Server 2022 on both pytds and pyodbc: the header now survives into Query Store and the plan cache for bare, dedented, CTE and parameterised statements, and none of them reach the parser.
❌ PR checklist incompleteThis PR cannot be merged until the following are addressed on its linked issue:
The fields live on the linked issue in the Shipping project (open the issue → right sidebar → Projects). After you set them, re-run this check (or push a commit) — issue/project changes do not re-trigger it automatically. Maintainers can bypass this check by adding the |
✅ TypeScript Types Auto-UpdatedThe generated TypeScript types have been automatically updated based on JSON schema changes in this PR. |
✅ Playwright Results — workflow succeededValidated commit ✅ 4599 passed · ❌ 0 failed · 🟡 6 flaky · ⏭️ 1 skipped · 🧰 0 lifecycle flaky PerformanceBlocking targets: ✅ met · Optimization targets: 🟡 in progress Shard-job maxima below are not the full workflow wall time; the linked run includes build, fixture, planning, and reporting. 🕒 Full workflow signal wall (to summary) 1h 5m 21s ⏱️ Max setup 4m 40s · max shard execution 20m 48s · max shard-job elapsed before upload 24m 11s · reporting 19s 🌐 122.64 requests/attempt · 2.23 app boots/UI scenario · 40.87% common-shard skew Optimization targets still in progress:
🟡 6 flaky test(s) (passed on retry)
How to debug locally# Download playwright-test-results-<shard> artifact and unzip
npx playwright show-trace path/to/trace.zip # view trace |
…r anchor T-SQL nests block comments. The regex that skipped leading comments stopped at the first `*/`, so for `/* outer /* inner */ AND condition */ SELECT ...` the anchor landed on `AND`, still inside the outer comment. Query Store discards the commented preamble along with the header, and the statement is reported as user usage and lineage again -- the same failure the enclosing change fixes for the simpler prefixes. Python's `re` cannot count nesting depth, so the skip is a short scan that tracks it instead. An unterminated block comment still yields no anchor: the header carries a `*/` that would close it. Verified against SQL Server 2022 with that exact statement -- previously the usage filter handed the row to the parser, now it excludes it. Statements whose leading comments are not nested are byte-identical to before.
Only OpenMetadata's own marker moved inside the statement, so only its filter needs to match anywhere in the text. dbt still sends its comment as a batch preamble; matching it anywhere recovers no dbt rows and drops user queries that merely quote the marker. Matches every other connector and the Vertica precedent, which loosened only the OpenMetadata pattern.
This reverts commit 5ffc2f4. Anchoring the dbt filter was wrong. dbt's query comment is not always a leading comment: query-comment.append moves it after the statement, and only an unanchored pattern excludes those queries from usage and lineage. Verified against SQL Server 2022 - with the comment appended, the anchored filter hands dbt's own query to the parser as user activity. StarRocks already unanchors both markers (starrocks/queries.py:81-82), so this is not a new convention either.
_executable_start scored 22 on Cognitive Complexity against a limit of 15 (SonarQube python:S3776). Most of that was the nesting penalty on the inner depth-tracking loop rather than the logic itself, so moving that loop into _past_block_comment drops the pair to 11 and 7 with no behavioural change: old and new agree on every string up to length 6 over '/*- \nS' and on 300k random inputs.
🚦 Removed from the merge queue —
|
🚦 Removed from the merge queue —
|
Code Review ✅ Approved 1 closed / 1 findings🟡 Medium risk · MSSQL and Azure SQL query tagging and usage filtering change ingestion behavior. Fixes SQL Server and Azure SQL connectors reporting their own queries as user usage and lineage by injecting the OpenMetadata marker after the first executable SQL token while correctly skipping whitespace, line comments, and nested block comments, and by making plan-cache and Query Store filters recognize markers anywhere in recorded statement text. The inline header injected inside a leading line comment issue has been resolved. ✅ 1 closed✅ Edge Case: Inline header injected inside a leading line comment (--)
OptionsDisplay: compact → Counting what did not apply, without listing it. Comment with these commands to change the behavior for this request:
Was this helpful? React with 👍 / 👎 | Powered by Gitar — free for open source |
|
|



SQL Server reports the connector's own reflection SQL back as user queries. There are two
independent defects behind it, one per code path, and fixing either alone changes nothing.
Query Store never sees the header
sys.query_store_query_textstores one row per statement, and a leading comment belongs tothe batch rather than to any statement, so SQL Server discards it:
No pattern can match a string that was never stored. Sending the header after the first token
keeps it inside the statement. This is why Vertica already has an override — though there the
history table drops the comment outright, which is a different failure with the same remedy.
Unlike Vertica's
statement.split(" "), the SQL Server version splits on the firstnon-whitespace token: 17 of the MSSQL query constants are
textwrap.dedentstrings that openwith a newline, and for those the naive split lands the comment between two whitespace runs —
still ahead of the first keyword, still discarded, and indistinguishable from success.
The plan cache keeps the header, but not at position 0
sp_executesqlputs the parameter declarations in front of the statement:The four usage, lineage and stored-procedure filters anchored the header at position 0, so they
missed every parameterised statement while working correctly for unparameterised ones. The
patterns now match the header anywhere in the text.
Azure SQL was never tagging anything
AzuresqlUsageSourceextendsMssqlUsageSourceand shares these queries, butazureSQLConnection.jsonnever declaredsupportsQueryComment, so thehasattr(connection, "supportsQueryComment")gate increate_generic_db_connectionskippedheader injection entirely. It advertised
supportsUsageExtractionwhile being unable toidentify its own queries. Declaring the field turns tagging on.
Verification
Against SQL Server 2022 on both pytds and pyodbc, the header now survives into Query Store and
the plan cache for bare, dedented, CTE and parameterised statements, and none of them reach the
parser. The integration test asserts this through the server rather than by inspecting strings —
a header at
"\n /* ... */ SELECT"looks correct, parses fine, and is still discarded.Metadata, usage, profiler and auto-classification pipelines were run end to end against a local
SQL Server for both connectors.
Cost
The unanchored predicate roughly doubles the cost of that one query — 338 ms to 662 ms measured
over 16 089 distinct query texts with a window covering all of them. It runs once per usage run
per database. No index is affected:
query_sql_textanddm_exec_sql_text.textarenvarchar(max), which SQL Server cannot use as an index key.Not covered
The dialect's own bootstrap queries (
fn_listextendedproperty,sys.system_views, theisolation-level probe) run on the raw DBAPI connection, outside
before_cursor_execute, so noinjection strategy can reach them. Excluding those needs a system-object deny-list, which is a
separate change. Query Store rows recorded before this fix stay untagged and age out with
retention.
The PR appears safe to merge; no new actionable failures remain, and both previously reported comment-handling defects are resolved.
Summary
This PR prevents MSSQL and Azure SQL connector queries from being reported as user usage or lineage.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[OpenMetadata SQL] --> B[Skip leading whitespace and comments] B --> C[Find first executable word] C --> D[Insert OpenMetadata marker after word] D --> E[SQL Server execution] E --> F[Plan cache or Query Store] F --> G{Marker present anywhere?} G -->|Yes| H[Exclude from usage and lineage] G -->|No| I[Process as user query]Reviews (9) · Last reviewed commit: "Merge branch 'main' into fix/mssql-azure..."