AIToday
Large Language ModelsHacker NewsPublished: Oct 6, 2026, 10:01 JST

schema-guard: AI SQL names checked vs real snapshot

schema-guard: AI SQL names checked vs real snapshot

3 Key Points

  1. What happened

    schema-guard, a tool by developer idk-arsh, checks AI-written SQL against a saved snapshot of real table and column names. In tests with Claude Haiku 4.5 and Sonnet 5, 0 of 24 files worked without it; 48 of 48 ran with it.

  2. Why it matters

    The snapshot stopped both models from inventing names that do not exist. Every one of the 48 files ran, so the wrong answers that remained were logic errors, not wrong names.

  3. What to watch

    The results come from a small, synthetic world with only 3 runs per cell, so the numbers hinge on a setup the author designed. Watch whether the snapshot readers for Snowflake and BigQuery get run against live accounts.

WHO IT HITSData and analytics engineers who let AI coding agents write SQL against a warehouse will see fewer failed runs and fewer broken dashboards from wrong column names. The team still needs one person with warehouse access to take the snapshot.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The problem schema-guard targets is not that coding agents write bad SQL syntax, but that they write against a schema they assume is real. As the author puts it, an agent bases its query on a README, an old query, or a naming convention, and the failure shows up later in CI, in a dashboard, or at 2am. Snowflake's own developer blog ran a post on this in September 2026, which suggests the vendor sees the same pattern among its users.

The tool's approach is to keep a snapshot of the actual tables and columns, names and types only, inside the repo, and to check the agent's SQL against that snapshot before it runs. The snapshot is taken once by someone with warehouse access, so the agent itself never needs that access. The author reports a Databricks reader run live on a Free Edition workspace on 2026-10-04, returning 9 tables and 277 columns in under 30 seconds on a cold warehouse.

The author is openly cautious about the evidence. The test world is small and synthetic, the stale README was designed in, and there are only 3 runs per cell; two grader references were even added after reading runs. A one-line rule in a project file gets a similar result if the agent follows it, while the hook does not depend on that. Whether this holds in a real warehouse with years of schema drift is the open question, and the readers for Snowflake and BigQuery have not yet been run against live accounts.

FAQ
How much does schema-guard cost to run?
The author reports the hook version cost about the same as not using it: $0.65 vs $0.58 for 12 Haiku runs, and $1.61 vs $1.61 for Sonnet.
Does schema-guard block valid SQL?
The author measured zero false blocks on 1,034 valid Spider queries and 960 defog queries. Its only 2 misses were inside correlated subqueries, where it deliberately gives the benefit of the doubt.
What does schema-guard not check?
It checks names, not meaning, so a query like SUM(gross_amount) when you wanted net_amount still passes. Dynamic SQL built from string pieces in application code is also not seen.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleHACH developer: split AI agent log retention, not one store