AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 1, 2026, 10:00 JST

Claude API checks pages against source of truth

Claude API checks pages against source of truth

3 Key Points

  1. What happened

    A developer built a checker that feeds source-of-truth JSON and a published HTML page to the Claude API, which returns only items that disagree in meaning. A plain diff caught wording variants like "email support" vs "support by email", burying real price gaps.

  2. Why it matters

    This means the checker's output is narrowed to genuine mismatches rather than noise, so the outdated prices and missing features that went unnoticed when they relied on someone happening to spot them are more likely to surface.

  3. What to watch

    The tool was never run against the live API because no ANTHROPIC_API_KEY was available in the build environment, so its mocked test only verifies the extractor to detect_drift to report wiring. Watch whether it behaves the same once a key is set.

WHO IT HITSDocumentation and pricing-page owners, plus the engineers who wire release CI, may benefit, since the tool outputs a Markdown report and exits with code 1 when mismatches appear.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The project addresses a routine operational gap: published pages are updated by hand while the real data sits elsewhere, so outdated prices or descriptions of removed features can keep facing customers. The developer chose not to rely on a plain text diff after testing it against sample data, where wording variants such as "email support" and "support by email" were flagged as differences and buried the price gaps and missing features that actually needed attention. Extraction on the page side avoids extra dependencies by using html.parser directly rather than BeautifulSoup, and Claude is prompted to return only a JSON array of semantic conflicts.

The developer also frames the problem as a familiar one in data work: comparing a master table with its output destination, such as a BI dashboard or an export. Deciding how much tolerance to allow before something counts as a drift is comparable to setting the acceptable error threshold in an aggregation pipeline, and a threshold that is too strict picks up too much noise while one that is too loose misses genuine gaps.

The outcome hinges on the one step that could not be exercised: because no ANTHROPIC_API_KEY was available in the environment, the live API was never called, and testing relies on a mock that returns fixed drift JSON to confirm the wiring from extractor to detect_drift to report. Whether the tool behaves as intended on real pages is likely to depend on how Claude handles the prompt against actual HTML, so anyone using it in CI should treat that first live run as the real test. The exit-code-1 behaviour and empty-array default suggest the intended home is an automated pipeline rather than a manual review.

FAQ
How does it decide what counts as a real mismatch?
A system prompt tells Claude to ignore wording and formatting differences, and to report only number differences, items present on one side only, or descriptions whose meaning changed.
What does it output when it finds a problem?
It emits a Markdown drift report as a JSON array with plan_id, field, source_value, published_value, severity and reason. If mismatches exist, it exits with code 1, staying silent with an empty array when they don't.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle unveils Gemini 4 Argon, limits early access