HTTP 200 Is a Lie: A 30-Line Schema Canary for Source Drift
The article discusses the issue of silent source drift in web scraping, where a scraper may return an HTTP 200 status but still produce incorrect data. It emphasizes the importance of monitoring not just the success of the request but also the integrity of the data being scraped. A proposed solution is to implement a contract that defines the expected shape of the data to catch discrepancies early in the process.
- ▪HTTP 200 indicates that the server responded, but does not guarantee the data is correct.
- ▪Silent source drift can lead to incorrect data being fed into a database without any error signals.
- ▪Implementing a contract for data shape can help identify issues with scraped data.
DEV.to (Top) files mainly under programming. We currently carry 4,924 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | DEV.to (Top) |
| Canonical URL | https://dev.to/0012303/http-200-is-a-lie-a-30-line-schema-canary-for-source-drift-4m74 |
| Publication time | Sat, 30 May 2026 12:12:29 +0000 |
| Retrieval time | 2026-05-30T12:29:36.408Z |
| Last seen | 2026-05-30T12:29:36.408Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | IfLFWcmJFDVN |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
try { if(localStorage) { let currentUser = localStorage.getItem('current_user'); if (currentUser) { currentUser = JSON.parse(currentUser); if (currentUser.id === 3831260) { document.getElementById('article-show-container').classList.add('current-user-is-article-author'); } } } } catch (e) { console.error(e); } Alex Spinov Posted on May 30 • Originally published at blog.spinov.online HTTP 200 Is a Lie: A 30-Line Schema Canary for Source Drift #api #dataengineering #python #webscraping A scraper that returns HTTP 200 is not a scraper that returns good data. Those are two different claims, and almost every monitoring setup I've seen conflates them. Here's the failure mode nobody writes code for. The source you scrape quietly changes.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at DEV.to (Top).