Hunting orphan objects: 45% off our ClickHouse storage bill and a near data-loss
Tinybird significantly reduced its cloud storage costs by 45% after cleaning up orphaned S3 objects. The company faced challenges with data loss during the cleanup process, but successfully recovered all data. This experience highlighted the importance of improving operational safety and recovery procedures.
- ▪Tinybird cleaned up orphaned S3 objects, leading to a 45% reduction in storage costs.
- ▪The cleanup process nearly resulted in the loss of legitimate data, which was ultimately recovered.
- ▪The company improved its garbage identification tooling to better manage orphaned objects.
Hacker News (Newest) files mainly under programming. We currently carry 5,306 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | Tinybird |
| Canonical URL | https://www.tinybird.co/blog/how-we-deal-with-cloud-orphan-objects |
| Publication time | Tue, 19 May 2026 09:58:43 +0000 |
| Retrieval time | 2026-05-19T10:04:57.575Z |
| Last seen | 2026-05-19T10:04:57.575Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | dkE0bnudW3KA |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
We were paying for petabytes of S3 objects nothing was reading. Last month we cleaned them up and our object storage bill dropped ~45%. We also almost lost real data along the way. Here's what happened.How do we deal with cloud orphan objects?At Tinybird, we run large-scale ClickHouse® clusters backed by object storage. Like many teams operating distributed storage systems at scale, we’ve spent a lot of time thinking about replication, consistency, and failure recovery.One issue that kept growing in the background was cloud storage garbage: objects that were no longer being used, but also never deleted.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Tinybird.