TLDR
A Durable Object alarm on Cloudflare rescheduled itself forever, ran up 6 trillion reads and writes, and produced a $10,811.41 bill that was paid in full. It enables any developer who pastes an alarm snippet into a worker without thinking about termination to learn the same lesson on their own card. The difference between this loop and a normal bug is that the retry machinery caps thrown errors at 6 attempts and does nothing at all about a handler that legitimately schedules the next alarm on every wake, and the billing side has usage alerts but no hard spending cap.
Caption: a thrown error gets 6 retries and then stops; a handler that calls setAlarm() before returning is scheduled forever, and every pass bills a request, duration, and storage rows.
What the bill is made of
Cloudflare publishes every unit the loop consumed, so the invoice decomposes without any inside information. Rows read cost $0.001 per million past the first 25 billion included each month. Six trillion reads is therefore $5,975. The rest of the $10,811.41, about $4,836, is the alarm’s compute duration at $12.50 per million GB-s, the requests each wake counts at $0.15 per million, and rows written at $1.00 per million. The arithmetic also bounds what the loop was not: six trillion rows written would have billed about six million dollars, so the 6 trillion figure from the victim’s post is read-dominant, an alarm that woke and scanned far more than it wrote.
Two numbers bracket the story. A pure read loop with instant wakes and no duration would have cost about $6,000. A write loop of the same magnitude would have cost about $6,000,000. The real bill sits slightly above the read floor at the price of days of machine wake time.
Caption: the same 6 trillion operations priced as reads, as the observed invoice, and as writes; the loop’s cost is set by which storage operation runs per pass and how long each wake keeps the object alive.
The mechanism: six retries versus forever
The Alarms API documents the boundary precisely. An alarm handler that throws is retried automatically with exponential backoff starting two seconds after the failure, at most six retries. That machinery contains a broken alarm to seven total runs and then the object goes quiet. Nothing in the same document caps a handler that exits cleanly after calling setAlarm(), because self-rescheduling is a legitimate and encouraged pattern: it is how you build queues, schedulers, and batchers on the platform. Each reschedule is a fresh alarm, each alarm invocation is a billable request, each wake bills duration for as long as the object stays active, and every setAlarm() writes one row. The loop in the victim’s project was the second kind: no error thrown, no retry cap consumed, six trillion storage operations accumulated until the invoice landed.
The billing side has exactly one guardrail and it fires after the fact. A Usage Based Billing notification can email an account holder when usage crosses a threshold, on pay-as-you-go accounts, from the Professional plan up. It is a smoke detector, not a circuit breaker; nothing in the Workers or Durable Objects billing model lets a customer set a number that turns the meters off. Community threads going back years ask for the hard cap and the answer has not changed.
The support epilogue
The follow-up thread records the support path that followed the invoice. The victim asked for a reduction through a support ticket and got bot replies in a loop, then reached Cloudflare staff on X and other channels and was told they could only take a look. The bill was paid in full on October 8 and the poster’s summary calls it the second most expensive lesson of their life, after which every project moves to self-hosted VPS. The vendor has no hard spending cap and left this customer with one support path, a bot, that replied in circles while the invoice came due. The bot itself ran on a competitor’s model, which the poster flagged with the driest line in the thread: you have the budget, why not use Claude.
The docket pattern
Andras Bacsai, who builds Coolify and collects these stories at serverlesshorrors.com, now has three Durable Object loop stories in six months: twin objects looping for $8,846.78 in August, a queue loop with unbatched DO writes reaching 16 billion operations for $36,000 in May, and this alarm for $10,811.41 in October. The same site holds a $46,485.99 Vercel bandwidth bill, a $100,000 single-day Firebase invoice after a DoS, and an AWS case where pausing a service still generated charges. The recurring shape is one bug, at usage prices, with meters that only face one direction, on infrastructure where the developer never sees a burn rate.
What to do about it
The guardrails fit in an afternoon:
- Termination condition inside the handler: the alarm body checks a deadline, a retry counter in storage, or a “work actually pending” flag, and returns without rescheduling when work is done. The vibe coded loop had no such check.
- A kill switch outside the loop: alarm schedules live in data the application controls, so a status flag checked before setAlarm() gives you a way to stop the loop without fighting the scheduler.
- The usage notification, set low: Cloudflare’s threshold alert on Workers usage is the tool that exists; a $50 threshold emails you on day one of a runaway instead of at invoice time.
- Load the meters locally: in staging, wrap the alarm in a budget that fails the test when storage operations exceed sane bounds for one handler.
A self-hosted server holding the same bug in a Redis-backed job queue reboots into a loop that costs a $40 VPS some CPU. The same logic on usage-billed serverless reads its price off a meter. That asymmetry, more than any single bug, is what the $10,811.41 invoice documents, and it is the pitch the person amplifying the story has been making with his own infrastructure product since before this loop existed.
Sources: original thread - follow-up - Andras Bacsai - ServerlessHorrors story page - Cloudflare DO pricing - Cloudflare Alarms API - Cloudflare notifications
Related on this site: Opus 5 Ultracode database wipe postmortem - Self-hosting AI: DGX Spark vs RTX vs Mac - Cloudflare Wallets and x402 agent payments - Local AI price tiers
Discussion
Be the first to commentStart a discussion
Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.