My Rate Limiter Caught Me First
The first client my canvas's abuse controls caught was me. Within ten hours of turning them on, they had throttled my zooming and shut out my own stress tool.
Infinite is a collaborative canvas I built in Go. Every browser holds one WebSocket, strokes live in tiles, and you can zoom in more or less forever. The plan for exposing it to the public internet set two rules early: validation and rate limiting run before the hub sees a message, using only the standard library.
16:42, day one
The protection commit landed at 16:42 on March 27 and added three layers. Every message gets validated. Every client gets token buckets for strokes, tile subscriptions, cursor moves, and raw bytes, and more than 50 rate-limited messages in a minute closes the socket. A connection guard caps the server at 500 connections total and 10 per IP, checked before the WebSocket upgrade.
The guard is small (trimmed):
func (g *ConnectionGuard) TryAcquire(ip string) bool {
g.mu.Lock()
defer g.mu.Unlock()
if g.total.Load() >= int64(g.maxTotal) {
return false
}
if g.perIP[ip] >= g.maxPerIP {
return false
}
g.perIP[ip]++
g.total.Add(1)
return true
}
18:25, my own zooming
Less than two hours later, the subscribe bucket went from a burst of 5 refilling at 10 per second to a burst of 20 at 30 per second. The commit gives the reason: fast pan and zoom at deep levels changes the viewport often enough to "hit the old limit too easily." On a canvas that zooms that deep, every pan or zoom can mean a new set of tiles to subscribe to. The fastest legitimate user was the one who wrote it.
02:34, my own stress tool
At 02:30 the next morning I committed a stress tool that simulates N concurrent clients, with ramp-up, stroke and cursor rates, and latency percentiles. Four minutes later the limits went to 5,000 total and 100 per IP. The commit says the old ones "blocked localhost stress testing entirely."
That part was guaranteed. The tool pointed at localhost by default, so every simulated client arrived from the same address. The usage line in that commit asks for 500 clients. The guard would admit 10 and hand the rest a 429 reading "too many connections from your IP."
May, a check that never fires
In May both constants became 50,000, in a commit whose whole message is "agents". The per-IP cap now equals the global cap. For one address to hit 50,000, the total has to be at least 50,000, and the total is checked first. So the per-IP line is still in the code, and it can never be the one that says no.
There's no reasoning behind 50,000 to report. I bumped it and forgot. A commit message that just says "agents" is what vibe-coded work looks like when nobody reviews the diff, and on this project, nobody is me.
What per-IP assumes
The 02:34 commit states the assumption outright: "each real user has a unique IP." That holds for one person at one desk. It fails for an office behind NAT, a school, or a phone carrier that puts many subscribers behind a shared address pool (carrier-grade NAT). A limit tuned for one person per address punishes all of them the way it punished my stress tool.
From the guard's side, my test harness and a school look the same: one address, a lot of sockets, all busy. Nothing in a per-IP counter can tell them apart.