sg

← Writing

What Unchanged Means

· performance · systems

Leaving a depth-22 view in Fractal and coming back, over a link shaped to 10 Mbps and 40 ms round trip, took 757.4 ms and pulled 867,735 bytes down the socket. Depth is the canvas's zoom level, from 0 to 24. The browser already held nearly all of those bytes. After the change in this post, the same return took 48.8 ms and 1,380 bytes.

The client keeps two caches. One holds cell images, the PNGs the server renders so a zoomed-out view can show art drawn deeper. The other holds vector tiles, the strokes the server sends for each tile the client subscribes to. Coming back in the same tab, with nothing edited in between, after those tiles had dropped out of the subscription, the client asked again and the server answered in full. I traced that repeated traffic and set the rules for when the server may skip a reply. The implementation was agent-written, under the two-session limit I keep.

Reusing a copy still downloaded it#

Images went first. An earlier change in the first performance round had already stopped the browser decoding the same PNG twice: when a reply matched a bitmap it still held, it kept the bitmap. The reply still came. On a depth-13 return at a device pixel ratio (DPR) of 2, that was 196 images, every one a copy of something already in memory.

The cheapest fix was rejected in the first plan in one line. Skipping the request because the client has a cached copy "misses mutations while away." Someone else can draw, delete, or clear while this tab is looking somewhere else. WebSocket compression was already on for large messages. It shrank the bytes, but the client still parsed and installed everything again.

So the browser now says what it holds, and the server answers "unchanged" where it can prove it. For images, on the same shaped link against a frozen 1,792-stroke corpus, three runs per variant:1

Depth 13, DPR 2 returnBeforeAfter
Ready2,905.1 ms403.2 ms
Downstream wire bytes3,308,361334,275
Upstream wire bytes2871,508
Cell-image JSON body bytes2,979,2486,780

Upstream grew, because the request now carries the list. Vectors were untouched: the vector JSON body on that depth-13 return was 1,189,809 bytes in both builds. The measured depth-22 view is drawn from vectors alone, so a deep return still downloaded every stroke. That's the 867,735 wire bytes in the opening, measured later on a larger corpus.

Asking before sending#

Both caches use the same shape. A request carries the client's identities for what it has. The server sends one short result listing the entries that still match, then ordinary full replies for everything else. For vector tiles, quoted as-is from the protocol notes, refresh names tiles to re-check even if already subscribed, and known is the subset the client holds a token for:

json
{"type":"subscribe","tiles":[{"z":0,"x":0,"y":0}],
 "requestId":"tv-12","refresh":[{"z":0,"x":0,"y":0}],
 "known":[{"z":0,"x":0,"y":0,"token":"v-812"}]}

{"type":"tile_validation_result","requestId":"tv-12","accepted":true,
 "matched":[{"index":0,"descendantStrokes":12,"maxDepth":7}]}

A missing match never means the tile is empty or deleted. It means a full body follows.

What an identity is#

An image's identity is three values from the snapshot that produced its PNG: the cell's revision, the mark (the last paint order baked into the image), and whether it's empty. Every applied add or remove of a stroke that could paint into a cell gives that cell a new revision. A delete changes the pixels without moving the mark, because removing a stroke never lowers the mark. Undoing an erase is the less obvious case: redo brings the eraser back at its original paint order, so neither step moves the mark, but both change the pixels. The server confirms a cell only when its current built image matches all three values exactly. A cell whose newest revision isn't built yet doesn't match.

Each canvas has a hub: one goroutine that owns every connection's subscriptions and applies mutations in order. The vector identity lives there, and it went through two designs. The first plan ruled out the paint-order high-water mark as a tile's version, because a delete changes what's drawn without advancing it. It proposed a runtime epoch plus a revision per tile, with tombstone revisions kept for tiles emptied during the epoch, and asked for a measurement of how far the tombstone map grew. The second round replaced that with an opaque token, v- plus a counter, held by the hub in an LRU of 4,096 tiles. An evicted token can't match, so the cost of eviction is one full reply and a fresh token. Empty tiles get ordinary tokens, so there's no tombstone map to grow.

Any applied add or remove drops the token of the tile it touched: drawing, delete, undo, redo, and both ends of a move. A clear empties the LRU but the counter keeps counting, so no token issued before a clear can equal one issued after it.

Unchanged here, new art deeper#

On a canvas this deep, a tile's reply is more than its own strokes. It also carries two numbers about everything below it: descendantStrokes and maxDepth. The client reads them to decide whether to ask for cell images at all, and the server only sends images for cells inside tiles whose descendant count is above zero. The client also subscribes to every ancestor of the current view, from depth 0 down.

A stroke drawn at depth 20 changes no depth-5 tile's own strokes. It does change that tile's descendant count, and the count of every ancestor up to depth 0.

That leaves two ways to get "unchanged" wrong. Put the descendant counts inside the token, and every stroke anywhere revokes the whole ancestor chain, including the depth-0 tile every view subscribes to. Leave them out and skip the whole reply, and the client keeps stale counts, so new art drawn deeper may never get an image requested.

The token covers a tile's own vectors only. Every match carries the current counts. The first plan put the reason plainly: this "avoids bumping every ancestor vector revision or adding versions to all live mutation broadcasts." The server-side check (trimmed):

go
func (ls *loopState) matchTileValidators(c *Client, msg protocol.SubscribeMsg) map[canvas.TileKey]bool {
	// ...
	for i, known := range msg.Known {
		key := canvas.TileKey{Z: known.Z, X: known.X, Y: known.Y}
		if ls.tileIdentities.current(key) != known.Token {
			continue // a full tile_data follows
		}
		matched[key] = true
		meta := ls.hub.canvas.GetMeta(key)
		results = append(results, protocol.TileValidationMatch{
			Index: i, DescendantStrokes: meta.DescendantStrokes, MaxDepth: meta.MaxDepth,
		})
	}
	c.Send(protocol.MarshalTileValidationResult(msg.RequestID, true, results))
	return matched
}

The identity read, the metadata read, and the queued result all happen in one turn of the hub's loop. A mutation after the match arrives later in the same queue and revokes the token normally. The client side had its own version of this mistake, caught in review: if a full snapshot for a tile arrived while its confirmation was still pending, the client skipped the confirmation's counts. It now applies them either way.

Strokes the server hasn't confirmed#

The client draws your stroke before the server acknowledges it. Those optimistic strokes sit in the same tile arrays as the server's, with IDs starting local-. A token certifies a complete server snapshot, and an array with a pending local stroke isn't one.

So the client records a token, along with its connection generation and the tile's stroke revision, only after installing a complete full reply with no local strokes in it. Any change to that tile's array revokes it first: an optimistic stroke, an acknowledgement, a remote stroke, pruning, or a clear.

A token only counts on the socket that delivered it, so a reconnect falls back to full bodies. Images follow the same rule. Their revisions are seeded from the server's clock at startup, which keeps them rising across a restart. That guard fails if the restarted server's clock is behind the last revision the old process issued: the new process could reissue a revision the client already holds for different content. Because reuse is scoped to one socket, that pair is never compared.

While a confirmation is in flight, a retained image can still draw as a fallback, but it doesn't count as current. Readiness and export wait for the server's answer.

Forty returns, zero repeated points#

Vector validation was measured on a frozen 6,481-stroke corpus, comparing the build with image validation against the same build plus vector validation:2

LinkDepth / DPRReady beforeReady afterDownstream bytes beforeAfter
Loopback13 / 144.4 ms32.0 ms405,3101,661
Loopback22 / 176.0 ms2.0 ms867,2451,345
Loopback13 / 249.4 ms34.4 ms411,6888,074
Loopback22 / 279.5 ms2.0 ms867,2451,345
10 Mbps / 40 ms13 / 1413.5 ms71.8 ms405,5201,731
10 Mbps / 40 ms22 / 1757.4 ms48.8 ms867,7351,380
10 Mbps / 40 ms13 / 2401.2 ms99.4 ms411,8988,109
10 Mbps / 40 ms22 / 2765.7 ms57.8 ms867,7351,380

Before, a depth-13 return carried 36,703 stroke points and a depth-22 return 78,194. Across all 40 returns after, the count was zero. Every run produced identical artwork.

A cold visit gets nothing#

A first visit has nothing cached to name, so every reply is a full body. On the image change, a fresh-page control on the shaped link went from 3,019.4 to 3,045.2 ms. Reuse is scoped to the socket that delivered each copy, and the copies live in the tab's memory, so a reload starts from nothing as well. Cold loads were what the coarse preview layer aimed at, and it didn't ship.

Across the whole round, a depth-22 cold load at DPR 1 came out slightly slower: source-ready median 93.1 to 96.4 ms, main-thread CPU median 58.088 to 63.245 ms.3 I can't say what the difference is. It stays in the record, and I didn't keep sampling until it went away.

Where the limits land#

The design is bounded, and the bounds trade hits for size. A vector request lists at most 128 tokens, an image request at most 256, and either message is capped at 16 KiB. Anything that doesn't fit gets a full reply. The hub keeps 4,096 vector identities and no per-client history. A rejected request disables validation for that connection and allows one delayed full refresh, not a retry loop. Every one of those limits fails toward the old behavior: a full reply, and a return as slow as it was before.

Footnotes

  1. Wire bytes are captured transport bytes in the stated direction, downstream or upstream, after WebSocket compression and including framing. JSON body bytes are message bodies before compression, so the two can't be added or compared directly. The image figures use a 1,792-stroke corpus; the vector figures below use a later 6,481-stroke one. ↩

  2. Medians of five returns per view, on one local machine, with headless Chrome and a FIFO proxy shaping the 10 Mbps / 40 ms link. Ready means the application had the vectors and current images it needed, not a compositor timestamp. ↩

  3. Five alternating pairs at DPR 1, a fresh browser against a warm server, comparing the build before vector validation with the build after the round's last performance change. Between them the round also shipped binary PNG frames, stroke JSON serialized once for many viewers, and new loading feedback, so the difference can't be assigned to validation alone. CPU is browser main-thread task time. Earlier grouped samples of the vector change alone were also slower; the report treats this final check as superseding them. ↩