On a cold load, a coarse preview layer got Fractal's canvas to 90% covered 37.4% sooner. Final detail arrived 46.8% later, so I removed it.
When you zoom, the server has to render the new level's 256×256 PNG cells and send them down. Until they arrive, the client fills the gap with whatever coarser images it already holds.
The first round of performance work went after the rasterizer's wasted scans. For this round I directed two independent experiments from a written plan, each with pass/fail gates set before agents wrote any of the code (the division of labor). Only one experiment survived, and a threshold from a later experiment turned out to have no reason behind its number.
The idea: one coarser level first#
The candidate added a single auxiliary coverage level. For every range of final cells the client asked for, the server also derived the range one level up, where each image covers a 2×2 group of final cells. Coverage images were ordinary cells with the same revision and cache semantics. The server prioritized them, dispatching at most four coverage builds while final work waited before serving one final build. Final images themselves did not change.
The hope was simple. A blurry image of the right region beats a hole.
The gates came first#
Coverage had to clear three conditions to ship:
- Median uncovered area-time on the primary cold-load and zoom route had to drop at least 20%.
- Final-detail readiness p95 could not regress more than 10%.
- Any CPU, wire, or memory increase over 20% failed the candidate.
The plan also says: "do not silently accept duplication for a first-pixel-only improvement." That sentence was written before any result existed to argue with.
How it was measured#
Both builds ran against the same frozen corpus of 6,481 strokes, in a 1024×1024 CSS-pixel viewport at DPR 2 (two device pixels per CSS pixel), through a FIFO proxy shaped to 10 Mbps and 40 ms round trip. Each run got a fresh server and a fresh clone of the data, loaded the overview until ready, then drove the same eight-second continuous zoom from depth 0 to depth 22 and back. Five baseline/candidate pairs alternated, and the plan ruled out rerunning to chase a better result. The baseline already included the raster change described further down, so none of that gain counts toward coverage. Everything ran on one local machine.
"Covered" and "uncovered" come from the renderer's per-frame source plan, which records whether each part of the viewport has an image to draw from. They are not compositor timestamps, and vectors may still draw over an uncovered region. Area-time is in viewport-milliseconds: uncovered area integrated over time, normalized to the viewport. CPU is browser main-thread task time. Every row is a median of five runs, except the p95 rows.1 A blank Gate cell means the row had no gate.
| Metric | Baseline | Coverage | Change | Gate |
|---|---|---|---|---|
| Cold 50% covered (ms) | 2,809.9 | 1,840.5 | -34.5% | |
| Cold 90% covered (ms) | 4,654.7 | 2,912.8 | -37.4% | |
| Cold uncovered area-time (viewport-ms) | 1,999.4 | 1,177.5 | -41.1% | Pass (needs 20% drop) |
| Cold final-ready (ms) | 5,118.3 | 7,513.6 | +46.8% | |
| Cold final-ready p95 (ms) | 5,202.9 | 7,571.0 | +45.5% | Fail (10% limit) |
| Motion uncovered area-time (viewport-ms) | 899.2 | 2,763.4 | +207.3% | Fail (needs 20% drop) |
| Final-ready wait after motion (ms) | 10,577.1 | 14,389.4 | +36.0% | |
| Final-ready wait after motion p95 (ms) | 11,653.6 | 15,429.4 | +32.4% | Fail (10% limit) |
| Downstream TCP bytes (MB) | 27.9 | 35.9 | +28.4% | Fail (20% limit) |
| Cold main-thread CPU (ms) | 545.1 | 802.4 | +47.2% | Fail (20% limit) |
| Zoom route and drain main-thread CPU (ms) | 2,349.1 | 2,609.6 | +11.1% | Pass (20% limit) |
The number I would have put on a slide#
The top three rows are real, and they are the rows I would have shown someone. On a cold load the canvas reached half coverage about a second sooner, and 90% coverage 1.7 seconds sooner.
The rest of the table is the cost. Coverage images competed with final detail for the same workers and the same link. Cold final-ready went from 5.1 to 7.5 seconds. Its p95 and the post-motion p95 both rose well past the 10% limit. During continuous zoom, the feature meant to reduce holes tripled uncovered area-time. Bytes and cold CPU both crossed the 20% line. Only CPU across the zoom route stayed under it.
The motion ranges don't overlap. Every candidate run (2,753.9 to 3,111.6 viewport-ms) was worse than every baseline run (887.9 to 2,745.6). The closest they came was 8.3 viewport-ms, the gap between the worst baseline run and the best candidate run.
Diagnostics also showed more PNG data arriving for cells the client no longer wanted: 13.5 MB of PNG bodies outside the current view in the baseline, 17.2 MB with coverage.2 They point at obsolete delivery during motion, but they don't say whether the server queue, the proxy, or network timing caused it, so I haven't acted on them.
The candidate was correct. It passed 100 scratch checks, an independent review found no confirmed defect, and all ten runs produced byte-identical final artwork. It was removed completely. The experiment had bumped the wire protocol to version 2; the shipped protocol is still version 1. The follow-ups that were skipped, including frame captures and multi-viewer resource runs, are not counted as passes. They never ran.
The half that shipped#
The same plan carried a second experiment with its own gate: an exact active-edge sweep in the rasterizer. The rasterizer splits each pixel row into thin strips and asks every polygon and disk in a stroke's outline what it covers at each strip's height. The first round's bounds check and scratch reuse made that cheaper, but every strip height still visited geometry that couldn't touch it. In overview cells, under 8% of the polygons, edges, and disks scanned at a given strip height were active.
The sweep keeps a set of the shapes crossing the current height. As it moves down the cell, it adds shapes whose top it has reached and drops shapes it has passed. The arithmetic and its order stay the same, so the output is exact.
The gate asked for three things: a 10% median raster improvement, no deep-zoom regression over 5%, and no unexplained growth in allocation or peak memory. Each view rendered and PNG-encoded 64 cells serially, medians of five alternating pairs:
| View | Before (ms) | After (ms) | Change |
|---|---|---|---|
| Overview | 2,230.2 | 1,659.1 | -25.6% |
| Depth 13 | 1,167.1 | 508.8 | -56.4% |
| Depth 22 | 2,396.1 | 724.1 | -69.8% |
All 192 cell cases matched the old PNG hashes exactly.
It isn't free. The event arrays add 6 to 7% more allocated bytes per render, and peak process memory rose by under 2%. The arrays belong to one render and are released with it, so I took that trade. This is render time in isolation, not a 69.8% faster browser view.
A threshold that would have been right for the wrong reason#
Neither verdict above depended on where a gate sat. Coverage missed by wide margins: sample p95 up 45.5% against a 10% limit, and cold CPU up 47.2% against 20%. The sweep cleared its 10% bar at 25.6% to 69.8%. No plausible threshold flips either result. Thresholds only decide things near the line, as with two interval-sort variants at 8.84% and 8.99% against 10%, and a simplification candidate at 11.0% against 15%.
That last one came from a later experiment. The generator that produces Fractal's showcase artwork emits densely sampled polylines, and three candidates thinned redundant points. They cut offline compressed size by 5.2%, 7.0%, and 11.0%, and raster CPU by 4.2%, 9.5%, and 10.4%. The experiment started with a 15% minimum compressed-size gate. All three fall under it, so the gate would have rejected every candidate on size.
Nothing justified 15%, so I removed it. Any repeatable gain can now qualify if quality and the other measures hold. All three candidates were still rejected, for a different reason: native deep-zoom crops showed shapes moving. One curved pen contour became nearly straight, displaced several pixels at camera depth 20.


The cause is magnification. Each stroke's points are stored in coordinates local to the tile it was drawn in, where one unit is one tile width, and its tile depth is how deep that tile sits. An error in those local coordinates grows on screen with every level you zoom past it:
1024 × 2^(cameraDepth - strokeTileDepth) × DPR × localError
The 1,024 is the client camera's base size: at zoom 0 the whole canvas is 1,024 CSS pixels across, so a tile at depth d spans 1024 × 2^(zoom - d) CSS pixels, and DPR converts that to device pixels. It is not the window width.3 A tolerance of 0.00005 is 1/20,000 of a tile. At tile depth 10, viewed at depth 24 and DPR 1, that is 1024 × 2^14 × 1 × 0.00005, roughly 839 device pixels. That figure is how far an error equal to the tolerance would land, not a shift that was observed. The pen stroke above was drawn on a depth-14 tile, so at depth 20 the 0.0001 tolerance is 1024 × 2^6 × 1 × 0.0001, about 6.55 pixels at DPR 1, close to the shift in the crop.
The size gate would have produced the same verdict and hidden the real finding. Tolerances picked without regard to the zoom range didn't hold shape at native deep zoom. Read through the size gate, the lesson would have been "remove more points," which is the direction that moves shapes further. A candidate that accounts for maximum zoom magnification might work. It hasn't been built.
What stays in the record#
The project's performance report keeps a table of experiments that didn't ship, with the measured benefit next to the failed condition. Coverage is in it. So are the interval-sort variants rejected under the 10% payoff threshold. Those decisions stand as recorded. I haven't gone back to check whether 10% was any better justified than 15%.