Scrolling capture should fail loudly
Stitching screenshots together is easy until the content stops being distinctive. Here is why Flint Capture would rather refuse than hand you a quietly mangled image.
Scrolling capture looks like a party trick and behaves like a research problem. You scroll a window a bit at a time, take a screenshot of each state, and glue the results into one tall image. The gluing is where it gets interesting.
The naive version works on the demo
Take frame A and frame B. Slide B upward over A until the overlapping pixels match. The offset that produces the best match is how far the window scrolled, so you paste B below A at that offset and repeat. On a documentation page with headings, paragraphs and code blocks, this is nearly foolproof — there is so much structure that only one alignment is plausible.
Then someone points it at a long table of near-identical rows, or a page with a large flat background, and the whole thing quietly falls apart.
Featureless content has no single right answer
If the overlap between two frames is a plain grey area, then every alignment matches perfectly. The score function has no peak — it has a plateau. The algorithm still returns a number, because it always returns a number, and that number is arbitrary.
The failure is nasty precisely because it is invisible. You do not get a garbled image with obvious seams. You get a plausible-looking screenshot that is missing four rows in the middle, or repeats one. The user has no reason to doubt it, and they paste it into a bug report.
A tool that silently drops content is worse than a tool that admits it cannot do the job.
What Flint Capture does instead
The stitcher scores candidate alignments and then asks a second question: is this the best match, and is it decisively the best? A clear winner means a confident offset. A plateau of equally good candidates means the content in the overlap carries no positional information, and no amount of arithmetic will conjure it.
When that happens the capture stops and tells you. You lose the automatic stitch, which is annoying. You do not lose four rows of a table without knowing, which is the outcome that actually costs you something.
Multi-round accumulation
A long page needs more than two frames, so the process accumulates: stitch A and B into a taller canvas, scroll, capture C, align C against the bottom of what you already have. Errors compound in this arrangement — a slightly wrong offset early on shifts everything after it — which is another reason to be strict about confidence rather than forgiving.
This is exercised by the automated tests on every build, including the case where the content is deliberately featureless and the correct behaviour is to refuse. A test that asserts a failure happens is an odd thing to write, and it has caught real regressions.
The general principle
Most of the interesting decisions in a capture tool are about what to do when the answer is unknowable. Scrolling capture is the clearest case, but the same instinct shows up elsewhere: window shadows are synthesised rather than scraped from whatever was behind the window, because scraping produces a shadow with someone else's desktop baked into it.
Guessing is cheap and usually right. It is the "usually" that ruins screenshots.