Build log · LVGenStudio
The wrong optimisation, and the one line that was right
Product architecture, orchestrated with Claude Code
The target moved from a three-second clip to one continuous six-second shot, inside fifteen minutes and sixteen gigabytes. I gave Claude four routes explicitly ruled out first: no dropping the resolution, no lowering the playback frame rate, no repeating frames, and no stitching two shorter clips together and calling it one take.
That last constraint is the whole difficulty. Joining two three-second shots is easy, and it looks exactly as easy as it was. One continuous take is a different problem entirely.
Four failures, and the arithmetic behind them
Every attempt died the same way — the memory guard tripping mid-generation. The reason, once Claude worked it out, is unpleasant and structural: the cost of attention grows with the square of how many frames are in the sequence. Doubling the length of the clip very nearly quadrupled the single most expensive operation in the whole pipeline. The three-second version ran comfortably. The six-second version couldn't finish at all.
Claude wrote three separate memory reductions in response, and verified each one produced bit-identical output — not "close enough," the same numbers to the last bit. One of them alone cut a major layer's memory in half. And the run still died, because at the level of the whole computation that saving was sitting underneath something larger it never touched.
A real 48% saving on the wrong component buys nothing.
This was also the stretch where Claude's own analysis was worst all week. Three separate measurements, all wrong, and every one of them wrong in the direction of sounding like a discovery — a cost attributed to the wrong piece of code, a benchmark quietly holding memory while it measured its own subject, an "unexplained" 27% of runtime that turned out, measured properly in place, to not exist at all. That's the pattern I started watching for: the wrong numbers weren't random. They all pointed toward something exciting. That's exactly the direction to be suspicious of, and I said so.
The fix was a setting, not a rewrite
The actual cause was a single policy inside the framework itself: freed memory gets kept in a reuse pool rather than returned to the system, and its default ceiling on this machine was larger than the GPU's entire addressable budget. Claude measured, during a failing run, that pool holding eleven and a half gigabytes while barely more than one and a half was actually live — memory the program was finished with, kept on the chance it might be wanted again, while the operating system quietly swapped real work to disk to make room around it.
Capped that pool down hard, and the six-second shot completed. Eighty-five minutes, at the native resolution — bit-identical output to every earlier attempt at the settings that used to fail, and no measurable time cost from the cap itself. The pool had never been buying speed here. Only holding ground.
Three exactly-correct optimisations hadn't finished the run. One allocator setting did.
Eighty-five minutes is completed. It isn't acceptable.
My target was fifteen. So I sent Claude back to a technique from the earlier project's own notes: render at half the linear resolution, then enlarge with a separate upscaler. A quarter of the pixels through the expensive part, a cheap pass to restore the size after.
Six times faster, end to end, and inside the memory budget the full-resolution path never was.
My own read on quality, watching it side by side with the full-resolution clip: nearly the same, with a slight colour tint in the hair and eyes. I'm recording that plainly rather than letting Claude paper over it with a blanket colour filter — a fix like that would shift correctly-coloured parts of the frame to hide a defect that's only local. I told Claude to leave it as a known, unrepaired flaw rather than mask it.
Two ways of telling the model who's in the shot, both rejected
A prompt alone can't specify a particular person. That needs a reference photo, so I had Claude test both approaches the engine offers.
Single reference image first — one photo, motion generated around it. It ran, and I rejected it on sight: visible artefacts in anything that moves.
Multiple reference images next, on a completely different piece of the model built for exactly that. Claude had estimated this as days of work — a separate architecture with its own conditioning path. It turned out to be a forty-four-line shim aliasing module paths. The estimate was wrong by two orders of magnitude, and the thing Claude had been calling a porting problem was a naming problem the whole time.
It ran too, and I rejected it for the reason that actually matters most: the generated face doesn't match the reference photos at all. Identity is the entire point of giving a model multiple references. Failing at that isn't a rough edge on an otherwise working feature. It's a failure at the one job the feature exists to do.
Both backends work, in the narrow sense of producing a video without crashing. Neither output is acceptable to me. Those are two different sentences, and I made sure the record says both instead of collapsing them into one comfortable "it works."
Next What 26 GB going missing taught me about my own tools
This is a running log of building LVGenStudio, a local video generation studio for Mac. Written as I go, including the parts that didn't work. Everything here is dated, and anything I later find out was wrong gets a retraction rather than a quiet edit.