← LVGenStudio

Build log · LVGenStudio

What 26 GB going missing taught me about my own tools

Product architecture, orchestrated with Claude Code

Cleanup running in a parallel session removed model files that were still in active use. The multi-reference path stopped working, and recovering it meant Claude identifying exactly what had gone missing, re-downloading at the exact same version, and verifying the checksum against the source — not just trusting a file with the right name back into place.

So I had Claude build something to stop it happening again: a full dependency manifest. Fifty-six files, eighty-one gigabytes, each one with its path, size, checksum, upstream source, version, licence, and which part of the pipeline actually consumes it.

And the tool Claude wrote to find unused files had a bug that would have caused the exact same disaster on its own. Its check for whether a file was still referenced only tested in one direction, and it confidently listed a live 26.68 GB asset as safe to delete. I found that by feeding it a file I knew was missing, to see whether it noticed. A cleanup tool that's wrong in the deletion direction is worse than having no cleanup tool at all — Claude nearly shipped the precise failure it existed to prevent.

Re-running the same thing and getting a different answer

Re-running the recovered setup produced output that differed from the original. I told Claude not to blame the re-downloaded file without real evidence. Good instruction — the actual difference was somewhere else entirely, and jumping to "the new file must be wrong" would have sent it re-verifying something that was already checksum-verified and clean.

The recovered configuration, replayed. Output differed from the original — the cause was not the file everyone suspected first.

There's a correction I made Claude earn twice over in the same stretch: matching file sizes are not matching files. It had cited equal byte counts as evidence once already. The real evidence is a checksum, or a decoded-frame comparison. Later, when a proper test produced output matching a known baseline, I asked for the actual hash — SHA-256, byte for byte — and it matched. That's a claim worth making, because a weaker one had already been made once and shouldn't have been.

Re-verifying the decoder, and finding four of Claude's own bugs on the way

The corrected decoder had already been validated once against a reference implementation. I asked for the check to run again on real production output, with the decoder itself left untouched — bridging its weights into a completely different framework to compare against.

Three attempts, in order: a correlation of negative 0.19, from a layer that had simply never been transferred and was left randomly initialised. Then 0.953 — high enough to look finished, and wrong anyway, because one implementation normalises its input inside the decode step and the other does it outside, so the two were quietly being fed different numbers under an identical-looking call. Then 1.00000000, exact, once both sides actually received the same input.

A correlation of 0.953 is the dangerous one. Nothing about it looks broken. It's high enough that someone hoping to be finished would call it close and move on — which is exactly why I don't accept "close" as an answer to a correctness check.

There was a fourth bug threaded through all of it: weight tensors had been paired by alphabetical filename order instead of their actual position in the model — a hundred and two mismatches out of a hundred and six. And worth stating plainly, because it's easy to skip past: two decoders agreeing with each other doesn't prove the model feeding them is correct. Agreement between two implementations of the same specification only tells you they implement the same specification.

Nine seconds, and three failures with no way to tell them apart

With six seconds working, the next target I set was nine — a new scene entirely, a woman in a recreation room, camera moving from a full-body shot to head-and-shoulders, speaking to camera.

145 frames, 9.06 seconds. Completes in 22m35s — and fails on camera movement, mouth movement and general quality, all three at once.

It completed. And it failed on the camera move, on the requested mouth movement, and on general image quality — all three, at the same time, on a duration and a scene neither one had ever been tested at before. Three failures and two brand-new variables in one run is not something anyone can actually diagnose. It just tells you something's wrong.

Isolating one thing at a time

So I changed how I was asking for tests. Go back to the length and scene already known to work, change exactly one clause, and write the pass/fail rule down before running it, not after.

Camera first. Take the accepted portrait, change nothing but "camera completely still" to "the camera slowly pushes in toward his face," and measure the actual zoom between first and last frame rather than eyeballing it. The still version drifted 1.10× on its own — a decent floor for what "no movement" looks like measured, not assumed. The instructed push measured 1.40×. It works, and it immediately killed an idea Claude had proposed — a dedicated camera-control layer bolted on separately — as unnecessary. The prompt alone was already doing the job.

Then the same test at increasing length, same prompt, only the frame count moving. It holds at forty-nine frames and again at seventy-three. By eighty-one, the movement is gone — not faded, gone, and the video still looks completely fine otherwise. It generates, it looks good, and it silently ignores what was asked.

73 frames — the camera push is still working, full strength.
81 frames, same prompt. The camera instruction is gone. The clip still looks fine — it just stopped listening.

That's the dangerous kind of limitation. A crash tells you something's wrong. This doesn't. So I had Claude turn it into a warning the engine raises before a run, rather than leave it as a line in a document someone might not read in time.

Claude had a tidy explanation ready — these models are commonly trained around eighty-one frames, and instructions can degrade past a model's trained length in a way that would explain a cliff exactly there. I told it to write down in advance that if eighty-one frames itself failed, the explanation was wrong and had to be dropped. Eighty-one failed. It was dropped. Two things saved us from keeping a wrong answer: writing the falsifying condition down before running, not after, and Claude having already flagged that the number it was leaning on came from its own general knowledge, with nothing in this project actually confirming it.

Mouth movement, same design, camera held still in every version so nothing could get confused with anything else. Measured relative to eye movement rather than the whole frame, because a clause that just makes everything move more isn't mouth movement specifically. It passed clearly, against a threshold fixed before the run.

Mouth motion up sharply relative to the eyes. Real — and explicitly not lip-sync, which is a different claim entirely.

Worth being exact here: this is silent video, and it is not lip-sync. The mouth moves more. Nothing here says it matches any particular speech. I had that distinction put directly into the engine's own output, so it can't quietly turn into a marketing claim later by someone reading the result out of context.

The number I made Claude take back

For image quality Claude had been using a standard sharpness measurement, and had cited the nine-second clip's low score against the portrait's high one as evidence it was the worst output in the project.

That comparison was never valid — the number is scene-dependent, and a wide room shot isn't comparable to a close-up face on it at all. Worse, tested directly, it disagreed with what the actual footage looks like, twice: one frame that's visibly dark and distorted scored higher than a clean one. That comparison had been repeated more than once before anyone tested whether the number meant what it was assumed to mean. It didn't. It's recorded now as unusable for a quality verdict, and quality currently requires an actual human looking at the clip — meaning me. What survived the correction: the scene itself matters far more than the duration does, with a smaller duration-linked dip sitting on top of that.

The first thing that's actually a product feature

Everything up to this point ran from hand-typed commands, which means none of it was a feature yet — a script that only I can run by hand isn't a product, and I said so.

So I asked Claude for one entry point that takes a job, streams its progress, and returns a result — wrapping the existing, already-validated pipeline rather than reimplementing it, so nothing about the memory limits or the preserved baselines changes underneath it.

Submitted through the new job interface. SHA-256 identical to the hand-run baseline, byte for byte.

The test that actually mattered: submit the known-good baseline as a job, and compare the file it returns against the file the hand-typed command produced. Same hash. Not similar — the same file.

It refuses things properly too: an invalid frame count is rejected rather than silently rounded to something close, the frame rate can't be nudged to fake a duration it didn't earn, and the two modes I'd already rejected on quality — the single reference and the multi-reference paths — are gated behind an explicit override, so "the backend ran" can never quietly get read as "the output is good." Even the interface itself had a bug worth catching before anyone else did: its progress reporting extrapolated a finish time from zero completed steps and repeated its own last step. A real interface built on top of it would have shown an invented countdown to nobody.

Where it actually stands

Three seconds runs at six and a half minutes. Six seconds at just under fourteen. Nine seconds completes, and its quality is rejected. Camera movement works up to seventy-three frames and silently stops at eighty-one. Mouth movement works at the shorter length and hasn't been tested longer. Both reference-conditioning paths run and both are rejected on the thing they exist to do.

From twenty-one minutes of compute per second of video on the first day back, to roughly two and a fifth minutes now — while the length ceiling moved from three seconds to six. Nothing is wired into an actual application yet. The interface is callable, tested, and proven identical to the hand-run version. No screen calls it yet.

A real optimisation on the wrong component is worth nothing — three bit-identical memory savings, one of them cutting a layer in half, and the run still died until an allocator setting fixed it. Every exciting-sounding measurement Claude made this stretch turned out to be wrong, and every one of them was wrong in the direction of sounding like a discovery — which is exactly the direction I've learned to be most suspicious of. Writing the falsifying condition down before a run, not after, is the only reason a tidy, plausible, wrong explanation got dropped instead of defended. And a silent failure is worse than a loud one: asking for a camera move at ninety-seven frames gets you a good-looking video that simply didn't listen. The engine says so now, before it runs, instead of after.

Next To be written

This is a running log of building LVGenStudio, a local video generation studio for Mac. Written as I go, including the parts that didn't work. Everything here is dated, and anything I later find out was wrong gets a retraction rather than a quiet edit.