Things I made

Can a model find the part worth watching?

The moment might last a second. The reason it matters might begin ten seconds earlier.

“A ball rolls across a table and falls off the edge.” That could describe a dozen videos. It tells me what happens, but very little about why I might keep watching.

I made a small illustration of the difference. Both scenes below have the same ball, the same table, and the same outcome. One gets there directly. The other slows near the edge before falling. Neither version is automatically better. What interests me is how much a change in timing can change the question in my head.

A / Straight to the event

B / A pause near the edge

Two original animations. Motion starts only when you choose Play.

This is a timing illustration, not a physics simulation or a model prediction. No video footage is used.

The model thought almost everything was funny enough

In one early experiment, I took a ranker trained on narrated stories and tried it on funny videos. It gave high scores to almost everything—even examples with very different observed performance.

The similarity search was doing its job: these were recognizably funny-video-shaped videos. The ranker was barely separating them. My interpretation was that it had picked up features common to the format, without learning what distinguished a good example within it.

That was a useful failure. Finding the right neighborhood of videos was much easier than finding the one worth stopping for.

I wanted the shape of a story to show up in the math

One hypothesis I particularly liked was to treat a video as a path through its hook, setup, and payoff. How far does it travel? Does it change direction? Does the ending return to something introduced at the beginning?

It felt like there might be a little geometry hiding inside a satisfying sequence. A setup establishes something; a payoff changes it. Perhaps those movements would help distinguish a story from a collection of related pictures.

In the small experiment I ran, the added trajectory features did not beat the simpler baseline. I liked the idea more than the evidence justified.

That does not settle whether timing matters. It tells me that those coarse measurements, in that experiment, did not capture it usefully. Three snapshots of a joke can miss the pause that makes it land.

More machinery did not automatically help

I also tried different combinations of visual and audio features and training objectives. The more elaborate alternatives did not consistently displace the simpler checkpoint.

It was tempting to interpret every disappointing result as a request for a better representation. But the target itself was slippery. A video’s observed popularity includes its audience, distribution, and time online. A content-only model sees none of that directly.

So “predict which video performed better” was only an imperfect way into the question I cared about: what in the video made someone want to watch?

The interesting part might include the waiting

Those runs were mostly about ranking whole videos. Finding an excerpt is the next question, and it changes what a useful answer looks like.

In the original animation above, locating the fall is easy. Choosing where the excerpt should begin is more interesting. The approach to the edge might be dead time, or it might be the setup. Cut it all away and you may remove the reason to care.

I want to test that directly: keep the event fixed, vary its lead-in, and see which portions people choose to keep. Then ask whether a model can learn those choices on unfamiliar scenes. That experiment is still ahead of me.

The moment might last a second. The reason it matters might begin ten seconds earlier.

That is the part I keep coming back to. A system can recognize the subject, find similar videos, and still miss why this particular stretch of time is worth watching.

The findings come from small exploratory ranking experiments; the animation illustrates timing and is not a trained model output.

A personal observation, and a question to come back to.