The pileup was still there

I was testing a flight route in a space game under development. The stretch had asteroids, larger structures in the background, and an effect that changed how the ship's sensors read the scene. Automated checks said the objects stayed within the density limit, and the rendered sequence had passed too. In motion, I could still see a pileup ahead. Some particles looked like translucent squares, and the effect I expected to notice was barely visible.

The automation was not useless. It had checked real properties: density, selected views, render time. The mistake was treating those measurements as an answer to the question I was asking from the cockpit: can I read this composition while playing?

When a test passes and someone still points at the same defect, the useful question is not who is right. It is what each of them is seeing.

Another screenshot of the flight route, with the ship, asteroids, background structures, and the effect indicator on the display.
Another automated capture along the route. The displayed value defines the scene; these images alone do not establish final human playtest acceptance.

Counting centers is not seeing silhouettes

The density check counted object positions in world space. From the cockpit, though, shapes far apart in the world can overlap after projection through the camera. Comparing rocks with one another was not enough either: one could sit directly on a much larger form in the background. Their centers satisfied the spacing rule while their visible edges told another story.

The investigation began measuring gaps between projected edges and including large background forms in the check. It also stopped relying on a nearby frame instead of the point reported during flight. A few steps of motion could change the alignment enough to hide the exact thing I had seen.

The old test still mattered for limiting quantity. It simply could not certify visual composition by itself.

One percentage, two different scenes

Another difference hid inside a metric's name. The automated capture called a raw simulation value by the effect's displayed name. The cockpit showed that value after attenuation. The image labeled as the point I had reported actually displayed a lower percentage to the player. We were comparing two different scenes as if they were the same one.

Review aligned sampling with the value shown in the cockpit. Only then did it make sense to keep a capture at that exact point and check whether the forms still touched. The internal number can be correct and still be the wrong axis for reproducing a human observation.

It also explains how an apparently complete capture sequence can leave a hole in the route. Mark frames against the internal scale, and the visible scale advances at a different pace.

What was measured, and what still needed eyes

The review record reports new captures at the cockpit's displayed value, separated silhouettes, and a guard against overlap with large background forms. Automated checks passed again. That gives the regression suite a better chance of catching this composition if it compresses along the same route later.

The last word still did not belong to an edge counter. The record left another human flight as a pending validation step. That makes sense. Automation can preserve a visual condition we have learned to measure; it cannot decree that a scene is readable, interesting, or pleasant to cross.

What stayed with me is that the camera, the display, and motion are also part of the test specification. If that is where I describe a defect, that is where verification has to begin.