For the val loss did you just use a single denoiser step error or was it more like a eval loss where the denoiser was run multiple times and the final loss after all denoiser steps was used?
Also, do think that the no correlation could be totally explained by multimodal nature of the task? What do you think would happen if we train a vey simple task, lets says move in a straigth line from a random spawn point to a random goal point?