Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Veo's cutting-edge latent diffusion transformers reduce the appearance of these inconsistencies, keeping characters, objects and styles in place, as they would in real life.

How is this achieved? Is there temporal memory between frames?



Probably similar to Sora, a patchified vision transformer, you sample a 3d patch (third dimension is time) instead of a 2d patch




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: