https://zhuang2002.github.io/FlashVSR/ This only implements the projection model and the VAE, which seems to be enough for upscaling. This does NOT implement any of the streaming and sparse attention code.