https://zhuang2002.github.io/FlashVSR/ This only implements the projection model and the VAE, which seems to be enough for upscaling. This does NOT implement any of the streaming and sparse attention code.
The model is available in fp16 only though