The model really isn't that good at this task, unfortunately. However, I did update the generate text with llama vision node to not need the special tokens, and I forgot to update the example.
396 KiB
2202x884px
396 KiB
2202x884px