Possibly it is because with things like Stable diffusion we give it a lot of passes when things don't exactly right. Images just have to be close enough.
Text however if it is only a single word out, the whole meaning and readability can change. It needs a significantly larger data set to ensure clearer readability.
Text however if it is only a single word out, the whole meaning and readability can change. It needs a significantly larger data set to ensure clearer readability.
Just a shot from the hip response on this one.