Not sure why this post is getting down voted? It's spot on - the limitation lies in relatively small text encoders within stable diffusion/mid journey. Larger language models contribute to better image generation (see: DeepFloyd). Google's CapPa also recently showed an alternative to the contrastive learning of CLIP at any param count.
Until larger models are released, there are "hacks" like this available: https://bair.berkeley.edu/blog/2023/05/23/lmd/ - GPT4-generated bounding boxes guiding stable diffusion