Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Not sure why this post is getting down voted? It's spot on - the limitation lies in relatively small text encoders within stable diffusion/mid journey. Larger language models contribute to better image generation (see: DeepFloyd). Google's CapPa also recently showed an alternative to the contrastive learning of CLIP at any param count.

Until larger models are released, there are "hacks" like this available: https://bair.berkeley.edu/blog/2023/05/23/lmd/ - GPT4-generated bounding boxes guiding stable diffusion



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: