Seems like we need more actor crific models that evaluate theae outputs on real world physical modeling and accuracy and not just quality of artistic output or token similarity.
But also Midjourney in particular seems trained to be more evocative / stylistic rather than photorealistic or precise.
But also Midjourney in particular seems trained to be more evocative / stylistic rather than photorealistic or precise.