Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It is not believable that someone would pay humans to label 400mln or 5bln images/samples to train a model on them, but I guess if you argument is "everything is possible" then gotcha


If it's done in a reCAPTCHA like way, it can be done fairly efficiently and for cheap. In fact Scale AI does just this, they do manual labor operations such as captioning images, as an API. Here's their product for image labeling: https://scale.com/rapid.

Unstable Diffusion is also doing their captioning like how I mentioned, with groups of volunteers as well as hired individuals.


Scale seems to do, for example, image classification but not captioning as it would be hard to compare the results with others people to verify the quality (when you have a discrete number of classes is really straightforward), also can you report where you read about the Unstable Diffusion plan for manually labeling image datasets? I want to dig deeper



> We are releasing Unstable PhotoReal v0.5 trained on thousands of tirelessly hand-captioned images

They seem to have created a much smaller dataset than LAION's, it would not work to train a generative model on such a small amount of images (obviously the images here do not have a single domain).


You seem to be confusing "possibility" with your personal opinion on what you think would be done by others.


As a human being I know human limitations, explicitly labeling 400mln/5bln images for a particular task seems absurd to me, but if you think it is realistically possible perhaps you can give an example.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: