Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It makes me wonder if long-neglected archives of old newspapers, internal corporate documentation, and government publications could have newfound economic value as training material. I know that the absolute volume is small compared to e.g. large scale document scraping from the Web, but intuitively I would guess that one old Bureau of Standards report written with a high degree of literacy has more value than 1000 SEO-chasing "how to make pancakes" guides. A lot of old documents were never digitized simply because people didn't think they had value. Is it time for a reinvigorated Google Books with a wider mission?


"A lot of old documents were never digitized simply because people didn't think they had value. Is it time for a reinvigorated Google Books with a wider mission?"

We have to be very careful with the "factual" content of old non-fiction works. So much of that, from history, to medicine, to biology, etc, turned out to be straight from their writers' imaginations. If an LLM considers such books to be no different from modern books on these subjects it would get a very skewed view of the world.

Imagine asking an LLM for a medical diagnosis and it responding with something about the humors.


That's a good point. I wonder how LLMs currently deal with the passage of time, if they can at all. Do they know that Satya Nadella is the current CEO of Microsoft and that Steve Ballmer no longer is simply because there are more training documents reflecting the present state of things, or is there an explicit time component that helps to resolve conflicting facts from different years? Based on what I've read so far about LLMs (and not actually working on any of these models myself), I wouldn't have thought they model time or facts in a way that you could expect them to resolve conflicting claims from a 2010 document and a 2020 document. Yet they're unreasonably effective at question answering anyway.


Something I've been thinking a lot about lately is the potential historical usefulness of LLMs which have been selectively trained on material from before a certain date, presumably resulting in a model exhibiting attitudes of the time in question. It would be extremely interesting to be able to ask for a 1950s take on a modern idea.

For this, large amounts of very old text would be extremely valuable.


I suppose the relative low value for low volumes of text makes this not a problem, but...

What about data that's sitting around and isn't supposed to be public? If training data gets scarce, does a market for small-medium sized data emerge? Like old homework papers, internal company documents, etc?


Not if you're trying to churn out "How to make pancakes" guides with your LLM




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: