Yep totally agreed. One of the things I'm excited to see going forward is multimodal models (trained e.g. on text + video + audio + images).
I'm sure there's a lot more to it than this, but maybe one factor that makes humans a lot more data efficient is the multimodal input we receive.
If that's the case, imagine how much better things could get when we train with all the videos, podcasts, radio etc in the world, in addition to all the text out there!
I'm sure there's a lot more to it than this, but maybe one factor that makes humans a lot more data efficient is the multimodal input we receive.
If that's the case, imagine how much better things could get when we train with all the videos, podcasts, radio etc in the world, in addition to all the text out there!