Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yep totally agreed. One of the things I'm excited to see going forward is multimodal models (trained e.g. on text + video + audio + images).

I'm sure there's a lot more to it than this, but maybe one factor that makes humans a lot more data efficient is the multimodal input we receive.

If that's the case, imagine how much better things could get when we train with all the videos, podcasts, radio etc in the world, in addition to all the text out there!



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: