Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

One argument (that I don't necessarily buy) is that inputs from sensory perception play a major role in building a language model in humans. A human might not read TB of books to learn language, but if you put them in a sensory deprivation tank that only displayed a massive series of Unicode characters it would take forever for us to learn language. If we did the same for an ML model, perhaps those non-language training materials would help.

(That said, DL architectures are obviously wildly different from how the human brain works. E.g. backprop is physically impossible.)



Yep totally agreed. One of the things I'm excited to see going forward is multimodal models (trained e.g. on text + video + audio + images).

I'm sure there's a lot more to it than this, but maybe one factor that makes humans a lot more data efficient is the multimodal input we receive.

If that's the case, imagine how much better things could get when we train with all the videos, podcasts, radio etc in the world, in addition to all the text out there!


Nature is good reference point but it's not by any means optimal. One example may be human eye - we can make cameras orders of magnitude better than eyes created by evolution.

You can look at it the other way around - dopamine and other neuro transmitters as poor approximation of backpropagation. It has many flaws for example tight harmful loops ie. addictions.

Majority of brain work is ignoring irrelevant information (attention) and small scale hallucinations (we don't see world as is but slightly hallucinated to keep it stable - ie. they way brain processes blinking <<turns off>>, you can peek at those nuances with ie. optical illusions etc).

One of missing bits in neural nets may be reusing its output as input (embedded in inference itself, not poor mans re-prompting).

Once it's sorted out I'd argue the performance will skyrocket and give opportunity to massive optimisations ie. embedding things like known functions - imagine brain which has known, very narrowed, available functions at its disposal - all mathematical functions on numbers, logic, optimal sorting etc. Imagine if as thinking human you'd have access to accurate functions - the sky is a limit.


“Nature is good reference point but it's not by any means optimal. One example may be human eye - we can make cameras orders of magnitude better than eyes created by evolution.”

This is only true when considering single performance axis like pixel resolution. When you consider the corpus of power efficiency, jitter resolution enhancement, dynamic contrast, performance per volume, etc. We aren’t close to building something as capable.


> all mathematical functions on numbers, logic, optimal sorting etc

Bayes formula as a built in primitive. You don't need to know much of statistics to see how limited humans are at processing information because estimating posterior updates is so expensive for them.

Thinking in terms of raw probabilities would be very alien to most humans, but could easily be technically superior for making plans.


Helen Keller managed to do pretty well without sight or hearing.


Wasn’t she ‘boot-strapped’ by having had both for the first 19 months of her life? - Presumably this gave her an internal model of the world on to which she could map touch perceptions and language concepts.

https://en.wikipedia.org/wiki/Helen_Keller


Touch is pretty important though...


[flagged]


1. You're mischaracterizing the parent's comment, to such a degree that it seems willful.

2. This tendency to police other people's behavior, telling them what they should or should not do, does not help your cause. It puts people on the defensive. No one, especially adults, like to be told what, or what not, to do unless it's coming from a well-qualified lawyer.


Agreed, i think the "telling other people what to do" thing is something i haven't seen brought up enough. Seems like the whole current social media is built on defending some group by telling other people that they're bad people for speaking the way they always have and that they must change themselves to fit your opinions, otherwise they're forever doomed to be "bad people". The bad person part is implicit in the command. When did people get so comfortable commanding others and demeaning their thoughts and agency?


> When did people get so comfortable commanding others and demeaning their thoughts and agency?

When people stopped having "friends" and started accruing "followers."

You can bully anybody about anything when your gang is bigger than theirs.


I can't believe you are serious. It is pretty clear what the context is here and I didn't mince my words. You are policing MY behavior and you should stop telling people how to be anti-ableist.


If anything I read the original comment as being the opposite of what you seem to have pulled out of it. HN actually has a guidelines entry on this:

"Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."

See: https://news.ycombinator.com/newsguidelines.html


Talking about an aspect of something or someone, is not the same as reducing something or someone to that aspect.

Nor was anyone disparaged or disrespected.


She was brought up in respect to a human in a sensory deprivation tank. You've made the connection to LLMs.


Having a model of the world (predominantly from vision, but also hearing and touch) surely does wonders for language acquisition as it provides the learner a backdrop against which to infer the meaning of speech. When an infant hears “look, a dog” and “look, a cat” in two separate instances, their eyes alone provide enough information to infer, at least at a high level, the meanings of “look, a”, “dog”, and “cat”.

It's pretty clear that humans, unlike LLMs, use external sensory data (the only external data we have, when you cut through it all) when they produce speech, as evidenced by the fact that they don't speak falsehoods that don't mesh with their internal data model. LLMs have such a weak model of reality -- it's whatever “sounds right” -- that they speak falsehoods all the time. The only way to give an LLM sensory data would be to encode every sensory experience people have into text.


> It’s pretty clear that humans, unlike LLMs, use external sensory data […] when they produce speech, as evidenced by the fact that they don’t speak falsehoods that don’t mesh with their internal data model.

I don’t have access to any other humans internal data model, but the indirect evidence I do have suggests that they do, in fact, speak falsehoods that don’t mesh with their internal data model for a variety of strategic purposes.


It's all a combination of both explicit (traditional software that's faster/has better specific accuracy but is brittle) and implicit (machine learning/etc where fuzziness is better) optimized hardware (wetware?)/software all up and down the biological stack, so to speak.

And it's been optimized over billions of years.


It's telling that there's always an argument like this in response to everything critical about llms and they have no coherence.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: