Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You can imagine that the embedding for the token “boat” includes wood, steel, water and so forth in some set of probabilistic paths as learnt by autorgressive training since the words appear together in the past N tokens. So they are directly in frame. A question is how to connect out of frame elements and are overlapping tokens sufficient to do “that”. Specifically is the token “that” sufficiently trained to reveal what it refers to? I think this depends on the fine tuning q/a task which adds in a layer of alignment rather than being an emergent property of the LLM in general.

Still alignment tasks are autoregressive (I think)… they could be masked or masked part of speech potentially.. but if autoregressive then I suspect you’re looking at regularities in positioning to identify things.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: