One area that needs a closer look if we are to achieve AGI is the ability to extract concepts from language and have a clean separation between what is learned and how to put it into words.
Yes we already have some of that today, but it’s very nascent and crude relative to what humans are able to do. Concepts and their language representation is more intermingled than we’d like in the model. More layers isn’t going to fix this either, at least not efficiently.