LocalLLaMA

11 readers

4 users here now

Community to discuss about Llama, the family of large language models created by Meta AI.

founded 2 years ago

MODERATORS

communick@poweruser.forum

Why didn't gpt4 work at first and how did they "fix it"? (alien.top)

submitted 2 years ago by Amgadoz@alien.top to c/localllama@poweruser.forum

11 comments fedilink hide all child comments

According to this tweet,

when gpt4 first finished training it didn’t actually work very well and the whole team thought it’s over, scaling is dead…until greg went into a cave for weeks and somehow magically made it work

So gpt-4 was kind of broken at first. Then greg spent a few weeks trying to fix it and then it somehow worked.

So why did it not work at first and how did they fix it?
I think this is an important question to the OSS community,

you are viewing a single comment's thread
view the rest of the comments

[–] wojtek15@alien.top 1 points 2 years ago (1 children)

According to https://openai.com/research/gpt-4 they were able to predict GPT4 performance while still training it, so this is contradiction to this tweet.

[–] dogesator@alien.top 1 points 2 years ago

Predicting the loss is very different from predicting real world abilities, they are able to top the former, not the latter.

Predicting the future loss once you’re already 10% into training is fairly trivial. Predicting the actual abilities though is not.