LocalLLaMA

14 readers

1 users here now

Community to discuss about Llama, the family of large language models created by Meta AI.

founded 2 years ago

MODERATORS

communick@poweruser.forum

Why didn't gpt4 work at first and how did they "fix it"? (alien.top)

submitted 2 years ago by Amgadoz@alien.top to c/localllama@poweruser.forum

11 comments fedilink hide all child comments

According to this tweet,

when gpt4 first finished training it didn’t actually work very well and the whole team thought it’s over, scaling is dead…until greg went into a cave for weeks and somehow magically made it work

So gpt-4 was kind of broken at first. Then greg spent a few weeks trying to fix it and then it somehow worked.

So why did it not work at first and how did they fix it?
I think this is an important question to the OSS community,

you are viewing a single comment's thread
view the rest of the comments

[–] maizeq@alien.top 1 points 2 years ago

Much more likely: the story is apocryphal, or is at least, highly exaggerated.