LocalLLaMA

3 readers

1 users here now

Community to discuss about Llama, the family of large language models created by Meta AI.

founded 1 year ago

MODERATORS

communick@poweruser.forum

Why Starling RLAIF has better results that Zephyr DPO ? (alien.top)

submitted 1 year ago by Puzzleheaded_Mall546@alien.top to c/localllama@poweruser.forum

1 comments fedilink hide all child comments

Based on this image:

https://preview.redd.it/z5vf03e8r54c1.png?width=648&format=png&auto=webp&s=0a652e76ab2489135ed2327e8156029eacf274b7

Starling has better results than Zephyr DPO in all the metrics, Why ?

Shouldn't DPO be better than RLHF/RLAIF ?

you are viewing a single comment's thread
view the rest of the comments

[–] fediverser@alien.top 1 points 1 year ago

This post is an automated archive from a submission made on /r/LocalLLaMA, powered by Fediverser software running on alien.top. Responses to this submission will not be seen by the original author until they claim ownership of their alien.top account. Please consider reaching out to them let them know about this post and help them migrate to Lemmy.

Lemmy users: you are still very much encouraged to participate in the discussion. There are still many other subscribers on !localllama@poweruser.forum that can benefit from your contribution and join in the conversation.

Reddit users: you can also join the fediverse right away by getting by visiting https://portal.alien.top. If you are looking for a Reddit alternative made for and by an independent community, check out Fediverser.