Voozh

VOOZH

URL: https://huggingface.co/alvarobartt/mistral-orpo-mix

⇱ alvarobartt/mistral-orpo-mix · Hugging Face

ORPO fine-tune of Mistral 7B v0.1 with DPO Mix 7K

👁 image/jpeg

Stable Diffusion XL "A capybara, a killer whale, and a robot named Ultra being friends"

This is an ORPO fine-tune of mistralai/Mistral-7B-v0.1 with alvarobartt/dpo-mix-7k-simplified.

⚠️ Note that the code is still experimental, as the ORPOTrainer PR is still not merged, follow its progress at 🤗trl - ORPOTrainer PR.

Reference

ORPO: Monolithic Preference Optimization without Reference Model

Downloads last month: 6

Safetensors

Model size

7B params

Tensor type

BF16

·

Model tree for alvarobartt/mistral-orpo-mix

Base model

mistralai/Mistral-7B-v0.1

Finetuned

(937)

this model

Quantizations

Dataset used to train alvarobartt/mistral-orpo-mix

Space using alvarobartt/mistral-orpo-mix 1

Collection including alvarobartt/mistral-orpo-mix

Contains some information and experiments fine-tuning LLMs using 🤗 `trl.ORPOTrainer` • 7 items • Updated Mar 2 • 5

Paper for alvarobartt/mistral-orpo-mix

Paper • 2403.07691 • Published Mar 12, 2024 • 73