## Model overview

`prometheus-13b-v1.0` is an alternative to GPT-4 for fine-grained evaluation of language models. Developed by [prometheus-eval](https://aimodels.fyi/creators/huggingFace/prometheus-eval), it uses the [Llama-2-Chat](https://huggingface.co/meta-llama/Llama-2-13b-chat-hf) model as a base and fine-tunes it on 100K feedback samples from the [Feedback Collection](https://huggingface.co/datasets/kaist-ai/Feedback-Collection) dataset. This specialized fine-tuning allows `prometheus-13b-v1.0` to outperform GPT-3.5-Turbo and Llama-2-Chat 70B, and perform on par with GPT-4 on various benchmarks. In contrast to GPT-4, `prometheus-13b-v1.0` is a more affordable and customizable evaluation model that can be tuned to assess language models based on specific criteria like child readability, cultural sensitivity, or creativity.

## Model inputs and outputs

### Inputs
- **Instruction**: The task or prompt to be evaluated
- **Response**: The text response to be evaluated
- **Reference answer**: A reference answer that would receive a score of 5 
- **Score rubric**: A set of criteria and descriptions for scoring the response on a scale of 1 to 5

### Outputs
- **Feedback**: A detailed assessment of the response quality based on the provided score rubric
- **Score**: An integer between 1 and 5 indicating the quality of the response, as per the score rubric

## Capabilities

`prometheus-13b-v1.0` excels at fine-grained evaluation of language model outputs. It can provide detailed feedback and scoring for responses across a wide range of criteria, making it a powerful tool for model developers and researchers looking to assess the performance of their language models. The model's specialized fine-tuning on human feedback data enables it to identify and react appropriately to the emotional context of user inputs, a key capability for providing empathetic and nuanced evaluations.

## What can I use it for?

`prometheus-13b-v1.0` can be used as a cost-effective alternative to GPT-4 for evaluating the performance of language models. It is particularly well-suited for assessing models based on customized criteria, such as child readability, cultural sensitivity, or creativity. The model can also be used as a reward model for Reinforcement Learning from Human Feedback (RLHF) approaches, helping to fine-tune language models to align with human preferences and values.

## Things to try

One interesting use case for `prometheus-13b-v1.0` is to provide detailed feedback on the outputs of large language models, helping to identify areas for improvement and guide further model development. Researchers and developers could use the model to evaluate their models on a wide range of benchmarks and tasks, and then use the detailed feedback to inform their fine-tuning and training processes. Additionally, the model could be used to assess the safety and appropriateness of language model outputs, ensuring that they align with ethical guidelines and promote positive behavior.

[](#links-for-reference)Links for Reference
-------------------------------------------

*   **Homepage: In Progress**
*   **Repository:[https://github.com/prometheus-eval/prometheus-eval](https://github.com/prometheus-eval/prometheus-eval)**
*   **Paper:[https://arxiv.org/abs/2405.01535](https://arxiv.org/abs/2405.01535)**
*   **Point of Contact:[seungone@cmu.edu](mailto:seungone@cmu.edu)**

[](#tldr)TL;DR
==============

Prometheus 2 is an alternative of GPT-4 evaluation when doing fine-grained evaluation of an underlying LLM & a Reward model for Reinforcement Learning from Human Feedback (RLHF). [![plot](/prometheus-eval/prometheus-7b-v2.0/resolve/main/finegrained_eval.JPG)](/prometheus-eval/prometheus-7b-v2.0/blob/main/finegrained_eval.JPG)

Prometheus 2 is a language model using [Mistral-Instruct](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2) as a base model. It is fine-tuned on 100K feedback within the [Feedback Collection](https://huggingface.co/datasets/prometheus-eval/Feedback-Collection) and 200K feedback within the [Preference Collection](https://huggingface.co/datasets/prometheus-eval/Preference-Collection). It is also made by weight merging to support both absolute grading (direct assessment) and relative grading (pairwise ranking). The surprising thing is that we find weight merging also improves performance on each format.

[](#model-details)Model Details
===============================

[](#model-description)Model Description
---------------------------------------

*   **Model type:** Language model
*   **Language(s) (NLP):** English
*   **License:** Apache 2.0
*   **Related Models:** [All Prometheus Checkpoints](https://huggingface.co/models?search=prometheus-eval/Prometheus)
*   **Resources for more information:**
    *   [Research paper](https://arxiv.org/abs/2405.01535)
    *   [GitHub Repo](https://github.com/prometheus-eval/prometheus-eval)

Prometheus is trained with two different sizes (7B and 8x7B). You could check the 8x7B sized LM on [this page](https://huggingface.co/prometheus-eval/prometheus-2-8x7b-v2.0). Also, check out our dataset as well on [this page](https://huggingface.co/datasets/prometheus-eval/Feedback-Collection) and [this page](https://huggingface.co/datasets/prometheus-eval/Preference-Collection).

[](#prompt-format)Prompt Format
-------------------------------

We have made wrapper functions and classes to conveniently use Prometheus 2 at [our github repository](https://github.com/prometheus-eval/prometheus-eval). We highly recommend you use it!

However, if you just want to use the model for your use case, please refer to the prompt format below. Note that absolute grading and relative grading requires different prompt templates and system prompts.

### [](#absolute-grading-direct-assessment)Absolute Grading (Direct Assessment)

Prometheus requires 4 components in the input: An instruction, a response to evaluate, a score rubric, and a reference answer. You could refer to the prompt format below. You should fill in the instruction, response, reference answer, criteria description, and score description for score in range of 1 to 5.

Fix the components with {text} inside.

    ###Task Description:
    An instruction (might include an Input inside it), a response to evaluate, a reference answer that gets a score of 5, and a score rubric representing a evaluation criteria are given.
    1. Write a detailed feedback that assess the quality of the response strictly based on the given score rubric, not evaluating in general.
    2. After writing a feedback, write a score that is an integer between 1 and 5. You should refer to the score rubric.
    3. The output format should look as follows: \"Feedback: (write a feedback for criteria) [RESULT] (an integer number between 1 and 5)\"
    4. Please do not generate any other opening, closing, and explanations.
    
    ###The instruction to evaluate:
    {orig_instruction}
    
    ###Response to evaluate:
    {orig_response}
    
    ###Reference Answer (Score 5):
    {orig_reference_answer}
    
    ###Score Rubrics:
    [{orig_criteria}]
    Score 1: {orig_score1_description}
    Score 2: {orig_score2_description}
    Score 3: {orig_score3_description}
    Score 4: {orig_score4_description}
    Score 5: {orig_score5_description}
    
    ###Feedback: 
    

After this, you should apply the conversation template of Mistral (not applying it might lead to unexpected behaviors). You can find the conversation class at this [link](https://github.com/lm-sys/FastChat/blob/main/fastchat/conversation.py).

    conv = get_conv_template("mistral")
    conv.set_system_message("You are a fair judge assistant tasked with providing clear, objective feedback based on specific criteria, ensuring each assessment reflects the absolute standards set for performance.")
    conv.append_message(conv.roles[0], dialogs['instruction'])
    conv.append_message(conv.roles[1], None)
    prompt = conv.get_prompt()
    
    x = tokenizer(prompt,truncation=False)
    

As a result, a feedback and score decision will be generated, divided by a separating phrase `[RESULT]`

### [](#relative-grading-pairwise-ranking)Relative Grading (Pairwise Ranking)

Prometheus requires 4 components in the input: An instruction, 2 responses to evaluate, a score rubric, and a reference answer. You could refer to the prompt format below. You should fill in the instruction, 2 responses, reference answer, and criteria description.

Fix the components with {text} inside.

    ###Task Description:
    An instruction (might include an Input inside it), a response to evaluate, and a score rubric representing a evaluation criteria are given.
    1. Write a detailed feedback that assess the quality of two responses strictly based on the given score rubric, not evaluating in general.
    2. After writing a feedback, choose a better response between Response A and Response B. You should refer to the score rubric.
    3. The output format should look as follows: "Feedback: (write a feedback for criteria) [RESULT] (A or B)"
    4. Please do not generate any other opening, closing, and explanations.
    
    ###Instruction:
    {orig_instruction}
    
    ###Response A:
    {orig_response_A}
    
    ###Response B:
    {orig_response_B}
    
    ###Reference Answer:
    {orig_reference_answer}
    
    ###Score Rubric:
    {orig_criteria}
    
    ###Feedback: 
    

After this, you should apply the conversation template of Mistral (not applying it might lead to unexpected behaviors). You can find the conversation class at this [link](https://github.com/lm-sys/FastChat/blob/main/fastchat/conversation.py).

    conv = get_conv_template("mistral")
    conv.set_system_message("You are a fair judge assistant assigned to deliver insightful feedback that compares individual performances, highlighting how each stands relative to others within the same cohort.")
    conv.append_message(conv.roles[0], dialogs['instruction'])
    conv.append_message(conv.roles[1], None)
    prompt = conv.get_prompt()
    
    x = tokenizer(prompt,truncation=False)
    

As a result, a feedback and score decision will be generated, divided by a separating phrase `[RESULT]`

[](#license)License
-------------------

Feedback Collection, Preference Collection, and Prometheus 2 are subject to OpenAI's Terms of Use for the generated data. If you suspect any violations, please reach out to us.

[](#citation)Citation
=====================

If you find the following model helpful, please consider citing our paper!

**BibTeX:**

    @misc{kim2023prometheus,
        title={Prometheus: Inducing Fine-grained Evaluation Capability in Language Models},
        author={Seungone Kim and Jamin Shin and Yejin Cho and Joel Jang and Shayne Longpre and Hwaran Lee and Sangdoo Yun and Seongjin Shin and Sungdong Kim and James Thorne and Minjoon Seo},
        year={2023},
        eprint={2310.08491},
        archivePrefix={arXiv},
        primaryClass={cs.CL}
    }
    

    @misc{kim2024prometheus,
        title={Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models},
        author={Seungone Kim and Juyoung Suk and Shayne Longpre and Bill Yuchen Lin and Jamin Shin and Sean Welleck and Graham Neubig and Moontae Lee and Kyungjae Lee and Minjoon Seo},
        year={2024},
        eprint={2405.01535},
        archivePrefix={arXiv},
        primaryClass={cs.CL}
    }

## Model Overview

The `prometheus-7b-v2.0` is a language model developed by the team at [prometheus-eval](https://aimodels.fyi/creators/huggingFace/prometheus-eval). It is an alternative to GPT-4 for fine-grained evaluation of language models and as a reward model for Reinforcement Learning from Human Feedback (RLHF).

The model is based on the [Mistral-Instruct](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2) base model and has been fine-tuned on 100K feedback within the [Feedback Collection](https://huggingface.co/datasets/prometheus-eval/Feedback-Collection) and 200K feedback within the [Preference Collection](https://huggingface.co/datasets/prometheus-eval/Preference-Collection) datasets. It supports both absolute grading (direct assessment) and relative grading (pairwise ranking), and surprisingly, the weight merging process used to support both formats also improves performance on each.

Similar models include the [prometheus-13b-v1.0](https://aimodels.fyi/models/huggingFace/prometheus-13b-v10-prometheus-eval) and [prometheus-13b-v1.0](https://aimodels.fyi/models/huggingFace/prometheus-13b-v10-kaist-ai) models, which use different base models and training approaches.

## Model Inputs and Outputs

The `prometheus-7b-v2.0` model is a language model that can be used for text-to-text generation tasks. It requires different prompt formats for absolute grading (direct assessment) and relative grading (pairwise ranking).

### Inputs
- An instruction (might include an input)
- A response to evaluate
- A reference answer that gets a score of 5
- A score rubric representing an evaluation criteria

### Outputs
- A detailed feedback assessing the quality of the response based on the given score rubric
- An integer score between 1 and 5 referring to the score rubric

## Capabilities

The `prometheus-7b-v2.0` model excels at fine-grained evaluation of language models, outperforming GPT-3.5-Turbo and on par with GPT-4 on various benchmarks. It can be used to evaluate LLMs with customized criteria, such as child readability, cultural sensitivity, or creativity. Additionally, it can be used as a reward model for Reinforcement Learning from Human Feedback (RLHF).

## What can I use it for?

The `prometheus-7b-v2.0` model can be leveraged for a variety of applications, particularly in the field of language model evaluation and development. It can be used to assess the performance of other language models, providing detailed feedback and scoring to help improve their capabilities.

Additionally, the model can be employed as a reward model in Reinforcement Learning from Human Feedback (RLHF) workflows, helping to fine-tune language models to better align with human preferences and values.

## Things to try

One interesting aspect of the `prometheus-7b-v2.0` model is its ability to perform well on both absolute grading (direct assessment) and relative grading (pairwise ranking) tasks, despite the weight merging process used to support both formats. Experimenting with different prompts and evaluation criteria could lead to insights into how the model achieves this performance.

Another area to explore is the potential for the `prometheus-7b-v2.0` model to be used in conjunction with other language models, either as a specialized evaluation tool or as part of a more comprehensive model development workflow. Combining the capabilities of this model with other state-of-the-art language models could yield interesting and powerful applications.