[](#model-card-for-codefuse-codellama-34b)Model Card for CodeFuse-CodeLlama-34B
===============================================================================

[![logo](/codefuse-ai/CodeFuse-CodeLlama-34B/resolve/main/LOGO.png)](/codefuse-ai/CodeFuse-CodeLlama-34B/blob/main/LOGO.png)

[\[\]](#chinese) [\[English\]](#english)

[](#model-description)Model Description
---------------------------------------

CodeFuse-CodeLlama-34B is a 34B Code-LLM finetuned by QLoRA of multiple code tasks600k instrunctions/answers on the base model CodeLlama-34b-Python. The context length of finetuning is 4K while it is able to be finetuned by 16k context if necessary.  

[](#news-and-updates)News and Updates
-------------------------------------

 2023-09-26 We are pleased to announce the release of the [4-bit quantized version](https://huggingface.co/codefuse-ai/CodeFuse-CodeLlama-34B-4bits) of CodeFuse-CodeLlama-34B. Despite the quantization process, the model still achieves a remarkable 73.8% accuracy (greedy decoding) on the HumanEval pass@1 metric.

 2023-09-11 CodeFuse-CodeLlama34B has achieved 74.4% of pass@1 (greedy decoding) on HumanEval, which is SOTA results for openspurced LLMs at present.

  

[](#code-community)Code Community
---------------------------------

**Homepage**:  [https://github.com/codefuse-ai](https://github.com/codefuse-ai) (**Please give us your support with a Star + Fork + Watch**)

*   If you wish to fine-tune the model yourself, you can visit [MFTCoder](https://github.com/codefuse-ai/MFTCoder)
    
*   If you wish to deploy the model yourself, you can visit [FasterTransformer4CodeFuse](https://github.com/codefuse-ai/FasterTransformer4CodeFuse)
    
*   If you wish to see a demo of the model, you can visit [CodeFuse Demo](https://github.com/codefuse-ai/codefuse)
    

[](#performance)Performance
---------------------------

Model

HumanEval(pass@1)

Date

**CodeFuse-CodeLlama-34B**

**74.4%**

2023.9

WizardCoder-Python-34B-V1.0

73.2%

2023.8

GPT-4(zero-shot)

67.0%

2023.3

PanGu-Coder2 15B

61.6%

2023.8

CodeLlama-34b-Python

53.7%

2023.8

CodeLlama-34b

48.8%

2023.8

GPT-3.5(zero-shot)

48.1%

2022.11

OctoCoder

46.2%

2023.8

StarCoder-15B

33.6%

2023.5

LLaMA 2 70B(zero-shot)

29.9%

2023.7

  

[](#requirements)Requirements
-----------------------------

*   python>=3.8
*   pytorch>=2.0.0
*   transformers==4.32.0
*   Sentencepiece
*   CUDA 11.4  
    

[](#inference-string-format)Inference String Format
---------------------------------------------------

The inference string is a concatenated string formed by combining conversation data(system, human and bot contents) in the training data format. It is used as input during the inference process. Here is an example format of the concatenated string:

    """
    <|role_start|>system<|role_end|>System instruction
    <|role_start|>human<|role_end|>Human 1st round input
    <|role_start|>bot<|role_end|>Bot 1st round output</s>
    <|role_start|>human<|role_end|>Human 2nd round input
    <|role_start|>bot<|role_end|>Bot 2nd round output</s>
    ...
    ...
    ...
    <|role_start|>human<|role_end|>Human nth round input
    <|role_start|>bot<|role_end|>{Bot output to be genreated}</s>
    """
    

When applying inference, you always make your input string end with "<|role\_start|>bot<|role\_end|>" to ask the model generating answers.

[](#quickstart)Quickstart
-------------------------

    pip install -r requirements.txt
    

    import torch
    from transformers import (
        AutoTokenizer, 
        AutoModelForCausalLM,
    )
    tokenizer = AutoTokenizer.from_pretrained(mode_name_or_path, trust_remote_code=True, use_fast=False, legacy=False)
    tokenizer.padding_side = "left"
    tokenizer.pad_token_id = tokenizer.convert_tokens_to_ids("<unk>")
    tokenizer.eos_token_id = tokenizer.convert_tokens_to_ids("</s>")
    # try 4bit loading if cuda memory not enough
    model = AutoModelForCausalLM.from_pretrained(mode_name_or_path,
                                                 trust_remote_code=True,
                                                 load_in_4bit=False,
                                                 device_map="auto",
                                                 torch_dtype=torch.bfloat16)
    model.eval()
    
    HUMAN_ROLE_START_TAG = "<|role_start|>human<|role_end|>"
    BOT_ROLE_START_TAG = "<|role_start|>bot<|role_end|>"
    
    text = f"{HUMAN_ROLE_START_TAG}write a python function of quick sort.{BOT_ROLE_START_TAG}" 
    inputs = tokenizer(text, return_tensors='pt', padding=True, add_special_tokens=False).to("cuda")
    outputs = model.generate(
            inputs=inputs["input_ids"],
            attention_mask=inputs["attention_mask"],
            max_new_tokens=512,
            top_p=0.95,
            temperature=0.1,
            do_sample=True,
            eos_token_id=tokenizer.eos_token_id,
            pad_token_id=tokenizer.pad_token_id
        )
    gen_text = tokenizer.batch_decode(outputs[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
    print(gen_text)
    

[](#md5)MD5
-----------

We notice that the file may be corrupted during transfer process. Please check MD5 value before use.

Model File

MD5 Value

pytorch\_model-00001-of-00007.bin

8d544b1bcb3449934184d4141137329c

pytorch\_model-00002-of-00007.bin

9d5dbb30911e48a42fb6d0fcabb322a4

pytorch\_model-00003-of-00007.bin

b0d4aecee0457d9332005a187e1fffed

pytorch\_model-00004-of-00007.bin

5c7e002de5eab77d0194a2b0f6de0c24

pytorch\_model-00005-of-00007.bin

d22a511aa26b5b17117b665a877490ab

pytorch\_model-00006-of-00007.bin

a5c28ac277fac07d16dd66537e54d109

pytorch\_model-00007-of-00007.bin

a967e2c6195477b7407089c0bffa2d53

[](#citation)Citation
---------------------

If you find our [work](https://arxiv.org/abs/2311.02303) useful or helpful for your R&D works, please feel free to cite our paper as below.

    @article{mftcoder2023,
          title={MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning}, 
          author={Bingchang Liu and Chaoyu Chen and Cong Liao and Zi Gong and Huan Wang and Zhichao Lei and Ming Liang and Dajun Chen and Min Shen and Hailian Zhou and Hang Yu and Jianguo Li},
          year={2023},
          journal={arXiv preprint arXiv},
          archivePrefix={arXiv},
          eprint={2311.02303}
    }
    

[](#)
-------------

CodeFuse-CodeLlama34B-MFT QLoRACodeLlama-34b-Python4k16k  

[](#)
---------

 CodeFuse-CodeLlama34B-MFTHumanEval pass@174.4%, SOTA

  

[](#)
-------------

****  [https://github.com/codefuse-ai](https://github.com/codefuse-ai) ** Star+ Fork + Watch**

*    [MFTCoder](https://github.com/codefuse-ai/MFTCoder)
    
*    [FasterTransformer4CodeFuse](https://github.com/codefuse-ai/FasterTransformer4CodeFuse)
    
*    [CodeFuse Demo](https://github.com/codefuse-ai/codefuse)
    

[](#)()
-------------------



HumanEval(pass@1)



**CodeFuse-CodeLlama-34B**

**74.4%**

2023.9

WizardCoder-Python-34B-V1.0

73.2%

2023.8

GPT-4(zero-shot)

67.0%

2023.3

PanGu-Coder2 15B

61.6%

2023.8

CodeLlama-34b-Python

53.7%

2023.8

CodeLlama-34b

48.8%

2023.8

GPT-3.5(zero-shot)

48.1%

2022.11

OctoCoder

46.2%

2023.8

StarCoder-15B

33.6%

2023.5

LLaMA 2 70B(zero-shot)

29.9%

2023.7

  

[](#requirements-1)Requirements
-------------------------------

*   python>=3.8
*   pytorch>=2.0.0
*   transformers==4.32.0
*   CUDA 11.4  
    

[](#)
-----------------

prompt

    """
    <|role_start|>system<|role_end|>System
    <|role_start|>human<|role_end|>1
    <|role_start|>bot<|role_end|>1</s>
    <|role_start|>human<|role_end|>2
    <|role_start|>bot<|role_end|>2</s>
    ...
    ...
    ...
    <|role_start|>human<|role_end|>n
    <|role_start|>bot<|role_end|>{}</s>
    """
    

prompt"<|role\_start|>bot<|role\_end|>"

[](#)
-------------

    from transformers import (
        AutoTokenizer, 
        AutoModelForCausalLM,
    )
    tokenizer = AutoTokenizer.from_pretrained(mode_name_or_path, trust_remote_code=True, use_fast=False, legacy=False)
    tokenizer.padding_side = "left"
    tokenizer.pad_token_id = tokenizer.convert_tokens_to_ids("<unk>")
    tokenizer.eos_token_id = tokenizer.convert_tokens_to_ids("</s>")
    # 
    model = AutoModelForCausalLM.from_pretrained(mode_name_or_path,
                                                 trust_remote_code=True,
                                                 load_in_4bit=False,
                                                 device_map="auto",
                                                 torch_dtype=torch.bfloat16)
    model.eval()
    
    HUMAN_ROLE_START_TAG = "<|role_start|>human<|role_end|>"
    BOT_ROLE_START_TAG = "<|role_start|>bot<|role_end|>"
    
    text = f"{HUMAN_ROLE_START_TAG}C++n{BOT_ROLE_START_TAG}" 
    inputs = tokenizer(text, return_tensors='pt', padding=True, add_special_tokens=False).to("cuda")
    outputs = model.generate(
            inputs=inputs["input_ids"],
            attention_mask=inputs["attention_mask"],
            max_new_tokens=512,
            top_p=0.95,
            temperature=0.1,
            do_sample=True,
            eos_token_id=tokenizer.eos_token_id,
            pad_token_id=tokenizer.pad_token_id
        )
    gen_text = tokenizer.batch_decode(outputs[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
    print(gen_text)
    

[](#md5-1)MD5
-------------

MD5



MD5

pytorch\_model-00001-of-00007.bin

8d544b1bcb3449934184d4141137329c

pytorch\_model-00002-of-00007.bin

9d5dbb30911e48a42fb6d0fcabb322a4

pytorch\_model-00003-of-00007.bin

b0d4aecee0457d9332005a187e1fffed

pytorch\_model-00004-of-00007.bin

5c7e002de5eab77d0194a2b0f6de0c24

pytorch\_model-00005-of-00007.bin

d22a511aa26b5b17117b665a877490ab

pytorch\_model-00006-of-00007.bin

a5c28ac277fac07d16dd66537e54d109

pytorch\_model-00007-of-00007.bin

a967e2c6195477b7407089c0bffa2d53

## Model overview

The `CodeFuse-CodeLlama-34B` is a 34 billion parameter code-focused large language model (LLM) developed by [codefuse-ai](https://aimodels.fyi/creators/huggingFace/codefuse-ai). This model is a fine-tuned version of the [CodeLlama-34b-Python](https://aimodels.fyi/models/huggingFace/codellama-7b-python-hf-codellama) model, trained on 600k instructions and answers across various programming tasks. It achieves state-of-the-art performance of 74.4% pass@1 on the HumanEval benchmark, outperforming other open-source models like [WizardCoder-Python-34B-V1.0](https://aimodels.fyi/models/huggingFace/codellama-7b-hf-codellama) and [GPT-4](https://aimodels.fyi/models/huggingFace/codellama-7b-hf-codellama) on this metric.

## Model inputs and outputs

### Inputs
- The model accepts a concatenated string of conversation data in a specific format, including system instructions, human messages, and bot responses.

### Outputs
- The model generates text continuations in response to the input prompt.

## Capabilities

The `CodeFuse-CodeLlama-34B` model is highly capable at a variety of code-related tasks, including code completion, infilling, and following programming instructions. It demonstrates strong performance on benchmarks like HumanEval, indicating its ability to synthesize and understand code. The model is also a Python specialist, making it well-suited for tasks involving the Python programming language.

## What can I use it for?

The `CodeFuse-CodeLlama-34B` model can be used for a wide range of applications that involve code generation, understanding, and assistance. Some potential use cases include:

- Building intelligent code editors or IDEs that can provide advanced code completion and suggestion capabilities.
- Developing chatbots or virtual assistants that can help programmers with coding tasks, answer questions, and provide code examples.
- Automating the generation of boilerplate code or repetitive programming tasks.
- Enhancing existing ML/AI systems with code-generation capabilities, such as automated machine learning pipelines or data processing workflows.

## Things to try

One interesting thing to try with the `CodeFuse-CodeLlama-34B` model is to provide it with open-ended programming challenges or tasks, and observe how it approaches and solves them. The model's strong performance on benchmarks like HumanEval suggests it may be able to tackle a variety of programming problems in creative and novel ways. Developers could also experiment with fine-tuning or adapting the model for their specific use cases, leveraging the provided tools and resources from the [codefuse-ai](https://aimodels.fyi/creators/huggingFace/codefuse-ai) team.

[](#model-card-for-codefuse-deepseek-33b)Model Card for CodeFuse-DeepSeek-33B
=============================================================================

[![logo](/codefuse-ai/CodeFuse-DeepSeek-33B/resolve/main/LOGO.jpg)](/codefuse-ai/CodeFuse-DeepSeek-33B/blob/main/LOGO.jpg)

[\[\]](#chinese) [\[English\]](#english)

[](#model-description)Model Description
---------------------------------------

CodeFuse-DeepSeek-33B is a 33B Code-LLM finetuned by QLoRA on multiple code-related tasks on the base model DeepSeek-Coder-33B.

  

[](#news-and-updates)News and Updates
-------------------------------------

 2024-01-12 CodeFuse-DeepSeek-33B has been released, achieving a pass@1 (greedy decoding) score of 78.65% on HumanEval.

 2024-01-12 CodeFuse-Mixtral-8x7B has been released, achieving a pass@1 (greedy decoding) score of 56.1% on HumanEval, which is a 15% increase compared to Mixtral-8x7b's 40%.

 2023-11-10 CodeFuse-CodeGeeX2-6B has been released, achieving a pass@1 (greedy decoding) score of 45.12% on HumanEval, which is a 9.22% increase compared to CodeGeeX2 35.9%.

 2023-10-20 CodeFuse-QWen-14B technical documentation has been released. For those interested, please refer to the CodeFuse article on our WeChat official account via the provided link.([https://mp.weixin.qq.com/s/PCQPkvbvfxSPzsqjOILCDw](https://mp.weixin.qq.com/s/PCQPkvbvfxSPzsqjOILCDw))

 2023-10-16 CodeFuse-QWen-14B has been released, achieving a pass@1 (greedy decoding) score of 48.78% on HumanEval, which is a 16% increase compared to Qwen-14b's 32.3%.

 2023-09-27 CodeFuse-StarCoder-15B has been released, achieving a pass@1 (greedy decoding) score of 54.9% on HumanEval, which is a 21% increase compared to StarCoder's 33.6%.

 2023-09-26 We are pleased to announce the release of the 4-bit quantized version of CodeFuse-CodeLlama-34B. Despite the quantization process, the model still achieves a remarkable 73.8% accuracy (greedy decoding) on the HumanEval pass@1 metric.

 2023-09-11 CodeFuse-CodeLlama-34B has achieved 74.4% of pass@1 (greedy decoding) on HumanEval, which is SOTA results for openspurced LLMs at present.

  

[](#code-community)Code Community
---------------------------------

**Homepage**:  [https://github.com/codefuse-ai](https://github.com/codefuse-ai) (**Please give us your support with a Star + Fork + Watch**)

*   If you wish to fine-tune the model yourself, you can visit [MFTCoder](https://github.com/codefuse-ai/MFTCoder)
    
*   If you wish to see a demo of the model, you can visit [CodeFuse Demo](https://github.com/codefuse-ai/codefuse)
    

  

[](#performance)Performance
---------------------------

### [](#code)Code

Model

HumanEval(pass@1)

Date

**CodeFuse-DeepSeek-33B**

**78.65%**

2024.01

**CodeFuse-Mixtral-8x7B**

**56.10%**

2024.01

**CodeFuse-CodeLlama-34B**

74.4%

2023.9

**CodeFuse-CodeLlama-34B-4bits**

73.8%

2023.9

**CodeFuse-StarCoder-15B**

54.9%

2023.9

**CodeFuse-QWen-14B**

48.78%

2023.10

**CodeFuse-CodeGeeX2-6B**

45.12%

2023.11

WizardCoder-Python-34B-V1.0

73.2%

2023.8

GPT-4(zero-shot)

67.0%

2023.3

PanGu-Coder2 15B

61.6%

2023.8

CodeLlama-34b-Python

53.7%

2023.8

CodeLlama-34b

48.8%

2023.8

GPT-3.5(zero-shot)

48.1%

2022.11

OctoCoder

46.2%

2023.8

StarCoder-15B

33.6%

2023.5

Qwen-14b

32.3%

2023.10

### [](#nlp)NLP

[![NLP Performance Radar](/codefuse-ai/CodeFuse-DeepSeek-33B/resolve/main/codefuse-deepseek-33b-nlp.png)](/codefuse-ai/CodeFuse-DeepSeek-33B/blob/main/codefuse-deepseek-33b-nlp.png)

  

[](#requirements)Requirements
-----------------------------

*   python>=3.8
*   pytorch>=2.0.0
*   transformers>=4.33.2
*   Sentencepiece
*   CUDA 11.4  
    

[](#inference-string-format)Inference String Format
---------------------------------------------------

The inference string is a concatenated string formed by combining conversation data(system, human and bot contents) in the training data format. It is used as input during the inference process. Here are examples of prompts used to request the model:

**Multi-Round with System Prompt:**

    """
    <s>system
    System instruction
    <s>human
    Human 1st round input
    <s>bot
    Bot 1st round output<endofsentence>
    <s>human
    Human 2nd round input
    <s>bot
    Bot 2nd round output<endofsentence>
    ...
    ...
    ...
    <s>human
    Human nth round input
    <s>bot
    """
    

**Single-Round without System Prompt:**

    """
    <s>human
    User prompt...
    <s>bot
    
    """
    

In this format, the system section is optional and the conversation can be either single-turn or multi-turn. When applying inference, you always make your input string end with "<s>bot" to ask the model generating answers.

For example, the format used to infer HumanEval is like the following:

    <s>human
    # language: Python
    from typing import List
    def separate_paren_groups(paren_string: str) -> List[str]:
        """ Input to this function is a string containing multiple groups of nested parentheses. Your goal is to
        separate those group into separate strings and return the list of those.
        Separate groups are balanced (each open brace is properly closed) and not nested within each other
        Ignore any spaces in the input string.
        >>> separate_paren_groups('( ) (( )) (( )( ))')
        ['()', '(())', '(()())']
        """
    <s>bot
    

Specifically, we also add the Programming Language Tag (e.g. "`# language: Python`" for Python) used by CodeGeex models.

[](#quickstart)Quickstart
-------------------------

    import torch
    from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
    
    model_dir = "codefuse-ai/CodeFuse-DeepSeek-33B"
    
    def load_model_tokenizer(model_path):
        tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
        tokenizer.eos_token = "<endofsentence>"
        tokenizer.pad_token = "<endofsentence>"
        tokenizer.eos_token_id = tokenizer.convert_tokens_to_ids(tokenizer.eos_token)
        tokenizer.pad_token_id = tokenizer.convert_tokens_to_ids(tokenizer.pad_token)
        tokenizer.padding_side = "left"
        
        model = AutoModelForCausalLM.from_pretrained(model_path, device_map='auto',torch_dtype=torch.bfloat16, trust_remote_code=True)
        return model, tokenizer
    
    
    HUMAN_ROLE_START_TAG = "<s>human\n"
    BOT_ROLE_START_TAG = "<s>bot\n"
    
    text_list = [f'{HUMAN_ROLE_START_TAG}Write a QuickSort program\n#Python\n{BOT_ROLE_START_TAG}']
    
    model, tokenizer = load_model_tokenizer(model_dir)
    inputs = tokenizer(text_list, return_tensors='pt', padding=True, add_special_tokens=False).to('cuda')
    input_ids = inputs["input_ids"]
    attention_mask = inputs["attention_mask"]
    generation_config = GenerationConfig(
            eos_token_id=tokenizer.eos_token_id,
            pad_token_id=tokenizer.pad_token_id,
            temperature=0.1,
            max_new_tokens=512,
            num_return_sequences=1,
            num_beams=1,
            top_p=0.95,
            do_sample=False
    )
    outputs = model.generate(
            inputs= input_ids,
            attention_mask=attention_mask,
            **generation_config.to_dict()
    )
    gen_text = tokenizer.batch_decode(outputs[:, input_ids.shape[1]:], skip_special_tokens=True)
    print(gen_text[0])
    

[](#)
-------------

CodeFuse-DeepSeek-33B QLoRADeepSeek-Coder-33B  

[](#)
---------

 2024-01-12 CodeFuse-DeepSeek-33BHumanEval pass@178.65% ()

 2023-11-10 CodeFuse-CodeGeeX2-6BHumanEval pass@1(greedy decoding)48.12%, CodeGeeX29.22%HumanEval

 2023-10-20 CodeFuse-QWen-14BCodeFuse[https://mp.weixin.qq.com/s/PCQPkvbvfxSPzsqjOILCDw](https://mp.weixin.qq.com/s/PCQPkvbvfxSPzsqjOILCDw)

 2023-10-16CodeFuse-QWen-14BHumanEval pass@1(greedy decoding)48.78%, Qwen-14b16%HumanEval

 2023-09-27CodeFuse-StarCoder-15BHumanEval pass@1(greedy decoding)54.9%, StarCoder21%HumanEval

 2023-09-26 [CodeFuse-CodeLlama-34B 4bits](https://modelscope.cn/models/codefuse-ai/CodeFuse-CodeLlama-34B-4bits/summary)HumanEval pass@173.8% ()

 2023-09-11 [CodeFuse-CodeLlama-34B](https://modelscope.cn/models/codefuse-ai/CodeFuse-CodeLlama-34B/summary)HumanEval pass@174.4% (), SOTA

  

[](#)
-------------

****  [https://github.com/codefuse-ai](https://github.com/codefuse-ai) **Star + Fork + Watch**

*    [MFTCoder](https://github.com/codefuse-ai/MFTCoder)
    
*    [CodeFuse Demo](https://github.com/codefuse-ai/codefuse)
    

  

[](#)
-------------

### [](#)



HumanEval(pass@1)



**CodeFuse-CodeLlama-34B**

74.4%

2023.9

**CodeFuse-CodeLlama-34B-4bits**

73.8%

2023.9

WizardCoder-Python-34B-V1.0

73.2%

2023.8

GPT-4(zero-shot)

67.0%

2023.3

PanGu-Coder2 15B

61.6%

2023.8

CodeLlama-34b-Python

53.7%

2023.8

CodeLlama-34b

48.8%

2023.8

GPT-3.5(zero-shot)

48.1%

2022.11

OctoCoder

46.2%

2023.8

StarCoder-15B

33.6%

2023.5

Qwen-14b

32.3%

2023.10

**CodeFuse-StarCoder-15B**

54.9%

2023.9

**CodeFuse-QWen-14B**

48.78%

2023.8

**CodeFuse-CodeGeeX2-6B**

45.12%

2023.11

**CodeFuse-DeepSeek-33B**.

**78.65%**

2024.01

### [](#nlp-1)NLP

[![NLP Performance Radar](/codefuse-ai/CodeFuse-DeepSeek-33B/resolve/main/codefuse-deepseek-33b-nlp.png)](/codefuse-ai/CodeFuse-DeepSeek-33B/blob/main/codefuse-deepseek-33b-nlp.png)

[](#requirements-1)Requirements
-------------------------------

*   python>=3.8
*   pytorch>=2.0.0
*   transformers>=4.33.2
*   Sentencepiece
*   CUDA 11.4  
    

[](#)
-----------------

prompt. 

**System:**

    """
    <s>system
    System instruction
    <s>human
    Human 1st round input
    <s>bot
    Bot 1st round output<endofsentence>
    <s>human
    Human 2nd round input
    <s>bot
    Bot 2nd round output<endofsentence>
    ...
    ...
    ...
    <s>human
    Human nth round input
    <s>bot
    """
    

**System:**

    """
    <s>human
    User prompt...
    <s>bot
    
    """
    

Systemprompt"<s>bot\\n"

HumanEval

    <s>human
    # language: Python
    from typing import List
    def separate_paren_groups(paren_string: str) -> List[str]:
        """ Input to this function is a string containing multiple groups of nested parentheses. Your goal is to
        separate those group into separate strings and return the list of those.
        Separate groups are balanced (each open brace is properly closed) and not nested within each other
        Ignore any spaces in the input string.
        >>> separate_paren_groups('( ) (( )) (( )( ))')
        ['()', '(())', '(()())']
        """
    <s>bot
    

CodeGeeXPython"`# language: Python`"

[](#)
-------------

    import torch
    from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
    
    model_dir = "codefuse-ai/CodeFuse-DeepSeek-33B"
    
    def load_model_tokenizer(model_path):
        tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
        tokenizer.eos_token = "<endofsentence>"
        tokenizer.pad_token = "<endofsentence>"
        tokenizer.eos_token_id = tokenizer.convert_tokens_to_ids(tokenizer.eos_token)
        tokenizer.pad_token_id = tokenizer.convert_tokens_to_ids(tokenizer.pad_token)
        tokenizer.padding_side = "left"
        
        model = AutoModelForCausalLM.from_pretrained(model_path, device_map='auto',torch_dtype=torch.bfloat16, trust_remote_code=True)
        return model, tokenizer
    
    HUMAN_ROLE_START_TAG = "<s>human\n"
    BOT_ROLE_START_TAG = "<s>bot\n"
    
    text_list = [f'{HUMAN_ROLE_START_TAG}\n#Python\n{BOT_ROLE_START_TAG}']
    
    model, tokenizer = load_model_tokenizer(model_dir)
    inputs = tokenizer(text_list, return_tensors='pt', padding=True, add_special_tokens=False).to('cuda')
    input_ids = inputs["input_ids"]
    attention_mask = inputs["attention_mask"]
    generation_config = GenerationConfig(
            eos_token_id=tokenizer.eos_token_id,
            pad_token_id=tokenizer.pad_token_id,
            temperature=0.2,
            max_new_tokens=512,
            num_return_sequences=1,
            num_beams=1,
            top_p=0.95,
            do_sample=False
    )
    outputs = model.generate(
            inputs= input_ids,
            attention_mask=attention_mask,
            **generation_config.to_dict()
    )
    gen_text = tokenizer.batch_decode(outputs[:, input_ids.shape[1]:], skip_special_tokens=True)
    print(gen_text[0])

## Model overview

The `CodeFuse-DeepSeek-33B` is a 33B parameter Code-LLM (Large Language Model) that has been fine-tuned by QLoRA (Quantized Low-Rank Adaptation) on multiple code-related tasks using the base model [DeepSeek-Coder-33B](https://aimodels.fyi/models/huggingFace/deepseek-coder-33b-base-deepseek-ai). The model has achieved a pass@1 (greedy decoding) score of 78.65% on the HumanEval benchmark, showcasing its strong performance in generating high-quality code. 

The model is part of the CodeFuse suite of code-focused AI models developed by the [codefuse-ai](https://aimodels.fyi/creators/huggingFace/codefuse-ai) team. Similar models in the CodeFuse lineup include [CodeFuse-Mixtral-8x7B](https://aimodels.fyi/models/huggingFace/codefuse-mixtral-8x7b-codefuse-ai), [CodeFuse-CodeGeeX2-6B](https://aimodels.fyi/models/huggingFace/codefuse-codegex2-6b-codefuse-ai), and [CodeFuse-QWen-14B](https://aimodels.fyi/models/huggingFace/codefuse-qwen-14b-codefuse-ai), all of which have shown significant improvements over their base models in code generation capabilities.

## Model inputs and outputs

### Inputs
- **Code-related prompts**: The model takes in text-based prompts related to coding tasks, such as algorithm descriptions, function stubs, or high-level specifications.
- **Natural language instructions**: The model can also accept natural language instructions for tasks like code generation, code completion, and code explanation.

### Outputs
- **Generated code**: The primary output of the `CodeFuse-DeepSeek-33B` model is high-quality, contextually relevant code in a variety of programming languages.
- **Explanations and insights**: The model can also generate natural language explanations and insights about the code, such as describing the purpose, functionality, or potential improvements.

## Capabilities

The `CodeFuse-DeepSeek-33B` model has demonstrated state-of-the-art performance on code generation tasks, outperforming many other open-source language models. It is particularly adept at tasks like algorithm implementation, code completion, and code refactoring. The model's deep understanding of programming concepts and syntax allows it to generate code that is both functionally correct and idiomatic.

## What can I use it for?

The `CodeFuse-DeepSeek-33B` model can be leveraged for a wide range of applications in the software development and AI research domains. Some potential use cases include:

- **Automated programming assistance**: Integrate the model into IDEs, code editors, or developer tools to assist programmers with tasks like code generation, code completion, and code explanation.
- **AI-powered coding tutorials**: Create interactive coding tutorials or educational content that leverage the model's ability to generate code and provide explanations.
- **Accelerated prototyping and experimentation**: Use the model to quickly generate code prototypes or explore different algorithmic approaches, speeding up the R&D process.
- **Intelligent code refactoring**: Leverage the model's understanding of code structure and semantics to suggest refactoring opportunities and optimize code quality.

## Things to try

To get the most out of the `CodeFuse-DeepSeek-33B` model, you can experiment with providing the model with detailed prompts or instructions that capture the specific requirements of your coding tasks. Additionally, you can explore fine-tuning or adapting the model further on your own dataset or use case to further improve its performance.